Accurate fetal movement (FM) detection is essential for assessing prenatal health, as abnormal movement patterns can indicate underlying complications such as placental dysfunction or fetal distress. Traditional methods, including maternal perception and cardiotocography (CTG), suffer from subjectivity and limited accuracy. To address these challenges, we propose Contrastive Ultrasound Video Representation Learning (CURL), a novel self-supervised learning framework for FM detection from extended fetal ultrasound video recordings. Our approach leverages a dual-contrastive loss, incorporating both spatial and temporal contrastive learning, to learn robust motion representations. Additionally, we introduce a task-specific sampling strategy, ensuring the effective separation of movement and non-movement segments during self-supervised training, while enabling flexible inference on arbitrarily long ultrasound recordings through a probabilistic fine-tuning approach. Evaluated on an in-house dataset of 92 subjects, each with 30-minute ultrasound sessions, CURL achieves a sensitivity of 78.01
BACKGROUND AND PURPOSE:Predicting the final location and volume of lesions in acute ischemic stroke is crucial for clinical management. While CTP is routinely used for estimating lesion outcomes, conventional threshold-based methods have limitations. We developed specialized outcome-prediction deep learning models that predict infarct core in successful reperfusion cases and the combined core-penumbra region in unsuccessful reperfusion cases. MATERIALS AND METHODS:We developed single-modal and multimodal deep learning models using CTP parameter maps to predict the final infarct lesion on follow-up DWI. Using a multicenter data set from multiple sites, we developed deep learning models and evaluated them separately for patients with complete recanalization (successful reperfusion [CR], n = 350) and no recanalization (unsuccessful reperfusion [NR], n = 138) after treatment. The CR model was designed to predict the infarct core region, while the NR model predicted the expanded, hypoperfused tissue encompassing both the core and penumbra regions. Five-fold cross-validation was performed for robust evaluation. RESULTS:The multimodal 3D nnU-Net model demonstrated superior performance, achieving mean Dice scores of 35.36% in patients with CR and 50.22% in those with NR. This model substantially outperformed the current clinically used method, providing more accurate outcome estimates than the conventional single-technique threshold-based measures, which yielded Dice scores of 15.73% and 39.71% for CR and NR groups, respectively. CONCLUSIONS:Our approach offered both successful reperfusion and unsuccessful reperfusion estimations for potential treatment outcomes, enabling clinicians to better evaluate treatment eligibility for reperfusion therapies and assess potential treatment benefits. This advancement facilitates more personalized treatment recommendations and has the potential to substantially enhance clinical decision-making in acute ischemic stroke management by providing more accurate tissue outcome predictions than conventional single-technique threshold-based approaches.
Skeleton-based human activity recognition (HAR) has achieved strong empirical performance, yet most existing models remain black boxes and difficult to interpret. In this work, we introduce a neurosymbolic formulation of skeleton-based HAR that reframes action recognition as concept-driven first-order logical reasoning over motion primitives. Our framework bridges representation learning and symbolic inference by grounding first-order logic predicates in learnable spatial and temporal motion concepts. Specifically, we employ a standard spatio-temporal skeleton encoder to extract latent motion representations, which are then mapped to interpretable concept predicates via a spatio-temporal concept decoder that explicitly separates pose-centric and dynamics-centric abstractions. These concept predicates are composed through differentiable first-order logic layers, enabling the model to learn human-readable logical rules that govern action semantics. To impose semantic structure on the learned concepts, we align skeleton representations with LLM-derived descriptions of atomic motion primitives, establishing a shared conceptual space for perception and reasoning. Extensive experiments on NTU RGB+D 60/120 and NW-UCLA demonstrate that our approach achieves competitive recognition performance while providing explicit, interpretable explanations grounded in logical structure. Our results highlight neurosymbolic reasoning as an effective paradigm for interpretable spatio-temporal action understanding.
Interpretability is essential for trustworthy medical image diagnosis. However, existing concept-driven interpretable methods have key limitations: Concept Bottleneck Models (CBMs) require scoring all predefined concepts at inference time and for manual intervention, imposing a substantial burden on clinicians, while rationale-based generative approaches often select concepts by class discriminability, which can drift from diagnostic ontologies. To address these issues, we propose Neuro-Symbolic Rule Distillation (NeRD), a framework that produces efficient, ontology-grounded reasoning chains that are sufficient yet non-redundant, without manually crafting diagnostic rules. Experiments on two skin datasets demonstrate strong diagnostic performance and interpretability, and blinded expert evaluation confirms the clinical plausibility of NeRD rationales. Our method further enables a first expert-in-the-loop study for Multimodal Chain-of-Thought-based diagnosis, achieving efficient and effective concept-level intervention.
Accurate seizure detection is vital for effective epilepsy management, often relying on multimodal data, such as video, EEG, and ECG, to capture comprehensive diagnostic information. However, integrating diverse modalities poses challenges, including handling missing data, aligning disparate formats, and achieving seamless fusion. This study focuses on utilizing non-invasive, privacy-preserving modalities, optical flow and pose (body, face, hand) extracted from video, alongside ECG recordings, to differentiate between Generalized Tonic-Clonic Seizures (GTCS) and Psychogenic NonEpileptic Seizures (PNES). To address these challenges, we propose two novel approaches: the Pose Attention Graph (PAG), a symmetric graph model for analyzing patient movement, and the Modalities Relational Graph (MRG) for dynamic coordination of modality interactions. Evaluated on our in-house multimodal (ECG+Video) dataset, the model achieves an 80.64for seizure detection using 10-second snippets, 80.11Furthermore, it maintains robust performance with a precision of 77.72even with missing modalities, addressing key challenges in multimodal seizure detection. Our code is available at: GitHub .
Epilepsy is a chronic neurological disorder requiring multi-faceted management, including seizure detection, syndrome diagnosis, prognostication, antiseizure medication recommendation, epileptogenic zone localization, and surgical outcome prediction. Although numerous deep learning approaches have been developed for individual tasks, these models are typically siloed and modality-specific (e.g., EEG for seizure detection, MRI for localization), failing to reflect the multidisciplinary nature of real-world epilepsy care, where epileptologists, neuroradiologists, neurosurgeons, neuropsychologists and neuropsychiatrists jointly interpret heterogeneous evidence to guide decisions. In this work, we propose a clinical guideline-grounded hybrid multi-agent framework for holistic epilepsy management. Heterogeneous patient data is processed through modality-specific discriminative and generative models, where textual interpretations from generative agents are combined with structured predictions from discriminative models as auxiliary guidance. This aggregated evidence is passed to a central orchestrating agent grounded in international epilepsy guidelines, which evaluates multi-modal findings within structured clinical pathways and performs iterative cross-agent coordination for evidence-informed decision-making. We evaluate our framework across two datasets spanning six epilepsy management tasks and also introduce a publicly available multi-modal, multi-task epilepsy benchmark. Results demonstrate that integrating discriminative evidence with guideline-grounded generative coordination yields more reliable and comprehensive decisions compared to conventional LLM-based and task-specific baselines. Our dataset and code is available at https://github.com/khoapham154/epi_guide.git}. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This study did not receive any funding ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: This study was approved by the Alfred Hospital Ethics Committee (Ethics number 437/21). The Ethics Committee is constituted according to NHMRC guidelines and reports to the Alfred Health Executive Committee, which in turn reports to the Alfred Health Board. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The data using for training model can be found at: https://github.com/khoapham154/epi_guide.git
Personalized Large Language Models (PLLMs) aim to tailor responses to individual users, where a central challenge is preference reasoning: inferring query-relevant preferences from user-related history. Existing benchmarks, however, largely assume that such preference can be retrieved from semantically related history. We study an underexplored but practically important regime, profile-preference conceptual misalignment (PRCM), where observable profile cues and query-specific preferences lie in different concept spaces, making semantic retrieval inconsistent for personalization. We introduce VIBE-Bench, a benchmark with two psychology-grounded tasks, 3,504 personas and 12,239 dialogues, including a manually verified gold test set, and requires cross-concept preference reasoning beyond surface semantic overlap. Experiments with several personalization methods show that current PLLMs largely rely on shallow semantic correlations and fail to acquire robust cross-concept mappings. These findings establish PRCM as a distinct failure regime in PLLMs and position VIBE-Bench as a focused testbed for advancing preference reasoning beyond semantic matching.
Intravascular ultrasound (IVUS) lumen and external elastic membrane (EEM) segmentation is important for quantitative coronary plaque burden assessment. Errors in lumen or EEM delineation directly propagate to plaque area, plaque burden and geometric measurements. However, standard methods prioritising overlap scores often suffer from boundary drift and topology errors, leading to inaccurate clinical measurements. We present GeoCat, a geometry-consistent network that processes 5-frame IVUS clips using dual Cartesian-polar encoders with cross-domain attention and temporal fusion. A differentiable geometry consistency loss directly supervises clinically relevant descriptors including diameters, orientations, and cross-sectional areas. The model is trained on 12,242 annotated frames from 146 patients acquired with two commercial IVUS systems. We evaluate performance using both segmentation accuracy and plaque-relevant clinical metrics, including Dice/IoU, boundary measures(95HD (mm), ASSD), topology violation rate, and clinical geometry errors (dmax/dmin, angles, and areas). On our dataset, GeoCat achieves a Dice of 0.93, reduces 95HD to 0.14 mm, and lowers topology violations to 1.0
With the rapid expansion of the Internet of Things (IoT), secure authentication has become paramount for safeguarding the digital ecosystem against spoofing attacks and privacy breaches. Optical and acoustic modalities currently dominate biometric authentication; however, they remain inherently vulnerable to environmental interference. Here, we report ultrathin (<50 nm), substrate-free gold nanoplatelet skins that conform intimately to finger-knuckle wrinkles, enabling dynamic and nontransferable wearable authentication. Unlike light- or sound-based systems, these nanocrystal skins generate unique resistive signatures that are insensitive to optical and acoustic perturbations. Fabricated from additive-free, polystyrene-capped gold nanoplatelets, the conformal films establish gapless skin-electronic interfaces that replicate knuckle microtopography with high fidelity. This intimate coupling is the key to converting bending-release motions into finger-specific electrical signatures that are unattainable with the corresponding substrate-supported, nonconformal system. Integrated with a deep-learning framework, the two-finger nanoplatelet skin achieves near-perfect authentication accuracy. Our findings indicate that substrate-free nanocrystal skins could enable next-generation wearable biometric authentication, advancing hardware-level security for the digital world.
BACKGROUND:Despite advances in epilepsy treatment options, selecting the appropriate therapy for an individual with epilepsy is a process of trial and error. Machine learning holds the potential to support clinical decision making. We aimed to provide an overview of the role of machine learning in epilepsy management and discuss future directions. METHODS:In this systematic review and meta-analysis, we searched Embase, MEDLINE, Scopus, and Web of Science from database inception to March 31, 2025, for human-only randomised controlled trials, cohort studies, and case-control studies predicting antiseizure medication outcomes, drug-resistant epilepsy, epilepsy surgery outcomes, and epilepsy surgery candidacy in populations with clinician-confirmed diagnosis of epilepsy. Studies on diagnosis, epilepsy classification, seizure prediction, engineering, and technical aspects of electroencephalograms, neuroimaging, or seizure detection by electrocardiogram or wearable devices were excluded to ensure clinical relevance. Summary data were extracted from published reports. Reporting quality and risk of bias were assessed with TRIPOD+AI and PROBAST, respectively. Meta-analysis was conducted by pooling the area under the receiver-operator characteristic curves (AUCs) of the best model of each study when CIs were available to assess performance. This study was registered with PROSPERO (CRD42023442156). FINDINGS:A total of 16 771 studies were identified, and 135 were included in the systematic review (33 [24%] that predicted antiseizure medication outcomes, 12 [9%] that predicted the development of drug resistance, 79 [59%] that predicted epilepsy surgery outcomes, nine [7%] that predicted epilepsy surgery candidacy, and two [1%] that predicted both antiseizure medication and epilepsy surgery outcomes). Only ten (7%) studies satisfied 70% or more of the subitems in the TRIPOD+AI reporting guidelines checklist, reflecting an overall inadequacy of most of the studies. All the included studies were rated high for overall risk of bias. The pooled AUC of the best-performing models in each study with available data was 0·82 (95% CI 0·77-0·88) in predicting antiseizure medication outcomes, 0·82 (0·76-0·88) in predicting epilepsy surgery outcomes, and 0·94 (0·92-0·96) in predicting epilepsy surgery candidacy. Studies had very high or high heterogeneity (studies predicting antiseizure medication outcomes I2=99·48%, p<0·0001; studies predicting epilepsy surgery outcomes I2=98·94%, p<0·0001; studies predicting epilepsy surgery candidacy I2=85·98%, p<0·0001). The AUCs of the models predicting drug-resistant epilepsy ranged from 0·76 to 0·99 for internal validation. INTERPRETATION:Although machine learning shows promise in predicting epilepsy treatment outcomes, the high heterogeneity and bias-particularly in small sample sizes, handling of missing data, and scarcity of studies with external validation-limit its clinical applicability. Future research should focus on larger, diverse datasets and standardised minimum reporting. Prospective trials are needed to evaluate machine learning models in real-world settings. FUNDING:Australian Government National Health and Medical Research Council.
Video-based seizure detection is essential for the management of epilepsy patients, offering a non-invasive complement to electroencephalography. While several deep learning approaches have been developed for video-based seizure detection, none are inherently interpretable, limiting their adoption and translation into clinical practice. We present, to our knowledge, the first exploration of a neurosymbolic framework for video-based seizure detection that directly addresses this gap. Our approach (1) extracts patient-centric skeleton sequences from epilepsy monitoring units via a prompt-guided foundation model, (2) predicts binary spatio-temporal concept activations grounded in clinical motor semiology guidelines, and (3) composes them via differentiable logic into interpretable Boolean rules with auditable contributions. Furthermore, to mitigate false positives arising from the traditional binary formulation (seizure vs. non-seizure), we sub-classify non-seizure segments into clinically relevant normal activities, providing the model with fine-grained discriminative supervision. Evaluated on two public seizure video benchmarks, our framework achieves 89.78
Lesion segmentation in non-contrast CT (NCCT) images is crucial for stroke management. However, stroke annotation on NCCT requires extensive time and expertise, resulting in data scarcity and thereby limiting segmentation model performance. While diffusion models have become widely used in medical image generation to mitigate data scarcity, significant challenges remain in generating high-quality 3D volumes and incorporating semantic constraints. To address these challenges, we propose a novel semantic-guided latent diffusion model (LDM) for NCCT synthesis. Our Semantic-guided 3D CT Generation (SCTG) model leverages semantic features extracted from both images and annotations using DINOv2 to guide the diffusion process, generating 3D NCCT images that correspond to specified stroke lesions. We also present a Spatial-Semantic Feature Interaction Module (SSFIM) that integrates semantic features into LDM, incorporating adapters to facilitate semantic information in guiding the generation process. Experimental results show that (1) SCTG outperforms existing conditional generation methods qualitatively and quantitatively for image synthesis, and (2) incorporating our generated data improves NCCT-based stroke lesion segmentation performance by up to 22% in Dice score.
Recent advancements in deep learning have shown significant potential for classifying retinal diseases using color fundus images. However, existing works predominantly rely exclusively on image data, lack interpretability in their diagnostic decisions, and treat medical professionals primarily as annotators for ground truth labeling. To fill this gap, we implement two key strategies: extracting interpretable concepts of retinal diseases using the knowledge base of GPT models and incorporating these concepts as a language component in prompt-learning to train vision-language (VL) models with both fundus images and their associated concepts. Our method not only improves retinal disease classification but also enriches few-shot and zero-shot detection (novel disease detection), while offering the added benefit of concept-based model interpretability. Our extensive evaluation across two diverse retinal fundus image datasets illustrates substantial performance gains in VL-model based few-shot methodologies through our concept integration approach, demonstrating an average improvement of approximately 5.8
PURPOSE:To develop and validate an automated lens cortex and nuclear opacity quantification method based on swept-source anterior segment optical coherence tomography (AS-OCT). METHODS:This cross-sectional study included 504 cataract surgery candidates. Lens images were captured using swept-source AS-OCT (CASIA-2; Tomey Corporation). Based on nnUNet framework, two artificial intelligence (AI) segmentation models were independently trained to quantify opacity in the lens cortex and nucleus. Data from 275 and 229 individuals were used for lens nucleus model training and external testing, respectively. The corresponding numbers for lens cortex model were 100 and 38. Five-fold cross-validation was employed for model selection. The performance of the auto-segmentation, as well as the mean pixel intensity values within the area of interest, were evaluated against the human-generated labels. RESULTS:The AI models demonstrated good segmentation accuracy for the lens cortex and nucleus (mean intersection over union [MIoU] = 0.959, 95% CI: 0.957 to 0.961 for cortex; MioU = 0.928, 95% CI: 0.925 to 0.931 for nucleus), and high agreement in the opacity quantification (intraclass correlation coefficient [ICC] = 0.9933, 95% CI: 0.9872 to 0.9965 for the cortex; ICC = 0.9939, 95% CI: 0.9921 to 0.9953 for the nucleus), compared to manual measurements by ophthalmologists. CONCLUSIONS:The AI model is capable of accurately and objectively quantifying the opacity of both the lens cortex and nucleus based on swept-source AS-OCT images, thereby offering a method that is more precise, objective, and rapid for quantification in both clinical practice and research settings.
Continuous video-based seizure detection remains clinically challenging owing to occlusions, environmental variations, and subtle seizure manifestations. We introduce a privacy-centric, non-invasive video-based seizure detection system that leverages dense surface normals to encode geometric features. This approach achieves superior generalization as the features remain invariant to patient appearance. This work explores the first application of surface normal analysis to seizure detection, demonstrating that geometry-based features not only preserve privacy but also outperform traditional pose-based methods. Through a rigorous evaluation on 821 clinical video clips from 7 patients, we systematically compare surface normals against pose estimation, semantic segmentation, and multi-modal fusion approaches under both patient-dependent (5-fold CV) and patient-independent (LOPO CV) validation protocols. Raw surface normals achieve 89.4% accuracy, significantly outperforming pose estimation ( 82.8% ) with a 26.8% relative improvement in F1-score. Critically, it maintains exceptional robustness in scenarios where semantic methods catastrophically fail, including multi-person interactions and severe occlusion, where its F1-score is robustly maintained at 0.833 while the pose-based F1-score drops to 0.0. Our proposed geometry-based, privacy-centric approach enables continuous monitoring in both clinical and home settings without compromising patient privacy.
Survival analysis holds a crucial role across diverse disciplines, such as economics, engineering and healthcare. It empowers researchers to analyze both time-invariant and time-varying data, encompassing phenomena like customer churn, material degradation and various medical outcomes. Given the complexity and heterogeneity of such data, recent endeavors have demonstrated successful integration of deep learning methodologies to address limitations in conventional statistical approaches. However, current methods typically involve cluttered probability distribution function (PDF), have lower sensitivity in censoring prediction, only model static datasets, or only rely on recurrent neural networks for dynamic modelling. In this paper, we propose a novel survival regression method capable of producing high-quality unimodal PDFs without any prior distribution assumption, by optimizing novel Margin-Mean-Variance loss and leveraging the flexibility of Transformer to handle both temporal and non-temporal data, coined UniSurv. Extensive experiments on several datasets demonstrate that UniSurv places a significantly higher emphasis on censoring compared to other methods.
Epilepsy affects over 50 million people worldwide, with antiseizure medications (ASMs) as the primary treatment for seizure control. However, ASM selection remains a “trial and error” process due to the lack of reliable predictors of effectiveness and tolerability. While machine learning approaches have been explored, existing models are limited to predicting outcomes only for ASMs encountered during training and have not leveraged recent biomedical foundation models for this task. This work investigates ASM outcome prediction using only patient MRI scans and reports. Specifically, we leverage biomedical vision-language foundation models and introduce a novel contextualized instruction-tuning framework that integrates expert-built knowledge trees of MRI entities to enhance their performance. Additionally, by training only on the four most commonly prescribed ASMs, our framework enables generalization to predicting outcomes and effectiveness for unseen ASMs not present during training. We evaluate our instruction-tuning framework on two retrospective epilepsy patient datasets, achieving an average AUC of 71.39 and 63.03 in predicting outcomes for four primary ASMs and three completely unseen ASMs, respectively. Our approach improves the AUC by 5.53 and 3.51 compared to standard report-based instruction tuning for seen and unseen ASMs, respectively. Our code, MRI knowledge tree, prompting templates, and TREE-TUNE generated instruction–answer tuning dataset are available at the link .
Multimodal Large Language Models (MLLMs) have shown promise in visual-textual reasoning, with Multimodal Chain-of-Thought (MCoT) prompting significantly enhancing interpretability. However, existing MCoT methods rely on rationale-rich datasets and largely focus on inter-object reasoning, overlooking the intra-object understanding crucial for image classification. To address this gap, we propose WISE, a Weak-supervision-guided Step-by-step Explanation method that augments any image classification dataset with MCoTs by reformulating the concept-based representations from Concept Bottleneck Models (CBMs) into concise, interpretable reasoning chains under weak supervision. Experiments across ten datasets show that our generated MCoTs not only improve interpretability by 37% but also lead to gains in classification accuracy when used to fine-tune MLLMs. Our work bridges concept-based interpretability and generative MCoT reasoning, providing a generalizable framework for enhancing MLLMs in fine-grained visual understanding.
Concept Bottleneck Models (CBMs) decompose image classification into a process governed by interpretable, human-readable concepts. Recent advances in CBMs have used Large Language Models (LLMs) to generate candidate concepts. However, a critical question remains: What is the optimal number of concepts to use? Current concept banks suffer from redundancy or insufficient coverage. To address this issue, we introduce a dynamic, agent-based approach that adjusts the concept bank in response to environmental feedback, optimizing the number of concepts for sufficiency yet concise coverage. Moreover, we propose Conditional Concept Bottleneck Models (CoCoBMs) to overcome the limitations in traditional CBMs' concept scoring mechanisms. It enhances the accuracy of assessing each concept's contribution to classification tasks and feature an editable matrix that allows LLMs to correct concept scores that conflict with their internal knowledge. Our evaluations across 6 datasets show that our method not only improves classification accuracy by 6% but also enhances interpretability assessments by 30%.
BACKGROUND:The surge in AI models for diagnosing skin lesions through image analysis is notable, yet their clinical implementation faces challenges. Common limitations include an over reliance on dermoscopy, lack of real-world applicability when only binary output (e.g. benign/malignant) is offered and low accuracy when faced with rare skin conditions. OBJECTIVES:To address these common constraints associated with limited diagnostic output, and applicability to real-world settings. METHODS:We developed an All-In-One Hierarchical-Out of Distribution-Clinical Triage (HOT) AI model for skin lesion analysis. Trained on a large dataset of ~208,000 lesion images, our HOT AI model generates three outputs: a hierarchical three-level prediction, an alert for out-of-distribution (OOD) images and a recommendation for dermoscopy to improve diagnostic prediction. RESULTS:Our hierarchical prediction output provides a binary level 1 prediction (benign/malignant), Level 2 prediction of eight possible categories (e.g. melanocytic and keratinocytic) and a more definitive Level 3 prediction from 44 lesion categories. The model produced high sensitivity for Level 1 prediction (88.14% CI: 87.42-88.51); however, significantly lower for Level 3 prediction (63.90%, CI: 62.27-65.61). By relying on all three prediction levels for consensus, Level 1 false-positives were reduced by 20-25%, and false-negatives were decreased by 11-13% of cases. OOD detection was benchmarked against previous landmark models and outperformed comparative models. Lastly, 44% of images were recommended for dermoscopy, and with additional image input, Level 3 sensitivity increased from 48.13% (CI:45.08-49.57) to 52.54% (CI:50.25-55.04). CONCLUSIONS:Our HOT-AI model attempts to address common challenges in existing models by combining three tasks in one model to increase accuracy and clinical utility. By providing a more nuanced prediction, and alert for OOD, the model output provides greater explainability of the AI decision process. Prospective clinical testing is required to measure how this additional output impacts user trust, and how the model performs in a real-world setting.