
In the past decade, virtual reality (VR) and augmented reality (AR) have seen a second wave of activity due to consumer-level availability of devices. In the medical realm, use of VR and AR has been investigated extensively and successfully for training, teaching, rehabilitation, and therapy. However, extended reality (XR) and its intraoperative use in the operating room has yet been underexplored and underdeveloped. Intelligent XR environments that integrate advanced visualization, intuitive interaction, and AI-assisted decision support have the potential to transform clinical care. Responsive, context-aware spaces could enhance medical teams’ capabilities by enabling real-time access to critical information, enhancing situation awareness, and supporting remote collaboration and access to clinical expertise across geographical boundaries—all while preserving the natural flow of procedures. But to realize the potential benefits of XR in the operating room requires a significant and focused effort by an interdisciplinary community. It is the intention of this paper to help to catalyze such an effort. In this paper, we identify pressing issues in clinical care, the potential of XR to address them, and the further research that is needed to realize that potential. We look at a broad range of areas where XR can play an important role, including direct support for the surgeon, support for the surgical team, support for the patient, and remote collaboration. We also consider ways in which XR can be integrated with AI to provide enhanced information and interaction. Discussions are underpinned by considerations of methods to evaluate contributions and progress. This paper is the result of a Dagstuhl Seminar on Extended Reality for the Operating Room (XR4OR), held at the Leibniz Center for Informatics, Schloss Dagstuhl, Germany, in February 2025. It gathered twenty-four researchers working in the areas of virtual reality, augmented reality, medical informatics, human-computer interaction, and surgery.
The rapid evolution of Generative Artificial Intelligence (GenAI) presents significant opportunities to transform healthcare, particularly in generating personalized treatment recommendations. This systematic literature review explores the current state of GenAI language models applications in various medical domains, assessing their effectiveness, applicability, and limitations. The review addresses nine specific research questions to understand the potential and challenges of integrating GenAI into clinical practice. We use the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. From a pool of 3237 studies, 42 were selected based on inclusion and exclusion criteria. These studies were analyzed to evaluate the use of generative language models, such as GPT-3 and GPT-4, in various medical domains including oncology, cardiovascular, gastrointestinal, and ophthalmological care. The analysis revealed that most GenAI applications in healthcare rely on general purpose LLMs to provide treatment recommendations. Fine-tuning with domain-specific data and prompt engineering were found to significantly improve output quality and reliability. However, persistent challenges include lack of clinical validation, ethical concerns such as bias, and issues related to transparency and regulatory compliance. While GenAI demonstrates strong potential to support clinical decision-making, real-world deployment remains limited due to unresolved ethical and validation issues. Future research should prioritize the development of interpretable, domain-specific models and rigorous clinical trials to ensure safe and effective integration into healthcare settings.
Autism Spectrum Disorder (ASD) screening often relies on structured questionnaires. Yet these data are not always easy to interpret or use in practice. There is a need for screening frameworks that support risk estimation, case comparison, and pattern exploration in a clear and accessible way. In this work, we present NeuroCypher ASD, a framework that combines tabular machine learning with a knowledge graph for ASD screening, interpretation, and data exploration. The tabular model uses questionnaire responses and demographic features to estimate ASD risk. The knowledge graph organizes the data and supports case comparison, pattern exploration, and natural language querying. The system is evaluated under a leakage-aware protocol using tabular, graph-based, and hybrid models. The results show strong performance for the tabular model and clarify the role of graph-based representations within the overall workflow. In this setting, the knowledge graph supports structured analysis of screening data by helping users examine relationships between cases, inspect shared features, and interpret predictions in context. To strengthen interpretability, the framework includes SHAP-based analysis and approximate mappings between embedding dimensions and original questionnaire features. It also includes an anomaly detection module that highlights atypical cases for closer review. Overall, NeuroCypher ASD provides a practical framework for combining prediction, interpretation, and data exploration in ASD screening. The results show the value of integrating machine learning, graph-based context, and natural language interaction in more transparent and accessible screening workflows, particularly in school-based environments.
Biomedical ontologies such as the Unified Medical Language System (UMLS) play a crucial role in supporting practical applications such as predicting gene-disease associations, drug-drug interactions, and biomedical question answering. However, automatically maintaining and updating them remains a critical barrier as authoring high-quality textual definitions is both time-consuming and labor-intensive. As a result, UMLS ontology updates significantly lag behind the pace of scientific discovery, leaving many concepts without definitions or with outdated descriptions. Despite the pressing need for automated generation of UMLS concept definitions, this problem has largely remained unaddressed in prior research. To address this, we propose a novel diffusion-based approach for generating high-quality definitions of biomedical concepts. A key innovation in our approach is the adaptive semantic modulation mechanism, which dynamically adjusts the influence of semantic priors throughout the denoising process, enabling the model to flexibly balance local lexical fidelity with global semantic alignment. Extensive experiments on the largest available biomedical ontology dataset demonstrate that the proposed approach both significantly and consistently outperforms strong baseline algorithms across multiple evaluation metrics.
Health literacy is a multidimensional construct essential for health decision-making, yet its computational assessment from naturalistic online discourse remains limited by categorical classifications that fail to capture its latent, continuous, and context-sensitive nature. To address this gap, we introduce an uncertainty-aware Bayesian deep learning framework that probabilistically infers latent health literacy from social media text while systematically quantifying predictive uncertainty. Using a large corpus of English-language health forum posts (N = 342 k), we operationalized five theoretical dimensions of health literacy—Functional, Communicative, Critical, Digital, and Expressed—through validated NLP features. A Bayesian Variational Autoencoder with Monte Carlo Dropout models health literacy as a continuous latent variable and provides epistemic uncertainty estimates. The framework recovers a robust three-factor latent struc- ture: Core Integrated Proficiency (merging Critical, Communicative, and Expressed dimensions), Digital Proficiency (exhibiting an inverse association with Functional Literacy), and Applied Functional Literacy. The model achieves strong reconstruction performance (MSE = 0.109) with uncertainty estimates reliably correlated to prediction error (r = 0.438). From the continuous latent representations, we derive three distinct user profiles—Balanced, Specialized, and Transitional—revealing heterogeneous patterns of health literacy expression and adaptive communication behavior. This work advances computational health literacy assessment by providing a probabilistic, uncertainty-aware framework that moves beyond static categorization, with direct implications for personalized public health communication and hybrid human-AI assessment systems.
The rapid advancement of Generative AI (GenAI) has transformed human-computer interaction, particularly in domains requiring personalized guidance, such as nutritional planning. While GenAIs show promise in generating contextual nutritional recommendations, their integration into social interactions and behavioral patterns (which are key aspects of the Internet of Behavior - IoB) remains inadequately understood. This study investigates how social media discussions can inform the development of more effective human-centered GenAI systems for dietary support. Through a comprehensive analysis of 10,219 user-generated contents across Reddit, YouTube, and Bluesky social media platforms, we examine social interaction patterns, behavioral trends, and user experiences with GenAI-generated nutritional content. Our mixed-method approach combines computational analysis of big social data with qualitative assessment of user feedback, revealing crucial insights on the benefits and pitfalls of GenAIs, ranging from their contribution to users’ health and wellness to their inconsistent outputs with human experts. Based on these insights, we recommend three conceptual frameworks that suggest the integration of GenAIs on social media with human experts’ involvement and social media conversational engagements. Furthermore, our user study experiment involving one of the conceptual frameworks, FR3, revealed significant positive effects in users’ attitude, dietary behavior intentions, and outcomes expectancy (p<.001), alongside moderate trust in GenAIs. These findings contribute to the field through: (1) a structured framework for enhancing human-AI collaboration in nutritional guidance, (2) a validated machine learning model for sentiment classification in dietary AI discourse, and (3) empirical, qualitative insights, and evidence-based recommendations for integrating GenAIs for dietary guidance in social spaces. This research advances our understanding for developing socially responsive and human-centered GenAI systems that effectively promote sustainable healthy eating behaviors.
Electronic Health Records (EHRs) provide rich opportunities for developing risk prediction tools to support clinical decision-making, yet they are inherently incomplete because data are recorded selectively during routine care. Such missingness may be informative, reflecting clinical judgment and patient status, and missing data patterns can shift between model development and real-world deployment. These challenges limit the reliability and transportability of predictive models in healthcare settings. We propose an imputation-free framework that jointly trains Conditional Variational Autoencoders with deep survival models to enable risk prediction directly from incomplete EHR data. We demonstrate the approach using the deep survival model DeSurv and evaluate its performance through simulation studies and two retrospective cohorts from the Clinical Practice Research Datalink primary care database. The proposed framework consistently outperforms conventional missing data methods, achieving superior performance on ground-truth metrics in simulations and improved calibration-based survival metrics in real-world cohorts. It also demonstrates increased robustness to unseen missingness patterns and distributional shifts. By providing a unified strategy for handling missing data across development, validation, and deployment, this work advances methodological robustness in healthcare informatics and supports more reliable clinical risk prediction in practice.
Large language models (LLMs) have revolutionized medical reasoning tasks, yet single-agent systems often falter on complex, interdisciplinary problems requiring robust handling of uncertainty and conflicting evidence. Multi-agent systems (MAS) leveraging LLMs enable collaborative intelligence, but prevailing centralized architectures suffer from scalability bottlenecks, single points of failure, and role confusion in resource-constrained environments. Decentralized MAS (D-MAS) promise enhanced autonomy and resilience via peer-to-peer interactions, but their application to high-stakes healthcare domains remains underexplored. We introduce MediHive, a novel decentralized multi-agent framework for medical question answering that integrates a shared memory pool with iterative fusion mechanisms. MediHive deploys LLM-based agents that autonomously self-assign specialized roles, conduct initial analyses, detect divergences through conditional evidence-based debates, and locally fuse peer insights over multiple rounds to achieve consensus. Empirically, MediHive outperforms single-LLM and centralized baselines on MedQA and PubMedQA datasets, attaining accuracies of 84.3
Recognition of daily activities is a critical element for effective Ambient Assisted Living (AAL) systems, particularly to monitor the well-being and support the independence of older adults in indoor environments. However, developing robust activity recognition systems faces significant challenges, including intra-class variability, inter-class similarity, environmental variability, camera perspectives, and scene complexity. This paper presents a multi-modal approach for the recognition of activities of daily living tailored for older adults within AAL settings. The proposed system integrates visual information processed by a 3D Convolutional Neural Network (CNN) with 3D human pose data analyzed by a Graph Convolutional Network. Contextual information, derived from an object detection module, is fused with the 3D CNN features using a cross-attention mechanism to enhance recognition accuracy. This method is evaluated using the Toyota SmartHome dataset, which consists of real-world indoor activities. The results indicate that the proposed system achieves competitive classification accuracy for a range of daily activities, highlighting its potential as an essential component for advanced AAL monitoring solutions. This advancement supports the broader goal of developing intelligent systems that promote safety and autonomy among older adults.
This study presents a hybrid ontology-based framework for clinical concept extraction from narrative EHR discharge summaries using large language models (LLMs) and standardized biomedical terminologies. The framework integrates multiple NLP components in sequence: SparkNLP for chunk detection and named entity recognition (NER), SentenceBERT embeddings for semantic similarity candidate generation, zero-shot inference with LLaMA3-8B and Mistral-7B for concept selection, and UMLS REST API normalization to CUIs and SNOMED CT terms. This coordinated integration of linguistic, semantic, and ontological modules forms a flexible architecture rather than a single-model comparison. We applied the framework to ten MIMIC-III discharge summaries spanning Chief Complaint, Brief Hospital Course, and History of Present Illness sections. Clinicians labeled extracted concepts as correct, partial, incorrect, missing, or spurious to assess model performance. LLaMA3-8B achieved the highest F1 score (0.77) and lowest false positive rate (3.04%), outperforming both Mistral-7B and cTAKES. While cTAKES demonstrated high precision, it had low recall and a significantly higher FPR (29.95%), indicating frequent misclassification. Mistral-7B offered faster processing for shorter notes, while LLaMA3-8B delivered higher accuracy for more detailed sections.LLMs outperformed traditional rule-based systems by more effectively handling context, modifiers, abbreviations, and multi-word expressions. Prompt refinement and semantic similarity embedding enhanced extraction quality. SparkNLP supported chunking but introduced errors related to spacing and abbreviation handling. We presented a flexible, context-aware framework for clinical concept extraction using LLMs, offering key advantages over rule-based tools. Future work should incorporate full ontology mapping, integrate assertion detection, and validate performance across diverse clinical datasets and domain-adapted LLMs.
Many colorectal cancer (CRC) survivors who have undergone resection surgery experience persistent bowel dysfunction that significantly affects their quality of life, highlighting the need for defecatory function rehabilitation in survivorship care. Although mobile health (mHealth) applications are increasingly recognized as promising tools for supporting post-surgical recovery, few are specifically designed to address the complex needs of CRC survivors. This study explores these challenges and identifies design requirements through ExerCompass—a prototype mHealth application designed to support recovery by offering a guided exercise program, bowel and dietary tracking, and condition-aware feedback. We conducted a formative study with 22 CRC survivors using task-based interaction sessions and semi-structured interviews. The prototype was developed by incorporating clinical knowledge specific to CRC, provided by healthcare experts, along with established design principles for digital health tools. We captured participants’ overall experiences, uncovered usability challenges and gathered feedback on the prototype’s capacity to support recovery. Participants valued the unified delivery of guided exercise, symptom tracking, and condition-aware feedback, noting its therapeutic relevance to their recovery. However, they reported difficulties in logging variable bowel symptoms, interpreting dietary trends, and sustaining motivation over time. While some interface-level issues were mentioned, most emphasized the need for flexible, personalized, and emotionally supportive features. This study highlights the importance of designing mHealth tools that address condition-specific needs of post-surgical CRC survivors and offers design implications for digital interventions that support both symptom management and psychosocial recovery.
Camera-based systems offer a comprehensive and inconspicuous approach to monitoring the well-being of individuals within the comfort of their homes. This study introduces a vision-based, fully autonomous pipeline for assessing eating behaviors and detecting musculoskeletal changes. The system captures eating activities and provides detailed insights such as hand-to-mouth motion duration and bite count. These indicators are vital for understanding behavioral and physiological influences on food consumption and their associated changes. The system integrates pose estimation and a temporal action localization network to classify actions and generate behavior profiles. Evaluated on the EatSense dataset and a supplementary test set, the system achieves strong performance, including a mean average precision (mAP) of 74% at 0.10 IoU for micro-action detection and a posture anomaly detection accuracy of over 76%. These results demonstrate the system's ability to detect subtle trends such as slower hand movements under increased wrist weights and changes in chewing behavior. Additionally, comparisons against Gemini-2.5-Pro, a state-of-the-art multimodal model, reinforces the system's accuracy. So, by successfully capturing trends aligned with ground truth data, the pipeline shows promise for long-term health monitoring, early detection of musculoskeletal decline, and behavioral changes in dietary habits-offering potential applications in elderly care and remote health assessment. The new test dataset is released on https://groups.inf.ed.ac.uk/vision/DATASETS/EATSENSE/.
Synthetic data generation using large language models (LLMs) demonstrates substantial promise in addressing biomedical data challenges and shows increasing adoption in biomedical research. This study systematically reviews recent advances in synthetic data generation for biomedical applications and clinical research, focusing on how LLMs address data scarcity, utility, and quality issues with different modalities. We conducted a scoping review following PRISMA-ScR guidelines and searched literature published between 2020 and 2025 through PubMed, ACM, Web of Science, and Google Scholar. A total of 59 studies were included based on relevance to synthetic data generation in biomedical contexts. Among the reviewed studies, the predominant data modalities were unstructured texts (78.0%), tabular data (13.6%), and multimodal sources (8.4%). Common generation methods included LLM prompting (74.6%), fine-tuning (20.3%), and specialized models (5.1%). Evaluations were heterogeneous: intrinsic metrics (27.1%), human-in-the-loop assessments (44.1%), and LLM-based evaluations (13.6%). However, limitations and key barriers persist in data modalities, domain utility, resource and model accessibility, and standardized evaluation protocols. Future efforts may focus on developing standardized, transparent evaluation frameworks and expanding accessibility to support effective applications in biomedical research.
Understanding the heterogeneity of neurodegeneration in Alzheimer's disease (AD) and identifying distinct progression pathways is critical for improving diagnosis, treatment, prognosis, and prevention. Motivated by this need, this study aimed to identify disease progression subphenotypes among patients with mild cognitive impairment (MCI) and AD using electronic health records (EHRs). We developed a novel approach that combines a graph neural network (GNN)-based framework with time series clustering to characterize progression subphenotypes from MCI to AD. We applied the proposed framework to a real-world cohort of 2,525 patients (61.66% female; mean age 76 years), of whom 64.83% were Non-Hispanic White, 16.48% Non-Hispanic Black, 2.53% were of other races, and 10.85% were Hispanic. Our model identified four distinct progression subphenotypes, each exhibiting characteristic clinical patterns, with average MCI-to-AD progression times ranging from 805 to 1,236 days. These findings indicate that AD does not follow a uniform progression trajectory but instead manifests heterogeneous pathways, and the proposed framework provides an explainable, data-driven approach for delineating AD progression subphenotypes, offering actionable insights for healthcare informatics research and the clinical management of patients at risk for AD.
Large Language Models (LLMs) have shown significant promise for clinical applications, yet their application to triage remains underexplored. In this study, we systematically investigate the capabilities of LLMs in emergency department triage through two key dimensions: (1) robustness to distribution shifts and missing data, and (2) intersectional biases across sex and race. We assess multiple LLM-based approaches, ranging from continued pre-training to in-context learning, as well as conventional machine learning (ML) approaches. First, we demonstrate that LLMs exhibit superior robustness compared to traditional ML, which is promising due to their ability to provide explanatory rationales. Second, we show that the most effective LLM-based methods are those that select similar examples from prior patient cases, whereas reasoning capabilities in LLMs offer little benefit for triage. Lastly, we identify critical gaps in LLM preferences that emerge at the intersections of sex and race. LLMs exhibit sex-based differences, and they are more pronounced in certain racial groups, suggesting that LLMs encode preferences that emerge in specific clinical contexts and combinations of characteristics. We perform this audit through counterfactual analysis, providing a systematic way to identify such biases before real-world integration.
Uncertainty Quantification (UQ) has gained traction in an attempt to improve the interpretability and robustness of machine learning predictions. Specifically (medical) biosignals such as electroencephalography (EEG), electrocardiography (ECG), electrooculography (EOG), and electromyography (EMG) could benefit from good UQ, since these suffer from a poor signal-to-noise ratio, and good human interpretability is pivotal for medical applications. To determine how uncertainty estimation can be used for biosignal tasks, we investigate current methods, use cases, applications, evaluations, and uncertainty measures. In this paper, we systematically review the state of the art of applying Uncertainty Quantification to Machine Learning tasks in the biosignal domain. All works from Web of Science, Scopus, IEEE XPlore and PsycINFO that discuss uncertainty in Machine Learning on one of the aforementioned biosignals is included. We present various methods, shortcomings, uncertainty measures and theoretical frameworks that currently exist in this application domain based on the 53 reviewed papers and related literature. We address misconceptions in the field, provide recommendations for future work, and discuss gaps in the literature in relation to diagnostic implementations as well as control for prostheses or brain-computer interfaces. Overall it can be concluded that promising UQ methods are available, but that research is needed on how people and systems may interact with an uncertainty-model in a (clinical) environment.
The reuse of clinical health data holds immense promise for advancing medical research, yet remains constrained by complex legal, technical, and organisational barriers. This article examines these challenges through the case study of TumorScope, a Belgian interdisciplinary initiative developing a secure, multimodal data environment for glioblastoma research. Drawing on five years of practical experience integrating imaging, genetic, tissue-based, and clinical datasets, the study identifies key legal, ethical, technical, and operational obstacles to effective data access, linkage, and reuse. Technical issues included fragmented data flows, pseudonymisation complexities, and limited interoperability, while legal and ethical barriers arose from strict interpretations of the General Data Protection Regulation, medical secrecy obligations, and intellectual property constraints. These were compounded by operational challenges such as unclear governance structures, resource limitations, and the limited capacity of Medical Research Ethics Committees to assess data-driven research. The analysis further considers the European Health Data Space Regulation (EHDS) as a potential enabler of responsible secondary data use, while noting uncertainties in its national implementation. Overall, the study demonstrates that meaningful health data reuse requires more than regulatory compliance, it depends on robust governance frameworks, institutional coordination, and sustained investment in infrastructure and expertise. The findings contribute to ongoing debates in healthcare informatics on how to translate the vision of the EHDS into practical, ethically grounded data reuse for patient benefit.
This article presents the design, implementation, and evaluation of MarIA, a GPT-3.5-powered virtual assistant integrated into a messaging platform to support patients with type 2 diabetes mellitus (DM). MarIA employs a multi-agent architecture that enables varying dialogue styles and degrees of personalization. In a 3-month longitudinal study involving 35 participants, personalized interactions increased engagement by 26
Deep learning has been extensively applied to medical imaging tasks over the past years, achieving outstanding results. However, the obscure reasoning of the models and the lack of supportive evidence causes both clinicians and patients to distrust the models’ predictions, hindering their adoption in clinical practice. In recent years, the research community has focused on developing explanations capable of revealing a model’s reasoning. Among various types of explanations, example-based explanations emerged as particularly intuitive for medical practitioners. Despite the intuitiveness and wide development of example-based explanations, no work provides a comprehensive review of existing example-based explainability works in the medical image domain. In this work, we review works that provide example-based explanations for medical imaging tasks, reflecting on their strengths and limitations. We identify the absence of objective evaluation metrics, the lack of clinical validation and privacy concerns as the main issues that hinder the deployment of example-based explanations in clinical practice. Finally, we reflect on future directions contributing towards the deployment of example-based explainability in clinical practice.
Pancreatic cancer remains one of the deadliest malignancies, primarily because of its subtle CT appearance and frequent late-stage diagnosis. We introduce MiniGPT-Pancreas, a lightweight multimodal large language model (MLLM) that interprets natural-language queries within an interactive ChatGPT-style interface, as well as computed tomography images, and returns precise bounding-box predictions for the pancreas and associated tumors. A cascaded fine-tuning strategy was applied to MiniGPTv2, a multi-task general-purpose MLLM, with a focus on pancreas and tumor detection, using the National Institute of Health (NIH) and Medical Segmentation Decathlon (MSD) pancreas datasets. Pancreas detection achieved an average intersection over Union (IoU) of 0.57 on NIH and MSD datasets, outperforming the base MiniGPT-Pancreas model and more recent MLLMs like GLM-4.1V-9B-Base (general-purpose) and UMIT (specific to the biomedical domain). Tumor observation on MSD yielded an accuracy, precision, recall, and F1 score of all about 0.87, surpassing MiniGPT-v2, GLM-4.1V-9B-Base, and UMIT. For tumor localization, the IoU was 0.28, higher than UMIT (IoU=0.07), but lower than GLM-4.1V-9B-Base (IoU=0.48). On multi-organ detection on the AbdomenCT-1k dataset, MiniGPT-Pancreas outperformed GLM-4.1V-9B-Base and UMIT in all organs, with an IoU of 0.50 on pancreas vs. 0.43 and 0.03, respectively. MiniGPT-Pancreas was rated highly by an international group of 10 expert general surgeons (Italy, Singapore, and the UK) as a potential training tool, especially for verification (4.5/5.0), and training of young specialists (4.5/5.0) on a 5-point Likert scale. While operating on 2D slices limits volumetric context, MiniGPT-Pancreas demonstrates that compact MLLMs can rival specialized vision networks in pancreas imaging, offering an intuitive, language-driven tool for AI-assisted radiology. The code is publicly available at https://github.com/elianastasio/MiniGPTPancreas .