BackgroundObjective Structured Clinical Examinations (OSCEs) are used as an evaluation method in medical education, but require significant pedagogical expertise and investment, especially in emerging fields like digital health. Large language models (LLMs), such as ChatGPT (OpenAI), have shown potential in automating educational content generation. However, OSCE generation using LLMs remains underexplored. ObjectiveThis study aims to evaluate 3 GPT-4o configurations for generating OSCE stations in digital health: (1) standard GPT with a simple prompt and OSCE guidelines; (2) personalized GPT with a simple prompt, OSCE guidelines, and a reference book in digital health; and (3) simulated-agents GPT with a structured prompt simulating specialized OSCE agents and the digital health reference book. MethodsOverall, 24 OSCE stations were generated across 8 digital health topics with each GPT-4o configuration. Format compliance was evaluated by one expert, while educational content was assessed independently by 2 digital health experts, blind to GPT-4o configurations, using a comprehensive assessment grid. Statistical analyses were performed using Kruskal-Wallis tests. ResultsSimulated-agents GPT performed best in format compliance and most content quality criteria, including accuracy (mean 4.47/5, SD 0.28; P=.01) and clarity (mean 4.46/5, SD 0.52; P=.004). It also had 88% (14/16) for usability without major revisions and first-place preference ranking, outperforming the other configurations. Personalized GPT showed the lowest format compliance, while standard GPT scored lowest for clarity and educational value. ConclusionsStructured prompting strategies, particularly agents’ simulation, enhance the reliability and usability of LLM-generated OSCE content. These results support the use of artificial intelligence in medical education, while confirming the need for expert validation.
Large language models have emerged as potential tools to support hypertension care, including diagnosis, treatment decision-making, and patient education. However, evidence regarding their validity, performance, and clinical applicability remains limited. The objective is to map current applications of large language models in hypertension care, with emphasis on model optimization strategies, evaluation approaches, and reported limitations. We conducted a Preferred Reporting Items for Systematic Reviews and Meta-Analyses extension for Scoping Reviews-compliant scoping review of primary studies published between 2023 and 2025 evaluating large language models in hypertension. Thirty-three studies were included. Data were charted on clinical use cases, model optimization techniques, evaluation metrics, data sets, and limitations. Applications were categorized into clinical decision support systems, patient education, medical education, research support, and administrative functions. GPT-based models predominated (82%). Model optimization was limited: 89% relied exclusively on prompt engineering. Most applications focused on patient education (52%) and clinical decision support systems (24%). In clinical decision support systems, reported accuracy ranged from 65% to 100%, reaching 87% to 91% for ambulatory blood pressure monitoring interpretation. Patient education applications showed accuracy between 80% and 90%, but frequent issues included excessive language complexity and occasional unsafe outputs. Across domains, evaluation methods were heterogeneous, reproducibility was inconsistently assessed, and safety concerns, including hallucinations and outdated knowledge, were commonly reported. Current evidence suggests that large language models may support selected tasks in hypertension care; however, their clinical reliability remains uncertain. The limited methodological rigor, minimal use of advanced optimization techniques, and narrow scope of evaluated applications preclude conclusions regarding routine clinical use. Further rigorously designed studies are required before broader implementation can be considered.
Background:Social media apps are widely used by health care professionals despite security and regulatory risks. Identifying factors associated with this use is important for developing effective risk-reduction strategies. Objective:This study aimed to investigate how medical residents use 6 popular social media apps in professional tasks and to identify factors influencing their adoption in health care, using the validated Unified Theory of Acceptance and Use of Technology 2 (UTAUT2) model. Methods:An anonymous web-based survey was conducted between June 2024 and November 2024 among medical residents in France. Participants reported demographic characteristics, frequency, and professional contexts of use for 6 apps (Facebook, Instagram, LinkedIn, Messenger, TikTok, and WhatsApp) and completed UTAUT2-based items. The model was adapted by adding a technology trust construct. Descriptive analyses were performed for all apps. With a sufficient sample size, partial least squares structural equation modeling was conducted for WhatsApp to identify factors associated with behavioral intention and use behavior. Results:A total of 137 residents (n=87, 63.5% female participants) across 40 specialties completed the survey. WhatsApp was the most widely and professionally used app (n=127, 92.7%), with 75.9% (n=104) using it at least many times per week. It was primarily used for patient care, including written transmissions (n=86, 62.8%), case discussions (n=76, 55.5%), and specialist advice (n=86, 62.8%), as well as for professional networking (n=62, 45.3%). Messenger was used by 46.7% (n=64) of participants for similar purposes. Facebook (n=35, 25.6%) and LinkedIn (n=20, 14.6%) were mainly used for education and networking, whereas Instagram (n=11, 8%) was rarely used, and TikTok was not used for professional purposes. Regarding adoption factors, WhatsApp had the highest overall scores, including the highest performance expectancy (mean 5.4, SD 1.12), behavioral intention (mean 5.28, SD 1.15), and use behavior (mean 5.91, SD 1.30), with high effort expectancy (mean 6.82, SD 0.55) and facilitating conditions (mean 6.07, SD 0.85). LinkedIn showed the highest social influence (mean 5.05, SD 1.06), whereas Instagram showed the highest hedonic motivation (mean 6.61, SD 0.51). Technology trust scores were low across all apps, ranging from a mean of 2.23 (SD 1.16) for Facebook to a mean of 3.72 (SD 1.39) for LinkedIn. In the partial least squares structural equation modeling analysis for WhatsApp, habit was the only significant predictor of behavioral intention (β=.53; P<.001) and use behavior (β=.45; P<.001). Conclusions:WhatsApp dominates professional use among residents despite low trust in its security, and its use is mainly driven by habit. Secure alternatives with features similar to popular social media apps, supported by institutional policies and digital professionalism training, are needed to encourage physicians to better consider safety when using social media.
Rare disease gene discovery is limited by small cohorts and the frequent absence of matched controls. We present the Case-Only Burden Test (COBT), a gene-based burden test for case-only designs accounting for multiple variants per individual and additive effects. COBT uses a Poisson model to test for excess variants in a gene relative to expectations from population mutation rates. Simulations show high power and competitive performance versus case-control burden tests. Validation on 1000 Genomes data demonstrated good model fit and low false-positive rates. Applied to 478 ciliopathy patients, COBT re-identified known causal genes and highlighted candidate variants in unsolved cases.
BACKGROUND AND OBJECTIVES:To address the multidimensional nature of health-related questions, advances in health research often require integrating information from various data sources within statistical analyses. When complementary information pertaining to the same set of individuals are distributed across different institutions, vertical methods make it possible to obtain analysis results without sharing or pooling individual-level data. To guide stakeholders toward a transparent and rigorous use of vertical methods with sensitive health data, this study aims to (1) Identify existing vertical methods enabling statistical inference (confidence interval estimation and hypothesis testing); and (2) Characterize the methodological properties of these methods and the current extent of their use with health data. METHODS:We conducted a scoping review following PRISMA-ScR using four interdisciplinary databases. We then systematically extracted the characteristics of identified vertical methods with respect to comparability with the pooled analysis, efficiency of communication schemes and confidentiality. We additionally screened studies that cited included articles to identify applications on vertically partitioned real-world health data. RESULTS:Among 2887 articles initially screened, 30 were included in the review, of which a majority mentioned health analytics. Inference for the linear and the logistic regression framework were the most frequent statistical inference tasks undertaken in proposed methods. Equivalence with the pooled analyses was not systematically addressed and most methods required multiple communications between participating parties. Almost all articles described their approach as privacy-preserving, although a minority provided privacy assessments. Very few published health studies were found to report the use these methods. CONCLUSION:The scope of existing approaches enabling statistical inference for vertically partitioned data is still relatively limited. Most existing methods do not concurrently achieve results equivalent to centralized analyses, high communication efficiency, and guaranteed protection of individual-level data.
Rare diseases affect over 300 million people worldwide and are characterized by complex care pathways, limited clinical expertise, and substantial unmet communication needs throughout the long patient journey. Recent advances in large language models (LLMs) offer new opportunities to support patient education and communication, yet their application in rare diseases remains unclear. We conducted a scoping review of studies published between January 2022 and March 2026 across major databases, identifying 12 studies on LLM-based rare disease patient education and communication. Data were extracted on study characteristics, application scenarios, model usage, and evaluation methods, and synthesized using descriptive and qualitative analyses. The literature is highly recent and dominated by general-purpose models, particularly ChatGPT. Most studies focus on patient question answering using curated question sets, with limited use of real-world data or longitudinal communication scenarios. Evaluations are primarily centered on accuracy, with limited attention to patient-centered dimensions such as readability, empathy, and communication quality. Multilingual communication is rarely addressed. Overall, the field remains at an early stage. Future research should prioritize patient-centered design, domain-adapted methods, and real-world deployment to support safe, adaptive, and effective communication in rare diseases.
Abstract Health analytics increasingly relies on variables held by different entities, such as clinical, laboratory, environmental, and genomic data. Due to legal, ethical, and social acceptability constraints, these vertically partitioned data often cannot be shared across organizations holding them. Conducting statistical analyses in such settings requires methods that protect privacy. We introduce VALORIS (Vertically partitioned Analytics under the LOgistic Regression model for Inference in Statistics), a novel method that enables lossless statistical inference (equivalent to the pooled analyses) under a logistic regression model without disclosing any individual-level data–including the outcome variable. VALORIS is a practical, one-shot algorithm that requires no third-party coordinator. The privacy-preserving properties of VALORIS were mathematically assessed, and a privacy-aware setting-dependent framework was provided to ensure individual-data privacy. We demonstrate the accuracy and feasibility of VALORIS through the investigation of potential factors associated with kidney failure among pediatric patients with chronic kidney disease using real health data from Necker-Enfants Malades Hospital. We further validate the proposed algorithm on a larger scale with a reproducible application using the MIMIC-IV database.
Background: Timely, uncertainty-aware forecasting from irregular electronic health records (EHR) can support critical-care decisions, yet most approaches either impute to a grid or sacrifice interpretability. We introduce StructGP, a continuous-time multi-task Gaussian process that couples process convolutions with differentiable structure learning to uncover a sparse, ordered directed acyclic graph (DAG) of inter-variable dependencies while preserving principled uncertainty. We further propose LP-StructGP, which augments StructGP with latent pathways-shared, temporally shifted trajectories inferred via subject-specific coupling filters and a softmax gating mechanism-to capture cross-patient progression patterns. Both models are trained under sparsity and acyclicity constraints (augmented Lagrangian, Adam) using scalable low-rank updates. Results: In simulations, the approach reliably recovers ground-truth graphs (Structural Hamming Distance approaching 0 as cohorts grow) and pathway assignments (high Adjusted Rand Index). On a MIMIC-IV septic shock cohort (n=1,008; norepinephrine, creatinine, mean arterial pressure), StructGP improves short-horizon (6 h) forecasting over independent-task baselines (average RMSE 0.68 [95
Abstract Extracting temporal information from unstructured clinical narratives is a foundational step toward automated patient timeline generation, a capability that has been proposed as having potential for rare disease diagnosis and care coordination, though prospective clinical validation remains future work. We present a comprehensive framework for temporal relation extraction from French clinical text, addressing a critical gap in non-English clinical NLP resources. We developed specialized annotation guidelines tailored to French medical language and created an annotated corpus of 490 clinical reports from Necker Hospital with 12,464 entity-relation pairs, achieving strong inter-annotator agreement (F1 $$\ge$$ 0.94 for core entities). Our comparative evaluation of modern AI approaches—including transformer-based models, large language models, and parameter-efficient fine-tuning (PEFT)-demonstrates that PEFT with CamemBERT-bio-base achieves the strongest temporal relation extraction performance (F1=0.82–0.87 for major relation types), significantly outperforming traditional approaches and matching few-shot large language models with greater computational efficiency. Entity consolidation substantially improves named entity recognition across all methods (DATE F1=0.96). This work provides validated temporal relation extraction methods as a technical foundation for future patient timeline generation systems. We discuss the pathway toward clinical integration, including deployment requirements, governance considerations, and the prospective validation studies needed to confirm clinical utility—particularly for rare genetic disease populations where automated temporal pattern recognition could support earlier diagnosis.
Abstract Objectives Rare diseases often require longitudinal monitoring to characterise progression, yet much clinical information remains locked in unstructured electronic health records (EHRs). Efficient recovery of such data is critical for accurate prognostic modelling and clinical trial preparation. We aimed to develop and evaluate a small language model (SLM)-based pipeline for extracting longitudinal information from French clinical notes of patients with rare kidney diseases. Methods As a use case, we focused on serum creatinine, a key biomarker of kidney function. We analyzed 81 clinical notes comprising 200 measurements (triplet of date, value and unit). Four open-source SLMs (Mistral-7B, Llama-3.2-3B, Qwen3-4B, Qwen3-8B) were systematically tested with different prompting strategies in French and English. Outputs were post-processed to standardize formats and resolve inconsistencies, and performance was assessed across model size, prompting, language, and robustness to text duplication. Results All SLMs extracted structured triplets, with F1-scores ranging from 0.519 to 0.928 (Qwen3-8B), outperforming the rule-based baseline. Larger models generally performed better, while prompting strategy and language had modest effects across models. SLMs also showed variable robustness to duplicated content common in real-world EHR notes. Discussion Lightweight, locally deployable language models can accurately extract longitudinal biomarkers from unstructured clinical notes. Our findings highlight their practicality for rare diseases where data scarcity often limits task-specific model training. Conclusion SLMs provide a privacy-preserving and resource-efficient solution for recovering longitudinal biomarker trajectories from unstructured notes, offering potential to advance real-world research and patient care in rare kidney diseases. 1) What is already known? Longitudinal monitoring is essential in rare kidney diseases, yet key biomarker data are often locked in unstructured clinical notes. Large language models (LLMs) have shown strong performance in clinical text processing tasks but face major challenges related to privacy, computational cost, and implementation feasibility in healthcare settings. Small language models (SLMs) are emerging as lightweight, locally deployable alternatives whose potential for clinical applications is increasingly recognized. 2) What does this paper add? This study provides the first real-world evaluation of SLMs for extracting longitudinal biomarker measurements in rare kidney disease cohorts. It introduces and validates an efficient extraction pipeline that combines document preselection, SLM prompting, and post-processing to accurately retrieve biomarker measurements from French clinical notes. The findings show that SLM-based extraction can help mitigate data scarcity in rare diseases, thereby improving prognosis modeling and supporting clinical research.
We develop and evaluate a structure learning algorithm for clinical time series. Clinical time series are multivariate time series observed in multiple patients and irregularly sampled, challenging existing structure learning algorithms. We assume that our times series are realizations of StructGP, a k-dimensional multi-output or multi-task stationary Gaussian process (GP), with independent patients sharing the same covariance function. StructGP encodes ordered conditional relations between time series, represented in a directed acyclic graph. We implement an adapted NOTEARS algorithm, which based on a differentiable definition of acyclicity, recovers the graph by solving a series of continuous optimization problems. Simulation results show that up to mean degree 3 and 20 tasks, we reach a median recall of 0.93 while keeping a median precision of 0.71 edges. We further show that the regularization path is key to identifying the graph. With StructGP, we proposed a model of time series dependencies, that flexibly adapt to different time series regularity, while enabling us to learn these dependencies from observations.
Life sciences research increasingly relies on variables held by different entities, such as clinical, laboratory, environmental, and genomic data. Due to legal, ethical, and social acceptability constraints, these data often cannot be shared across organizations holding them. As a result, they cannot be pooled, and analyses must be conducted within the framework of vertically partitioned data. Supporting such analyses requires methods that protect privacy. However, the mere fact that line-level data are not exchanged should not be mistaken for true privacy protection. We introduce VALORIS (Vertically partitioned Analytics under the LOgistic Regression model for Inference in Statistics), a novel method that enables statistical inference under a logistic regression model without disclosing any individual-level data—including the outcome variable. VALORIS is a practical, communication-efficient algorithm that requires no third-party coordinator. Most importantly, it includes a novel framework for evaluating privacy, allowing users to distinguish among different levels of privacy preservation.
OBJECTIVE:Individual patients' data sharing requires interoperability, security, ethical, and legal compliance. The aim was to assess the landscape and sharing capacities between endocrine researchers. DESIGN:A standardized survey (SurveyMonkey®) with 67 questions was sent to European Network for the Study of Adrenal Tumors centers. METHODS:Answers were counted as absolute numbers and percentages. Comparisons between inclusiveness target countries (ITC) and non-ITC (defined by Cooperation in Science & Technology Action) were performed using Fisher's exact test. RESULTS:Seventy-three centers from 34 countries answered the survey. Electronic health record (EHR) systems are now the main source of data (90%). However, significant variability was reported, entailing >35 EHR providers, and variable data collected. Variable stakeholders' implication for enabling data sharing was reported, with more lawyers (P = .023), patient representatives (P < .001), ethicists (P = .002), methodologists (P = .023), and information technology experts (P < .001) in non-ITC centers. Implication of information technologies experts for data collection and sharing was underwhelming (33%). Funding for clinical research was higher in non-ITC than in ITC for clinical trials (P = .01) and for registry-based and cohort studies (P = .05). However, for retrospective studies addressing a specific clinical question, the funding was either very low (<10%) or nonexistent for both ITC and non-ITC (37% and 46%, respectively), with no dedicated funding for information technology (86%) and ethical and regulatory aspects (88%). CONCLUSIONS:In the absence of dedicated funding for retrospective research, current requirements for data sharing are obstacles.
The integration of big data and artificial intelligence (AI) has revolutionized biomedicine, enhancing our understanding of diseases and health care practices. Although AI has shown remarkable success in some medical fields, its application in nephrology faces challenges because of the complex disease mechanisms and intricate physiology. These obstacles are further compounded in rare diseases, affecting <1 in 2000 people, where data scarcity and clinical complexities create additional challenges for AI in accurate disease characterization and prediction. Rare kidney diseases encompass >150 different conditions, with significant clinical and genetic heterogeneity, posing unique challenges for AI applications. Embracing AI for rare kidney diseases is essential, not only for driving the discovery of novel genes, pathways, and mechanisms relevant to both rare and common diseases, but also for shortening the diagnostic odyssey faced by patients with rare conditions, a goal regarded as the most urgent and transformative need in rare disease care. Recent reviews highlight AI applications in nephrology, focusing on big data sources, decision support systems, imaging data, multi-omics integration, and genotype-phenotype analysis. This review explores the current landscape of AI in rare genetic kidney diseases, examining key challenges and advancements in disease characterization and clinical decision support, with an emphasis on hypothesis generation using unsupervised methods and generative AI. It shows how AI can empower physicians to interpret complex data sets, identify patterns, and generate insights that can lead to improved patient outcomes and innovative medical research for rare genetic kidney conditions.
BACKGROUND:Pediatric emergency departments face overcrowding, often driven by non-urgent consultations. Telephone triage, supported by clinical decision support systems (CDSSs), offers a potential solution to improve decision accuracy and reduce unnecessary visits. However, pediatric-specific CDSSs are scarce and underexplored. OBJECTIVE:This study aimed to evaluate the impact of a pediatric-specific CDSS, PED-IA, on decision-making accuracy, confidence, and response time. METHODS:PED-IA is an ontology-based CDSS featuring a rule-driven inference engine and a dynamic interface that guides practitioners through structured clinical reasoning. A crossover study was conducted with 51 practitioners who had to answer clinical cases with and without the CDSS. Decision accuracy, confidence, and response times were measured, and satisfaction was assessed through questionnaires. RESULTS:The CDSS significantly improved decision accuracy from 52.9 % to 76.0 % (+23.1 %, p < 0.01) and increased confidence levels by 0.65 points on a 10-point scale (p < 0.01). Residents benefited the most, with an improved accuracy (odds ratio of 3.70 [2.15, 6.36]). Response times increased by an average of 261.8 s per case (p < 0.01). Practitioners expressed high satisfaction, with 88.2 % finding the system useful for decision-making and 84.3 % believing it could reduce stress in clinical practice. CONCLUSION:The PED-IA CDSS significantly enhances triage decision accuracy and user confidence, making it a promising system for clinical practice and medical education. Practitioners viewed the system positively and identified its long-term time-saving potential. Future works should focus on refining system ergonomics and exploring hybrid models that combine data-driven and logic-based approaches to improve usability and adaptability.
Objectives:To evaluate large language models (LLMs) for extracting temporal relations from pediatric rare disease clinical reports to enable automated patient timeline creation. Materials and Methods:We developed a temporal relation extraction framework for electronic health records, using 25 clinical reports from a pediatric rare disease hospital. We implemented few-shot prompting with 3 different LLMs in secure environments. Results:Our findings reveal that binary classification significantly outperforms multi-class approaches for temporal relation extraction, with best F1 scores reaching 0.70 for simpler relations while more complex relations remain challenging (F1: 0.03-0.40). Mistral 22B emerged as the strongest overall performer, though model superiority varied by relation type. Discussion:The dramatic performance improvement from reducing cognitive load (binary vs multi-class classification) demonstrates that task formulation critically impacts LLM effectiveness in specialized clinical domains. Our few-shot approach successfully enables temporal relation extraction from French pediatric texts while maintaining data privacy through local deployment, offering a viable methodology for healthcare institutions with strict data governance requirements. Conclusion:Our few-shot prompting approach demonstrated promising results in secure environments. This methodology allows technique sharing without exposing sensitive data, advancing research possibilities for clinical natural language processing in restricted settings.
Introduction: General Practitioners (GPs) play a key role of gatekeeper, as they coordinate patients' care. However, most of them reported having difficulty to refer patients to hospital, especially in semi-urgent context. To facilitate the referral of semi-urgent patients, we implemented an e-referral platform, named SIPILINK, within 4 wards from a large public French hospital (internal medicine, diabetology, gynaecological surgery and oncology wards). Here, we aimed to evaluate the SIPILINK e-referral platform after 2 years of implementation. Methods: The evaluation included a multidimensional assessment based on the RE-AIM framework with the analysis of implementation, requests, health professionals' satisfaction, and estimated hospital payment. Results: Over 2 years of implementation, GPs sent 113 requests to hospital. Hospital respected the time of response requested by GPs in 93 % of cases and proposed a consultation or hospitalization in respectively 40.7 % and 10.6 % of cases. 100 % of GPs and 78 % of Hospital Practitioners (HPs) were satisfied with the quality of exchanges. 77 % of HPs and 100 % of Care Pathway Managers (CPMs) found that patient care pathways were improved. Nearly all practitioners would recommend this platform for patient referrals. Discussion: SIPILINK shows promise in streamlining the referral process, enhancing communication, and improving patient care pathways. Further studies including the impact on the quality of care, are needed to assess its effectiveness and sustainability in healthcare settings.
Background: If hospital Clinical Data Warehouses are to address today's focus in personalized medicine, they need to be able to track patients longitudinally and manage the large data sets generated by whole genome sequencing, RNA analyses, and complex imaging studies. Current Clinical Data Warehouses address neither issue. This paper reports on methods to enrich current systems by providing provenance data allowing patient histories to be followed longitudinally and managing the linking and versioning of large data sets from whatever source. The methods are open source and applicable to any clinical data warehouse system, whether data schema it uses. Method: We introduce gITOMMIx, an approach that overcomes these limitations, and illustrate its usefulness in the management of medical omics data. gITOMMIx relies on (i) a file versioning system: git, (ii) an extension that handles large files: git-annex, (iii) a provenance knowledge graph: PROV-O, and (iv) an alignment between the git versioning information and the provenance knowledge graph. Results: Capabilities inherited from git and git-annex enable retracing the history of a clinical interpretation back to the patient sample, through supporting data and analyses. In addition, the provenance knowledge graph, aligned with the git versioning information, enables querying and browsing provenance relationships between these elements. Conclusion: gITOMMIx adds a provenance layer to CDWs, while scaling to large files and being agnostic of the CDW system. For these reasons, we think that it is a viable and generalizable solution for omics clinical studies.
BACKGROUND:Telephone triage could limit admissions to emergency departments. However, telephone triage is challenging in pediatrics due to nonspecific symptoms, reliance on parental description, and emotional distress. Clinical decision support systems (CDSSs) could improve the accuracy and quality of telephone triage. Despite proven benefits, current CDSSs are not well suited to the nuances of pediatrics. This study aims to develop a CDSS for pediatric emergency telephone triage. METHODS:We developed a formal knowledge base (KB) for pediatric telephone triage inspired by the ontology model and implemented a generic medical reasoning system that mimics the clinical reasoning used in pediatric emergency triage. The CDSS is built in three layers (a knowledge layer, a Python-based decision layer, and a web interface layer) and provides real-time recommendations. We assessed its accuracy on 96 fictitious clinical cases. RESULTS:The CDSS uses an ontology-oriented KB that includes 303 concepts and 1780 axioms and a generic algorithm that provides recommendations based on user input, exploring and updating decisions continuously. It demonstrated 100 % internal validity compared to written recommendations and 77.1 % accuracy compared to a trio of experts. The 22.9 % discrepancies were due to experts using additional elements not documented in the written recommendations (11.5 %) or experts making different decisions despite consistent rules in the textual recommendations (10.4 %), emphasizing the challenges of standardized guidelines in this narrow but complex field. DISCUSSION/CONCLUSION:The CDSS provides explainable and interpretable recommendations designed to alleviate healthcare professionals' cognitive load so that they can focus on complex clinical situations. Future improvements involve enriching the KB, enhancing user interaction with patient-friendly language, and combining this knowledge-based approach with data-driven approaches.