
Pregnancy complications are a leading cause of maternal and neonatal mortality worldwide. Understanding their underlying mechanisms is hindered by dispersed evidence across thousands of studies and complex biological interactions. We present PregBase, a comprehensive knowledge base for pregnancy research comprising automated extraction, validation, and exploration. PregBase was constructed by utilising large language models (LLMs) to extract relationships from literature; 13 LLMs were benchmarked across seven prompting strategies, with ensemble shuffle prompting outperforming single-strategy alternatives. A three-tier validation pipeline combining ontology mapping, graph neural networks, and statistical analysis produced PregKG, a knowledge graph containing 64,087 associations between 13,303 biomedical entities spanning 155 semantic types and 50 vocabularies across 8 relationship types from 8420 articles. Link prediction validated PregBase’s inference capability beyond extracted knowledge, recovering established clinical interventions, reconstructing canonical hormonal pathways across maternal-placental-fetal compartments, and identifying novel biomarker candidates for preterm birth, gestational diabetes, and preeclampsia, supported by genetic and expression databases. An interactive web interface (https://pregknowledgebase.com) has been created to enable further exploration via conversational queries. This work provides a scalable foundation for systematic discovery in maternal health research.
Deciphering herb-symptom associations (HSAs) and their underlying mechanisms is essential for the modernization of Traditional Chinese Medicine. However, current methods primarily focus on graph topology modeling, which often relies on coarse-grained structures that overlook the heterogeneous contributions of individual proteins and are susceptible to spurious correlations. To address these limitations, we proposed CIGMA, a novel framework explicitly designed for out-of-distribution generalization in biological network association prediction. First, we introduced a gated graph matching module integrated with multi-view distillation, which captured fine-grained, context-aware interactions by enforcing consistency across local, global, and interaction-based representations. Second, to mitigate spurious correlations, we implemented a causal representation invariance learning mechanism that employed invariant risk minimization to prioritize invariant and robust pathways over observational biases. Experiments on the benchmark dataset demonstrated that CIGMA achieved average AUPRC improvements of 2.84% and 2.26%, and AUROC improvements of 3.02% and 1.83% under cross validation and independent validation, respectively. Furthermore, case studies on classic formulas like Siwu Decoction and Xuefu Zhuyu Decoction demonstrated the model’s ability to provide plausible mechanistic deconstruction and identify stable and environment-invariant protein pathways. CIGMA provides a powerful, generalizable, and interpretable framework for HSA prediction, advancing the systematic discovery of herb-based therapeutic mechanisms.
Endometriosis is a widespread gynecological disorder causing severe pain and infertility, with diagnosis currently relying on slow, costly, and risky laparoscopy. This highlights the critical need for non-invasive imaging diagnostics using transvaginal ultrasound (TVUS) and magnetic resonance imaging (MRI). A key challenge is that patients typically receive only one scan modality in practice, despite TVUS and MRI offering differing diagnostic strengths for endometriosis signs like Pouch of Douglas (POD) obliteration and bowel nodules (BN). Previous work partially addressed this challenge by leveraging unpaired multi-modal data for detecting a single marker: Pouch of Douglas (POD) obliteration. However, this is restrictive because endometriosis signs, such as POD obliteration and bowel nodules (BN), often provide correlated diagnostic cues. Capturing these correlations is essential for accurate detection of endometriosis imaging signs, particularly when combined with multi-modal learning, as each modality offers complementary strengths for different signs. To overcome these limitations, we propose EndoFusion, a novel unpaired multi-modal, multi-label learning framework that enables the detection of POD and BN from TVUS and MRI. Our approach introduces three key innovations: (1) label-based pairing, mixup, and cross-modal feature exchange for robust single-modality inference; (2) Dynamic Mutual Knowledge Distillation (DMKD), which adaptively selects teachers using a worst-student-oriented strategy for effective cross-modal transfer; and (3) label correlations modeling with multi-head attention and a specialized loss to handle imbalance and boost accuracy. This design ensures that knowledge from the superior modality and from co-occurring signs is effectively transferred, mitigating modality-specific weaknesses and improving robustness in imaging sign detection. Experiments on our endometriosis dataset show that our method significantly outperforms all comparison methods, achieving an average AUC of 0.827 (95% CI: 0.790-0.861) when evaluated using single-modality inference. These results represent an initial proof-of-concept toward multi-modal, non-invasive assessment of selected endometriosis imaging signs from MRI and TVUS.
Clinical large language models (LLMs) are increasingly used for documentation, diagnosis, and decision support, but their opaque reasoning can limit clinician trust, regulatory assessment, and safe deployment. Explainability research has expanded rapidly, yet existing reviews largely address traditional machine learning or general-domain LLMs. We conducted a PRISMA-ScR scoping review to map explainability approaches for decoder-only clinical LLMs with over one billion parameters, searching PubMed, Scopus, Web of Science, ACM Digital Library, and arXiv through early 2026. Among 69 included studies, LLM-native generative and interactive methods dominated (58.0%, n = 40), spanning chain-of-thought rationales, retrieval-augmented evidence citation, and agentic decomposition. Intrinsic by-design methods accounted for 24.6% (n = 17); post-hoc XAI methods accounted for 17.4% (n = 12). General medicine was the most represented clinical domain, and diagnosis was the dominant task. Proprietary models were used in 75.4% of studies, yet every mechanistic analysis relied on open-source models, revealing a transparency asymmetry: most deployed models are the least transparent. Although 59.4% of studies quantitatively evaluated explanations, metrics remain non-standardized and rarely assess faithfulness. Local explanations predominated, and no study prospectively evaluated explanations in live clinical workflows. These findings show that clinical LLM explainability has shifted toward fluent generative rationales, but evidence that such explanations reflect model reasoning remains limited. To support trustworthy deployment, we highlight three regulatory priorities: prioritizing explanations that enable independent verification or logic auditing over plausibility-only rationales; preferring inspectable models where regulatory documentation is required; and prospectively validating explanations in clinical workflows before scaling.
Skin cancer diagnosis relies on visual interpretation of lesion morphology, yet most deep learning models function as black boxes with limited clinical interpretability. This study presents an explainable diagnostic framework that integrates a fine-tuned Xception network with a Convolutional Block Attention Module (CBAM) for skin lesion classification and a domain-adapted vision-language model (VLM) for generating clinically meaningful explanations. The classifier was trained on a preprocessed and augmented HAM10000 dataset, while Grad-CAM++ was employed to highlight discriminative lesion regions. These heatmaps were manually annotated with expert morphological descriptions to construct a novel image-text corpus comprising 347 samples, which was subsequently used to parameter-efficiently fine-tune MedGemma-4B using Low-Rank Adaptation (LoRA). The proposed classifier achieved a test accuracy of 85.66% with an AUC of 0.9392 across seven lesion classes. Furthermore, the fine-tuned VLM improved the average explanation quality score from 62.14 to 87.86, representing a 25.72-point gain over the baseline according to automated LLM-based evaluation. The resulting framework provides clinicians with classification predictions, visual attention maps, and expert-aligned textual rationales, offering a practical step toward trustworthy and explainable AI for dermatological diagnosis.
Alzheimer’s disease (AD) is a prevalent neurodegenerative disorder where early diagnosis is pivotal for effective intervention, yet it is hindered by subtle pathological feature differences among AD subtypes and severe class imbalance in medical imaging datasets. Existing structural magnetic resonance imaging (sMRI) and resting-state functional MRI (rs-fMRI) multimodal fusion methods for AD diagnosis mostly adopt simple concatenation or summation without fine-grained cross-modal alignment and interaction. To address these issues, we propose a Deep Cross-Branch Multi-Modal Fusion Network (DCMFNet) for early AD diagnosis. We first preprocess sMRI and rs-fMRI to extract ROI-based features, then perform dimension unification and normalization to realize cross-modal feature alignment. A novel Deep Cross-branch Multi-modal Feature Fusion (DCMF) module with three parallel branches and a dual-pathway cross-modal branch is designed to fully mine complementary and correlated cross-modal information, and the fused features are input into a Transformer encoder for classification. Moreover, we introduce the Logit Adjustment Cross-Entropy (LACE) loss to mitigate class imbalance by correcting decision boundaries based on class prior probabilities, enhancing the recognition of minor classes. The model is evaluated on a private clinical dataset. Experimental results show that DCMFNet outperforms traditional machine learning methods and state-of-the-art deep learning models in six binary AD subtype classification tasks, with the LACE loss and DCMF module effectively alleviating class imbalance and improving cross-modal feature representation. This work provides a reliable multimodal fusion framework for early AD diagnosis and reduces the diagnostic burden on healthcare professionals.
Diagnostic accuracy in Large Language Models (LLM) is an increasing concern as physicians employed LLMs into their medical practice. We measured whether diagnostic correctness in LLMs changed after a certainty challenge or clinician specialty framing. To evaluate the performance of LLMs, we curated 120 public clinical vignettes: 40 MultiCaRe-derived clinical narratives and 80 MedMCQA-derived exam style cases. Ten proprietary and open-weight LLMs were tested across four prompt conditions with two passes per condition, yielding 9,600 responses. Diagnostic correctness was evaluated using an LLM-as-a-judge approach. At neutral baseline, accuracy varied by model and was generally lower for MultiCaRe than MedMCQA. GPT-5 had the highest baseline and post-challenge accuracy (75.0% and 76.7%); Claude Sonnet 4 had the highest accuracy changing flip rate (AcFR; 57.5%) and largest post-challenge accuracy loss (-37.5 percentage points). Across all neutral model-case pairs, corrected accuracy decreased from 50.6% to 41.7% after 'Are you sure?' (AcFR = 20.4\%). Correct-to-incorrect transitions (n = 176) exceeded incorrect-to-correct transitions (n = 69). Specialty context prompts produced smaller changes. Adjacent specialty prompts increased overall accuracy by 2.9 percentage points, and differential specialty prompts decreased accuracy by 1.5-1.8 percentage points. The 'Are you sure?' challenge was the strongest source of diagnostic instability in the study; case source also strongly affected accuracy. Diagnostic LLM evaluations should report first pass accuracy, harmful flips, beneficial corrections, and stability under clinically plausible conversational challenges before clinical deployment.
The complex spatial organization of the tumor microenvironment (TME) plays a critical role in colorectal cancer progression, yet its quantitative characterization from histopathological images remains challenging and often relies on subjective pathological assessment. In this study, we develop a CAMuTILS-based TME scoring framework for quantitative analysis and prognostic evaluation in colorectal cancer. First, we propose CAMuTILS, a panoptic segmentation network for histopathological images. The model adopts a dual-branch U-shaped architecture to simultaneously segment tissue regions and cell nuclei at 1.0 MPP and 0.5 MPP resolutions. A cross-channel Transformer module is introduced for multi-scale feature alignment, together with an iterative attention-based (iAFF) cross-branch interaction mechanism and histology-informed spatial constraints that allow tissue regions to guide nuclear classification. Based on the segmentation results generated by CAMuTILS, TME features were quantified across five biological themes including epithelial architecture, stromal characteristics, tumor-infiltrating lymphocytes, necrosis, and spatial interactions. A total of 45 prognostic features were selected to construct a computational risk score termed CRS. Experimental results demonstrate that CAMuTILS achieves superior segmentation performance on the PanopTILs dataset compared with existing methods. Survival analysis on the TCGA colorectal cancer cohort shows that CRS is an independent predictor of progression-free survival and provides prognostic information complementary to conventional clinicopathological variables. External validation on an independent CPTAC cohort further supports the robustness and generalizability of the proposed framework across institutions. These findings highlight the potential of computational pathology to enable quantitative TME characterization and support precision prognostic assessment in colorectal cancer.
In recent years, generative models have shown remarkable capabilities in synthesizing realistic human motion, with applications ranging from animation to virtual reality. However, their potential in clinical and rehabilitation settings remains underexplored. In this work, we introduce a conditional diffusion-based generative framework for rehabilitation-oriented motion synthesis, which directly operates on joint-angle representations of full-body movement. Unlike most existing approaches that rely on joint positions, our method generates motion in a clinically meaningful space that explicitly encodes joint range of motion, aligning the generation process with how motor performance is assessed in rehabilitation practice. This design enables subject-independent modeling while improving the interpretability of the generated movements from a clinical perspective. We propose a comprehensive evaluation protocol by combining qualitative and quantitative metrics, including simulation visualizations, similarity analysis, and automated assessment of simulations adherence to users input. Experiments based on cross-subject and leave-one-combination-out settings demonstrate the model’s ability to generate plausible, contextually accurate motion sequences, with improved generalization when using joint angle representations, achieving superior performance compared to a position-based approach. Despite limitations due to dataset size and gesture diversity, results support the feasibility of generating rehabilitation-oriented motion simulations, motivating future investigation in personalized rehabilitation scenarios.
This work proposes a voice-based clinical decision support tool that detects a generalized behavioral health condition, defined by Major Depressive Disorder (MDD) and Generalized Anxiety Disorder (GAD), from spontaneous speech. We compare Random Forest, XGBoost, and fully connected deep neural network classifiers trained on high-dimensional voice representations, including X-vector, Wav2Vec2, TRILLsson, and HuBERT embeddings. Models are trained using validated symptom instruments, including the PHQ-8 and GAD-7. To enhance clinical safety, we incorporate an ‘uncertain’ classification category for low-confidence samples. Our model achieves a sensitivity of 0.68 and specificity of 0.80, demonstrating that a consolidated vocal biomarker metric provides a scalable, non-invasive mechanism to identify health conditions in both in-clinic and remote settings.
The design and evaluation of Retrieval-Augmented Generation (RAG) pipelines remain fragmented, with system components often tailored to specific datasets and tasks. In this paper, we propose STRAGMED, a simple yet effective RAG pipeline to provide a standardized baseline for system engineering and retrieval evaluation. Our experiments show that a hybrid pipeline combining a retriever, reranking at each step, and Reciprocal Rank Fusion (RRF), yields consistent improvements across medical datasets, outperforming retrieval-only and individually reranked outputs. These results highlight a robust and generalizable baseline configuration for medical RAG systems, enabling researchers to focus on task-specific optimizations such as query augmentation, model selection, and retrieval refinement.
This systematic review evaluates the effectiveness of Artificial Intelligence (AI) in preventing cardiovascular disease (CVD). We conducted a systematic review using Joanna Briggs Institute (JBI) methodology, with the protocol registered under CRD42022369548 on PROSPERO. We included studies on all populations who receive or provide care and used AI for risk prediction and prevention of CVD. We searched seven databases from 1999 till October 2022. Two reviewers independently screened the identified records and extracted data from the included studies. We conducted a subgroup analysis on studies that compare AI to standard care. We assessed the risk of bias using a Modified IJMEDI and PROBAST tool, and data are synthesized in narrative form, considering various study aspects. This review adheres to PRISMA guidelines for systematic reviews and meta-analyses. After screening 7,505 identified records, 266 articles were included in the review. Risk prediction and early disease diagnosis were the most common prevention modalities. The subgroup analysis of 16 studies that compared these interventions with standard care included 51,687,627 participants. The Framingham Risk Score was the most used comparator, reported in 8 of the 16 studies (50
Accurate localization of the SubThalamic Nucleus (STN) during Deep Brain Stimulation (DBS) surgery is critical for therapeutic efficacy and is commonly supported by intraoperative MicroElectrode Recordings (MERs). While deep learning approaches have shown promising performance in automatic STN identification, their limited transparency hinders their clinical adoption. In this work, we present an interpretable deep learning pipeline for MERs classification coupled with a structured validation framework aimed at assessing the clinical relevance of the model reasoning. To this aim, the classification output from a patch-based convolutional neural network with self-attention is combined with Grad-CAM relevance maps. Alignment between relevance maps and manual annotations from 3 expert neurologists with 27, 20, and 11 years of experience was quantified using overlap-based metrics, while perceived transparency, usefulness, and trustworthiness were assessed through Likert-scale questionnaires. Results show competitive classification performance (0.92 ± 0.07 AUC) and consistent agreement between automatic explanations and expert reasoning, with a maximum Dice Score of 0.77 ± 0.11, alongside high clinician acceptance and perceived interpretability. These findings suggest that structured, expert-centred validation of XAI can provide a meaningful contribution toward trustworthy AI-assisted decision support systems in intraoperative neurosurgery.
Myocardial Infarction (MI) remains one of the leading causes of death worldwide. Electrocardiogram (ECG) is a reliable diagnostic tool for MI and its interpretation requires cardiology expertise. Advancement of multimodal large language models (MLLMs) attracted attempts to interpret ECG images, but prior studies showed low accuracies. To address the issue, we investigated a few strategies for pretrained MLLMs: (1) inclusion of domain diagnostic prompt instruction, (2) few-shot technique by providing annotated examples in the prompt, and (3) adjustment of model hyperparameters. On a corpus of 928 annotated 12-lead ECG images, our results utilizing Google Gemini 2.5 Pro and OpenAI GPT-4o for MI versus non-MI classification showed noticeable performance improvements. With 30-shot prompting, Gemini achieved 76.08 ∼ 30
With rapidly growing medical data, some of the biomedical knowledge graphs are often incomplete and unreliable due to scattered evidence, which limits clinical discovery. We developed an explainable link prediction framework that combines diverse models such as DistMult (symmetric relations), ComplEx (asymmetric relations), and SimplE (diverse head–tail interactions). To reduce the bias of individual model ranking, we aggregate model scores using an ensemble strategy and rank candidate links. We retrieve supporting graph paths and quantify edge contributions using Shapley-based attribution, and validate predictions against biomedical literature using a domain-specific NLI model, which enables expert review. Moreover, experiments on 9,102 test triples showed improved ranking performance, with correct links appearing higher and more frequently among top results (MRR 0.299 vs. 0.269; Hits@10 0.484 vs. 0.443 for ComplEx). This provides a reliable and verifiable pipeline for biomedical knowledge graph completion.
Kidney transplantation (KT) remains the optimal treatment for end-stage renal disease, yet severe organ scarcity necessitates maximally efficient donor-recipient matching (DRM) strategies. Current allocation systems rely on rule-based scoring mechanisms that cannot predict individualized treatment effects (ITE) for specific donor-recipient combinations. Moreover, transplant registries exhibit substantial systematic biases from allocation policies that introduce confounding, which standard machine learning models tend to reproduce rather than correct. This paper addresses the challenge of transforming counterfactual treatment estimation into counterfactual treatment optimization for DRM in KT. Building on our Confounding-Adjusted Model (CAM) for predicting outcomes of alternative donor-recipient pairings, this work introduces a specialized search strategy to identify optimal matches from a high-dimensional treatment space. We leverage gradient-based optimization combined with Logit masking and Gumbel-Softmax relaxation techniques to efficiently identify the optimal counterfactual recipients for a given donor ensuring that the recommendations satisfy clinical donor-recipient compatibility criterion. Applied to a comprehensive dataset of 186,000 kidney transplants from the Scientific Registry of Transplant Recipients (2000–2023), our framework demonstrates substantial improvements over baseline prediction models with nearest-neighbor retrieval validation revealing that our counterfactual optimization corresponds to 394 additional days of observed post-transplant survival (p < 0.001).
Acute coronary syndromes (ACS), including ST-elevation myocardial infarction (STEMI) and non-ST-elevation myocardial infarction (NSTEMI), remain leading causes of mortality worldwide. Despite advances in diagnosis and treatment, single-omics approaches have proven insufficient to capture the molecular complexity underlying ACS pathophysiology. Consequently, we adopted a multilayer network approach to construct phenotype-specific networks for STEMI and NSTEMI, alongside a control multilayer network derived from patients with stable angina pectoris (SAP). The multilayer networks were constructed using data collected from 200 patients within the CardioSCOPE project, integrating one metabolomics layer and one microRNA layer. Nodes represented molecular features, while edges were defined based on Pearson correlation coefficients. Network analyses included interlayer connection investigation, hub identification, and community detection, whose results were used to select a compact panel of discriminative features that achieved a cross-validated AUC of 0.84 (95
Hepatocellular carcinoma (HCC) incidence is rising alongside the global type 2 diabetes (T2D) pandemic, with alcohol use disorder (AUD) as a major risk factor. However, a substantial population with undiagnosed subclinical AUD (sAUD) remains overlooked in current estimates. Using XGBoost and One-Class SVM, we identified hidden sAUD populations whose clinical profiles resemble those of diagnosed AUD patients, suggesting that the alcohol-related burden of HCC in T2D patients may be underestimated by 10–30