
Multi-organ segmentation in medical imaging remains challenging due to its reliance on costly manual annotations. While semi-supervised learning (SSL) mitigates this by utilizing unlabeled data, current methods lack effective integration of medical domain knowledge and struggle to balance local details with global context. To address these, we propose a novel SSL framework that integrates structured, template-based medical text priors with a complementary dual-branch architecture. Our method synergizes a CNN-based V-Net and a Transformer-based SwinUNETR, effectively capturing both local anatomical details and long-range contextual dependencies. We further introduce a task-aware attention-guided text-visual fusion (ATVF) module that dynamically aligns medical text embeddings with visual features at the voxel level, enhancing semantic discrimination. A discrepancy-aware fusion mechanism is also designed to focus learning on uncertain regions by leveraging inter-branch prediction inconsistencies. Extensive experiments across three public datasets (SegTHOR, MM-WHS, MyoPS) under varying annotation ratios demonstrate competitive performance, especially under low annotation regimes. With only 10% labeled data, it achieves a mean Dice score of 70.86% on SegTHOR, surpassing the second-best solution by 5.8%, and demonstrating notable gains on challenging structures such as the esophagus and right atrium. This work provides a robust and annotation-efficient solution for multi-organ segmentation using structured text prompts, demonstrating strong potential for further evaluation in clinical settings.
Objective Temporal lobe epilepsy (TLE) is increasingly recognized as a network disorder that extends beyond the hippocampus. Among the extra-hippocampal regions, the cerebellum has emerged as a remote yet functionally connected node showing consistent metabolic alterations. However, its mechanistic contribution to seizure propagation and disease progression remains largely unclear. The aim of this study was to elucidate the contribution of cerebellar microglial activation to aberrant glucose metabolism and to determine whether metabolic-inflammatory coupling within the cerebellum drives seizure generalization and cognitive decline in TLE. Methods We integrated [18F]FDG PET analysis from 55 patients with TLE and a lithium chloride-pilocarpine rat model of TLE to delineate cerebellar metabolic abnormalities across disease stages. Longitudinal [18F]FDG and [18F]DPA-714 microPET/CT imaging was performed to assess glucose metabolism and microglial activation. Histological and flow cytometric analyses characterized microglial phenotypes. Cerebellar mRNA sequencing with KEGG pathway enrichment analysis and hub gene identification using Cytoscape identified c-Fos as a central regulator of MAPK signaling, which was functionally validated via stereotactic cerebellar administration of the c-Fos inhibitor T-5224. Results Patients with TLE exhibited bilateral cerebellar hypermetabolism. In rats with chronic TLE, elevated [18F]FDG and [18F]DPA-714 uptake in the cerebellum was confirmed, accompanied by Purkinje cell loss and increased proinflammatory cytokines (IL-6, TNF-α, IFN-γ, IL-12p70). Pharmacological inhibition with minocycline or T-5224 normalized cerebellar PET signals, attenuated M1-like microglial polarization, reduced seizure frequency, and improved spatial and recognition memory. Conclusion These findings identify cerebellar metabolic-inflammatory coupling, driven by c-Fos-dependent microglial activation, as a key mechanism promoting epileptic network progression. Targeting this cerebellar axis provides a promising therapeutic avenue for precision neuromodulation in TLE.
Learning-based methods for medical image translation, which synthesize missing modalities from existing ones, offer significant clinical value by overcoming the limitations of individual imaging techniques and reducing patient radiation exposure. Unsupervised methods have garnered significant attention for their capacity to train without paired datasets, thereby addressing the challenges of data acquisition and registration. While GAN-based unsupervised methods have shown notable progress, they are often hampered by issues such as mode collapse and limited adversarial learning capacity. Diffusion models, renowned for their explicit likelihood formulation and superior image generation quality, typically rely on paired data and demand substantial computational resources. To overcome these limitations, we introduce Cycle-LVDM, an unsupervised medical image translation framework based on a latent variational diffusion model that enables fully unpaired image synthesis. Our approach integrates a cycle-consistency loss with the powerful generative capabilities of variational diffusion models to enhance synthetic image quality in an unsupervised setting. Furthermore, we perform diffusion training and sampling in a compressed latent space and employ the Mamba-based DiM architecture as the model backbone, providing favorable scalability characteristics for high-resolution medical image synthesis while reducing computational burden compared with conventional pixel-space diffusion frameworks. Extensive quantitative and qualitative evaluations on a multi-contrast MRI dataset and a cross-modality MRI-CT dataset demonstrate that Cycle-LVDM achieves competitive performance compared with representative unsupervised GAN- and diffusion-based methods. A downstream segmentation evaluation further demonstrates the practical utility of the synthesized images in clinically relevant analysis.
Purpose The differentiation grade of pancreatic ductal adenocarcinoma (PDAC) is a crucial determinant of its aggressiveness and patient prognosis. Recently, advanced diffusion models have been increasingly applied to assess the grading of many tumors. However, the value of these models in discriminating PDAC differentiation grades remains unclear. Methods A retrospective analysis was conducted on 71 patients with pathologically confirmed PDAC between July 2022 and March 2025. Clinicopathological data and conventional imaging features were collected. Eleven diffusion parameters were derived from four advanced diffusion models via multi-model diffusion imaging (DXI) technology: intravoxel incoherent motion (IVIM: f, D∗, D), diffusion kurtosis imaging (DKI: D, K), fractional order calculus (FROC: D, β, μ), and continuous-time random walk (CTRW: D, β, α). Based on postoperative histological differentiation, patients were categorized into a low-grade group (n = 38) and a high-grade group (n = 33). The clinicopathological features, imaging characteristics, and diffusion parameters were compared between the two groups. Receiver operating characteristic (ROC) curves were utilized to evaluate the diagnostic performance of these indicators in predicting the differentiation grade of PDAC, and the DeLong test was employed to compare differences in the area under the curve (AUC). Parameter stability was assessed using bootstrap internal validation with 1000 resamples. Results The high-grade PDAC group exhibited significantly lower values in IVIM-f, DKI-D, and CTRW-α metrics (all p < 0.0045). Among all metrics, CTRW-α had the numerically highest AUC of 0.823. A CTRW-α cutoff of <0.856 diagnosed high-grade PDAC with 81.8% sensitivity and 71.1% specificity. Lymph node status differed significantly between the low- and high-grade groups (p = 0.024). Conclusions DXI technology provides valuable advanced diffusion models and derived parameters for non-invasively assessing PDAC differentiation grade, providing a novel tool for clinical assessment.
Learning meaningful representation constitutes a pivotal problem in constructing foundation models. Nevertheless, the complex anatomical patterns and the random distribution of lesions in medical images pose significant challenges to understanding and disentangling useful representations. Contrastive learning has demonstrated remarkable success in decoupling representations, but measuring the distance in a high-dimensional feature space is still hard. In this paper, we propose a mutual information-based mechanism for quantifying the representation distance. However, collecting millions of samples and constructing a huge positive-negative sample bank for conducting effective contrastive learning is impractical in the medical domain. To address such an issue, we introduce a constrained multiview learning paradigm. Specifically, we conduct a dynamic representation reranking and selection process to enhance the quality of the positive and negative sample pairs. Our method benefits both the continuous MI estimating and the representation significance measuring, enhancing the contrastive learning process and semantic comprehension. Our proposed framework was rigorously evaluated using publicly accessible CT-captured lung lesion segmentation datasets and compared against influential baseline models with either pure CNN modules or transformer modules. The statistical results under the four metrics demonstrate that our proposed framework proficiently optimizes the multi-view contrastive learning process and improves MI maximization-driven representation learning.
The joint analysis of macro-level imaging and micro-level genetic information facilitates a holistic understanding of the pathological processes underlying Alzheimer’s disease (AD). Although recent studies have made advances, most current approaches struggle to exploit the critical features embedded in imaging genetics data and fail to effectively reveal biological interactions. To remedy these shortcomings, this paper proposes a multi-level learning and interactive fusion framework based on large foundation models (LFMs), and accordingly develops an algorithm termed MLLIFA for the diagnosis of AD and the extraction of etiology. Specifically, MLLIFA first employs two LFMs to construct high-quality representations of brain region and gene features. Then, two sparse attention mechanisms are applied to extract key information from the constructed features. Finally, interaction learning is utilized to explore the latent relationships between features within biological contexts, guiding the effective fusion of multi-omics information. Experimental results on the ADNI dataset demonstrate that MLLIFA achieves an AD classification accuracy of 91.22%, outperforming state-of-the-art methods. Moreover, the proposed MLLIFA successfully identifies disease-related brain regions, risk genes, and brain-gene associations. These findings not only provide strong support for the precise diagnosis of AD but also offer new insights into the extraction of etiology.
Purpose: The lack of biomarkers hinders precision use of anti-TNF-alpha agents in Crohn's disease (CD). In this real-world study, we aimed to develop and validate a machine learning model based on ACC-focused multiparametric MRI to predict patient suitability for anti-TNF-alpha agents. Materials and methods: In this prospective dual-center study, 109 CD patients underwent baseline ACC-focused multiparametric MRI, MR enterography, ileocolonoscopy, and serum biomarker assessment. Using 11 machine learning algorithms, an ACC neurophenotype model was built from 584 features to stratify patients into high- or low-risk groups. Outcomes after 25 +/- 4 months were compared between anti-TNF-alpha agents and conventional therapy using propensity-score matching (PSM) and inverse-probability-of-treatment weighting (IPTW), with subsequent validation in a DSS-induced colitis mouse model where stratified mice received anti-TNF-alpha agents or conventional treatment. Results: The ACC model showed strong neurophenotype discrimination (AUC(training) = 0.921; AUC(test) = 0.861, both P < 0.050). Higher ACC neuro-scores correlated with increased blood-brain-barrier-permeability (elevated S100 beta), psychological distress, reductions in key serum neurotransmitters (tryptophan and histidine), and worse abdominal conditions (abdominal pain scores, intestinal inflammation) (r = 0.189-0.578; all P < 0.050). Critically, high-risk patients derived significantly longer progression-free survival from anti-TNF-alpha agents versus conventional therapy (Plog-rank = 0.049), a benefit robust to PSM (Plog-rank = 0.029) and IPTW (Plog-rank = 0.049) adjustments, while low-risk patients showed no benefit. Similarly, in DSS-colitis mice, high-risk mice showed significantly greater improvement with anti-TNF-alpha agents than conventional therapy across brain (neuronal loss, S100 beta, microtubule-associated protein 2), systemic (IL-6), and intestinal (alpha-1-antitrypsin) metrics (all P < 0.050), whereas low-risk mice showed no such benefit. Conclusions: An ACC-focused neurophenotype model derived from multiparametric MRI assessed brain-gut dysregulation, allowing for patient stratification to optimize selection of anti-TNF-alpha agents therapy in CD.
Purpose: Intravoxel incoherent motion (IVIM) and diffusion kurtosis imaging (DKI) are well-established diffusionweighted imaging (DWI) techniques, but both typically require multi-b-value acquisitions and computationally expensive parameter estimation.This study is aimed to develop a physics-informed deep learning-empowered hybrid IVIM-DKI framework (PI-DL-IVIM-DKI) for efficient parameter estimation and preoperative identification of microvascular invasion (MVI) in hepatocellular carcinoma (HCC). Materials and Methods: PI-DL-IVIM-DKI consists of two modules: (1) an optimal b-value design module combining Fisher information matrix analysis and Monte Carlo simulation; and (2) an end-to-end neural network module that estimates hybrid IVIM-DKI parameters based on the optimized b-value set. Parameter fitting and time efficiency were evaluated in both synthetic simulations and in vivo cohort (74 subjects; 56.1 +/- 10.6 years), using conventional nonlinear least squares (NLLS) fitting as the reference. The diagnostic performance of PI-DLIVIM-DKI-derived parameters for preoperative MVI identification was further assessed. Results: The 5-b-value protocol B5 (b = 0, 40, 200, 800, 2000 s/mm2) was selected as the optimal scheme. The deep neural network (DNN) achieved the most accurate parameter estimation under B5 compared with a convolutional network (ConvNet) and a Transformer, and was therefore used as the backbone of PI-DL-IVIM-DKI. PIDL-IVIM-DKI yielded higher fitting accuracy (NRMSE = 0.156) and substantially improved computational efficiency (up to 1.2 & times; 105-fold speed-up) relative to NLLS. In vivo, PI-DL-IVIM-DKI parameters agreed well with the conventional "full-b-value-NLLS" strategy (Pearson r = 0.812-0.993). PI-DL-IVIM-DKI-derived parameters further demonstrated good performance for identifying MVI, with AUC up to 0.765. Conclusion: PI-DL-IVIM-DKI enables accurate and efficient hybrid IVIM-DKI parameter estimation using only four non-zero b-values. PI-DL-IVIM-DKI-derived parameters show potential as imaging biomarkers for MVI in HCC.
Quantum Artificial Intelligence (QAI) has emerged at the nexus of quantum computing and AI, promising to redefine computational frontiers. This survey critically synthesizes the state-of-the-art through 2024, elucidating the profound bidirectional synergy between these fields. We analyze how classical machine learning is accelerating quantum hardware control, circuit optimization, and error correction. Conversely, we assess the potential quantum advantage of algorithms, including variational and kernel-based methods, across domains such as drug discovery, financial modeling, and cybersecurity. Our analysis reveals a critical trade-of between the utility of near-term Noisy Intermediate-Scale Quantum (NISQ) devices and the long-term promise of fault-tolerant architectures. We identify fundamental obstacles to QAI's advancement, including hardware decoherence, algorithmic barren plateaus, and data-encoding bottlenecks. While QAI's potential is transformative, achieving practical quantum advantage requires a concerted effort to overcome these core challenges at the hardware-software interface. This work provides a roadmap for navigating the current landscape and prioritizing future research in this rapidly evolving discipline.
Purpose: The peritumoral region may harbor prognostically relevant information. Radiomics enables extraction of subtle imaging features beyond visual assessment, potentially improving outcome prediction in nasopharyngeal carcinoma (NPC).To investigate whether integrating peritumoral radiomic features-alongside intratumoral and clinical data-enhances prediction of overall survival (OS), progression-free survival (PFS), and distant metastasis-free survival (DMFS) in NPC. Materials and methods: We retrospectively analyzed 252 NPC patients who received definitive chemoradiotherapy (2010-2019). Gross tumor volumes (GTVs) were manually contoured, and peritumoral regions were created by expanding the GTV boundary (5-20 % of in-slice diameter). Radiomic features from both intra- and peritumoral regions were extracted using PyRadiomics. Clinical variables included age, gender, T/N stage, overall stage, and Epstein-Barr virus (EBV) DNA level. Feature selection used intraclass correlation coefficient (ICC), univariate Cox regression, and recursive feature elimination with cross-validation (RFECV). Prognostic models were built using multivariate Cox regression and evaluated with Harrell's C-index. Results: In EBV-included models, the clinical + intratumoral + peritumoral radiomics model yielded the best performance for OS (C-index = 0.776 +/- 0.089; test = 0.729) and DMFS (0.768 +/- 0.056; test = 0.654), outperforming clinical-only models (OS: 0.709 +/- 0.112; DMFS: 0.733 +/- 0.070; p < 0.05). For PFS, clinical + intratumoral radiomics was optimal. Without EBV, peritumoral-inclusive models still enhanced OS and DMFS prediction, while PFS prediction remained reliant on clinical + intratumoral features. Conclusion: Integrating peritumoral radiomics significantly improved NPC prognostication, especially for OS and DMFS, even in the absence of EBV, underscoring its potential for refining risk stratification.
Purpose To evaluate the impact of a deep learning reconstruction (DLRecon) algorithm on the image quality and scar quantification in cardiac magnetic resonance (CMR) late gadolinium enhancement (LGE) imaging for patients with ventricular arrhythmias (VAS) Materials and methods Seventy-two patients with suspected or known cardiomyopathy were prospectively scanned with 3.0T scanner. Short-axis LGE images were reconstructed using conventional reconstruction (ConRecon) and DLRecon, respectively. 4-point Likert score, contrast-to-noise radio (CNR) were used to assess the image quality of LGE and were compared between ConRecon and DLRecon, separately. Scar size (LGE extent) was quantified using the standard deviation (SD) thresholding and full width at half maximum (FWHM) methods and compared between ConRecon and DLRecon images. Results Sixty-four patients (27 VAS, 37 non-VAS) were included. DLRecon images received significantly higher Likert scores than ConRecon images overall, and within both VAs and non-VAs subgroups. DLRecon significantly improved CNR among scar, myocardium, and blood pool, particularly in patients with high heart rate (>75 bpm) or low left ventricular ejection fraction (≤35%). Significantly greater LGE extent was detected using DLRecon with the 5SD method in VAS patients with high heart rate, leading to a higher proportion being classified as high SCD risk. No significant differences were found between DLRecon and ConRecon for LGE quantification using FWHM or the resulting SCD risk stratification (based on LGE extent, with ≥15% defining high risk). Conclusion DLRecon significantly improves LGE image quality and enhances myocardial scar detection in VAS patients, particularly those with high heart rates. This technique shows potential to aid in scar assessment and SCD risk stratification, warranting further clinical investigation.
Purpose: This study aims to investigate the X-ray manifestations of hypotrophic new bone formation in tibial bone transport, propose a classification system, and establish standard treatment protocols. Materials and methods: A retrospective analysis was conducted on 53 out of 378 cases of hypotrophic distraction osteogenesis in tibial bone transport, performed from January 2012 to December 2023. The cohort included 34 males and 19 females, aged 18-71 years (mean age: 37.8 years). Distraction sites comprised 31 cases in the proximal tibia, 7 in the tibial shaft, and 15 in the distal tibia, with defect lengths ranging from 3.3 cm to 22.4 cm (average: 6.3 cm). X-ray imaging categorized hypotrophic bone formation into four types: longitudinal shape (Type A), transverse shape (Type B), worm-bitten shape (Type C), and complete shape (Type D). The treatment protocol included assessment and management of the general condition, callus stimulation through adjustments in transport rate or direction, and surgical interventions. The external fixation index (EFI) assessed healing and mineralization, while limb function was evaluated using the Paley method. Results: Follow-up data over an average of 33.71 +/- 11.7 months indicated that 2 cases required amputation, while 51 achieved bone union, restoring mobility in the transported leg. The EFI ranged from 1.47 to 2.73 months/cm, averaging 1.78 +/- 0.32 months/cm. Outcomes were classified as excellent in 34 cases, good in 11, fair in 3, and poor in 3, resulting in an overall excellent and good rate of 84.9 %. Conclusion: The X-ray characteristics of hypotrophic bone formation in tibial transport can be effectively categorized. A systematic evaluation followed by tailored interventions leads to favorable treatment outcomes.
Recent years have seen an increase in the development of foundation models for chest X-ray (CXR) analysis. Such foundation models provide robust, generalizable feature extraction abilities and allow adaptation for a wide range of downstream tasks. However, there remains no survey paper in the literature that compiles the technologies of foundation models for CXR analysis. This survey aims to fill this gap by providing a comprehensive review of both vision foundation models (VFMs) and vision-language foundation models (VLFMs) for CXR analysis. Specifically, we compiled a list of recently-developed, high-performance CXR foundation models, discussed the commonly-used pretraining techniques and datasets, compared model architectures and parameters, analyzed downstream adaptation methods and corresponding tasks, and summarized model performance on common datasets for downstream tasks. Based on this thorough summary of CXR foundation models, we further highlighted their limitations, open challenges, and development trends. Finally, we discussed interesting future directions, from reasoning capability and agentic workflow to efficiency and interpretability, on CXR foundation model research, which would spawn next-generation artificial intelligence (AI) models for computer-aided CXR interpretation.
Generative Artificial Intelligence (AI) models have demonstrated strong potential in radiology report generation, but their clinical adoption depends on physician trust. In this pilot study, we conducted a radiology-focused Turing test to evaluate how well attendings and residents distinguish AI-generated reports from those written by radiologists, and how their confidence and decision time reflect trust. We developed an integrated web-based platform for report evaluation. Using the web-based platform, eight participants (4 attendings and 4 residents) evaluated 48 anonymized X-ray cases, each paired with two reports from three comparison groups: radiologist vs. AI model 1, radiologist vs. AI model 2, and AI model 1 vs. AI model 2. Participants were asked to select the AI-generated report, rate their confidence, and indicate report preference. Results show that attendings outperformed residents in identifying AI-generated reports (49.9% vs. 41.1%) and exhibited longer decision times, suggesting more deliberate judgment. Both groups took more time when both reports were AI-generated. Our findings highlight the role of clinical experience in AI acceptance and the need for design strategies that foster trust in clinical applications.
Background Irritable bowel syndrome (IBS) is a prevalent disorder characterized by abnormal brain-gut interactions. Notable, psychosocial stressors have been consistently observed to increase susceptibility to additional physical symptoms of IBS. However, the neuropsychological mechanism underlining this disordered top-down control by the central nervous system remains unclear. Methods To address this issue, we investigated gray matter characteristics in the brain and spinal cord of 53 IBS patients and 35 healthy controls (HCs) using simultaneous three-dimensional T1-weighted cortico-spinal imaging. All participants had their IBS symptom severity and pain-related negative emotions quantified using a comprehensive clinical and psychosocial symptom questionnaire battery. We used partial least-squares correlation analysis to assess the associations from complex, heterogeneous variable sets. Results Our results showed that gut-brain dysregulation may be related to abnormal changes in brain gray matter volumes, particularly in superior frontal gyrus, middle frontal gyrus and precentral gyrus. Mediation analyses indicated that pain catastrophizing is associated with these gray matter abnormalities, which may be linked to more severe physical symptoms of IBS. Additionally, the gray matter properties in segments C5 to C6 of the spinal cord serve as serial mediators in the relationship between the top-down brain control and gut function, influencing abdominal pain severity. Conclusions Our study provides novel insights into the top-down control mechanisms contributing to IBS symptoms, emphasizing the potential benefits of early intervention for pain catastrophizing.
Purpose To evaluate the effectiveness of deep learning image reconstruction (DLIR) with metal artifact reduction (MAR) (DLIR-MAR) for improving carotid dual-energy CT angiography (DECTA) in patients with dental hardware. Methods This retrospective study included 49 patients with dental hardware who underwent carotid DECTA. Virtual monochromatic images (VMIs) were reconstructed with DLIR-MAR, adaptive statistical iterative reconstruction-Veo (ASIR-V), and ASIR-V with MAR (ASIRV-MAR) at 40-keV and 50-keV energy levels. We quantitatively compared image noise, CT attenuation, contrast-to-noise ratio (CNR), signal-to-noise ratio (SNR), and artifact index (AI) among different reconstructions, and performed a qualitative assessment of internal carotid artery (ICA) visualization affected by metal artifacts and overall image quality using a 5-point scale. Statistical analyses used repeated measures ANOVA or Friedman tests, with paired t-tests or Wilcoxon signed-rank tests as appropriate. Results Both ASIRV-MAR and DLIR-MAR reduced metal artifacts compared with ASIR-V, with lower AI and higher vascular visualization scores (P < 0.001). Compared to ASIRV-MAR, DLIR-MAR reduced image noise at both keV levels (P < 0.001). DLIR-MAR 40-keV VMIs exhibited comparable image noise (P > 0.05), higher CT attenuation, SNR, CNR, and overall image quality compared with ASIRV-MAR 50-keV VMIs (P < 0.05). Although DLIR-MAR 40-keV VMIs had higher AI values than ASIRV-MAR 50-keV VMIs (P < 0.001), this did not significantly affect ICA visualization (P > 0.05). Conclusions DLIR reconstructed with MAR improves carotid DECTA images in patients with dental hardware, reducing image noise in lower-keV VMIs to capture their contrast enhancement benefits.
Purpose Effective tools for risk stratification of intestinal disease progression (IDP) in Crohn's disease (CD) patients remain limited. Electronic medical records (EMRs) contain detailed information describing disease characteristics. Large language model (LLM) excels at summarizing and analyzing these data. Materials and Methods We developed an LLM-agent using EMRs to predict IDP. In this retrospective study, we collected EMRs from 563 patients across six centers, built a human knowledge base incorporating literature-derived and EMRs-derived knowledge. Three levels of prompts were developed through a step-by-step superposition strategy: basic task description, Chain-of-Thought technique, and "Self-Reflection" module. The third level further integrated the human knowledge base. Three prompt versions (V1-3) were paired with three LLMs (ChatGPT-4omini, Gemini-1.5-Pro, ChatGPT-4o), creating nine LLM-agents. We compared their predictive performance for IDP to select the optimal model and contrasted the readability and semantics of LLM-generated reports with physician-written summaries. Results After evaluating 304 initial factors using a human-in-the-loop strategy with LLM and 13 CD specialists, we selected 98 IDP-related factors to construct the knowledge base. GPT-4o_V3 outperformed eight competing models in training set (accuracy, 84.4% vs. 53.9%-75.8%; F1 score: 0.867 vs. 0.663-0.808) and test set (accuracy, 83.5% vs. 59.8%-72.2%; F1 score, 0.882 vs. 0.748-0.814). Disease-progression-free survival of high-risk patients identified by GPT-4o_V3 was significantly shorter than low-risk patients (P<0.001). Prediction reports generated from GPT-4o_V3 showed intermediate characteristics between Gemini-1.5-Pro_V3 and GPT-4omini_V3 in both readability and semantics. Conclusion In conclusion, GPT-4o_V3 outperforms other models and accurately predicts IDP risk for CD patients across multiple medical centers, assisting clinicians in clinical decision-making.
Purpose: Abdominal trauma detection via ultrasound, particularly through Focused Assessment with Sonography for Trauma (FAST), is a cornerstone of emergency medicine due to its portability, real-time imaging and applicability in unstable patients. Artificial intelligence (AI) models have been used to enhance diagnostic accuracy and reduce operator dependency. Conduct a systematic review and meta-analysis to evaluate the effectiveness of AI models in ultrasound-based detection of abdominal trauma. Methods: A search was carried out across MEDLINE, Embase, and Cochrane databases in accordance with the Preferred Reporting Items for Systematic Reviews and Meta-Analyses for Diagnostic Test Accuracy (PRISMA-DTA) 2018 guidelines. Diagnostic studies that assessed the precision of AI in ultrasound-based detection of abdominal trauma, particularly FAST, were included. Results: The analysis included five studies with 209,577 frames evaluated across all included studies. The pooled analysis of AI models demonstrated a sensitivity of 90.7 % (95 % CI: 79.1-96.2 %; I 2 = 99 %), specificity of 96.9 % (95 % CI: 95.1-98.0 %; I 2 = 89 %), and an area under the curve (AUC) of 0.979 (95 % CI: 0.964-0.983). Conclusion: This study suggests that AI models may demonstrate promising sensitivity and specificity in ultrasound-based detection of abdominal trauma, with an overall AUC of 0.979. However, given the limited number of studies and sample sizes included, these findings should be considered preliminary, highlighting the need for further validation in larger and more diverse cohorts.
The development of large-scale artificial intelligence (AI) models is influencing neuroscience research by enabling end-to-end learning from raw brain signals and neural data. In this paper, we review applications of large-scale AI models across four major neuroscience domains: neuroimaging and data processing, brain-computer interfaces and neural decoding, clinical decision support and translational frameworks, and disease-specific applications across neurological and psychiatric disorders. These models show great potential for addressing major computational neuroscience challenges, including multimodal neural data integration, spatiotemporal pattern interpretation, and the development of translational frameworks for clinical research. Moreover, the interaction between neuroscience and AI has become increasingly reciprocal, as biologically informed architectural constraints are now incorporated to develop more interpretable and computationally efficient models. This review highlights both the promise of such technologies and critical implementation considerations, with particular emphasis on rigorous evaluation frameworks, effective integration of domain knowledge, prospective clinical validation, and comprehensive ethical guidelines. Finally, we provide a systematic listing of critical neuroscience datasets used to develop and evaluate large-scale AI models across diverse research applications.
Background The methylation status of the O6-methylguanine-DNA methyltransferase (MGMT) promoter is a key predictive biomarker for chemotherapy response in glioblastoma (GBM). Current reliance on complex molecular testing necessitates the development of alternative predictive methods for postoperative decision support during the waiting period. Although both preoperative MRI and histopathological images contain valuable biological information, their combined potential for predicting MGMT status remains unexplored. We aimed to develop and validate a deep learning radiopathomics model (DLRPM) that integrates MRI and pathological images for predicting MGMT promoter methylation. Methods A retrospective collection of pathologically confirmed isocitrate dehydrogenase (IDH) wild-type GBM patients (n=207) from three centers was performed, all of whom underwent MRI scanning within 2 weeks prior to surgery. The pre-trained ResNet50 was used as the feature extractor. Features of 1024 dimensions were extracted from MRI and pathological images, respectively, and the features were screened for modeling. Then feature fusion was performed by calculating the normalized multimode MRI fusion features and pathological features, and prediction models of MGMT based on deep learning radiomics, pathomics, and radiopathomics (DLRM, DLPM, DLRPM) were constructed and applied to internal and external validation cohorts. Results In the training, internal and external validation cohorts, the DLRPM further improved the predictive performance, with a significantly better predictive performance than the DLRM and DLPM, with AUCs of 0.920 (95% CI 0.870–0.968), 0.854 (95% CI 0.702–1.000), and 0.840 (95% CI 0.625–1.000) and the corresponding accuracy was 83.2%, 82.1% and 80.0%, respectively. The AUCs of DLRM were 0.786, 0.771, and 0.600 respectively, with accuracy rates of 67.3%, 67.9%, and 60.0% respectively. The AUCs of DLPM were 0.864, 0.844, and 0.780 respectively, with accuracy rates of 79.6%, 78.6%, and 66.7% respectively. Conclusion By integrating MRI and histopathological images, the DLRPM narrows the diagnostic gap created by the waiting period for molecular testing, offering a practical approach to predicting MGMT methylation status. The model’s robust performance across validation cohorts demonstrates its potential as a practical clinical tool to supplement or reduce reliance on invasive tissue sampling, thereby aiding in personalized treatment planning for GBM patients. Future studies will focus on prospective validation and explore its utility in predicting other molecular markers and treatment outcomes.