
Chromium-compensated Gallium Arsenide (GaAs:Cr) has emerged as a promising sensor material for photon counting detectors (PCDs) due to its excellent resistivity and charge carrier mobility. However, due to the limitations of the Cr compensation process, the GaAs:Cr wafer thickness is less than 1 mm, which distorts its X-ray absorption efficiency and limits its application in clinical computed tomography (CT). This study proposes a GaAs detector with an edge-on structure for clinical photon counting CT (PCCT) applications. To evaluate the performance of the detector, we modeled the spectral response and analyzed material decomposition (MD) noise using a full-chain detector model. In imaging domain study, digital phantoms, such as low contrast phantom, Gammex phantom, and XCAT digital phantoms, were virtually scanned to assess the image uniformity, CT number accuracy, material decomposition, and virtual mono-energetic imaging (VMI) performance, with the detector system modeled to reflect actual PCCT prototypes. The results indicate that the GaAs:Cr detector exhibits excellent mono-energetic spectrum response and low material decomposition (MD) noise. Additionally, the GaAs:Cr detector provides lower noise, higher contrast-to-noise ratio (CNR) in VMI results, and more precise concentration measurements in quantitative analysis, compared to existing PCDs. These findings highlight the feasibility of developing GaAs:Cr detectors for clinical CT applications, offering potential advantages over existing PCDs.
Multi-instance learning (MIL) has significantly advanced AI-assisted cancer diagnosis using histochemically stained whole-slide images (WSIs). However, acquiring such WSIs is time-consuming, labor-intensive, and environmentally unfriendly. To address this, we propose using unlabeled autofluorescence (UAF) WSIs as a cost-effective alternative for MIL-based diagnosis. We introduce a dedicated UAF WSI dataset, LCUHI-UAF, along with a matched H&E-stained WSI dataset, LCUHI-H&E, derived from the same tissue sections for direct comparison. To tackle the low signal-to-noise ratio and blurred morphological details in UAF images for classification tasks, we propose a novel Multi-Granularity Graph-Mamba (MGGM) MIL framework. In this framework, each WSI is represented as a graph constructed from the spatial coordinates of tissue instances, allowing graph convolution to capture local spatial dependencies while the Mamba architecture models long-range relationships. A multi-granularity mechanism is further proposed to enable comprehensive representation of hierarchical relationships among cell clusters and microenvironments within the tissue. Experiments on the paired lung cancer datasets LCUHI-H& E and LCUHI-UAF show that MGGM-MIL achieves top-performing results across two pre-trained feature settings. These results simultaneously confirm the viability of UAF-stained WSIs as a cost-effective substitute for H&E in cancer diagnosis. Additional evaluation on the Camelyon16 breast cancer dataset further validates the generalizability of MGGM-MIL across tissue types. Code is available at https://github.com/JiuyangDong/MGGMMIL.
Accurate breast tumor segmentation in dynamic contrast-enhanced magnetic resonance imaging (DCE-MRI) is vital for diagnosis and treatment planning. Despite advances in deep learning, its performance remains constrained by the need for extensive voxel-wise annotations. To mitigate this burden, we propose an annotation-efficient framework that jointly optimizes data selection, unlabeled data utilization, and data augmentation under limited annotation budgets. A diversity-aware uncertainty query (DUQ) strategy guides the annotation by jointly modeling data representativeness and informativeness through a representative candidate selector (RCS) and an uncertainty-based decision maker (UDM), ensuring efficient and targeted labeling. To leverage unlabeled data, a cross-decoder consistency regularization (CDCR) mechanism enforces prediction consistency between two decoders with distinct attention mechanisms, enhancing robustness and confidence. Furthermore, a lesion transplant augmentation (LTA) technique synthesizes anatomically valid pseudo samples by transplanting lesion regions from labeled to unlabeled images, effectively expanding training diversity. Experiments were conducted on two DCE-MRI datasets with biopsy-proven breast cancers, one as internal dataset containing 676 subjects and the other as external dataset with 344 subjects. Comparative and ablation results demonstrate that our framework consistently outperforms state-of-the-art semi-supervised and active learning methods, providing a simple yet effective annotation-efficient solution for breast cancer segmentation in DCE-MRI. The code is publicly available at https: //github.com/zouquanling/DUQ_and_CDCR.
Interventional radiology (IR) requires joint reasoning over procedural images and domain-specific clinical knowledge. Existing medical retrieval-augmented generation (RAG) methods are mainly text-oriented or designed for general medical vision-language tasks, and therefore remain limited in retrieving fine-grained visual-textual evidence for IR scenarios. To address this limitation, we present Prototype-guided Retrieval for Interventional Medical Assistance (PRIMA), a multimodal RAG framework that jointly leverages multimodal imaging and clinical text to support IR decision-making. PRIMA constructs a multimodal IR knowledge index through anatomy-aware visual-textual alignment and modality-preserving representation learning. It then introduces domain-informed prototype learning to organize IR concepts, enabling prototype-guided retrieval that re-ranks evidence using both query similarity and prototype affinity. We conduct comprehensive evaluations on literature-curated and clinically collected IR datasets. Experimental results show that PRIMA consistently improves generation quality, question-answering accuracy and expert-rated clinical interpretability compared with existing RAG baselines. These findings demonstrate the effectiveness of clinically grounded prototype-guided retrieval for multimodal knowledge assistance in interventional radiology. Related resources are available at https://github.com/StonHamA/PRIMA.
Multi-energy CT (MECT) offers unique advantages in material decomposition, tissue characterization, and functional imaging, positioning it as a pivotal direction for next-generation CT. Currently, standardized scanning protocols for MECT have not yet been established. Considering growing public concern over X-ray radiation exposure, we propose a complementary sparse-view scanning protocol tailored for MECT, which reduces radiation dose while maximizing angular coverage. To reconstruct high-quality images from these sparse-view data and ensure algorithmic reliability in practical applications, we introduce an Online Adaptive Reconstruction (OA-Recon) framework that adapts robustly to varying acquisition settings through two designs. First, it adopts a Bayesian adaptation strategy for instance-specific optimization while preserving the learned prior. Second, it incorporates a Frequency-adaptive and Physics-informed Network (FaPiNet) for adaptive feature extraction and acquisition-conditioned feature modulation. In addition, it incorporates a spectral attention mechanism to fully exploit complementary information across energy channels. Experiments on simulated MECT and real mouse PCCT data show that FaPiNet-OA-Recon achieves better performance in suppressing streak artifacts, restoring image details, and maintaining CT-value accuracy. More importantly, OA-Recon demonstrates adaptability to changes in view, spectrum, and anatomy, providing a preliminarily feasible solution for clinical applications of MECT.
Nuclei segmentation is a fundamental but challenging task in computational pathology due to diverse morphologies, blurred boundaries, and staining variations. Despite remarkable progress, existing models often suffer from structural instability under morphological and staining variations. We attribute this instability to disrupted frequency-spatial consistency and address it through FreqPath-Net, which enforces frequency-spatial consistency for robust nuclei segmentation. By operating directly in the frequency domain, FreqPath-Net achieves morphology-invariant and stain-robust feature representations. The Spectral Wavelet Attention Module (SWAM) adaptively enhances high-frequency boundary cues while maintaining low-frequency consistency, addressing boundary blurring and detail loss. Furthermore, the Orthogonal Direction-Constrained Frequency Module (ODFM) captures global spectral patterns and enforces directional consistency, effectively preserving boundary orientation and structural integrity by leveraging frequency-spatial consistency. Extensive experiments on twelve nuclei segmentation benchmarks show that FreqPath-Net consistently outperforms state-of-the-art methods. On the multi-organ Pan-Nuke dataset, FreqPath-Net achieves an mIoU of 85.32%, outperforming the second-best method by 2.93%. Code: https://github.com/huangjin520/FreqPath-Net.
Magnetic resonance imaging (MRI) can estimate three-dimensional (3D) time-resolved blood velocity fields using 4D-flow MRI, providing rich relative pressure field information. Clinical alternatives—catheterization and Doppler echocardiography—provide one-dimensional pressure drops. While the accuracy of one-dimensional pressure drops from 4D-flow has been explored, the accuracy of 3D relative pressure field estimates requires further evaluation. This work analyzes three state-of-the-art relative pressure estimators: virtualWork-Energy Relative Pressure (vWERP), the Pressure Poisson Estimator (PPE), and the Stokes Estimator (STE). Spatiotemporal characteristics and noise sensitivity were determined in silico via comparison with CFD-derived relative pressure fields. Estimators were then validated using a type B aortic dissection (TBAD) flow phantom with varying tear geometry and twelve catheter pressure measurements. Finally, each estimator was evaluated across eight patient cases in a clinical feasibility analysis. Estimators were cross-compared and evaluated for coherence since no reference catheter pressure measurements were available. PPE outperformed STE in the in silico relative pressure analysis, while both outperformed vWERP in plane-to-plane assessment. High velocity gradients and low spatial resolution contributed most significantly to local variations in 3D relative pressure field errors. Low temporal resolution led to systematic underestimation of highly transient peak pressure events. In the flow phantom analysis, vWERP was the most accurate method, followed by STE and then PPE. Each relative pressure estimator was strongly correlated with ground truth relative pressure values, despite the tendency to underestimate peak relative pressure. Patient case results demonstrated that each relative pressure estimator could be feasibly integrated into a clinical workflow.
Functional connectivity networks (FCNs) derived from functional magnetic resonance imaging (fMRI) have been widely used to characterize topological alterations of brain networks in neuropsychiatric disorders (NDs). Given the frequent restrictions on direct multi-site fMRI data sharing, federated learning (FL) offers a collaborative modeling paradigm without exchanging raw neuroimaging data. However, conventional parameter-averaging FL approaches struggle under cross-site non-IID distributions. Prototype-based FL provides a promising alternative, yet existing designs implicitly rely on spatially structured image data and fail to capture the topology-centric semantics of FCNs. To bridge this gap, we propose ToPPFed, a Topological Prototype-Enhanced Personalized Federated Learning framework for multi-site classification between subjects with each studied disorder and normal controls (NCs). ToPPFed introduces a Graph Topological Prototype Learning module to extract discriminative topology-aware prototypes from FCNs and a Contrastive Mask-Induced Residual Scaling mechanism to adaptively integrate group-level priors into individual representations. By exchanging topology prototypes instead of raw data or full model parameters, ToPPFed supports cross-site collaboration while reducing direct data exposure. Experiments on multi-site fMRI datasets of three representative NDs show that ToPPFed improves accuracy (ACC) by 1.7-7.9 percentage points over the best-performing federated baseline on each dataset. Interpretability analyses indicate that ToPPFed highlights model-derived discriminative brain regions and functional connections. The topology-aware exchange of node and edge prototypes offers an effective framework for collaborative FCN modeling across imaging sites without centralizing neuroimaging data.
Accurate tissue point tracking in endoscopic videos is crucial for robotic-assisted surgical navigation and scene understanding, yet remains challenging due to complex tissue deformations, instrument occlusions, and the scarcity of dense trajectory annotations. Existing methods struggle with long-term tracking robustness under these conditions because they primarily rely on local motion cues that fail in textureless regions and lack mechanisms to quantify tracking reliability. Without high-level semantic context to differentiate homogeneous tissues and explicit uncertainty modeling to prevent error accumulation during occlusions, existing trackers inevitably drift over time. We present Endo-TTAP, addressing these challenges through a two-stage hybrid supervision approach. Our method introduces: (1) A Multi-Facet Guided Attention (MFGA) module that fuses multi-scale optical flow features, semantic embeddings, and motion patterns via guided attention to jointly predict point positions, occlusion states, and tracking uncertainty; (2) An Auxiliary Curriculum Adapter (ACA) enabling progressive domain adaptation from synthetic to surgical data through exponential scheduling; (3) A Pseudo Label Generator (PLG) that creates high-quality dense annotations from sparse surgical data. Our two-stage training strategy first initializes components using synthetic datasets with optical flow ground truth, then transitions to real surgical data through unsupervised flow consistency and semi-supervised pseudo-label learning. We further contribute the Endo-TTAPC5 dataset, comprising 250 video segments across five clinically meaningful challenges. Extensive validation on two public datasets (SurgT, STIR) and our Endo-TTAPC5 dataset demonstrates that Endo-TTAP achieves state-of-the-art performance in tissue point tracking, particularly in complex endoscopic scenarios. Code and video demo are available at https://adampc888.github.io/Endo_TTAP/.
The clinical translation of AI in medical imaging faces critical challenges including scare annotated data, long-tailed pathological distributions, and privacy constraints in virtual imaging trials. To address these limitations, we propose a Multi-conditional Diffusion framework with Texture Constraints (MDTC) for synthesizing clinically reliable lesions in CT images. The key innovation lies in the joint integration of anatomical mask guidance and a Gray-Level Co-occurence Matrix-based texture classifier into the diffusion network. This dual-constraint mechanism uniquely enforces structural fidelity and pathologically heterogeneous characteristics in synthetic lesions, effectively preserves pathological heterogeneity and simultaneously enhances the authenticity of the generated lesionst in existing methods. The experiments on hepatocellular carcinoma and pulmonary nodule datasets demonstrate that the proposed MDTC method achieves favorable performance in terms of texture fidelity, a significant improvement in PSNR, and notable optimization in FID compared to classic generation models. Furthermore, downstream classifiers trained on synthetic data remain capable of maintaining the vast majority of baseline classification performance, while data augmentation for rare pulmonary nodules significantly improved the classification efficacy. This work establishes a new paradigm for generating diagnostically meaningful synthetic data, effectively alleviating critical bottlenecks in virtual imaging trial. The experimental results demonstrate that MDTC-generated lesions preserve structural authenticity and have potential to mitigate texture homogenization, thereby enabling robust downstream diagnostic model training and offering a practical, privacy preserving solution to expand virtual imaging trials toward underrepresented disease groups and real world clinical workflows.
Patient motion remains a source of image degradation in brain MRI, leading to signal loss, blurring, and geometric distortion that compromise quantitative analysis. Existing deep learning methods for motion correction typically rely on paired clean-corrupted data or k-space acquisitions, which are rarely available in clinical settings. We propose SSRL-MAR, a motion artifact-aware unpaired representation learning framework for motion artifact reduction that requires neither paired training data nor explicit motion labels. SSRL-MAR employed a three-stage training strategy: (1) contrastive learning on 3D patches to extract motion representations by contrasting clean and synthetically corrupted images, (2) a motion artifact-aware synthesis network to generate motion artifacts from clean scans, and (3) a motion artifact-aware generator to restore clean volumes using the learned degrader for self-supervised supervision. On in-silico dataset, SSRL-MAR achieved PSNR 23.81dB, SSIM 91.55%, and NMSE 0.79%. On in-vivo MR-ART dataset, the pretrained model reduced motion distortion, and unsupervised domain adaptation further improved anatomical fidelity. Against a source-only supervised model trained on the same simulated pairs, SSRL-MAR improved PSNR by up to 2.0 dB on MR-ART after unsupervised domain adaptation, and remained within 0.25-0.47 dB of an oracle supervised model that requires real paired data unavailable in practice. At the milder motion level, volumetric error in structures such as the corpus callosum and ventricular system decreased by more than 50%, confirming improved neuroanatomical consistency. These results indicate that SSRL-MAR provides a robust and scalable image-domain solution for 3D brain MRI motion correction, enabling reliable structural quantification in large-scale neuroimaging studies without requiring prospectively acquired pairs or acquisition-specific calibration.
Computational analysis of breast histopathological images is critical for reliable computer-aided diagnosis and treatment planning. Owing to the ultra-high resolution of whole-slide images (WSIs), most existing WSI segmentation methods rely on patch-wise processing. However, independently processing isolated patches breaks spatial continuity and weakens global tissue context, ultimately limiting segmentation performance. To overcome these limitations, we propose CAT-WSI, a context-aware trajectory learning framework for breast pathology WSI segmentation. Rather than treating patches as unordered samples, CAT-WSI organizes them into structured transverse trajectories across each slide, thereby preserving long-range spatial dependencies while reducing the directional bias and boundary fragmentation inherent in conventional patch-based pipelines. To further enhance global positional awareness, CAT-WSI augments these trajectory representations with a paired downsampled whole-slide thumbnail, enabling explicit global-local contextual modeling over the entire slide. We evaluate CAT-WSI on the CAMELYON16 and Breast-HER2+ datasets across multiple magnification levels. Extensive experiments demonstrate that CAT-WSI achieves consistently strong performance across multiple magnification levels, attaining the best overall results on the evaluated benchmarks in our experimental setting.
Automatic medical image segmentation, as a prerequisite for clinical quantitative analysis, forms the basis of computer-aided diagnosis. However, blurry object boundaries caused by factors such as imaging quality and inherent physiological properties of tissues or lesions are the main causes of imprecise segmentation. This aligns with the common understanding that high uncertainty and misclassification tend to occur at boundaries in segmentation. To address the challenge, we investigate this phenomenon and explore the connection between uncertainty and tissue boundaries by analysing various tissues. Then an Evidential Uncertainty-Guided Boundary (EUGB) loss is further proposed to demonstrate that uncertainty information can indeed facilitate combating boundary segmentation errors. The proposed EUGB loss not only emphasizes challenging pixels along blurry boundaries using evidential uncertainty, but also introduces a regularization term that constrains uncertainty learning by penalizing incorrect predictions and reinforcing correct ones. The effectiveness of the proposed EUGB loss is verified in the public LIDC-IDRI, ISIC 2018, and OCTA-500 datasets with two classic medical image segmentation networks (U-Net and TransU-Net). Experimental results demonstrate that the proposed loss outperforms seven other segmentation loss functions in terms of boundary segmentation, while maintaining competitive region-level segmentation accuracy. Beyond introducing a new loss function, this paper provides empirical insights for selecting appropriate loss functions across different application scenarios. We systematically analyze the strengths and limitations of existing losses from multiple perspectives, including reliability and dataset characteristics. This analysis offers practical insights that enable researchers and practitioners to optimize segmentation performance based on specific data attributes.
Histopathological analysis of stained tissue remains central to biomedical research and clinical care. Virtual staining (VS) offers a promising alternative, with potential to reduce costs and streamline workflows, yet hallucinations pose serious risks to clinical reliability. Here, we formalize the problem of hallucination detection in VS and propose a scalable post-hoc baseline method: Neural Hallucination Precursor (NHP), which explores the generator’s latent space to preemptively identify hallucinations. Extensive experiments across diverse VS tasks show NHP is both effective and robust. Critically, we also find that models with fewer hallucinations do not necessarily offer better detectability, exposing a gap in current VS evaluation and underscoring the need for hallucination detection benchmarks.
Automatic tooth segmentation is essential for computer-aided diagnosis and treatment planning in dentistry. Proposal generation plays a pivotal role in accurately delineating instance boundaries and is critical for improving segmentation accuracy. However, existing methods suffer from dispersive offset bias, spatial information loss, and spillover-induced centroid shift, all of which are difficult to alleviate in Euclidean space. This limitation commonly leads to missed detections or excessive merging of instances, particularly in cases of missing or crowded teeth. To overcome these issues, we introduce a novel proposal generation framework based on Topological Morphological Clustering (TMC), which integrates edge logits fields with topological morphology. The key insight is that morphological erosion applied to tooth and edge masks enables the separation of individual tooth instances. To further enhance the continuity of tooth-tooth edges, we design an edge enhancement module that fuses the alpha-shape algorithm with our proposed KNN-based iterative contour approximation (KNN-ICA) method, thereby constructing a supplementary geometric edge logits field. Extensive experiments show that our method achieves superior detection performance over existing proposal generation techniques, and outperforms prior instance segmentation methods on two public benchmarks. Our code is available at https://github.com/jfzhuang314/TMELFNet.
Pretrained segmentation models for cardiac magnetic resonance imaging (MRI) often fail to generalize across imaging sequences due to substantial contrast variations. These variations arise from different imaging protocols, yet fundamentally, all contrasts are governed by the same underlying tissue properties, primarily captured by three components: the magnetization strength (M0), T1, and T2. Building on this insight, we introduce Reverse Imaging, a physics-driven framework for data augmentation and domain generalization in cardiac MRI. Our method infers tissue properties from observed MR images with annotation by solving an ill-posed nonlinear inverse problem, regularized by a generative prior. The prior is learned from the multiparametric saturation-recovery single shot acquisition (mSASHA) dataset for joint cardiac T1 and T2 mapping. In inference, we characterize imaging sequences as weak, moderate, or strong observations according to the physical information they provide and the degree of ill-posedness. This motivates an iterative prior-learning strategy that uses moderate T1-mapping observations to alleviate mSASHA data scarcity via pseudo tissue-property estimates. We further integrate MRI physics into posterior inference by expressing the sequence model as a likelihood term guiding the reverse diffusion process. For widely used but weak cine observations, we develop a sequence-specific ControlNet to improve efficiency and spatial consistency. Extensive experiments on eight unseen cardiac MRI sequences with markedly different contrast mechanisms show that Reverse Imaging yields plausible tissue-property estimates, supports synthesis of diverse yet physically consistent contrasts, and improves segmentation robustness under severe cross-sequence shifts.
Surgical full scene segmentation is essential for laparoscopic assistance but remains challenging due to the high visual similarity among anatomical structures, and illumination variations caused by single moving light source. Moreover, accurately segmenting thin, elongated instruments is still difficult, especially when they appear at oblique orientations. Although Mamba-based segmentation methods effectively model long-range spatial relationships through state-space updates, their 2D selective scan strategy is limited in capturing the oblique spatial distributions of surgical instruments. To address these challenges, a holistic vision Mamba block (HVMamba) is proposed with a Holistic Directional Selective Scan module (HSSD) module to integrate anisotropic spatial features from multiple directions while simultaneously modeling cross-channel dependencies to address ambiguous visual features and complex illumination. Specifically, HSSD comprises an Attention-Guided Holistic Directional Selective Scan (AHSD) for efficient integration of horizontal and oblique spatial features under the guidance of cross-coordinate relationships, and a Channel-aware Directional Selective Scan (CASD) to model bidirectional cross-channel dependencies and enhance responses in ambiguous regions. Based on HVMamba, a hierarchical hybrid network, SurgMamba, is further developed by combining an HVMamba branch with a convolutional neural network branch to jointly capture global and local representations. The two types of representations are adaptively fused by Local-Global Feature Coupling Units (LG-FCUs) in the encoder and an Attention-Aware Gating Mechanism (AGM) in the prediction head. Experiments on two public datasets demonstrate that SurgMamba achieves superior performance over state-of-the-art methods, particularly in challenging cases involving obliquely oriented thin instruments, strong specular highlights, and low-contrast tissue boundaries. Code is available at https://github.com/hailinhhh/SurgMamba.
Vision Transformer has achieved significant performance improvements in natural image segmentation tasks owing to its superior global modeling capabilities. However, applying vision Transformers to 3D medical image segmentation is challenging because of the quadratic computational complexity of the self-attention mechanism and their limited generalization on small-scale datasets. To address these limitations, we propose a hybrid CNN-Transformer architecture guided by channel attention, referred to as CACFormer, for 3D medical image segmentation. Specifically, we design a simple and effective channel attention module to guide the fusion of local and global features in each channel. This module adaptively assigns weights to each channel based on its semantic contribution to accurate segmentation. Meanwhile, we introduce a novel linear Transformer variant that integrates a linear attention mechanism with tanh activation. This design encourages the model to focus on the target regions and produce robust segmentation outcomes. The effectiveness and competitive generalization of the proposed framework are validated across five benchmark datasets. On AMOS2022, CACFormer achieves an average Dice score of 89.71%, outperforming 3D UX-Net (89.30%) while reducing inference time from 3.77 s to 2.49 s (a 33.95% reduction). On BraTS2021, CACFormer attains an average Dice score of 90.20%, comparable to TransBTS (90.33%), with 28.54% fewer parameters and 15.22% faster inference time (from 0.46 s to 0.39 s), demonstrating a favorable trade-off between performance and efficiency. Moreover, CACFormer demonstrates competitive cross-dataset generalization, achieving an average Dice score of 86.50% on BraTS2021 when trained on BraTS2019, significantly out-performing TransBTS (47.90%). Index Terms—3D Medical Image
Accurate T-staging is crucial for guiding personalized treatment strategies for laryngopharyngeal cancer. However, current clinical practice relies on invasive biopsy procedures, whereas CT-based staging remains challenging due to the complex patterns of tumor invasion. Recent computer-aided approaches face two key challenges: 1) Structural relationship modeling: existing methods underrepresent anatomically structured patterns of tumor invasion, as they either process whole CT volumes without tumor-specific anatomical constraints or rely on labor-intensive tumor segmentation. 2) Fine-grained cross-modal alignment: while radiology reports contain organ-specific invasion details, current methods that apply global feature fusion struggle to accurately align individual anatomical structures with their corresponding textual descriptions. To address these issues, we propose an anatomy-aware multimodal framework that integrates organ-level CT context and radiology reports into a unified representation for laryngopharyngeal T-staging. The framework first constructs an Anatomy-Structured Organ Graph (AOG) that captures invasion patterns between primary sites and surrounding organs, then performs Organ-Anchored Cross-Modal Alignment (OCA) so that each organ node aggregates textual evidence from the radiology report, and finally refines this graph representation by injecting organ-specific invasion cues extracted from the report via Report-Enhanced Graph-Refinement (REG), yielding a multimodal organ graph that combines spatial and textual evidence. Extensive experiments demonstrate that the proposed framework achieves superior performance in T-staging of laryngopharyngeal cancer.
Accurate MRI-based brain tumor analysis requires not only tumor subtype classification but also localization at an anatomical granularity that is consistent with radiology reports. Most vision-only methods address localization and classification as separate label-prediction tasks, and therefore provide limited alignment with the fine-grained anatomical semantics used in routine reporting. To address this limitation, we propose LIGHT (Learning Image-text Grounding for Hierarchical Tumor analysis), a 3D vision-language framework that formulates brain tumor localization and subtype classification as image-text retrieval in a shared embedding space. Given multimodal MRI inputs, LIGHT retrieves coarse anatomical regions, fine-grained subregions, and tumor subtypes in a unified coarse-to-fine procedure. The framework has three main components: (1) a hierarchical retrieval vocabulary containing 21 coarse-grained regions, 563 fine-grained subregions, and 5 tumor subtypes; (2) large-scale foundation pretraining on a curated patient-disjoint in-house dataset of 99,813 MRI-report pairs; and (3) grounded task fine-tuning that uses segmentation-guided tumor crops and multi-template prompt supervision to improve tumor-prompt alignment. Across 11,034 annotated 3D brain tumor MRI cases from in-house, external clinical, and public datasets, LIGHT achieved 72.1% accuracy for coarse-grained localization, 70.1% top-1 accuracy for fine-grained localization, and 76.0%, 83.7%, and 84.6% accuracy for subtype classification on three 3D datasets, respectively. Additional public 2D subtype-classification experiments provide further evidence of cross-setting subtype recognition. These results suggest that report-grounded image-text retrieval can provide interpretable anatomical and diagnostic outputs for brain tumor MRI analysis. Code: https://github.com/qiuzhaoyu/LIGHT.