
Proton therapy is a tumor treatment technique that uses the unique dose profile of protons to deliver high doses to the tumor while sparing surrounding tissue. This precision creates a need to actively verify the delivered dose during treatment. Many techniques for dose verification have been proposed; most focus on the spatial distribution of secondary particle emissions caused by proton-matter interactions. Among them are prompt gamma’s (PGs), i.e. high-energy photons emitted within 10 ps after a proton interacts with the tissue. Therefore, their emission is linked to the proton in time and space. Using a specialized measurement procedure – prompt gamma timing – and a tomographic-like reconstruction we can retrieve the PGs spatial and temporal distribution [1]. The combination of time and space enables the study of proton motion. Through non-linear regression with proton motion models, we then estimate the stopping power – a crucial material parameter to accurately plan and execute the treatment [2, 3]. Several challenges in instrumentation, reconstruction, modelling and postprocessing must be overcome to implement this unique approach. Still, we have used this novel combination of image computation techniques to estimate the stopping power of a homogeneous phantom with an accuracy of 2
Breast cancer screening in Asian populations faces significant challenges due to high breast density prevalence, which reduces mammographic sensitivity. Low-dose computed tomography (LDCT) scans acquired for lung cancer screening capture breast tissue and present an opportunity for opportunistic breast cancer risk assessment. This study develops and evaluates deep learning frameworks using multiple instance learning (MIL) for breast cancer risk stratification from LDCT scans combined with clinical features. Two complementary approaches were developed: an individual breast model using attention mechanisms for slice-level interpretability, and a bilateral model employing global pooling to capture asymmetry patterns. Evaluated on 60 patients, the individual model achieved 72.7
Radiomics extracts quantitative features from medical images, offering biomarkers for diagnosis, prognosis, and evaluation of treatment response. Yet, its broader application in research is limited by the absence of standardized, end-to-end workflows for multimodal imaging. We present an open-source Python-based pipeline that allows for interactive studies and series selection, as well as automated conversion, segmentation, and quantitative analysis of positron emission tomography (PET) / computed tomography (CT) DICOM images. Leveraging widely adopted segmentation models for PET analysis and CT organ delineation, the pipeline computes key radiomics, producing structured outputs for analysis. Its modular design facilitates reproducible, scalable, and clinically relevant radiomics studies, addressing a critical gap in medical image analysis infrastructure. The code is available under: https://github.com/Clinical-Computational-Medical-Imaging/MUSIQ
Brain aging is an inevitable process in adulthood, yet there remains a critical need for objective and accurate biomarkers to assess its progression. In this study, we develop a deep learning-based framework for brain age estimation using multiparameter MRI. Structural (T1 and T2 weighted) and diffusion-weighted images were acquired, from which we extracted cortical features, including gray matter volume, surface area, and thickness along with white matter integrity metrics such as fractional anisotropy, mean diffusivity, axial diffusivity, and radial diffusivity. To integrate these multimodal neuroimaging features, we propose MN-FNet (multimodal neurofeature fusion network), a dedicated regression architecture that effectively combines gray matter structure and white matter microstructure. Our model achieves accurate brain age prediction with low estimation error and identifies key neuroanatomical regions associated with aging, additionally providing evidence of hemispheric lateralization as a factor in brain aging. This approach offers a reliable and interpretable tool for brain age estimation, with potential applications in early detection of neurodegenerative conditions.
Foundation models provide general-purpose image representations that promise to capture structural and semantic information. However, their suitability for measuring similarity across image sequences has not been thoroughly examined. Conventional metrics such as the structural similarity index measure (SSIM) are commonly used to assess frame-to-frame consistency but are sensitive to motion, deformation, and intensity changes, which limits their usefulness for dynamic imaging. In this study, we compare embeddings from a variety of pretrained models, including DINOv2, ResNet50, CLIP, SAM, and LPIPS, to evaluate their ability to represent temporal and structural similarity in videos and their ability to assess the quality of image registration. We analyzed sensitivity to global and local motion in two medical imaging datasets. We focused on videofluoroscopic swallowing studies (VFSS) with global and local motion and the BAGLS dataset of vocal fold vibrations with mainly local motion. Our results indicate differences in how models maintain consistent similarity under motion, and suggest that some embedding-based approaches provide a more stable representation than SSIM without additional fine-tuning.
Accurate preoperative mapping of abdominal vasculature is essential in colorectal surgery to reduce intraoperative bleeding and postoperative ischemia. We developed deep learning models for automated 3D segmentation of arteries and veins from dual-phase contrast-enhanced CT scans. Our dataset included 55 patients with arterial and venous phase scans, where major vessels down to third-order branching were manually annotated. Three nnU-Net variants (standard 3D U-Net, residual encoder U-Net, and SkeletonRecall) were trained independently for arteries and veins using five-fold cross-validation. Segmentation performance was evaluated using Dice score, centerline Dice score, sensitivity, and precision. All models achieved comparable Dice and centerline Dice scores, with slightly better results for veins segmentation. SkeletonRecall showed the highest sensitivity and superior vessel continuity, despite lower precision. On the independent test set, Dice scores were lower due to incomplete ground truth, yet SkeletonRecall correctly captured vessels up to third-order branching. The trained models are publicly available at https://github.com/LERCO-FNO/Abdominal-Vessels-Segmentation.
Bladder cancer is among the most common malignancies, with early detection being critical for effective treatment. This work investigates AI-based lesion detection in cystoscopic images, leveraging both YOLO and visual transformer (VT) architectures. Multiple datasets, including newly collected and publicly available sources, were systematically combined to train and evaluate detection models. Results show that increasing diversity and volume of training data significantly improves detection performance. Pretraining with colonoscopic images further improved model accuracy, indicating similarities in the appearance of lesions across different organs. While VT initially performed better, advanced YOLO outperformed VT with enriched data. These findings highlight the importance of heterogeneous datasets and model selection to advance automated bladder cancer detection.
Accurate knowledge of patient and mobile C-arm system orientation is essential for intraoperative 3D scan acquisition, yet this information is currently entered manually by operating room staff, making the process timeconsuming and error-prone.We propose a deep learning approach for the joint classification of patient and C-arm orientation using only a pair of anteriorposterior and lateral projection images. The method builds on frozen DAX foundation model embeddings, combined with a task-specific head network trained on 633 clinical 3D scans. The developed model achieved a weighted mean F1-score of 89.9
The performance of deep learning models in medical image analysis critically depends on the access to large and high-quality datasets. However, ethical, legal, and privacy constraints often limit data availability. Generative models offer a promising solution by producing synthetic training data, yet their resource-efficient fine-tuning remains an open challenge. This study investigates whether stable diffusion (SD) can be adapted using low-rank adaptation (LoRA) with minimal data and computational resources to generate synthetic hand radiographs (X-rays). The aim is not perfect anatomical fidelity but the reproduction of key X-ray characteristics under restrictive conditions. Quantitative and qualitative comparisons of a generic and a medically pre-trained SD 1.4 model showthat both can produce visually plausibleX-rays despite anatomical imperfections. These findings demonstrate the potential of lightweight fine-tuning methods for medical imaging and underscore the need for systematic research on training efficiency, quality assessment of synthetic data, and integration of synthetic and real datasets in medical AI.
The proposed algorithm is a differentiable approximate truncation robust computed tomography (ATRACT) reconstruction algorithm for end-to-end trainable cone-beam CT reconstruction, offering enhanced robustness to truncated geometries and a significant reduction of truncation artifacts. The proposed method utilizes known operator learning to map the analytical reconstruction into a neural network. This approach preserves physical consistency while enabling data-driven optimization of redundancy weights under the 180◦ limited-angle condition. The experimental results demonstrate that the proposed frame work surpasses the Parker-weighted analytical reconstruction, achieving a 1.5
The loss of a lower limb severely affects mobility and quality of life, with the fit of the prosthesis being essential for user comfort and functionality. Techniques for manufacturing prosthetic sockets, such as plaster casting, are time-consuming and dependent on the prosthetist’s expertise, while advanced imaging modalities like computed tomography (CT) and magnetic resonance imaging (MRI) are costly and introduce positioning-related inaccuracies. Ultrasound (US) presents a promising alternative for capturing both external and internal limb structures in a non-invasive manner, though mostly limited to 2D. This study investigates a novel approach for 3D reconstruction of US scans using a N-shaped fiducial pattern integrated into a liner applied to the limb. The fiducials facilitate registration of US images to reconstruct a 3D model. Validation on a limb phantom yielded an average reconstruction accuracy of 1.5 mm for bone and 1.7 mm for skin on a partial volume. While our findings demonstrate feasibility of the method, future work aims to enhance reconstruction accuracy by refining image alignment techniques and expanding the scanning approach to the whole limb.
Accurate detection and segmentation of brain metastases (BM) are essential for stereotactic radiosurgery (SRS) planning. This study first systematically compares six loss functions across three representative 3D deep learning models. Results show Dice is a robust baseline, CE improves precision, and Focal has limited effectiveness. While JVSS substantially enhances small-lesion sensitivity, it reduces precision; combined losses achieve the most balanced and robust performance. Furthermore, we propose a novel inference strategy: local overlap fusion for subvolume merging (LOF_SM). By applying axis-wise local weighting in overlapping regions, LOF_SM significantly improves computational efficiency, reducing inference time by approximately 30
Anti-scatter grids enhance image contrast in digital radiography but can introduce gridline artifacts when grid and detector sampling frequencies misalign. We propose a dual-domain AI framework that integrates a frozen DINO Vision Transformer with a FiLM-conditioned U-Net, combining global semantic encoding with frequency-aware reconstruction. The DINO embeddings modulate U-Net activations via Feature-wise Linear Modulation, enabling anatomically consistent correction through the joint prediction of a spatial residual and a frequency-domain attenuation mask, which are adaptively fused to suppress gridline artifacts while preserving tone and structural detail. A physics-based synthetic dataset comprising 1,475 clean and 4,425 grid-contaminated radiographs was generated from real detector captures. On 555 test images with available clean references, the model achieved a mean PSNR of 36.9 dB and an SSIM of 0.98, demonstrating high-fidelity and structurally consistent reconstruction across varying grid frequencies and exposure conditions.