High-density metal implants introduce severe artifacts in computed tomography (CT) images that degrade diagnostic quality and obscure anatomical structures. Although deep learning-based metal artifact reduction (MAR) methods have shown progress, they often struggle to suppress globally distributed streaks while preserving fine anatomical details. In this work, we present a frequency-aware proximal optimization framework for CT metal artifact reduction, termed the Fourier-Enhanced Proximal Network. This architecture introduces a prior-driven iterative model with Fast Fourier Convolution to enable joint spatial and frequency-domain processing.
Patient motion is a significant problem in interventional C-arm computed tomography (CT), as the relatively long acquisition times increase the likelihood of patient motion.
Differentiable shift-variant filtered backprojection (SV-FBP) is a fast and efficient reconstruction algorithm for cone-beam computed tomography with non-circular trajectories. While this method is significantly faster than iterative methods, its reliance on trajectory-specific training of Radondomain weights limits its ability to generalise and reuse it across different scanning geometries. In this paper, we propose a novel training framework that facilitates trajectory-generalizable learning for SV-FBP. The fundamental concept underpinning this methodology is a preprocessing pipeline that standardises a set of CBCT trajectories by reordering projections based on geometric continuity. This standardization decouples the learned SV-FBP weights from the absolute trajectory configuration, thereby making them invariant to projection order, allowing a single model to be reused across multiple trajectories. Experimental results show that the proposed method enables a single SV-FBP model to generalize across four distinct non-circular CBCT trajectories, achieving consistent performance with an average MSE of 0.1207, PSNR of 34.68 dB, and SSIM of 0.9177, without the need for retraining. This confirms the effectiveness of the standardization framework in enhancing flexibility and scalability.
Breast cancer is the most frequently diagnosed malignancy among women globally and presents a growing challenge for healthcare systems. While early detection and accurate lesion assessment are vital, traditional radiology workflows are labor-intensive and prone to variability. Existing AI solutions in breast imaging are task-specific and lack the ability to generate comprehensive radiology reports. Vision-language models (VLMs), such as CLIP and BLIP, offer a promising direction for unifying visual understanding with natural language generation. We propose MammoBLIP, an end-to-end framework for automated mammography report generation based on the MedBLIP architecture. We curated a dataset of 81,076 images from five public sources-VinDR-Mammo, RSNA, CMMD, InBreast, and KAU—with standardized clinical annotations. Images are processed through a frozen EVA-CLIP Vision Transformer, followed by a lightweight transformer and projection layers to produce vision embeddings. These are aligned with text embeddings using a contrastive loss. A GPT-2-based BioMedLM model generates reports conditioned on these visual features. Only the transformer and projection heads are trained, keeping backbone models frozen to ensure computational efficiency. MammoBLIP achieves strong results across datasets, with an overall BLEU score of 65.36, ROUGE-1 of 0.75, BERT-F1 of 0.88, and SBERT similarity of 0.91. The generated reports are clinically coherent, with an average Flesch-Kincaid grade of ∼ 7. These findings provide initial insights into automated report generation and lay the groundwork for developing robust vision-language models tailored to clinical imaging tasks. Our results highlight MammoBLIP's potential to streamline workflows, enhance diagnostic consistency, and aid in radiologist training in future applications.
Interventional C-arm cone beam CT significantly reduces time-to-therapy for patients suffering from acute stroke. Extended acquisition times as compared to modern helical CT systems make rigid patient motion more likely, introducing a mismatch with the geometry alignment presupposed during reconstruction, leading to typical blurring or streaking artifacts. Different autofocus methods exist which aim to find a compensation trajectory and reestablish a motion-free reconstruction by optimizing a quality metric on the reconstructed images. However, depending on the type of implemented metric, these methods show diminished performance for larger motion amplitudes, often converging into local minima. To overcome these drawbacks, we present OrdinalNet, a deep learning model predicting a pseudo-binary score to assess motion artifact deterioration or improvement between two reconstructions resulting from slightly different motion trajectories. Our method eliminates the necessity for absolute artifact quantification in each image, along with establishing a relative ordering among data points which is sufficient for effective simplex-based optimization. The model is trained and evaluated on simulated motion applied to motion-free clinical acquisitions. When presented 1000 motion-affected pairs from previously unseen patients, it shows superior performance on a threshold-dependent binary classification task when compared to an entropy-based metric, resulting in areas under the curve of $\mathrm{AUC}_{\text {ONet }}=0.9295$ and $\mathrm{AUC}_{\mathrm{Ent}}=0.6366$
die Diagnostik maligner Schleimhautveränderungen im Bereich des oberen Aerodigestivtrakts zu verbessern
Objective. Multiple algorithms have been proposed for data driven gating (DDG) in single photon emission computed tomography (SPECT) and have successfully been applied to myocardial perfusion imaging (MPI). Application of DDG to acquisition types other than SPECT MPI has not been demonstrated so far, as limitations and pitfalls of current methods are unknown. Approach. We create a comprehensive set of phantoms simulating the influence of different motion artifacts, view angles, moving objects, contrast, and count levels in SPECT. We perform Monte Carlo simulation of the phantoms, allowing the characterization of DDG algorithms using quantitative metrics derived from the data and evaluate the Center of Light (COL) and Laplacian Eigenmaps methods as sample DDG algorithms. Main results. View angle, object size, count rate density, and contrast influence the accuracy of both DDG methods. Moreover, the ability to extract the respiratory motion in the phantom was shown to correlate with the contrast of the moving feature to the background, the signal to noise ratio, and the noise in the data. Significance. We showed that reporting the average correlation to an external physical reference signal per acquisition is not sufficient to characterize DDG methods. Assessing DDG methods on a view-by-view basis using the simulations and metrics from this work could enable the identification of pitfalls of current methods, and extend their application to acquisitions beyond SPECT MPI.
Abstract Funding Acknowledgements Type of funding sources: Private company. Main funding source(s): Research Support from Siemens Healthineers GmbH. Background In bSSFP sequences commonly used for cardiac MRI, signal modulation (e.g. banding artifacts) due to B0 inhomogeneity is often observed, especially at higher field strengths. The spatial position of these artifacts can be shifted by a frequency offset to reduce artifacts in a region of interest (ROI), e.g. the heart. To this end, frequency scout (FS) scans are acquired to visually select the optimal frequency offset [1,2]. In this work, we propose a fully automated image-based system for selecting the optimal frequency offset on FS images based on machine learning. Methods The proposed prototype system consists of four main steps (Fig.1). First, a pre-trained deep-learning-based whole heart segmentation network is applied on a four chamber-view FS image to localize the ROI where artifacts should be reduced. Second, high frequency components within the ROI (for each frequency offset in the FS series) are extracted by successive processing of Fourier transformation, high-pass filtering, inverse Fourier transformation and subtraction over series. and N images with the lowest high-frequency content are selected. Third, an adaptive weighting map for each FS image is generated which penalizes signal deviations from a pixel-wise median that is calculated based on the selected images [3]. By averaging the maps and selecting the frame with maximum percentage, the optimal frequency offset is selected. A total of 38 datasets, acquired on multiple clinical 3T MRI scanners (MAGNETOM Skyra, Vida, Prisma, Lumina; Siemens Healthcare, Erlangen, Germany), were used to evaluate the proposed system. All FS series were annotated manually and used to compare with the system output. The experts were allowed to select multiple possible optimal FS images within a FS series. In case of multiple annotations, the system output was labelled as correct when it selected one of the offsets chosen by the expert. Further, the generated weighting maps were visually evaluated. Results The proposed system achieved an accuracy of 92.1% compared to experts’ ground truth annotations. From the failed cases (n=3), the maximum difference was off by 2 frames. Based on the generated weighting maps, a reasonable decision on the selection of the optimal frequency offset is made. The algorithm successfully selects an FS image with minimized banding and flow artifacts within the ROI (Fig. 2a). Further, it reveals that the generated weighting map correctly suppress areas containing artifacts (Fig. 2b). Conclusions Initial results demonstrate the feasibility of the proposed system to automatically select the optimal frequency offset on FS scans. Therefore, it can improve the automation of a cardiac MRI workflow. An example of the result of each step
Abstract Funding Acknowledgements Type of funding sources: Private company. Main funding source(s): Research support from Siemens Healthineers GmbH. Background Mitral valve (MV) motion parameters, assessable using CMR [1, 2], have been shown to help the diagnosis of cardiac dysfunction. To extract valve motion parameters, we propose a fully automatic AI-based prototype system that tracks annulus and apex landmarks by the registration network on time-resolved two- and four-chamber CMR cine views. Parameters such as displacements, velocities, mitral annular plane systolic excursion (MAPSE), or longitudinal shortening (LS) are automatically extracted and evaluated on a large CMR dataset (N=11000). Methods The system consists of two sequential neural networks with a processing step in between (Fig. 1a) [3]. Initially, a 2D UNet is applied to localize both MV annulus insertion points as well as the apex. Based on these points, the image processing step consists of rotating, cropping, and interpolating the images, allowing a standardized image impression for both long axis views. Finally, the registration network (VoxelMorph framework [4]) is applied to the processed series and tracks the MV annulus insertion points and apex over the cardiac cycle by the deformation fields obtained by the network. The system was trained on (N=166) multivendor, multi-field strength, ground-truth annotated datasets [5]. A total of 11000 datasets, acquired on a 1.5T scanner (MAGNETOM Aera, Siemens Healthcare, Erlangen, Germany) from January 2016 to September 2017 [6], were used for parameter extraction. 200 of these datasets were additionally annotated semi-automatically for the performance evaluation of the system. Five motion parameters were automatically derived by the system that are defined as follows (Fig. 1b): (1) The atrioventricular plane displacement (AVPD) as the distance of the plane spanned by the MV annulus points relative to the first frame, (2) the atrioventricular plane velocity (AVPV) as the discrete temporal derivate of the AVPD, (3) the diameter of the annulus as the maximum distance between the MV annulus points, (4) the lateral/inferior and septal/superior MAPSE, as the maximum MV points’ excursion, and (5) the LS as the percentage size difference of the distance between the mid valvular point and the apex point at end-systole and end-diastole. Results The accuracy of the system resulted in deviations on the annotated dataset of 1.02 ± 0.87 mm, 0.01 ± 0.02 mm/s, 1.54 ± 1.21 mm, 2.30 ± 1.35 mm, 2.1 ± 1.8 mm for AVPD, AVPV, diameter, MAPSE, and LS respectively. Initial statistics on all datasets (Fig. 2) revealed a mean lateral/inferior, septal/superior MAPSE and LS of 8.7 ± 2.7 mm, 10.5 ± 3.2 mm and 16.3 ± 4.2 % for two-chamber and 9.6 ± 2.6 mm, 8.7 ± 2.6 mm and 15.5 ± 3.9 % for four-chamber views, respectively. Conclusions The results demonstrate the versatility of the proposed system for automatic extraction of various MV motion parameters. The proposed system enables automatic extraction of clinically relevant parameters and can improve the automation of MV-based analyses. System overview & Parameter of interestsAnalysis of the extracted parameters
BackgroundWhile MRI evaluation of joints has been primarily used to quantify inflammation at a cross-sectional and longitudinal level, less is known about the potential of MRI in distinguishing different patterns of inflammation in the various forms of arthritis.ObjectivesTo evaluate (i) whether deep learning using neural networks can be trained to distinguish between seropositive rheumatoid arthritis (RA+), seronegative RA (RA-), and psoriatic arthritis (PsA) based on structural inflammatory patterns on hand magnetic resonance imaging and (ii) to assess if psoriasis patients with subclinical inflammation fit into such patterns.MethodsResNet 3D [1] neural networks were trained to distinguish (i) RA+ vs. PsA, (ii) RA- vs. PsA and (iii) RA+ vs. RA- with respect to hand MRI data. Diagnosis of patients was determined using the following guidelines: ACR/EULAR 2010 [2] for RA and CASPAR [3] for PsA. Results from T1 coronal, T2 coronal, T1 coronal and axial fat suppressed contrast-enhanced (CE) and T2 fat suppressed axial sequences were used. The performance of such trained networks was analyzed by the area-under-the-receiver-operating-characteristic curve (AUROC) with and without imputation of demographic and clinical parameters (Figure 1A). Additionally, the trained networks were applied to psoriasis patients without clinical signs of PsA.Figure 1.(A) Neural network combining MR sequences with optional additional clinical data. The prediction for a single case is formed by averaging the prediction of all sequences and the clinical data. (B) Plot of the AUROC for increasing percentages (0.6 – 60%) of training data for the differentiation between RA+ and PsA by the neural network. The light blue area around the dark blue mean indicates the uncertainty measured using a 5-fold cross-validation.ResultsMRI scans from 649 patients (135 RA-, 190 RA+, 177 PsA, 147 psoriasis) were included (Table 1). The AUROC for differentiation between disease entities was 75% (SD 3%) for RA+ vs. PsA, 74% (SD 8%) for RA- vs. PsA, and 67% (6%) for RA+ vs. RA-. All MRI sequences were relevant for classification, however, when deleting CE sequences, the loss of performance was only marginal. The addition of patient-specific data to the networks did not provide significant improvements. Increasing amounts of training data demonstrated improved performance of the networks (Figure 1B). Psoriasis patients were mostly assigned to PsA by the neural networks, suggesting that PsA-like MRI pattern may be present early in the course of psoriatic disease.Table 1.Overview of demographic and clinical information.RA+RA-PsAPsoriasisTotal Number (N)649Number (N)190135177147Age (years), mean±SD56.9±12.660.5±10.356.3±12.049.6±13.8Sex (female/male)126/6493/4292/8571/76BMI (kg/m2), mean±SD26.6±10.527.6 ±9.329.1±11.326.7±6.9Disease duration (years), mean±SD2.6±4.91.3±2.30.8±2.34.2±5.1DAS28, mean±SD3.3±1.33.4±1.23.2±1.3-CRP (mg/L), mean±SD0.9±2.50.7±1.20.5±0.80.5±1.3HAQ, mean±SD0.8±0.60.9±0.80.6±0.60.3±0.4MedicationbDMARD88.46%83.87%81.32%35.01%csDMARD89.52%88.89%80.54%12.28%ConclusionDeep learning can be successfully applied to differentiate MRI inflammatory patterns related to RA+, RA-, and PsA. Early changes in psoriasis patients can be recognized by neural networks and are characterized by a pattern that allowed the networks to classify them as PsA.References[1]Kensho Hara, Hirokatsu Kataoka, and Yutaka Satoh 2018. Can Spatiotemporal 3D CNNs Retrace the History of 2D CNNs and ImageNet? In Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR) (pp. 6546–6555).[2]Aletaha D, Neogi T et al. 2010 Rheumatoid arthritis classification criteria: an American College of Rheumatology/European League Against Rheumatism collaborative initiative. Arthritis Rheum. 2010 Sep;62(9):2569-81.[3]Helliwell PS, Taylor WJ. Classification and diagnostic criteria for psoriatic arthritis. Annals of the Rheumatic Diseases 2005;64:ii3-ii8.AcknowledgementsThe study was supported by the Deutsche Forschungsgemeinschaft (DFG-FOR2886 PANDORA and the CRC1181 Checkpoints for Resolution of Inflammation). Additional funding was received by the Bundesministerium für Bildung und Forschung (BMBF; project MASCARA), the ERC Synergy grant 4D Nanoscope, the IMI funded projects HIPPOCRATES and RTCure, the Emerging Fields Initiative MIRACLE of the Friedrich-Alexander-Universität Erlangen-Nürnberg and the Else Kröner-Memorial Scholarship (DS, no. 2019_EKMS.27). Furthermore, infrastructural and hardware support was provided by the d.hip Digital Health Innovation Platform.Disclosure of InterestsNone declared
Deep learning-based speech enhancement has seen huge improvements and recently also expanded to full band audio (48 kHz). However, many approaches have a rather high computational complexity and require big temporal buffers for real time usage e.g. due to temporal convolutions or attention. Both make those approaches not feasible on embedded devices. This work further extends DeepFilterNet, which exploits harmonic structure of speech allowing for efficient speech enhancement (SE). Several optimizations in the training procedure, data augmentation, and network structure result in state-of-the-art SE performance while reducing the real-time factor to 0.04 on a notebook Core-i5 CPU. This makes the algorithm applicable to run on embedded devices in real-time. The DeepFilterNet framework can be obtained under an open source license.
Since the introduction of model-based iterative reconstruction for computed tomography (CT) by Thibault et al. in 2007, statistical weights play an important role in the problem formulation with the objective to improve image quality. Statistical weights depend on the variance of measurements. However, this variance is not known and therefore weights must be estimated. So far, the literature neither discusses how statistical weights should be estimated nor how accurate the estimation needs to be. Our submission aims at filling this gap in the literature. Specifically, we propose an estimation procedure for statistical weights and assess this procedure with real CT data. The estimated weights are compared against (clinically unpractical) sample weights obtained from repeated scans. The results show that the estimation procedure delivers reliable results for the rays that pass through the scanned object. Four imaging scenarios are considered; in each case the root mean square difference between the estimated and sample weights is below 5% of the maximum statistical weight value. When used for reconstruction, these differences are seen to have little impact: all voxel values within soft tissue (low contrast) regions differ by less than 1 HU. Our results demonstrate that the statistical weights can be sufficiently well estimated to closely approach the result that would be obtained if the weights were known.
Ziel Diese Pilotstudie zielte darauf ab, die Durchführbarkeit der intraoperativen Beurteilung der R0 Resektion mit der konfokalen Laserendomikroskopie (CLE) beim Oropharynxkarzinom zu evaluieren.
Background: Early diagnosis and reliable differentiation between rheumatic diseases (RMDs) are crucial to start an adequate therapy and prevent irreversible damage. Since finger joints are commonly affected in rheumatoid arthritis (RA) and psoriatic arthritis (PsA), imaging of the peripheral skeleton is an essential step of diagnosis at a rheumatologist. High resolution peripheral quantitative computed tomography (HR-pQCT) allows an even more detailed and three-dimensional (3D) illustration of the peripheral bone than conventional radiographs. Segmented scans contain further information, such as the density, microstructure, and shape of the bones, which can be further analyzed by neural networks. Objectives: We hypothesize that, based on the shape of the second metacarpophalangeal (MCP) joint from HR-pQCT images, a neural network can be trained to differentiate between RA, PsA, and healthy controls and to reveal regions in the bone shape characteristic for the diseases. Methods: HR-pQCT images of MCP joints from patients with classified CCP positive RA, classified PsA, and healthy controls with low motion artifacts and appropriate scan region were selected as reported previously [3]. Scans were performed as part of the clinical routine and patients gave their informed consent to use pseudonymized data (Ethics approval 334_16B). Based on the assumption that pathognomonic changes develop over time, only images were used, where the period between classification and imaging exceeded one year. Based on previous work [4], a pixel-wise mask of the second metacarpal bone was generated using a neural network based on the HR-pQCT scans of patients. Supervised auto-encoder [1] networks were used to predict the correct class given the bone mask only. For the neural network experiment, the patient scans were split on a patient-level into training (70%), validation (20%), and testing (10%). Guided backpropagation [2] was used as a method to investigate the regions influencing the class prediction most. Results: In total, images of 331 patients were included in the experiments. The evaluation of the model on the 33 test cases yielded a high accuracy for the healthy control with 94%, RA patients with 84%, and PsA patients with 89%. An area under the receiver operator curve of 91% could be achieved. The regions of the bone mask influencing the network´s decision most are highlighted exemplary in Figure 1. Figure 1. Visualization of the HR-pQCT slices with gradient maps. Higher values (red) represent regions that had a stronger contribution to the classification result. The HR-pQCT images are displayed for reference only. (a) Healthy patient, (b) RA diagnosed patient, and (c) PsA diagnosed patient. The first row shows the single slices with the highest values corresponding to the 3D bone masks in the second row. Conclusion: For the first time, a neural network-based approach successfully provides a differential diagnosis of RA and PsA based only on the shape of the second MCP in HR-pQCT images. The evaluation of the test set suggests that high curvatures of the bone surface in the joint region significantly influence the prediction of the network, suggesting an in-depth investigation of these regions for patients affected by RA and PsA. Based on these promising findings, we aim to extend the approach to seronegative RA as well as early RA and PsA. References: [1]Le, L. et al. (2018). Supervised autoencoders: Improving generalization performance with unsupervised regularizers. In Advances in Neural Information Processing Systems. [2]Springenberg, J. T. et al. (2015). Striving for simplicity: The all convolutional net. 3rd International Conference on Learning Representations, ICLR 2015 - Workshop Track Proceedings. [3]Simon, D. et al. (2017). Age- and Sex-Dependent Changes of Intra-articular Cortical and Trabecular Bone Structure and the Effects of Rheumatoid Arthritis. Journal of Bone and Mineral Research, 32(4), 722–730. [4]Folle, L. et al. (2021). Fully Automatic Bone Mineral Density Measurements using Deep Learning. Manuscript submitted for publication. Acknowledgements: This work was supported by the emerging field initiative (project 4 Med 05 “MIRACLE”) of the University Erlangen-Nürnberg and MASCARA - Molecular Assessment of Signatures Characterizing the Remission of Arthritis grant 01EC1903A. Disclosure of Interests: Lukas Folle: None declared, Chang Liu: None declared, David Simon Speakers bureau: Lilly, Novartis, Consultant of: Lilly, Novartis, Gilead, BMS, Abbvie, Grant/research support from: Lilly, Novartis, Timo Meinderink: None declared, Anna-Maria Liphardt Consultant of: Mylan/Meda Pharma, Grant/research support from: Novartis, Gerhard Krönke Speakers bureau: Lilly, Novartis, Consultant of: Lilly, Novartis, Gilead, BMS, Abbvie, Grant/research support from: Lilly, Novartis, Georg Schett Speakers bureau: Lilly, Novartis, Consultant of: Lilly, Novartis, Gilead, BMS, Abbvie, Grant/research support from: Lilly, Novartis, Andreas Maier: None declared, Arnd Kleyer Speakers bureau: Lilly, Novartis, Consultant of: Lilly, Novartis, Gilead, BMS, Abbvie, Grant/research support from: Novartis, Lilly
Tino Haderlein合作论文数Department of Computer Science, Friedrich-Alexander-Universität4