Purpose To develop and validate LION (Lesion Identification in Oncological Nuclear imaging), an open-source PET-only tumor segmentation pipeline for [ 18 F]FDG and PSMA-targeted PET/CT, and to investigate how training data characteristics influence segmentation performance. Materials and Methods In this retrospective multicenter study, 5,209 [ 18 F]FDG PET/CT scans spanning 19 disease types and 2,046 PSMA-targeted PET/CT scans were used to train PET-only segmentation models. Tumor segmentation incorporated organs with physiological uptake as auxiliary classes to enable PET-only inference. Tumor Occurrence Maps (TOMs) quantified tumor spatial diversity across the training data. For [ 18 F]FDG, disease-specific and mixed-disease models trained on progressively larger subsets were compared to test whether increasing spatial diversity improves generalization. Scanner-related domain shift was analyzed using DINOv2 embeddings. Models were evaluated on multicenter holdout cohorts (616 [ 18 F]FDG across 4 diseases; 443 PSMA-targeted prostate cancer scans) and compared with three open-source tools. Results Organ context improved median Dice from 0.62 to 0.71 for [ 18 F]FDG and from 0.75 to 0.83 for PSMA. Spatial diversity measured by TOMs was strongly associated with Dice (Spearman ρ = 0.87, P < 0.001). A mixed-disease model trained on 500 patients matched the performance of a lymphoma specialist model trained on 3,031 cases. DINOv2 embeddings revealed scanner-induced domain shift, between same-disease cohorts. LION achieved median Dice scores of 0.71 ([ 18 F]FDG) and 0.85 (PSMA) and outperformed other open-source approaches on common holdout patients. Conclusion LION enables PET-only automated segmentation for [ 18 F]FDG and PSMA-targeted PET. Training data composition, particularly spatial diversity quantified by TOMs, was strongly associated with segmentation performance.
We present a large whole-body and total-body curated dataset of dual-modality 2-deoxy-2-[18F]fluoro-D-glucose (FDG)-Positron Emission Tomography/Computed Tomography (PET/CT) studies, consisting of 1,683 PET/CT images and the corresponding CT-derived segmentations of 130 target regions. This multi-center dataset includes images from individuals without overt disease and patients with a range of malignant and inflammatory pathologies, including arthritis, lymphoma, and melanoma, as well as cancers of the lung, head-neck, and genito-urinary tract. Target regions were first automatically segmented from CT images using an in-house software and subsequently verified and corrected by physicians-in-training. In total, the segmented regions encompass 130 volumes, including abdominal organs, muscles, bones, cardiac subregions, vessels, adipose tissue, and skeletal muscle around the third lumbar vertebra. PET/CT images and corresponding CT-derived segmentations are provided in anonymized NIfTI format. The dataset can be used for deep learning training, validation, or multi-modality image analysis and thus fills an important gap in available resources to advance the use of PET/CT data in clinical management.
Advancements of deep learning in medical imaging are often constrained by the limited availability of large, annotated datasets, resulting in underperforming models when deployed under real-world conditions. This study investigated a generative artificial intelligence (AI) approach to create synthetic medical images taking the example of bone scintigraphy scans, to increase the data diversity of small-scale datasets for more effective model training and improved generalization. We trained a generative model on 99mTc-bone scintigraphy scans from 9,170 patients in one center to generate high-quality and fully anonymized annotated scans of patients representing two distinct disease patterns: abnormal uptake indicative of (i) bone metastases and (ii) cardiac uptake indicative of cardiac amyloidosis. A blinded reader study was performed to assess the clinical validity and quality of the generated data. We investigated the added value of the generated data by augmenting an independent small single-center dataset with synthetic data and by training a deep learning model to detect abnormal uptake in a downstream classification task. We tested this model on 7,472 scans from 6,448 patients across four external sites in a cross-tracer and cross-scanner setting and associated the resulting model predictions with clinical outcomes. The clinical value and high quality of the synthetic imaging data were confirmed by four readers, who were unable to distinguish synthetic scans from real scans (average accuracy: 0.48
[18F]-Flurpiridaz is an emerging radiopharmaceutical with significant potential to enhance myocardial perfusion imaging. With a myocardial extraction fraction comparable to [15O]-H₂O, Flurpiridaz offers superior image quality, higher diagnostic accuracy, and reduced radiation exposure compared to traditional SPECT tracers. Unlike short half-life PET tracers, Flurpiridaz does not require an on-site cyclotron, making it more accessible for clinical use. Its ability to accurately quantify myocardial blood flow and coronary flow reserve could improve ischemia detection and patient management. However, challenges remain regarding regulatory approval, economic feasibility, and infrastructure adaptation, particularly in Italy, where SPECT remains dominant. Despite these hurdles, Flurpiridaz represents a significant advancement in nuclear cardiology, with the potential to improve diagnostic precision, guide revascularization decisions, and optimize personalized treatment strategies for coronary artery disease. [18F]-Flurpiridaz is an emerging radiopharmaceutical with significant potential to enhance myocardial perfusion imaging. With a myocardial extraction fraction comparable to [15O]-H₂O, Flurpiridaz offers superior image quality, higher diagnostic accuracy, and reduced radiation exposure compared to traditional SPECT tracers. Unlike short half-life PET tracers, Flurpiridaz does not require an on-site cyclotron, making it more accessible for clinical use. Its ability to accurately quantify myocardial blood flow and coronary flow reserve could improve ischemia detection and patient management. However, challenges remain regarding regulatory approval, economic feasibility, and infrastructure adaptation, particularly in Italy, where SPECT remains dominant. Despite these hurdles, Flurpiridaz represents a significant advancement in nuclear cardiology, with the potential to improve diagnostic precision, guide revascularization decisions, and optimize personalized treatment strategies for coronary artery disease.
Explainability is a leading solution offered to address the challenge of AI’s black boxing. However, a lot can go wrong when trying to apply explainability, and its success is far from certain. Moreover, there is insufficient empirical data regarding the effectiveness of concrete explainability efforts. We examined an explainability scenario for an AI decision support tool under development for the early detection of cancer-related cachexia, a potentially fatal metabolic syndrome. We conducted 13 interviews with clinicians who deal with cachexia, and asked about their prior experience with AI tools, their views on explainability, and presented an explainability scenario based on the Shapley Additive Explanations (SHAP) method. Most clinicians we interviewed had limited prior experience with AI tools, and a majority of them believed that the explainability of such an AI system for the early detection of cachexia is essential. When presented with the SHAP explainability scheme, they had limited familiarity with the features that contributed to the tool’s ruling, and only a minority of the clinicians (nuclear medicine experts) stated that they could utilize these features in a meaningful manner. Paradoxically, it is the clinicians who come in contact with patients who cannot make use of this specific SHAP explanation. This study highlights the challenges of offering a hyper-selective explainability tool in clinical settings. It also shows the challenge of developing explainable-by-design AI systems.
PurposeIn this short communication, we consider the need for explainable AI from the perspective of a large multi-disciplinary research project for predicting cachexia in cancer patients.Materials and methodsIn a series of meetings, comprising expertise from medicine, data science, sociology, and philosophy, project participants discussed the need for explainability.ResultsWe distinguish between contexts in which a black box AI tool undertakes tasks that users can perform or validate themselves and contexts in which this is not the case.ConclusionWe conclude that explanations are likely required when a black box AI tool undertakes tasks that users cannot perform or validate themselves. If the user can verify outputs manually, documented reliability and accuracy may suffice, but explainability can still add value when outputs are uncertain or errors occur. More generally, close collaboration among physicians, AI developers, and other stakeholders is crucial to ensure that AI tools are trustworthy and useful in clinical practice.
Combined PET/CT imaging provides critical insights into both anatomic and molecular processes, yet traditional single-tracer approaches limit multidimensional disease phenotyping; to address this, we developed the PET Unified Multitracer Alignment (PUMA) framework-an open-source, postprocessing tool that multiplexes serial PET/CT scans for comprehensive voxelwise tissue characterization. Methods: PUMA utilizes artificial intelligence-based CT segmentation from multiorgan objective segmentation to generate multilabel maps of 24 body regions, guiding a 2-step registration: affine alignment followed by symmetric diffeomorphic registration. Tracer images are then normalized and assigned to red-green-blue channels for simultaneous visualization of up to 3 tracers. The framework was evaluated on longitudinal PET/CT scans from 114 subjects across multiple centers and vendors. Rigid, affine, and deformable registration methods were compared for optimal coregistration. Performance was assessed using the Dice similarity coefficient for organ alignment and absolute percentage differences in organ intensity and tumor SUVmean Results: Deformable registration consistently achieved superior alignment, with Dice similarity coefficient values exceeding 0.90 in 60% of organs while maintaining organ intensity differences below 3%; similarly, SUVmean differences for tumors were minimal at 1.6% ± 0.9%, confirming that PUMA preserves quantitative PET data while enabling robust spatial multiplexing. Conclusion: PUMA provides a vendor-independent solution for postacquisition multiplexing of serial PET/CT images, integrating complementary tracer data voxelwise into a composite image without modifying clinical protocols. This enhances multidimensional disease phenotyping and supports better diagnostic and therapeutic decisions using serial multitracer PET/CT imaging.
Background/Objectives: Artificial Intelligence (AI) is becoming increasingly important in Medicine. The aim of this review is to summarize its use in the field of Nuclear Cardiology. Methods: First, we provide a short description of how AI works. Then we performed a review of the literature focusing on the articles in which AI is used for image interpretation for diagnostic or prognostic purposes. Results: AI has been applied according to various approaches for both diagnosis and prognosis. The achieved gains have been so far relatively limited as compared to traditional methodologies. However, promising results have been reported, including interesting perspectives for the explainability of AI results and their potential integration in clinical routine. Conclusions: AI is soon going to play an important role in Nuclear Cardiology, but further improvements are needed to reach significant gains in terms of diagnostic accuracy, and prospective studies on its prognostic capabilities are still lacking. Furthermore, several important issues must be solved, such as availability and feasibility within the processing workflow, explainability, liability, and ethics of its application in clinical decision-making.
Not many years ago, there was still a debate about the best sequence for performing myocardial perfusion imaging. The traditional rest-stress approach was preferred by many centers for its simplicity and the advantage of acquiring the stress part with the higher dose in case of single day protocol [1]. With time, however, it became apparent that an increasing percentage of patients did not show stress-induced perfusion defects, so that after a normal stress study rest images were not required [2].
BACKGROUND:Cancer-associated cachexia (CAC) is a metabolic syndrome contributing to therapy resistance and mortality in lung cancer patients (LCP). CAC is typically defined using clinical non-imaging criteria. Given the metabolic underpinnings of CAC and the ability of [18F]fluoro-2-deoxy-D-glucose (FDG)-positron emission tomography (PET)/computer tomography (CT) to provide quantitative information on glucose turnover, we evaluate the usefulness of whole-body (WB) PET/CT imaging, as part of the standard diagnostic workup of LCP, to provide additional information on the onset or presence of CAC. METHODS:This multi-centre study included 345 LCP who underwent WB [18F]FDG-PET/CT imaging for initial clinical staging. A weight loss grading system (WLGS) adjusted to body mass index was used to classify LCP into 'No CAC' (WLGS-0/1 at baseline prior treatment and at first follow-up: N = 158, 51F/107M), 'Dev CAC' (WLGS-0/1 at baseline and WLGS-3/4 at follow-up: N = 90, 34F/56M), and 'CAC' (WLGS-3/4 at baseline: N = 97, 31F/66M). For each CAC category, mean standardized uptake values (SUV) normalized to aorta uptake () and CT-defined volumes were extracted for abdominal and visceral organs, muscles, and adipose-tissue using automated image segmentation of baseline [18F]FDG-PET/CT images. Imaging and non-imaging parameters from laboratory tests were compared statistically. A machine-learning (ML) model was then trained to classify LCP as 'No CAC', 'Dev CAC', and 'CAC' based on their imaging parameters. SHapley Additive exPlanations (SHAP) analysis was employed to identify the key factors contributing to CAC development for each patient. RESULTS:The three CAC categories displayed multi-organ differences in . In all target organs, was higher in the 'CAC' cohort compared with 'No CAC' (P < 0.01), except for liver and kidneys, where in 'CAC' was reduced by 5%. The 'Dev CAC' cohort displayed a small but significant increase in of pancreas (+4%), skeletal-muscle (+7%), subcutaneous adipose-tissue (+11%), and visceral adipose-tissue (+15%). In 'CAC' patients, a strong negative Spearman correlation (ρ = -0.8) was identified between and volumes of adipose-tissue. The machine-learning model identified 'CAC' at baseline with 81% of accuracy, highlighting of spleen, pancreas, liver, and adipose-tissue as most relevant features. The model performance was suboptimal (54%) when classifying 'Dev CAC' versus 'No CAC'. CONCLUSIONS:WB [18F]FDG-PET/CT imaging reveals groupwise differences in the multi-organ metabolism of LCP with and without CAC, thus highlighting systemic metabolic aberrations symptomatic of cachectic patients. Based on a retrospective cohort, our ML model identified patients with CAC with good accuracy. However, its performance in patients developing CAC was suboptimal. A prospective, multi-centre study has been initiated to address the limitations of the present retrospective analysis.
To recognize patients at high risk of refractory disease, the identification of novel prognostic parameters improving stratification of newly diagnosed Hodgkin Lymphoma (HL) is still needed. This study investigates the potential value of metabolic and texture features, extracted from baseline 18F-FDG Positron Emission Tomography/Computed Tomography (PET) and Contrast-Enhanced Computed Tomography scan (CECT), together with clinical data, in predicting first-line therapy refractoriness (R) of classical HL (cHL) with mediastinal bulk involvement. We reviewed 69 cHL patients who underwent staging PET and CECT. Lesion segmentation and texture parameter extraction were performed using the freeware software LIFEx 6.3. The prognostic significance of clinical and imaging features was evaluated in relation to the development of refractory disease. Receiver operating characteristic curve, Cox proportional hazard regression and Kaplan-Meier analyses were performed to examine the potential independent predictors and to evaluate their prognostic value. Among clinical characteristics, only stage according to the German Hodgkin Group (GHSG) classification system significantly differed between R and not-R. Among CECT variables, only parameters derived from second order matrices (gray-level co-occurrence matrix (GLCM) and gray-level run length matrix (GLRLM) demonstrated significant prognostic power. Among PET variables, SUVmean, several variables derived from first (histograms, shape), and second order analyses (GLCM, GLRLM, NGLDM) exhibited significant predictive power. Such variables obtained accuracies greater than 70% at receiver operating characteristic analysis and their PFS curves resulted statistically significant in predicting refractoriness. At multivariate analysis, only HISTO_EntropyPET extracted from PET (HISTO_EntropyPET ) and GHSG stage resulted as significant independent predictors. Their combination identified 4 patient groups with significantly different PFS curves, with worst prognosis in patients with higher HISTO_EntropyPET values, regardless of the stage. Imaging radiomics may provide a reference for prognostic evaluation of patients with mediastinal bulky cHL. The best prognostic value in the prediction of R versus not-R disease was reached by combining HISTO_EntropyPET with GHSG stage.
Purpose The purpose of this study was to create 123 I-FP-CIT reference values for ultra-high-resolution fan beam collimators (UHR-FB) from a sample of subjects without dopaminergic degeneration and to compare them to a normal database -PPMI database- of a commercial software (DaTQUANT) obtained using high-resolution parallel-hole collimators (HR-PH). Methods A striatal phantom study was performed to compare UHR-FB with HR-PH and to obtain a correction factor between collimators. Normal 123 I-FP-CIT studies from 177 subjects acquired using UHR-FB were retrospectively selected on the basis of visual and semi-quantitative analysis as well as of the neurological follow-up (range of 2–9 years). SPECT images were reconstructed using the same parameters of DaTQUANT normal database and SBR values were obtained for striatal structures. Correction factor was applied to the UHR-FB database to test differences against DaTQUANT database. Results Correction factor obtained from the phantom study was 0.84. Uncorrected SBR values of the local database were significantly higher than PPMI database values, but no significant differences were found using corrected values. Coefficients of variations of SBR values were significantly lower in a local database than PPMI database (15% vs 20%). Significant effects of age on SBR were observed in both databases with a reduction rate for a decade of 6% in the PPMI database and 4.5% in the local database. In the latter, women had slightly higher SBR values and a steeper decline with advancing age compared to men, whereas no significant gender differences were found in the PPMI database. Conclusion The SBR values obtained using UHR-FB have an age-related distribution comparable to that of healthy subjects but with lower variability. The reduction rate per decade was similar between the two databases but the gender effect was found only in the local database, probably related to the better performance of UHR-FB.
BACKGROUND: The extent of myocardial bone tracer uptake with technetium pyrophosphate, hydroxymethylene diphosphonate, and 3,3-diphosphono-1,2-propanodicarboxylate in transthyretin amyloid cardiomyopathy (ATTR-CM) might reflect cardiac amyloid burden and be associated with outcome. METHODS: Consecutive patients with ATTR-CM who underwent diagnostic bone tracer scintigraphy with acquisition of whole-body planar and cardiac single-photon emission computed tomography (SPECT) images from the National Amyloidosis Centre and 4 Italian centers were included. Cardiac uptake was defined according to the Perugini classification: 0=absent cardiac uptake; 1=mild uptake less than bone; 2=moderate uptake equal to bone; and 3=high uptake greater than bone. Extent of right ventricular (RV) uptake was defined as focal (basal segment of the RV free wall only) or diffuse (extending beyond basal segment) on the basis of SPECT imaging. The primary outcome was all-cause mortality. RESULTS: Among 1422 patients with ATTR-CM, RV uptake accompanying left ventricular uptake was identified by SPECT imaging in 100% of cases at diagnosis. Median follow-up in the whole cohort was 34 months (interquartile range, 21 to 50 months), and 494 patients died. By Kaplan-Meier analysis, diffuse RV uptake on SPECT imaging (n=936) was associated with higher all-cause mortality compared with focal (n=486) RV uptake (77.9% versus 22.1%; P<0.001), whereas Perugini grade was not associated with survival (P=0.27 in grade 2 versus grade 3). On multivariable analysis, after adjustment for age at diagnosis (hazard ratio [HR], 1.03 [95% CI, 1.02-1.04]; P<0.001), presence of the p.(V142I) TTR variant (HR, 1.42 [95% CI, 1.20-1.81]; P=0.004), National Amyloidosis Centre stage (each category, P<0.001), stroke volume index (HR, 0.99 [95% CI, 0.97-0.99]; P=0.043), E/e' (HR, 1.02 [95% CI, 1.007-1.03]; P=0.004), right atrial area index (HR, 1.05 [95% CI, 1.02-1.08]; P=0.001), and left ventricular global longitudinal strain (HR, 1.06 [95% CI, 1.03-1.09]; P<0.001), diffuse RV uptake on SPECT imaging (HR, 1.60 [95% CI, 1.26-2.04]; P<0.001) remained an independent predictor of all-cause mortality. The prognostic value of diffuse RV uptake was maintained across each National Amyloidosis Centre stage and in both wild-type and hereditary ATTR-CM (P<0.001 and P=0.02, respectively). CONCLUSIONS: Diffuse RV uptake of bone tracer on SPECT imaging is associated with poor outcomes in patients with ATTR-CM and is an independent prognostic marker at diagnosis.
Background The diagnosis of cardiac amyloidosis can be established non-invasively by scintigraphy using bone-avid tracers, but visual assessment is subjective and can lead to misdiagnosis. We aimed to develop and validate an artificial intelligence (AI) system for standardised and reliable screening of cardiac amyloidosis-suggestive uptake and assess its prognostic value, using a multinational database of Tc-99m-scintigraphy data across multiple tracers and scanners. Methods In this retrospective, international, multicentre, cross-tracer development and validation study, 16 241 patients with 19 401 scans were included from nine centres: one hospital in Austria (consecutive recruitment Jan 4, 2010, to Aug 19, 2020), five hospital sites in London, UK (consecutive recruitment Oct 1, 2014, to Sept 29, 2022), two centres in China (selected scans from Jan 1, 2021, to Oct 31, 2022), and one centre in Italy (selected scans from Jan 1, 2011, to May 23, 2023). The dataset included all patients referred to whole-body Tc-99m-scintigraphy with an anterior view and all Tc-99m-labelled tracers currently used to identify cardiac amyloidosis-suggestive uptake. Exclusion criteria were image acquisition at less than 2 h (Tc-99m-3,3-diphosphono-1,2-propanodicarboxylic acid, Tc-99m-hydroxymethylene diphosphonate, and Tc-99m-methylene diphosphonate) or less than 1 h (Tc-99m-pyrophosphate) after tracer injection and if patients' imaging and clinical data could not be linked. Ground truth annotation was derived from centralised core-lab consensus reading of at least three independent experts (CN, TT-W, and JN). An AI system for detection of cardiac amyloidosis-associated high-grade cardiac tracer uptake was developed using data from one centre (Austria) and independently validated in the remaining centres. A multicase, multireader study and a medical algorithmic audit were conducted to assess clinician performance compared with AI and to evaluate and correct failure modes. The system's prognostic value in predicting mortality was tested in the consecutively recruited cohorts using cox proportional hazards models for each cohort individually and for the combined cohorts. Findings The prevalence of cases positive for cardiac amyloidosis-suggestive uptake was 142 (2%) of 9176 patients in the Austrian, 125 (2%) of 6763 patients in the UK, 63 (62%) of 102 patients in the Chinese, and 103 (52%) of 200 patients in the Italian cohorts. In the Austrian cohort, cross-validation performance showed an area under the curve (AUC) of 1 center dot 000 (95% CI 1 center dot 000-1 center dot 000). Independent validation yielded AUCs of 0 center dot 997 (0 center dot 993-0 center dot 999) for the UK, 0 center dot 925 (0 center dot 871-0 center dot 971) for the Chinese, and 1 center dot 000 (0 center dot 999-1 center dot 000) for the Italian cohorts. In the multicase multireader study, five physicians disagreed in 22 (11%) of 200 cases (Fleiss' kappa 0 center dot 89), with a mean AUC of 0 center dot 946 (95% CI 0 center dot 924-0 center dot 967), which was inferior to AI (AUC 0 center dot 997 [0 center dot 991-1 center dot 000], p=0 center dot 0040). The medical algorithmic audit demonstrated the system's robustness across demographic factors, tracers, scanners, and centres. The AI's predictions were independently prognostic for overall mortality (adjusted hazard ratio 1 center dot 44 [95% CI 1 center dot 19-1 center dot 74], p<0 center dot 0001). Interpretation AI-based screening of cardiac amyloidosis-suggestive uptake in patients undergoing scintigraphy was reliable, eliminated inter-rater variability, and portended prognostic value, with potential implications for identification, referral, and management pathways. Copyright (c) 2024 The Author(s). Published by Elsevier Ltd. This is an Open Access article under the CC BY 4.0 license.