Objective.Spectral computed tomography (CT) data from photon-counting CT (PCCT) enables material decomposition. Mechanistic approaches such as maximum likelihood estimation are noise sensitive. Deep learning alternatives mitigate this issue, but their accuracy remains limited due to lack of incorporation of underlying physics principles and lack of ground truth data. This study aims to develop and validate a physics-informed deep-learning model, trained on validated simulated data, to decompose spectral CT images into density (ρ)and effective atomic number (Zeff) maps.Methods.The training dataset included simulated abdominal PCCT scans from 32 human models with corresponding ground truth. The scans were obtained at two clinical dose levels, four detector energy thresholds, different iodinated contrast agent concentrations and reconstructed using three clinically-used kernels. A generative adversarial network (GAN) was trained with and without a physics-informed regularization loss to estimateρandZeffmaps. Model performance was evaluated on 16 computational phantoms and validated on 6 clinical cases. A reader study was performed on 30 image slices to assess the comparative performance ofρandZeffmaps to multi-rendered virtual monochromatic images (VMIs) for assessing liver lesion conspicuity.Main results.With physics-informed regularization, NRMSE of 1.29% and 0.68%, SSIM of 0.99 and 0.99, and PSNR of 29.8 dB and 29.04 dB were achieved. A maximum RMSE of 5.45% was achieved on clinical data. Reader study results showedρandZeffimages had higher conspicuity scores compared to VMIs (median: 4.52 vs 4.13; 95% CIs: [4.19, 4.52] vs [4.01, 4.31]). The study showed equivalent conspicuity between VMIs and material images within a ±0.5 margin, though the small sample limits generalization.Significance.This study demonstrates the feasibility of material decomposition using a physics-informed GAN model trained on realistic simulated data. The maps provided equivalent conspicuity under a clinically acceptable margin, with a significantly small number of images for interpretation.
BACKGROUND:The adoption of virtual imaging trials (VITs) is rapidly expanding, offering a cost-effective and ethically viable alternative to large-scale clinical trials for imaging system evaluation. However, differences in demographic composition between virtual phantom populations and real-world clinical cohorts can introduce bias in imaging performance assessments, particularly for underrepresented populations. Such discrepancies, if unaddressed, can limit the translational relevance of VIT findings by misrepresenting diagnostic performance across diverse patient groups. PURPOSE:To address this limitation, we introduce DISTINCT (Distributional Subsampling for Covariate-Targeted Alignment), a statistical framework for selecting demographically aligned subsamples from large clinical datasets to support robust comparisons with virtual cohorts. METHODS:We applied DISTINCT to the National Lung Screening Trial (NLST) and a companion virtual trial dataset (VLST). The algorithm jointly aligned typical continuous (age, BMI) and categorical (sex, race, ethnicity) variables by constructing multidimensional bins based on discretized covariates. For a given target size, DISTINCT samples individuals to match the joint demographic distribution of the reference population. We evaluated the demographic similarity between VLST and progressively larger NLST subsamples using Wasserstein and Kolmogorov-Smirnov (K-S) distances to identify the maximal subsample size with acceptable alignment. After demographic alignment, we evaluated lung cancer risk prediction performance by applying two established NLST risk scores to the aligned subsamples and assessing their stability with receiver operating characteristic (ROC) analysis. RESULTS:The DISTINCT algorithm identified a maximal demographically aligned NLST subsample of 9974 participants that preserved similarity to the VLST population. To assess whether such aligned subsets were sufficient for downstream applications, we applied two established NLST lung cancer risk scores and evaluated their performance using ROC analysis. Area under the curve (AUC) estimates stabilized once subsample sizes exceeded approximately 6000 participants, demonstrating that moderately sized aligned subsets provide reliable predictive model evaluation. Stratified analyses revealed demographic-specific variations in AUC, underscoring the importance of covariate alignment for fair and representative comparisons. CONCLUSION:DISTINCT provides a statistically rigorous and scalable approach for covariate alignment between real and virtual imaging cohorts based on demographic factors of variability. Although demonstrated for lung cancer screening with low-dose CT, the framework is broadly applicable to other imaging modalities and diseases, and across wide ranges of factors of variability. By enabling fair and representative performance assessments, DISTINCT advances the integration of VITs into imaging research and protocol optimization workflows.
Background Diagnostic reference levels (DRLs) and achievable doses (ADs) are intended to reflect contemporary CT practice, but national benchmarks are updated infrequently. Purpose To update U.S. national adult CT DRLs and ADs derived from 2025 data and to demonstrate a scalable, acquisition-level, size-normalized benchmarking framework for generating timely and clinically relevant national benchmarks. Materials and Methods This retrospective observational study analyzed adult CT acquisitions performed between January 1 and December 31, 2025, across U.S. facilities using the Imalogix radiology optimization platform. The 10 most common adult CT examination categories, matching the 2014 benchmark study, were mapped using standardized RadLex Playbook identifiers. Dose metrics (volume CT dose index [CTDIvol] and dose-length product [DLP]) were extracted from dose structured reports and normalized at the acquisition level. Patient size was quantified using attenuation-based water-equivalent diameter from axial images with automated truncation correction, and size-specific dose estimates (SSDE) were calculated using standardized American Association of Physicists in Medicine task group methods. ADs and DRLs were calculated as the 50th and 75th percentiles, respectively, of facility-level median dose distributions per International Commission on Radiological Protection Publication 135 and compared with 2014 U.S. benchmarks. Results A total of 5 234 285 CT acquisitions performed in 2 802 416 adults (mean age, 58.1 years ± 19.7 [SD], 1 575 029 female) at 592 U.S. facilities were analyzed, representing a nearly fourfold larger sample than the 2014 study. Relative to 2014, national CTDIvol ADs and DRLs decreased by 9.3% and 21.8%, respectively. The largest reductions were observed for chest examinations with contrast material (31.2%), whereas smaller reductions were observed for head examinations without contrast material (3.5%). Dose metrics demonstrated consistent size dependence across all torso examinations. Conclusion U.S. adult CT radiation doses have decreased substantially over the past decade. This acquisition-level, size-normalized methodology enables near-real-time generation of national DRLs and ADs for modern CT dose benchmarking. © RSNA, 2026 Supplemental material is available for this article. See also the editorial by McCollough in this issue.
PURPOSE:The purpose of this study was to demonstrate the use of integrated PET/CT virtual imaging trials (VITs) for evaluating PET/CT acquisition protocols and clinically relevant sources of PET quantification variability under controlled, repeatable conditions. METHODS:An integrated PET/CT simulation framework was developed by combining SimSET for PET and DukeSim for CT, enabling concordant multi-modality simulations from a single human model input, including CT-derived attenuation correction. Geometric concordance was verified using an in-silico PET/CT coregistration test. Validation employed a NEMA IEC body phantom by comparing simulated PET/CT images with clinical scans using standard uptake value (SUV) and contrast recovery coefficient (CRC) metrics. Application studies assessed respiratory phase mismatch between PET and CT acquisitions, the impact of CT protocol selection on attenuation correction for free-breathing PET, and CT dose-reduction strategies for attenuation correction. RESULTS:Geometric concordance was achieved with a maximum PET/CT centroid difference of 1.83 mm. Phantom validation demonstrated agreement between simulated and clinical PET quantification, with SUV measurements within 6% using CT-derived attenuation correction and within 12% using an ideal attenuation map. Application studies revealed that respiratory phase mismatch and CT protocol selection can produce clinically significant differences in PET quantification metrics. Dose-reduction studies showed minimal sensitivity of PET quantification to ultra-low-dose CT attenuation correction, with CRC deviations ≤ 0.23% across CT conditions for the same sphere size. CONCLUSION:The integrated PET/CT simulation framework enables virtual imaging trials for controlled and reproducible investigations of PET/CT protocols, supporting multi-modality imaging optimization and quantitative performance evaluation under clinically relevant conditions.
BACKGROUND:Photon-counting CT (PCCT) is the latest technology enabling imaging with reduced noise and inherent spectral separation, with the potential to directly calculate a more accurate tissue stopping power from spectral data. This potential benefit is difficult to quantify in practice and is currently evaluated mainly in phantoms with simplified geometries that only approximate real patient anatomy. PURPOSE:In this work, we proposed virtual imaging simulators as an alternative approach to experimental validation of beam range uncertainty in complex patient geometry using a computational model of a human head and a CT system. In addition, we validate the accuracy of stopping power ratio (SPR) calculations on a model of a PCCT scanner using a conventional stoichiometric calibration approach and a prototype software TissueXplorer. METHODS:A validated CT simulator (DukeSim) was used to generate PCCT projections of a computational head phantom, which were reconstructed with an open-source toolbox (ASTRA). The dose of 2 Gy was delivered through protons in a single fraction to target two different cases of nasal and brain tumors. The ground-truth treatment plan was made directly on the computational phantom using clinical treatment planning software (RayStation). This plan was then recalculated on the corresponding CT images for which SPR values were estimated using both the conventional method and the prototype software TissueXplorer. The resulting dose distributions were subsequently compared against the ground-truth plan to quantify dose differences arising from SPR estimation. RESULTS:The mean percentage difference in estimating the SPR with TissueXplorer in all head tissues inside the scanned volume was 0.28%. SPRs obtained with this method showed smaller dose distribution differences from the ground truth plan than the conventional stoichiometric calibration method on the computational head phantom. CONCLUSIONS:Virtual imaging offers an alternative approach to validation of the SPR prediction from CT imaging, as well as its effect on the dose distribution and thus downstream clinical outcomes. According to this simulation study, software solutions that utilize spectral information hold promise for more accurate prediction of the SPR than the conventional stoichiometric approach.
OBJECTIVE:Image quality evaluation in radiology is most relevant when reflects radiologists' performance. This study assessed how image quality measurement in terms of in vivo -characterized detectability index ( ) for low-contrast liver lesion assessment in CT is correlated with radiologists' performance across 2 different CT reconstructions. METHODS:Fifty-one contrast-enhanced abdominal studies for investigating colorectal liver metastases were prospectively performed using 2 radiation dose exposures and reconstructed with Filtered back projection (FBP) and deep learning image reconstruction (DL) algorithms for a total of 161 noncalcified hypoattenuating lesions for 3 lesion size (D) subsets (<6 mm, 6 to 10 mm, and >10 mm). Images were assessed by expert radiologists for hepatic lesion detection task and likelihood of malignancy across the 2 imaging conditions. All cases were also evaluated automatically in terms of in vivo as a metric of task-based performance, both using a conventional technique and a new formalism of an added frequency term in the internal noise component of to accommodate the nonlinearity of the DL reconstruction ( adj ). RESULTS:The study found conventionally defined d' well-reflective of radiologists' evaluation of FBP images but not well-aligned with that of DL images. The new formalism provided more consistent reflection of performance across reconstruction techniques. In particular, in the lesion group D <=6 mm, the difference between radiologists' accuracy in images acquired with DL and images acquired with FBP was -26%, and the related adj difference was -9%, whereas the was 34%. Analogously, for the lesion group 6 mm < D <=10 mm, the differences were -15%, -13%, and 29%, respectively. Lastly, for the lesion group D>10 mm, radiologists showed the same accuracy in both FPB and DL images, difference in adj was -11%, and difference in was 31%. CONCLUSION:The new formalism can robustly reflect CT systems clinical performance irrespective of reconstruction algorithm. The methodology can be more readily applied to assess the real-world performance of CT systems.
Virtual imaging trials rely on complex, multi-stage computational workflows spanning a large array of software systems for patient modeling, image simulation, and analysis. Integrating heterogeneous software modules into efficient, reproducible, and scalable workflows remains challenging: ad hoc conventions, manual handoffs between developers, and inconsistent file formats often lead to fragile, error-prone results. To address these challenges, we developed GEMSTONe, a Generic Engine for Modular Software Task Orchestration and Negotiation, which decouples domain-specific module computation from a centralized workflow manager through declarative configuration, automated parameter validation, standardized module interfaces, and centralized path management. We deployed three GEMSTONe example workflows: a refactor of a recently studied virtual imaging trial (i.e., Virtual Lung Screening Trial), a high-throughput CT simulation scheduler, and an interoperable workflow integrating two breast imaging pipelines (FDA VICTRE and the University of Pennsylvania OpenVCT). Key findings included a 3.5 × speedup through engine-level mechanisms alone and the feasibility of refactoring, integrating, validating, and combining VICTRE and OpenVCT in less than five weeks. The generic nature of GEMSTONe makes it applicable beyond medical imaging, establishing it as a general framework for complex, multi-stage computational workflows.
Spectral CT can acquire signal at multiple x-ray energy levels. This enables material quantification by exploiting differences in x-ray attenuation across energy levels, particularly for k-edge materials. This simulation study quantified the signal and separability of current and potential clinical contrast agents across a range of materials and energies. A validated CT simulation platform was used to simulate a clinical photon-counting CT scanner with two energy thresholds. A cylindrical phantom containing common biological materials, clinical contrast agents, candidate contrast agents and nanoparticles, and investigational materials was imaged with varying upper energy thresholds (50–90 keV). At each energy level, images were assessed for noise, each material was assessed for contrast, and each material pair was evaluated for separability. Material contrasts reached peak value at the closest threshold higher than their respective k-edge. The energy threshold that produced the highest separability for each pair was characterized. Selection of energy threshold was dependent on the materials of interest. Threshold values at or just above a material’s k-edge maximized material signal while separability was maximized by the threshold that best separated k-edge signals.
Objective: To reduce variability in chest radiography (CXR) from acquisition and post-processing, and assess whether harmonization improves image quality and down-stream diagnostic performance. Methods: A generative adversarial network (GAN) was trained exclusively on virtual images produced by a Virtual Imaging Trial (VIT). The model maps non-harmonized CXRs to noise-free, unprocessed references using a dual U-Net with contrastive loss. Evaluation spanned virtual, physical-phantom, and clinical data using generic (GIQMs) and chest-specific (CIQMs) metrics. A phantom benchmark compared the method with ComBat harmonization and a conventional denoising algorithm. Clinical generalization was tested on VinDr-CXRs, examining feature compactness via Uniform Manifold Approximation and Projection (UMAP). We also assessed NIH-CXRs, quantifying diagnostic accuracy with bootstrap uncertainty. Results: Compared with non-harmonized images, harmonized CXRs yielded approximately 89% lower NRMSE, 50% higher PSNR, and 86% higher SSIM. On phantom data, CIQM variability fell from 0.255 to 0.026 and was reduced more consistently than with ComBat or denoising. Clinical analyses on VinDr-CXRs showed tighter UMAP clusters for harmonized features, indicating suppression of acquisition-related variability. On NIH-CXRs, training and testing a classifier in the harmonized domain improved diagnostic accuracy over the non-harmonized domain and lowered cross-domain sensitivity; pleural and cardiac categories showed consistent gains, while texture-dependent labels exhibited task-dependent effects. Conclusion: VIT-guided training enables a GAN to harmonize CXRs toward a raw reference, improving quantification variability and harmonizing image representations across systems. Significance: The proposed virtual-to-clinical strategy is scalable and generalizable, offering a practical path to standardized CXR appearance and reliable downstream detection across institutions.
PURPOSE:Free-space phase contrast propagation coupled with a photon-counting detector enables CT imaging with improved contrast in soft tissues at lower radiation dose. In addition, photon-counting detectors enable inherent spectral separation that can be used to capture tissue contrast at different energy levels. The objective of this study was to (i) develop a novel split-beam method for spectral synchrotron-based imaging considering limitations of photon-counting technology and clinical requirements, and (ii) propose a redefined mathematical model to calculate contrast-to-noise ratio in spectral imaging applications. METHODS:Our novel approach was applied in a CT imaging setup using a custom-made breast phantom with tissue-equivalent inserts and compared to the more common setup utilizing monochromatic beams. To complement the traditional contrast-to-noise ratio metric, a new mathematical framework for spectral contrast-to-noise ratio was introduced as a composite metric that integrates the signal-to-noise performance across spectral channels. RESULTS:The results show that the split-beam method proposed in this study obtains a comparable spectral contrast-to-noise ratio at the same radiation dose. Relative spectral contrast-to-noise differences were 0.12 (polyethylene), 1.92 (polyamide), 1.19 (polymethylmethacrylate), -0.29 (polyoxymethylene), and -0.17 (polytetrafluoroethylene) when comparing spectral imaging with two monochromatic beams at energies of 24 keV and 38 keV against the split-beam method. CONCLUSION:The potential advantages of the split-beam method for spectral CT imaging are numerous - it avoids non-rigid deformations, is fast to implement, and enables optimization in the post-processing step. The model for contrast-to-noise ratio redefined in this study applies to new generation spectral CT scanners beyond synchrotron setups.
Purpose:Accurate airway measurement is critical for bronchitis quantification with computed tomography (CT), yet optimal protocols and the added value of photon-counting CT (PCCT) over energy-integrating CT (EICT) for reducing bias remain unclear. We quantified biomarker accuracy across modalities and protocols and assessed strategies to reduce bias. Approach:A virtual imaging trial with 20 bronchitis anthropomorphic models was scanned using a validated simulator for two systems (EICT: SOMATOM Flash; PCCT: NAEOTOM Alpha) at 6.3 and 12.6 mGy. Reconstructions varied algorithm, kernel sharpness, slice thickness, and pixel size. Pi10 (square-root wall thickness at 10-mm perimeter) and WA% (wall-area percentage) were compared against ground-truth airway dimensions obtained from the 0.1-mm-precision anatomical models prior to CT simulation. External validation used clinical PCCT ( n = 22 ) and EICT ( n = 80 ). Results:Simulated airway dimensions agreed with pathological references ( R = 0.89 - 0.93 ). PCCT had lower errors than EICT across segmented generations ( p < 0.05 ). Under optimal parameters, PCCT improved Pi10 and WA% accuracy by 26.3% and 64.9%. Across the tested PCCT and EICT imaging protocols, improvements were associated with sharper kernels (25.8% Pi10, 33.0% WA%), thinner slices (23.9% Pi10, 49.8% WA%), smaller pixels (17.0% Pi10, 23.1% WA%), and higher dose ( ≤ 3.9 % ). Clinically, PCCT achieved higher maximum airway generation ( 8.8 ± 0.5 versus 6.0 ± 1.1 ) and lower variability, mirroring trends in virtual results. Conclusions:PCCT improves the accuracy and consistency of airway biomarker quantification relative to EICT, particularly with optimized protocols. The validated virtual platform enables modality-bias assessment and protocol optimization for accurate, reproducible bronchitis measurements.
Purpose:The credibility of artificial intelligence (AI) models for medical imaging continues to be a challenge, affected by the diversity of models, the data used to train the models, and the applicability of their combination to produce reproducible results for new data. We aimed to explore whether emerging virtual imaging trial (VIT) methodologies can provide an objective resource to approach this challenge. Approach:We conducted this study for the case example of COVID-19 diagnosis using clinical and virtual computed tomography (CT) and chest radiography (CXR) processed with convolutional neural networks. Multiple AI models were developed and tested using 3D ResNet-like and 2D EfficientNetv2 architectures across diverse datasets. Results:Model performance was evaluated using the area under the curve (AUC) and the DeLong method for AUC confidence intervals. The models trained on the most diverse datasets showed the highest external testing performance, with AUC values ranging from 0.73 to 0.76 for CT and 0.70 to 0.73 for CXR. Internal testing yielded higher AUC values (0.77 to 0.85 for CT and 0.77 to 1.0 for CXR), highlighting a substantial drop in performance during external validation, which underscores the importance of diverse and comprehensive training and testing data. Most notably, the VIT approach provided an objective assessment of the utility of diverse models and datasets while offering insight into the influence of dataset characteristics, patient factors, and imaging physics on AI efficacy. Conclusions:The VIT approach enhances model transparency and reliability, offering nuanced insights into the factors driving AI performance and bridging the gap between experimental and clinical settings.
RATIONALE AND OBJECTIVE:Noise magnitude is one of the image quality indicators in computed tomography (CT) for assessing imaging performance, and the CMS (Centers for Medical and Medicaid Services) recently included a measure of noise magnitude in terms of "global noise". Despite its importance, no standard method currently exists to measure noise magnitude in patient images. Theoretically, the most accurate approach is to assess voxel value variation across repeated images, the so-called ensemble noise. Obviously, such a method is not ethically feasible in actual patients. To surmount this impasse, we deployed virtual imaging techniques to benchmark three noise magnitude calculation methods against two gold standard ensemble noise measures across 36 imaging conditions. MATERIALS AND METHODS:Over 1800 virtual image datasets were generated from imaging an American College of Radiology (ACR) phantom and Extended Cardiac-Torso (XCAT) human models using a validated, scanner-specific CT simulator (DukeSim). The ACR phantom was imaged under 36 different imaging conditions defined by combinations of chest and abdominopelvic protocols, three dose levels, three reconstruction kernels, and both Filtered Back Projection and Iterative Reconstruction algorithms. Under the same conditions, XCAT models were repeatedly imaged 50 times. Noise magnitudes in the ACR phantom were calculated in 5 circular regions of interest. In patients, noise was measured in the air surrounding the body and in soft tissues by applying HU < -900 and -300 ≤ HU ≤ 100 thresholds, respectively. For each imaging condition, measured noise magnitudes were compared against the ensemble noise in soft tissue, liver, and lungs. RESULTS:Across all imaging conditions, noise measurements in the ACR phantom and in the air surrounding the patient underestimated ensemble noise by approximately 60% and 50%, respectively. In contrast, soft tissue-based noise measurements were closer to the gold standard with median differences between -8% and +4%. CONCLUSION:This study introduced a virtual imaging-based framework to benchmark clinical CT noise metrics against ensemble noise measures. Virtual imaging enabled objective comparison of different noise magnitude calculation methods in large, realistic populations simulating clinical conditions. Noise measured in the air cannot represent soft tissue noise. The results validated soft tissue-based noise measurements as a reliable surrogate to inform protocol design, technology assessment, and equitable healthcare reimbursement.
With preparations underway for extended-duration crewed deep space missions, the health risks of solar particle events (SPEs) to astronauts are becoming increasingly pertinent. To address this hazard, the AstroRad vest, a personal radiation shielding garment providing targeted organ protection, was tested during Artemis I. Two anthropomorphic female phantoms, equipped with internal and external passive and active dosimeters, were flown aboard the Orion spacecraft: one unshielded and the other wearing AstroRad. Inner Van Allen belt transit active dosimeter measurements were extrapolated to simulate SPE scenarios, predicting effective dose reductions of ∼60% for an August 1972-like SPE and nearly 40% for an October 1989-like SPE, varying slightly with anatomical model. Such reductions could spare astronauts the equivalent of up to 193 and 131 days of deep space radiation exposure, respectively. These findings demonstrate that wearable shielding such as AstroRad could serve as a vital element for safe and sustainable human deep space exploration.
Congenital Heart Disease (CHD) is the most common birth defect worldwide, affecting millions of newborns yearly. Pulmonary artery stenosis is a narrowing around the pulmonary valve that restricts blood flow to the lungs and can lead to heart failure if untreated. Current treatment uses stents to support blood flow, but this lacks long-term effectiveness for children and newborns. The purpose of our work is to leverage a comprehensive database of pediatric phantoms based on biomechanical simulation to simulate stent placement preoperatively. We segmented patient anatomies using TotalSegmentator with additional refinements, then reconstructed them using the RFTA algorithm. This approach reduced geometric error by 73% on average compared to Marching Cubes (measured by Chamfer distance and Hausdorff distance) while maintaining fine details. For non-rigid deformation, we used the SyN algorithm in ANTs, which achieved high normalized cross-correlation values of 91.89% to 95.09% as well as a mutual information score of 45.93% to 79.13%, showcasing strong alignment between the warped template and patient images. Our resulting library of patient-specific phantoms enables preoperative stent placement simulation. This framework supports personalized planning for pediatric CHD interventions, aiming to improve stent performance and reduce restenosis.
Purpose To develop a physics-based image harmonization method that transforms images into a reference quality index of noise, spatial resolution, and lung volume and to evaluate its performance for improving reproducibility of lung density measurement. Materials and Methods This retrospective analysis of Genetic Epidemiology of Chronic Obstructive Pulmonary Disease (COPDGene) study data included participants who underwent chest CT with full-dose (200 mAs) and reduced-dose (40-80 mAs) protocols during the same visit between November 2014 and July 2017. A harmonization algorithm was developed to reduce variations in lung density by adjusting spatial resolution, noise, and lung volume in sequence, aligning scans to a reference condition based on full-dose, soft-kernel reconstruction. The percentage of lung voxels less than -950 HU (LAA-950) and the 15th percentile of the lung density histogram (Perc15) were calculated from scans. Improvements in bias, limits of agreement, and reproducibility coefficients were evaluated, and harmonizer performance was compared with that of two existing techniques: volume-adjusted lung density (VALD) and median filtering followed by VALD (MF-VALD). Results A total of 1159 participants (mean age ± SD, 65 years ± 9; 586 male) were studied. Across all imaging conditions, the harmonization technique improved the Perc15 reproducibility coefficient 4.8-fold, from 35.6 HU ± 0.7 to 7.4 HU ± 0.2. The harmonizer outperformed VALD and MF-VALD across the full-dose and reduced-dose scans, with reproducibility coefficients improving from 32.4 HU ± 0.7 before harmonization to 30.1 HU ± 0.3 with VALD, 10.3 HU ± 0.3 with MF-VALD, and 7.7 HU ± 0.2 with the harmonizer. Conclusion This physics-based technique harmonized CT images to a reference quality index, improving the reproducibility of lung density metrics. Keywords: CT, Thorax, Lung, Physics, Chronic Obstructive Pulmonary Disease Clinical trial registration no. NCT00608764 Supplemental material is available for this article. © RSNA, 2026.
Objective. The performance of photon-counting CT (PCCT) depends on how detector energy thresholds or bins (collectively referred to as energy settings) are defined. This study aimed to identify optimal energy settings for various PCCT technologies to enhance spectral separability and spatial detectability of materials relevant to diagnostic imaging. Approach. We modeled vendor-neutral PCCT scanners with cadmium telluride (CdTe), cadmium zinc telluride (CZT), and silicon (Si) based detectors using an in-silico imaging framework. The scanner models incorporated scanner geometries and spatio-energetic detector responses to account for inter-pixel and inter-energy crosstalk and noise correlation. Using these models, we imaged two sizes of a cylindrical phantom containing inserts with various concentrations of calcium, iodine, and gadolinium across various energy settings and two dose levels. Other settings were held constant, including tube voltage (120 kV), pitch (1), gantry rotation speed (0.5 rot s-1), and reconstruction settings. Image quality was quantified in terms of separability index (s ') and contrast-to-noise ratio. The scanner-specific optimal energy settings were determined by ranking the measured s ' values using a rank-sum method and evaluating them with the Friedman test. Main results. The optimal energy settings varied primarily with the phantom size and material pairs to separate (Ca-I, Ca-Gd, I-Gd) rather than the dose level. When aggregated across all imaging tasks and conditions, the optimal energy settings were 30-65 keV and 20-35-50-70 keV for two- and four-threshold CdTe-, 20-35-50-70 keV for four-threshold CZT-, and 5-35-50-80-120 keV for four-bin Si-based PCCT systems. Optimization of energy settings significantly improved material decomposition accuracy for Ca, I, and Gd across PCCT systems, particularly under low-concentration conditions, reducing percent differences from as high as 422.0% to within 17.0%. Significance. This work presents a simulation-based framework for optimizing energy settings of photon-counting detectors used in PCCT scanners, providing clinically translatable insights towards clinical adoption for accurate and standardized spectral imaging.
Objective.Lung nodule appearance may provide prognostic information, as the presence of spiculation increases the suspicion of a nodule being cancerous. Spiculations can be quantified using morphological radiomics features extracted from CT images. Radiomics features can be affected by the acquisition parameters and scanner technologies; thus, it is essential to identify imaging conditions that provide reliable measurements, particularly for emerging technologies like photon-counting CT (PCCT). This study aimed to systematically quantify the effect of imaging parameters on the radiomics measurements using a virtual imaging trial (VIT) platform, and further verify the findings with human clinical data.Approach.The VIT utilized nine virtual patients, each with three 6 mm nodules of varying spiculations. The virtual patients were run through a validated CT simulator (DukeSim) to acquire images at three dose levels (CTDIvol = 2.85, 5.69, and 11.38 mGy) with a clinical energy-integrating CT and a PCCT. The acquired projection images were reconstructed using multiple slice thicknesses, kernels, and matrix sizes. The reconstructed images were processed to extract morphological features using three segmentation methods. The features were clustered into three broad type categories. Features extracted from the acquired CT images were compared to their corresponding ground truth values, across all imaging conditions.Main results.Among all imaging conditions, slice thickness had the greatest effect on the radiomics measurements. When the thickest slices were used, the coefficient of variation increased by [1.19%-9.66%] in the energy integrating CT images, and [3.94%-24.43%] in the PCCT images. For both scanners, varying the kernel sharpness and dose affected the radiomics measurements insignificantly, while pixel size and segmentation method were observed to have stronger effects. Under varying imaging conditions, the trends and magnitude of radiomics features measurements were coherent with virtual trial results.Significance.The findings stress the importance of choosing optimal reconstruction settings for radiomics extraction to achieve precise feature quantifications.
Protocol optimization is critical in Computed Tomography (CT) for achieving desired diagnostic image quality while minimizing radiation dose. Due to the inter-effect of influencing CT parameters, traditional optimization methods rely on the testing of exhaustive combinations of these parameters. This poses a notable limitation due to the impracticality of exhaustive parameter testing. This study introduces a novel methodology leveraging Virtual Imaging Trials (VITs) and reinforcement learning to more efficiently optimize CT protocols. Computational phantoms with liver lesions were imaged using a validated CT simulator and reconstructed with a novel CT reconstruction Toolkit. The optimization parameter space included tube voltage, tube current, reconstruction kernel, slice thickness, and pixel size. The optimization process was done using a Proximal Policy Optimization (PPO) agent which was trained to maximize the Detectability Index (d') of the liver lesion for each reconstructed image. Results showed that our reinforcement learning approach found the absolute maximum d' across the test cases while requiring 79.7% fewer steps compared to an exhaustive search, demonstrating both accuracy and computational efficiency, offering a efficient and robust framework for CT protocol optimization. The flexibility of the proposed technique allows for use of varying image quality metrics as the objective metric to maximize for. Our findings highlight the advantages of combining VIT and reinforcement learning for CT protocol management.