Purpose:There are multiple commercially available, Food and Drug Administration (FDA)-cleared, artificial intelligence (AI)-based tools automating stroke evaluation in noncontrast computed tomography (NCCT). This study assessed the impact of variations in reconstruction kernel and slice thickness on two outputs of such a system: hypodense volume and Alberta Stroke Program Early CT Score (ASPECTS). Approach:The NCCT series image data of 67 patients imaged with a CT stroke protocol were reconstructed with four kernels (H10s-smooth, H40s-medium, H60s-sharp, and H70h-very sharp) and three slice thicknesses (1.5, 3.0, and 5.0 mm) to create 1 reference condition (H40s/5.0 mm) and 11 nonreference conditions. The 12 reconstructions per patient were processed with a commercially available FDA-cleared software package that yields total hypodense volume (mL) and ASPECTS. A mixed-effect model was used to test the difference in hypodense volume, and an ordered logistic model was used to test the difference in e-ASPECTS. Results:Hypodense volume differences from the reference condition ranged from - 14.6 to 1.1 mL and were significant for all nonreference kernels (H10s p = 0.025 , H60s p < 0.001 , and H70h p < 0.001 ) and for thinner slices (1.5 mm p < 0.001 and 3.0 mm p = 0.002 ). e-ASPECTS was invariant to the nonreference kernels and slice thicknesses, with a mean difference ranging from - 0.1 to 0.5. No significant differences were found for any kernel or slice thickness (all p > 0.05 ). Conclusions:Automated hypodense volume measured with a commercially available, FDA-cleared software package is substantially impacted by reconstruction kernel and slice thickness. Conversely, automated ASPECTS is invariant to these reconstruction parameters.
Computed tomography (CT) images are typically reconstructed along the scanner's z-axis, resulting in axial slices. While sufficient for most diagnostic tasks, axial slices are sub-optimal when the target anatomy is not aligned with the scanner axis, such as the vasculature. Currently, multi-planar and curved-planar reformation are used to resample axial volumes into oblique or curved slices, respectively. However, as a post-reconstruction interpolative process, reformation is limited by the sampling frequency of the underlying axial images, which can degrade the spatial resolution of the resulting slices. To mitigate this, we propose space curve reconstruction (SCR), a framework that enables the native reconstruction of planar slices along arbitrary curves in 3D space. Unlike reformation, SCR backprojects attenuation values directly from the sinogram, obviating the need for an intermediate axial volume. To evaluate SCR, we selected three representative anatomical structures: the lumbosacral spine, coronary vessels, and pancreatic duct. For each structure, equivalent reformation and SCR images were generated for visual inspection. Quantitative validation was performed by calculating the full width at half maximum (FWHM) from an oblique pixel intensity profile of a coronary calcium granule. In all cases, SCR produced qualitatively sharper images compared to reformation, particularly for off-axis structures. FWHM measurements confirmed these visual findings, showing a 14% reduction in FWHM compared to reformation for the analyzed calcium granule. This study suggests that SCR can enhance the visualization of complex anatomy and lays the groundwork for future rigorous investigations of non-standard projection sampling patterns.
To establish a public repository of computed tomography (CT) texture phantom images paired with objective 3D image quality measurements from multiple scanner models using a diverse range of imaging protocols, facilitating investigations into relationships between image quality and quantitative imaging features. Three specialized CT phantoms were scanned: (1) the Corgi® phantom for image quality assessment, (2) a radiomics liver phantom, and (3) an open-source 3D-printed texture phantom. Image quality assessment included measurement of contrast-to-noise ratio, 3D modulation transfer function, and 3D noise power spectrum. Data were acquired on four CT scanner models from two manufacturers at five CTDI v o l levels (2.05-17.11 mGy) and reconstructed with eight different kernels, yielding 160 total conditions (combinations of scanner, dose, kernel). Data were validated for integrity and completeness, resulting in the exclusion of six conditions. The paired texture phantom and image quality (PTP-IQ) dataset includes: (1) DICOM image series of two texture phantoms acquired across the 154 conditions, as well as (2) image quality metrics derived from each corresponding set of scanner, acquisition and reconstruction settings provided in HDF5 format. This dataset enables the development and validation of harmonization methods for multi-center quantitative imaging studies, investigation of protocol-dependent QIF variability, and optimization of acquisition protocols for radiomics applications. The controlled and systematic study design facilitates isolation of individual protocol effects on quantitative measurements.
BACKGROUND:The lack of standardized, reproducible, and publicly available reference objects limits the interpretability and reproducibility of radiomic imaging studies across institutions. Current handcrafted phantoms lack the precision, reproducibility, and accessibility required for large-scale, multi-center validation studies. PURPOSE:This work presents the design, manufacture, and initial evaluation of an open-source modular CT texture phantom intended as a reproducible, accessible platform for quantitative imaging research. METHODS:A modular phantom system was designed comprising an epoxy-filled cylindrical body and eight cavities for interchangeable 3D-printed texture inserts. The inserts utilized mathematically defined geometric primitives (e.g., periodic rod lattices, and random shape distributions) to provide standardized and reproducible printable objects. The design files for the phantom body and inserts were released in the public domain. Initial evaluation of the phantom body and inserts was done by manufacturing three copies of the phantom body and four sets of the inserts and then performing four assessments: (a) phantom body uniformity quantified through intra- and inter-phantom HU evaluations; (b) texture insert reproducibility based on an evaluation of the intraclass correlation coefficient (ICC) and coefficient of variation (CV) of radiomic feature values across insert replicates; (c) evaluating that the set of inserts provide non-redundant texture stimuli by performing a principal component analysis (PCA) on radiomic features extracted from each insert; and (d) evaluating the sensitivity of inserts to CT imaging protocol variation across two dose levels and three reconstruction kernels. RESULTS:The manufacturing pipeline demonstrated high intra- and inter-phantom uniformity across all four phantom bodies (<5 HU). The texture inserts demonstrated exceptional reproducibility of radiomic features across independent manufacturing runs with a ICC > 0.999 both across insert set prints and between slices. The PCA demonstrated that the set of inserts provided distinct texture stimuli. The protocol variation investigation demonstrated that the inserts showed anticipated variation in texture-based features across CT dose levels and reconstruction kernels. CONCLUSIONS:We present an open-source phantom design and manufacturing approach capable of producing high-fidelity, reproducible reference objects for radiomic imaging studies using consumer-grade equipment and low-cost materials. The public domain release of design files removes licensing barriers and enables any research group to replicate, extend, and share phantom designs, supporting standardized phantom-based investigation of radiomic feature behavior across institutions.
Harmonization is a critical step in mitigating variability in computed tomography (CT) scans caused by differences in CT parameters such as radiation dose, reconstruction kernel, and slice thickness, among others. Currently, several deep learning-based models exist that focus on minimizing voxel-based distances between pairs of scans rather than being optimized for specific downstream tasks. In this work, we chose emphysema scoring as our downstream task and demonstrated that incorporating task-specific information while training a harmonization model results in more consistent performance across different image reconstructions than when such information is not used. We generated 9 different variations in CT scans from raw sinogram data. Using a reference condition, we trained two models to map non-reference condition scans to the reference: (1) a generative adversarial network with spectral normalization (SNGAN) and (2) SNGAN with a masking module that generates a weighted density mask (SNGAN-WDM), emphasizing low attenuation regions in the scan. We assessed the performance difference between the models using voxel similarity metrics: peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). We also computed emphysema scores for the scans and assessed mean absolute deviation (MAD) in scores from the reference for all conditions before and after harmonization. The mean +/- standard deviation across conditions were 35.207 +/- 5.147 for PSNR, 0.921 +/- 0.049 for SSIM, and 0.062 +/- 0.029 for LPIPS with no harmonization; 40.972 +/- 4.664 for PSNR, 0.961 +/- 0.023 for SSIM, and 0.034 +/- 0.017 for LPIPS after SNGAN; and 41.891 +/- 2.657 for PSNR, 0.976 +/- 0.012 for SSIM, and 0.026 +/- 0.009 for LPIPS after SNGAN-WDM, respectively. The MAD in emphysema scores across conditions was 0.127 +/- 0.033 with no harmonization; 0.100 +/- 0.005 with SNGAN; and 0.068 +/- 0.003 with SNGAN-WDM.
Agentic Artificial Intelligence (AI) systems leveraging Large Language Models (LLMs) exhibit significant potential for complex reasoning, planning, and tool utilization. We demonstrate that a specialized computer vision system can be built autonomously from a natural language prompt using Agentic AI methods. This involved extending SimpleMind (SM), an open-source Cognitive AI environment with configurable tools for medical image analysis, with an LLM-based agent, implemented using OpenManus, to automate the planning (tool configuration) for a particular computer vision task. We provide a proof-of-concept demonstration that an agentic system can interpret a computer vision task prompt, plan a corresponding SimpleMind workflow by decomposing the task and configuring appropriate tools. From the user input prompt, "provide sm (SimpleMind) config for lungs, heart, and ribs segmentation for cxr (chest x-ray)"), the agent LLM was able to generate the plan (tool configuration file in YAML format), and execute SM-Learn (training) and SM-Think (inference) scripts autonomously. The computer vision agent automatically configured, trained, and tested itself on 50 chest x-ray images, achieving mean dice scores of 0.96, 0.82, 0.83, for lungs, heart, and ribs, respectively. This work shows the potential for autonomous planning and tool configuration that has traditionally been performed by a data scientist in the development of computer vision applications.
Consistent and reliable data in computed tomography (CT) is vital for developing generalizable AI algorithms for downstream tasks. However, differences in CT acquisition and reconstruction parameters lead to inconsistencies in image quality and characteristics, impacting AI model performance. We introduce CT-Norm, an open-source toolkit addressing CT variability through three primary modules: Data Characterization, Data Harmonization, and Robustness Analysis. Each module targets a distinct phase of the data understanding pipeline. The data characterization module identifies voxel-level and DICOM header-level variability for a given dataset and generates a summary report. The data harmonization module provides a solution to mitigate the identified variability. It offers a library of image-level harmonization methods, ranging from traditional image processing to convolutional neural networks (CNNs) and generative adversarial networks (GANs). This module is validated on an external test set from the TCIA low-dose CT (LDCT) dataset using peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS). Without the harmonization module, the observed PSNR, SSIM, and LPIPS were 19.77 +/- 0.368, 0.33 +/- 0.010, and 0.43 +/- 0.012, respectively, which increased to 26.51 +/- 0.281 and 0.57 +/- 0.011 for PSNR and SSIM, while LPIPS decreased to 0.30 +/- 0.004 after applying CNN-based harmonization. The robustness analysis module allows developers to evaluate the sensitivity of an existing deep-learning model across the identified variability within a dataset. This allows developers to identify the imaging conditions under which a model performs optimally and those where it may struggle, enabling them to optimize the model further and enhance its reliability. CT-Norm enables model developers and end users to understand data variability, harmonize inconsistencies, and optimize AI models for robust and generalizable performance.
Objective. The study aims to systematically characterize the effect of CT parameter variations on images and lung radiomic and deep features, and to evaluate the ability of different image harmonization methods to mitigate the observed variations. Approach. A retrospective in-house sinogram dataset of 100 low-dose chest CT scans was reconstructed by varying radiation dose (100%, 25%, 10%) and reconstruction kernels (smooth, medium, sharp). A set of image processing, convolutional neural network (CNNs), and generative adversarial network-based (GANs) methods were trained to harmonize all image conditions to a reference condition (100% dose, medium kernel). Harmonized scans were evaluated for image similarity using peak signal-to-noise ratio (PSNR), structural similarity index measure (SSIM), and learned perceptual image patch similarity (LPIPS), and for the reproducibility of radiomic and deep features using concordance correlation coefficient (CCC). Main Results. CNNs consistently yielded higher image similarity metrics amongst others; for Sharp/10%, which exhibited the poorest visual similarity, PSNR increased from a mean +/- CI of 17.763 +/- 0.492 to 31.925 +/- 0.571, SSIM from 0.219 +/- 0.009 to 0.754 +/- 0.017, and LPIPS decreased from 0.490 +/- 0.005 to 0.275 +/- 0.016. Texture-based radiomic features exhibited a greater degree of variability across conditions, i.e. a CCC of 0.500 +/- 0.332, compared to intensity-based features (0.972 +/- 0.045). GANs achieved the highest CCC (0.969 +/- 0.009 for radiomic and 0.841 +/- 0.070 for deep features) amongst others. CNNs are suitable if downstream applications necessitate visual interpretation of images, whereas GANs are better alternatives for generating reproducible quantitative image features needed for machine learning applications. Significance. Understanding the efficacy of harmonization in addressing multi-parameter variability is crucial for optimizing diagnostic accuracy and a critical step toward building generalizable models suitable for clinical use.
CT is crucial for diagnosing chest diseases, with image quality affected by spatial resolution. Thick-slice CT remains prevalent in practice due to cost considerations, yet its coarse spatial resolution may hinder accurate diagnoses. Our multicenter study develops a deep learning synthetic model with Convolutional-Transformer hybrid encoder-decoder architecture for generating thin-slice CT from thick-slice CT on a single center (1576 participants) and access the synthetic CT on three cross-regional centers (1228 participants). The qualitative image quality of synthetic and real thin-slice CT is comparable (p = 0.16). Four radiologists’ accuracy in diagnosing community-acquired pneumonia using synthetic thin-slice CT surpasses thick-slice CT (p < 0.05), and matches real thin-slice CT (p > 0.99). For lung nodule detection, sensitivity with thin-slice CT outperforms thick-slice CT (p < 0.001) and comparable to real thin-slice CT (p > 0.05). These findings indicate the potential of our model to generate high-quality synthetic thin-slice CT as a practical alternative when real thin-slice CT is preferred but unavailable.
To support access to CT projection data, DICOM-CT-PD was previously proposed, which encodes each projection as a separate DICOM file. However, this poses challenges with modern projection datasets, which often exceed 10,000 projections. We propose SF-CT-PD, a single-file derivative of the DICOM-CT-PD file format for CT projection data that stores projections within a single DICOM file, stores pixel data detector-row-major, and stores projection-specific parameters as ordered tables within the DICOM header. We compared the performance of SF-CT-PD against DICOM-CT-PD in read speed, file size, and network transfer. Cases were sampled from TCIA's "LDCT-and-Projection-data" dataset and encoded into DICOM-CT-PD and SF-CT-PD representations. Read tests were conducted for four programming languages on hard disk and solid-state drives. rsync-based network transfer analysis measured the Ethernet throughput for each format. The accuracy of the implementation was confirmed by analyzing offline reconstructions and transfer file checksums for each format. SF-CT-PD was generally more performant in read operations and file size. The network throughput was equivalent between the formats, with file-checksums indicating file integrity. Reconstruction accuracy was supported by difference image agreements. SF-CT-PD represents a viable extension of DICOM-CT-PD where a single file is preferred.
This study is an initial investigation into methods to harmonize quantitative imaging (QI) feature values across CT scanners based on image quality metrics. To assess the impact of harmonization on QI features, we: (1) scanned an image quality assessment phantom on three scanners over a wide range of acquisition and reconstruction conditions; (2) from those scans, assessed image quality for each scanner at each acquisition and reconstruction condition; (3) from these assessments, identified a set of parameters for each scanner that yielded similar image quality values ("harmonized condition"); (4) scanned a second phantom with texture (i.e., local variations in attenuation) under the same set of conditions; and (5) extracted QI features and compared values between non-harmonized and harmonized image quality conditions. Quantitative image quality assessments provided contrast to noise ratio (CNR) and modulation transfer function frequency at 50% (MTF f(50)) values for each scanner and each condition used. A set of harmonized conditions was identified across three CT scanners based on the similarity of CNR and MTF f(50). To provide a comparison, several non-harmonized condition sets were identified. From the texture phantom, the standard deviation of the QI feature values (intensity mean and variance, GLCM autocorrelation and cluster tendency, GLDM high and low gray level emphasis) across the three CT systems decreased between 72.8% and 81.1% between the unharmonized and harmonized groups (with exception of intensity mean which showed little difference across scanners). These initial results suggest that selecting protocols that produce similar quantitative image quality metric values across different CT systems can reduce the variance of QI feature values across those systems.
Purpose:To rule out hemorrhage, non-contrast CT (NCCT) scans are used for early evaluation of patients with suspected stroke. Recently, artificial intelligence tools have been developed to assist with determining eligibility for reperfusion therapies by automating measurement of the Alberta Stroke Program Early CT Score (ASPECTS), a 10-point scale with > 7 or ≤ 7 being a threshold for change in functional outcome prediction and higher chance of symptomatic hemorrhage, and hypodense volume. The purpose of this work was to investigate the effects of CT reconstruction kernel and slice thickness on ASPECTS and hypodense volume. Methods:The NCCT series image data of 87 patients imaged with a CT stroke protocol at our institution were reconstructed with 3 kernels (H10s-smooth, H40s-medium, H70h-sharp) and 2 slice thicknesses (1.5mm and 5mm) to create a reference condition (H40s/5mm) and 5 non-reference conditions. Each reconstruction for each patient was analyzed with the Brainomix e-Stroke software (Brainomix, Oxford, England) which yields an ASPECTS value and measure of total hypodense volume (mL). Results:An ASPECTS value was returned for 74 of 87 cases in the reference condition (13 failures). ASPECTS in non-reference conditions changed from that measured in the reference condition for 59 cases, 7 of which changed above or below the clinical threshold of 7 for 3 non-reference conditions. ANOVA tests were performed to compare the differences in protocols, Dunnett's post-hoc tests were performed after ANOVA, and a significance level of p < 0.05 was defined. There was no significant effect of kernel (p = 0.91), a significant effect of slice thickness (p < 0.01) and no significant interaction between these factors (p = 0.91). Post-hoc tests indicated no significant difference between ASPECTS estimated in the reference and any non-reference conditions. There was a significant effect of kernel (p < 0.01) and slice thickness (p < 0.01) on hypodense volume, however there was no significant interaction between these factors (p = 0.79). Post-hoc tests indicated significantly different hypodense volume measurements for H10s/1.5mm (p = 0.03), H40s/1.5mm (p < 0.01), H70h/5mm (p < 0.01). No significant difference was found in hypodense volume measured in the H10s/5mm condition (p = 0.96). Conclusion:Automated ASPECTS and hypodense volume measurements can be significantly impacted by reconstruction kernel and slice thickness.
We introduce a simple physics-based model of RA-950 emphysema scoring. Our model assumes that the lung is strictly composed of healthy tissue and emphysematous tissue, each described by a single attenuation value and contaminated with Gaussian noise. We show that when combined with curve-fitting, the model can accurately capture change in RA-950 score with respect to image noise and subject breath hold, and accurately compute “true” RA-950 scores (relative to a clinical reference scan) in a cohort of 16 patients. To validate the model, noise realizations of 10 lung screening subjects and 6 COPD patients were created using various combinations of reconstruction parameters and simulated reduced dose acquisitions. Least-squares curve fitting software was utilized to determine the amount of emphysema and the attenuation value of healthy lung tissue for each subject using the model. The derived model provided accurate emphysema scores (difference between model value and clinical reference of < 0.02) in all cases except one. Upon radiologist review of this case, the score derived from our model was deemed more appropriate than RA-950 from the clinical reference scan. The R-squared values were < 0.9 in all cases except one, and < 0.95 in 12 of 16 cases. The case with low R2 value was also reviewed by a radiologist and found to have substantial other disease that violated key model assumptions. The model appears to be robust to breath hold, image noise, and amount of emphysema present, factors that have been found to confound other approaches (such as denoising) to improving emphysema scoring.
Concerns over the risks of radiation dose from diagnostic CT motivated the utilization of low dose CT (LdCT). However, due to the extremely low X-ray photon statistics in LdCT, the reconstruction problem is ill-posed and noisecontaminated. Conventional Compressed Sensing (CS) methods have been investigated to enhance the signal-to-noise ratio of LdCT at the cost of image resolution and low contrast object visibility. In this work, we adapted a flexible, iterative reconstruction framework, termed Plug-and-Play (PnP) alternating direction method of multipliers (ADMM), that incorporated state-of-the-art denoising algorithms into model-based image reconstruction. The PnP ADMM framework is achieved by combining a least square data fidelity term with a regularization term for image smoothness and was solved through the ADMM. An off-the-shelf image denoiser, the Block-Matching 3D-transform shrinkage (BM3D) filter, is plugged in to substitute an ADMM module. The PnP ADMM was evaluated on low dose scans of ACR 464 phantom and two lung screening data sets and is compared with the Filtered Back Projection (FBP), the Total Variation (TV), the BM3D post-processing method, and the BM3D regularization method. The proposed framework distinguished the line pairs at 9 lp/cm resolution on the ACR phantom and the fissure line in the left lung, resolving the same or better image details than FBP reconstruction of higher dose scans with up to 18 times less dose. Compared with conventional iterative reconstruction methods resulting in comparable image noise, the proposed method is significantly better at recovering image details and improving low contrast conspicuity.
Screen failure rates in Alzheimer's disease (AD) clinical trial research are unsustainable, with participant recruitment being a top barrier to AD research progress. The purpose of this project was to understand the neuropsychological, psychiatric, and functional features of individuals who failed screening measures for AD trials. Previously collected clinical data from 38 patients (aged 50-83) screened for a specific industry-sponsored clinical trial of MCI/early AD (Biogen 221AD302, [EMERGE]) were analyzed to identify predictors of AD trial screen pass/fail status. Worse performance on non-memory cognitive domains like crystalized knowledge, executive functioning, and attention, and higher self-reported anxiety, was associated with failing the screening visit for the EMERGE AD clinical trial, whereas we were not able to detect a relationship between screening status and memory performance, self-reported depression, or self-reported daily functioning. By identifying predictors of AD trial screen passing/failure, this research may influence decision-making about which patients are most likely to successfully enroll in a trial, thereby potentially lowering participant burden, maximizing study resources, and reducing costs.
Iterative coordinate descent (ICD) is an optimization strategy for iterative reconstruction that is sometimes considered incompatible with parallel compute architectures such as graphics processing units (GPUs). We present a series of modifications that render ICD compatible with GPUs and demonstrate the code on a diagnostic, helical CT dataset. Our reference code is an open-source package, FreeCT ICD, which requires several hours for convergence. Three modifications are used. First, as with our reference code FreeCT ICD, the reconstruction is performed on a rotating coordinate grid, enabling the use of a stored system matrix. Second, every other voxel in the z-is updated direction simultaneously, and the sinogram data is shuffled to coalesce memory access. This increases the parallelism available to the GPU. Third, NS voxels in the xy-plane are updated simultaneously. This introduces possible crosstalk between updated voxels, but because the interaction between non-adjacent voxels is small, small values of NS still converge effectively. We find NS = 16 enables faster reconstruction via greater parallelism, and NS = 256 remains stable but has no additional computational benefit. When tested on a pediatric dataset of size 736x16x14000 reconstructed to a matrix size of 512x512x128 on a single GPU, our implementation of ICD can converge within 10 HU RMS in less than 5 minutes. This suggests that ICD could be competitive with simultaneous update algorithms on modern, parallel compute architectures.
PURPOSE:The computational burden associated with model-based iterative reconstruction (MBIR) is still a practical limitation. Iterative coordinate descent (ICD) is an optimization approach for MBIR that has sometimes been thought to be incompatible with modern computing architectures, especially graphics processing units (GPUs). The purpose of this work is to accelerate the previously released open-source FreeCT_ICD to include GPU acceleration and to demonstrate computational performance with ICD that is comparable with simultaneous update approaches.METHODS:FreeCT_ICD uses a stored system matrix (SSM), which precalculates the forward projector in the form of a sparse matrix and then reconstructs on a rotating coordinate grid to exploit helical symmetry. In our GPU ICD implementation, we shuffle the sinogram memory ordering such that data access in the sinogram coalesce into fewer transactions. We also update NS voxels in the xy-plane simultaneously to improve occupancy. Conventional ICD updates voxels sequentially (NS = 1). Using NS > 1 eliminates existing convergence guarantees. Convergence behavior in a clinical dataset was therefore studied empirically.RESULTS:On a pediatric dataset with sinogram size of 736 × 16 × 13860 reconstructed to a matrix size of 512 × 512 × 128, our code requires about 20 s per iteration on a single GPU compared to 2300 s per iteration for a 6-core CPU using FreeCT_ICD. After 400 iterations, the proposed and reference codes converge within 2 HU RMS difference (RMSD). Using a wFBP initialization, convergence within 10 HU RMSD is achieved within 4 min. Convergence is similar with NS values between 1 and 256, and NS = 16 was sufficient to achieve maximum performance. Divergence was not observed until NS > 1024.CONCLUSIONS:With appropriate modifications, ICD may be able to achieve computational performance competitive with simultaneous update algorithms currently used for MBIR.
PURPOSE With recent substantial improvements in modern computing, interest in quantitative imaging with CT has seen a dramatic increase. As a result, the need to both create and analyze large, high-quality datasets of clinical studies has increased as well. At present, no efficient, widely available method exists to accomplish this. The purpose of this technical note is to describe an open-source high-throughput computational pipeline framework for the reconstruction and analysis of diagnostic CT imaging data to conduct large-scale quantitative imaging studies and to accelerate and improve quantitative imaging research. METHODS The pipeline consists of two, primary "blocks": reconstruction and analysis. Reconstruction is carried out via a graphics processing unit (GPU) queuing framework developed specifically for the pipeline that allows a dataset to be reconstructed using a variety of different parameter configurations such as slice thickness, reconstruction kernel, and simulated acquisition dose. The analysis portion then automatically analyzes the output of the reconstruction using "modules" that can be combined in various ways to conduct different experiments. Acceleration of analysis is achieved using cluster processing. Efficiency and performance of the pipeline are demonstrated using an example 142 subject lung screening cohort reconstructed 36 different ways and analyzed using quantitative emphysema scoring techniques. RESULTS The pipeline reconstructed and analyzed the 5112 reconstructed datasets in approximately 10 days, a roughly 72× speedup over previous efforts using the scanner for reconstructions. Tightly coupled pipeline quality assurance software ensured proper performance of analysis modules with regard to segmentation and emphysema scoring. CONCLUSIONS The pipeline greatly reduced the time from experiment conception to quantitative results. The modular design of the pipeline allows the high-throughput framework to be utilized for other future experiments into different quantitative imaging techniques. Future applications of the pipeline being explored are robustness testing of quantitative imaging metrics, data generation for deep learning, and use as a test platform for image-processing techniques to improve clinical quantitative imaging.
PURPOSE:It is important to enhance image quality for low-dose CT acquisitions to push the ALARA boundary. Current state-of-the-art block-matching three-dimensional (BM3D) denoising scheme assumes white Gaussian noise (WGN) model. This study proposes a novel filtering module to be incorporated into the BM3D framework for ultra-low-dose CT denoising, by accounting for its specific power spectral properties.METHODS:In the current BM3D algorithm, the Wiener filtering is applied in the transform domain to a post-thresholding signal for enhanced denoising. However, unlike most natural/synthetic images, low-dose CTs do not obey the ideal Gaussian noise model. Based on the specific noise properties of ultra-low-dose CT, we derive the optimal transform-domain coefficients of Wiener filter based on the minimum mean-square-error (MMSE) criterion, taking the noise spectrum and the signal/noise cross spectrum into consideration. In the absence of ground-truth signal, the hard-thresholding denoising module in the previous stage is used as a plug-in estimator. We evaluate the denoising performance on thoracic CT image datasets containing paired full-dose and ultra-low-dose images simulated by a well-validated clinical engine (or pipeline). We also assess its clinical implication by applying the denoising methods to the emphysema quantification task. Our modified BM3D method is compared with the current one, using peak signal-to-noise ratio (PSNR) and emphysema scoring results as evaluation metrics.RESULTS:The noise in ultra-low-dose CT presented distinct non-Gaussian characteristics and was correlated with image intensity. Performance evaluation showed that the current Wiener filter in basic BM3D algorithm yielded little denoising enhancement on ultra-low-dose CT images. In contrast, the proposed Wiener filter achieved (1.46, 1.91) dB performance gain in mean and median peak signal-to-noise ratio (PSNR) for 5%-dose image denoising and (0.93, 0.95) dB improvement for 10% dose. A paired t-test of the PSNRs between denoising using the current and the proposed Wiener filters demonstrated statistically significant improvement, yielding P-values of 1.45E-12 and 1.34E-7 on 5% and 10%-dose images, respectively. In addition, emphysema quantification on the denoised images using the modified BM3D method also had statistically significant advantage over that using the current BM3D scheme, resulting in a P-value of 6.30E-5 with the commonly used measure.CONCLUSIONS:This work tailors the Wiener filter in BM3D algorithm to data statistics and demonstrates statistically significant performance improvement on ultra-low-dose CT image denoising and a subsequent emphysema quantification task. Such performance gain is more pronounced with a lower dose level. The development and rationale are generally enough for other image denoising tasks when the WGN assumption is violated.
Quantitative imaging in lung cancer CT seeks to characterize nodules through quantitative features, usually from a region of interest delineating the nodule. The segmentation, however, can vary depending on segmentation approach and image quality, which can affect the extracted feature values. In this study, we utilize a fully-automated nodule segmentation method – to avoid reader-influenced inconsistencies – to explore the effects of varied dose levels and reconstruction parameters on segmentation. Raw projection CT images from a low-dose screening patient cohort (N=59) were reconstructed at multiple dose levels (100%, 50%, 25%, 10%), two slice thicknesses (1.0mm, 0.6mm), and a medium kernel. Fully-automated nodule detection and segmentation was then applied, from which 12 nodules were selected. Dice similarity coefficient (DSC) was used to assess the similarity of the segmentation ROIs of the same nodule across different reconstruction and dose conditions. Nodules at 1.0mm slice thickness and dose levels of 25% and 50% resulted in DSC values greater than 0.85 when compared to 100% dose, with lower dose leading to a lower average and wider spread of DSC values. At 0.6mm, the increased bias and wider spread of DSC values from lowering dose were more pronounced. The effects of dose reduction on DSC for CAD-segmented nodules were similar in magnitude to reducing the slice thickness from 1.0mm to 0.6mm. In conclusion, variation of dose and slice thickness can result in very different segmentations because of noise and image quality. However, there exists some stability in segmentation overlap, as even at 1mm, an image with 25% of the lowdose scan still results in segmentations similar to that seen in a full-dose scan.