Medical foundation models show promise to learn broadly generalizable features from large, diverse datasets. This could be the base for reliable cross-modality generalization and rapid adaptation to new, task-specific goals, with only a few task-specific examples. Yet, evidence for this is limited by the lack of public, standardized, and reproducible evaluation frameworks, as existing public benchmarks are often fragmented across task-, organ-, or modality-specific settings, limiting assessment of cross-task generalization. We introduce UNICORN, a public benchmark designed to systematically evaluate medical foundation models under a unified protocol. To isolate representation quality, we built the benchmark on a novel two-step framework that decouples model inference from task-specific evaluation based on standardized few-shot adaptation. As a central design choice, we constructed indirectly accessible sequestered test sets derived from clinically relevant cohorts, along with standardized evaluation code and a submission interface on an open benchmarking platform. Performance is aggregated into a single UNICORN Score, a new metric that we introduce to support direct comparison of foundation models across diverse medical domains, modalities, and task types. The UNICORN test dataset includes data from more than 2,400 patients, including over 3,700 vision cases and over 2,400 clinical reports collected from 17 institutions across eight countries. The benchmark spans eight anatomical regions and four imaging modalities. Both task-specific and aggregated leaderboards enable accessible, standardized, and reproducible evaluation. By standardizing multi-task, multi-modality assessment, UNICORN establishes a foundation for reproducible benchmarking of medical foundation models. Data, baseline methods, and the evaluation platform are publicly available via unicorn.grand-challenge.org.
Constraining the major contributors to the ionisation of the early universe is an ongoing endeavour of high-redshift galaxy research. We measure the ionising photon production efficiency and Lyman Continuum escape fraction for a sample of 230 intermediate redshift (1.35. This control sample allows us to verify the correlations between ionising and spectral/physical properties suggested by previous studies. We find no significant correlations between the ionising photon production efficiency (ξ_ion) with the UV slope, M_UV, M_* or sSFR. We do find that ξ_ion correlates with [OIII]5007Å equivalent width (EW) (Spearman coefficient ρ =0.24; p< 4×10^-4) and Hα EW (ρ =0.63; p<< 1×10^-6) hold even at low EW albeit with more scatter. We also find that our novel approach to determining the ionising photon escape fraction f_esc results in values within theoretical ranges (0-10%) though vary substantially in comparison to the empirical results (median f_esc = 0.9%^+1.1_-0.5 including non-detections, median f_esc = 1.9%^+8.9_-1.8 above a 0.01% threshold). We find that this escape fraction method has consistently significant correlations with the redshift, SFR and M_UV and sample-dependent correlations with [OIII]5007Å EW,Hα EW and stellar mass.
The current BTS guidelines recommend evaluation of suspicious pulmonary nodules using [18F]FDG-PET/CT imaging, followed by Herder model risk stratification. However, it is based on limited imaging features, which may limit diagnostic accuracy. This study aims to develop a PET/CT-based deep learning (DL) model for malignancy probability estimation (AITO-PETCT-MP) and compare its performance to the Herder model and clinician performance. In a single-center retrospective study, we collected 533 indeterminate pulmonary nodules (268 malignant) with a mean diameter of 18.4 mm (SD ± 12.1) in 436 patients. Histopathological malignancy confirmation or a minimum 2-year benign national cancer registry follow-up served as the reference standard. Model diagnostic performance was compared against the Herder model and seven clinicians in a reader study on a test set of 161 nodules (80 malignant). AITO-PETCT-MP achieved an AUC of 0.78 [95
Background. Low-dose computed tomography (LDCT) screening reduces lung cancer mortality fundamentally through patient stratification, determining management intensity and screening intervals. Clinical guidelines such as Lung-RADS and NELSON make this assignment from nodule features, which can be measured by a radiologist or by a deep learning feature extractor. Emerging end-to-end deep learning approaches predict multi-year lung cancer (continuous) risk from a single LDCT, independent of nodule presence. These models are evaluated almost entirely by aggregate discrimination — AUC and the concordance index — which summarize ranking but do not localize where across the risk distribution a model is useful. They are not assessed as the actionable stratification tools they are meant to be. Methods. We propose an evaluation framework for end-to-end models that is aligned with clinical deployment and rooted in lung screening standards. It retains high-level ranking metrics (AUC, PR-AUC, concordance index) for comparability, and adds operating-point evaluation anchored to Lung-RADS: cancer capture from Lorenz curves, clinical utility from decision-curve analysis (DCA), calibration assessed both overall and within Lung-RADS categories. Because every operating point is a Lung-RADS operating point, the framework compares models with one another and with clinical standards such as Lung-RADS or NELSON. We applied it to three contemporary models that predict 6-year risk from a single LDCT on the National Lung Screening Trial (NLST) test cohort (6,223 scans / 2,199 patients), with retrospectively assigned Lung-RADS as the clinical standard and the non-imaging PLCOm2012 model as a baseline. Findings. The models have similar AUCs with overlapping confidence intervals (year-1 AUC 0.91–0.95). Calibration was good in aggregate but varied by category: the models over-predicted in the low-risk Lung-RADS tiers, were well calibrated at intermediate risk, and under-predicted in the high-risk tiers where workup is decided (11% predicted against 22% observed at the highest-risk tier). On cumulative-gains curves and decision-curve analysis, the stronger end-to-end models achieved at least on-par or higher cancer capture and greater net benefit than Lung-RADS, and the advantage was largest on baseline scans, where Lung-RADS has no prior CT for comparison, exceeding it by more than 25 percentage points in capture. Furthermore, the continuous scores were also effective at the extremes the categories cannot reach: they isolated very-high-risk patients at flag rates below the smallest Lung-RADS category, and recovered and ranked higher-risk patients placed in the benign Lung-RADS tiers, where the top score decile held about 70% of the cancers hidden in those tiers (6.7–7.4x enrichment). Implications. Evaluating end-to-end models as stratification tools, rather than by global discrimination, assesses them on the task they are designed for, and places them in the same frame as the clinical standards they would support: Lung-RADS, and feature-based deep learning models that classify by its rules. In that frame, end-to-end deep learning and Lung-RADS are not best understood as competitors. They derive risk differently and have complementary strengths, and our results indicate they are more useful combined than ranked against each other.
BACKGROUND:Computed Tomography (CT) scans allow opportunistic evaluation of body composition. We investigated whether body composition and change through time are associated with lung cancer incidence and all-cause/lung cancer-specific mortality in a lung cancer screening cohort. METHODS:A machine learning segmentation method was used in this retrospective cohort study to measure skeletal muscle area and density, and subcutaneous adipose tissue area (SAT) on repeated chest CTs from the Dutch-Belgian lung cancer screening trial. Hazard ratios by sex adjusted for age, smoking status, and smoking pack-years (aHR) were calculated for each outcome. FINDINGS:During median follow-up of 12.2 (interquartile range, 1.2) years, 4.1% of 6187 subjects (85.5% male, mean age ± SD, 58.6 ± 5.5 years, smoking pack-years 41.2 ± 18.3) developed lung cancer, 12.2% died, and 2.1% died due to lung cancer. For males, SAT loss was associated with lung cancer incidence (aHR 1.19, 95% CI 1.02-1.39) and lung cancer-specific mortality (aHR 1.26, 95% CI 1.03-1.55), and less baseline muscle and muscle loss with all-cause mortality (aHR 1.20, 95% CI 1.10-1.31 and 1.17, 1.07-1.27). For females, less baseline SAT and SAT loss was associated with all-cause mortality (aHR 1.44, 95% CI 1.06-1.97 and 1.48, 1.13-1.94) and lung cancer-specific mortality (aHR 2.85, 95% CI 1.50-5.39 and aHR 1.96, 1.11-3.44). Models improved by including body composition trends for all-cause mortality (males: p < 0.001; females: p = 0.012) and for lung cancer-specific mortality (males: p = 0.102; females: p = 0.005). INTERPRETATION:Body composition trends based on automated analysis of chest CT are associated with worse outcomes in participants screened for lung cancer. FUNDING:Dutch Cancer Society, Health Holland, Siemens Healthineers.
A screening-trained deep learning (DL) model (DL1) for pulmonary nodule malignancy probability estimation on CT previously demonstrated good discrimination on a single-centre dataset of incidental nodules. An updated DL model (DL2) was trained on both screening and clinical data. We aimed to test the performance of both models in a multicentre dataset of incidental nodules. A retrospective, multicentre, case-control dataset of incidental nodules was collected, sampled across size buckets (5–10 mm, 10–15 mm, 15–30 mm), aiming for 10 malignant and 20 benign nodules per bucket and centre, resulting in 270 nodules. Performance was assessed using AUCs and specificity at a fixed sensitivity. AUCs were compared using the DeLong method. Both DL models were compared with the Brock model. Consistent discrimination was investigated by centre-stratified analyses. The multicentre dataset contained 269 nodules (89 malignant) from 231 patients. DL1 and DL2 achieved AUCs of 0.74 and 0.72, respectively, versus 0.63 for Brock (both p < 0.01). Using a 10
To characterize the capabilities of CE-marked AI products for lung nodule analysis in lung cancer screening (LCS), quantify their coverage of tasks defined in nodule management recommendations, and assess their peer-reviewed evidence. Six core tasks in LCS (nodule detection, classification, measurement, growth assessment, malignancy risk estimation, and structured management) were derived from 4 nodule management recommendations: Lung-RADS 2022, British Thoracic Society (BTS) guidelines, European Union Position Statement (EUPS), and European Society of Thoracic Imaging (ESTI). Products were identified through www.healthairegister.com . Vendors confirmed capabilities using a standardized questionnaire; public documentation supplemented non-responders. Task coverage was calculated as the percentage of functional overlap (0–100
Foundation models are increasingly used as frozen feature extractors for CT classification tasks, yet the determinants of their downstream performance remain unclear. We assess how foundation model choice, feature aggregation strategy, and abnormality type are associated with performance in organ-level abnormality classification. We further evaluate whether more expressive aggregation strategies outperform simple pooling, and whether performance varies by abnormality type. We evaluate five state-of-the-art medical imaging foundation models on binary organ-level abnormality classification (normal vs. abnormal) across six abdominal organs. Local patch-level embeddings are aggregated using six strategies, including simple pooling methods (e.g., mean pooling) and attention-based multiple instance learning. A linear classifier is trained on feature embeddings, and performance is evaluated on an annotated test set of 200 CT scans. Among the evaluated models, 3D CT-native models generally outperformed 2D multi-modal models. None of the tested aggregation strategies significantly outperformed mean pooling, including attention-based multiple instance learning (best-performing aggregation vs. mean: Δ AUC = 0.008, 95 [-0.002, 0.022] ). In our organ-level classification setting, performance differed between abnormality type for all well-performing models (SPECTRE, TAP-CT, CT-FM), with lower AUCs observed for focal abnormalities compared to diffuse abnormalities (largest difference: Δ AUC = 0.108, 95
Background/Objectives: The objective of this study is to evaluate the performance of the traditional age/smoking criteria and existing risk prediction models in selecting high-risk populations for lung cancer screening from a Western European general population. Methods: Baseline data from the Dutch population-based Lifelines cohort, collected between 2006 and 2013, were linked to the Dutch cancer registry to confirm lung cancer diagnoses. Five-year lung cancer risk was estimated based on traditional age/smoking criteria (NLST, NELSON, SPSTF-2021) and risk prediction models (LLPv2, PLCOm2012, Hoggart, Bach and Shanghai-LCM). For every strategy, the number of individuals eligible was determined, and total lung cancer cases in the eligible groups versus the ineligible groups were calculated. Results: Among 139,120 participants (aged ≥18 years), 218 (0.2%) developed lung cancer within five years. Age/smoking criteria identified 2161-6295 (1.6-4.5%) participants as eligible, comprising 62-92 (28.4-42.2%) lung cancer cases. Risk prediction models identified 2372-4315 (1.7-3.1%) participants as eligible, comprising 40-85 (18.4-38.9%) lung cancer cases. Among lung cancers in ineligible groups, 46.2-59.6% occurred in individuals who formerly smoked, and 28.7-39.3% occurred in individuals who currently smoke. Additionally, 41.2-70.0% of lung cancer cases in ineligible groups were in individuals younger than 50, and 44.3-72.3% in individuals who had quit smoking > 15 years prior to diagnosis. Conclusions: In a Western European population, current lung cancer screening selection criteria resulted in identifying only 18-42% of lung cancer cases. Cases in ineligible groups predominantly concern individuals who currently smoke and are below the threshold age and individuals who quit smoking > 15 years ago, highlighting the opportunity for more personalized risk-based screening strategies to increase lung cancer detection.
Purpose To compare the performance of an artificial intelligence (AI) system with that of radiologists for estimating malignancy risk of indeterminate-size nodules (5-15 mm) at low-dose CT (LDCT) within a standardized and transparent evaluation framework. Materials and Methods Teams participating in the AI study had access to a public dataset of 555 malignant and 5608 benign nodules on 4069 baseline LDCT scans from the National Lung Screening Trial to develop AI systems. External testing was performed on 156 malignant and 312 benign size-matched nodules, all of indeterminate size, from 463 baseline scans collected from three large European lung cancer screening trials, and the best-performing AI system (based on area under the receiver operating characteristic curve [AUC]) was selected. An observer study was conducted in which radiologists assessed 300 randomly selected nodules (100 malignant, 200 benign) from the external test set. Radiologists categorized nodules as low, intermediate, or high risk, and the threshold of intermediate or greater risk (intermediate or high-risk) was used to define a positive test result. The selected AI system was compared with radiologists on this subset using the AUC. Results The selected AI system demonstrated superior performance to the 65 radiologists' mean (AUC, 0.78 [95% CI: 0.73, 0.84] vs 0.70 [95% CI: 0.65, 0.74]; P = .001). With use of the intermediate risk or greater threshold, the AI system correctly classified 12% more malignant nodules at matched specificity and yielded 20% fewer false-positive results at matched sensitivity. Conclusion The selected AI system was superior to radiologists in estimating malignancy risk of indeterminate lung nodules at LDCT. Keywords: CT, Thorax, Lung, Observer Performance, Screening, Supervised Learning, Lung Cancer Screening, Radiologists, Artificial Intelligence, Benchmarking, Pulmonary Nodule Malignancy Risk, Deep Learning Supplemental material is available for this article. © RSNA, 2026 See also commentary by Júdice de Mattos Farina and Szarf in this issue.
The ASTRO 3D Galaxy Evolution with Lenses (AGEL) Survey is an ongoing effort to spectroscopically confirm a diverse sample of gravitational lenses with high spatial resolution imaging, to facilitate a broad range of science outcomes. The AGEL systems span single galaxy-scale deflectors to groups and clusters, and include rare targets such as galaxy-scale lenses with multiple sources, lensed quiescent galaxies, and Einstein rings. We build on the 68 systems presented in Tran et al. (AGEL data release 1) to present a total of 139 lenses, and high-resolution Hubble Space Telescope images for 167 lenses from three programs (including one ongoing). Lens candidates were originally identified by convolutional neural networks in the Dark Energy Camera Legacy Survey imaging fields, and of the targets with follow-up spectroscopy, we find a high (96%) success rate. Compared with other spectroscopic lens samples, AGEL lenses tend to have both higher redshift deflectors and sources. We briefly discuss the common causes of false-positive candidates, and suggest strategies for mitigating false-positives in next-generation lens searches. Lastly, we present the newly measured redshifts for six (five confirmed strong lenses, one probable) galaxy-scale double-source plane lenses, targets which are useful for cosmological analyses. With next-generation telescopes and surveys such as Euclid, Vera Rubin’s Legacy Survey of Space and Time, Keck Observatory’s KAPA program, and 4MOST’s 4SLSLS surveys on the horizon, the AGEL survey represents a pathfinder for refining automated candidate search methods and identifying and triaging candidates for follow-up based on scientific potential.
Early detection of lung cancer through low-dose CT lung cancer screening in a high-risk population has proven to reduce lung cancer-specific mortality. Nodule management plays a pivotal role in early detection and further diagnostic approaches. The European Society of Thoracic Imaging (ESTI) has established a nodule management recommendation to improve the handling of pulmonary nodules detected during screening. For solid nodules, the primary method for assessing the likelihood of malignancy is to monitor nodule growth using volumetry software. For subsolid nodules, the aggressiveness is determined by measuring the solid part. The ESTI-recommendation enhances existing protocols but puts a stronger focus on lesion aggressiveness. The main goals are to minimise the overall number of follow-up examinations while preventing the risk of a major stage shift and reducing the risk of overtreatment. Question Assessment of nodule growth and management according to guidelines is essential in lung cancer screening. Findings Assessment of nodule aggressiveness defines follow-up in lung cancer screening. Clinical relevance The ESTI nodule management recommendation aims to reduce follow-up examinations while preventing major stage shift and overtreatment.
Incidental airway tumors are rare and can easily be overlooked on chest CT, especially at an early stage. Therefore, we developed and assessed a deep learning-based artificial intelligence (AI) system for detecting and localizing airway nodules. At a single academic hospital, we retrospectively analyzed cancer diagnoses and radiology reports from patients who received a chest or chest–abdomen CT scan between 2004 and 2020 to find cases presenting as airway nodules. Primary cancers were verified through bronchoscopy with biopsy or cytologic testing. The malignancy status of other nodules was confirmed with bronchoscopy only or follow-up CT scans if such evidence was unavailable. An AI system was trained and evaluated with a ten-fold cross-validation procedure. The performance of the system was assessed with a free-response receiver operating characteristic curve. We identified 160 patients with airway nodules (median age of 64 years [IQR: 54–70], 58 women) and added a random sample of 160 patients without airway nodules (median age of 60 years [IQR: 48–69], 80 women). The sensitivity of the AI system was 75.1
Incidental pulmonary nodules are very frequently found on CT imaging and may represent (early stage) lung cancers without any signs or symptoms. These incidental findings can be solid lesions or ground glass lesions that may be solitary or multiple. Careful, and systematic evaluation of these findings in imaging is needed to determine the risk of malignancy, based on imaging characteristics, patient factors like smoking habits, prior cancers or family history, and growth rate preferably determined by volume measurements. Once the risk of malignancy is increased, minimal invasive image guided biopsy is warranted, preferably by navigation bronchoscopy. We present two cases to illustrate this clinical workup: one case with a benign solitary pulmonary nodule, and a second case with multiple ground glass opacities, diagnosed as synchronous primary adenocarcinomas of the lung. This is followed by a review of the current status of computer and artificial intelligence aided diagnostic support and clinical workflow optimization.
Lung cancer is the leading cause of cancer-related mortality in adults worldwide. Screening high-risk individuals with annual low-dose CT (LDCT) can support earlier detection and reduce deaths, but widespread implementation may strain the already limited radiology workforce. AI models have shown potential in estimating lung cancer risk from LDCT scans. However, high-risk populations for lung cancer are diverse, and these models' performance across demographic groups remains an open question. In this study, we drew on the considerations on confounding factors and ethically significant biases outlined in the JustEFAB framework to evaluate potential performance disparities and fairness in two deep learning risk estimation models for lung cancer screening: the Sybil lung cancer risk model and the Venkadesh21 nodule risk estimator. We also examined disparities in the PanCan2b logistic regression model recommended in the British Thoracic Society nodule management guideline. Both deep learning models were trained on data from the US-based National Lung Screening Trial (NLST), and assessed on a held-out NLST validation set. We evaluated AUROC, sensitivity, and specificity across demographic subgroups, and explored potential confounding from clinical risk factors. We observed a statistically significant AUROC difference in Sybil's performance between women (0.88, 95
Low-dose CT screening for lung cancer reduces the risk of death from lung cancer by at least 21
To assess changes in peer-reviewed evidence on commercially available radiological artificial intelligence (AI) products from 2020 to 2023, as a follow-up to a 2020 review of 100 products. A literature review was conducted, covering January 2015 to March 2023, focusing on CE-certified radiological AI products listed on www.healthairegister.com . Papers were categorised using the hierarchical model of efficacy: technical/diagnostic accuracy (levels 1–2), clinical decision-making and patient outcomes (levels 3–5), or socio-economic impact (level 6). Study features such as design, vendor independence, and multicentre/multinational data usage were also examined. By 2023, 173 CE-certified AI products from 90 vendors were identified, compared to 100 products in 2020. Products with peer-reviewed evidence increased from 36
We present the formation histories of 19 massive (≳3 × 10 10 M ⊙ ) quiescent (specific star formation rate, sSFR < 0.15 Gyr −1 ) galaxy candidates at z ~ 3.0–4.5 observed using JWST/NIRSpec. This completes the spectroscopic confirmation of the 24 K -selected quiescent galaxy sample from the ZFOURGE and 3DHST surveys. Utilizing Prism 1–5 μ m spectroscopy, we confirm that all 12 sources that eluded confirmation by ground-based spectroscopy lie at z > 3, resulting in a spectroscopically confirmed number density of ~1.4 × 10 −5 Mpc −3 between z ~ 3 and 4. Rest-frame U − V versus V − J color selections show high effectiveness in identifying quiescent galaxies, with a purity of ~90%. Our analysis shows that parametric star formation histories (SFHs) from FAST++ and binned SFHs from Prospector on average yield consistent results, revealing diverse formation and quenching times. The oldest galaxy formed ~6 × 10 10 M ⊙ by z ~ 10 and has been quiescent for over 1 Gyr at z ~ 3.2. We detect two galaxies with ongoing star formation and six with active galactic nuclei (AGNs). We demonstrate that the choice of stellar population models, stellar libraries, and nebular or AGN contributions does not significantly affect the derived average SFHs of the galaxies. We demonstrate that extending spectral fitting beyond the rest-frame optical regime reduces the inferred average star formation rates (SFRs) in the earliest time bins of the SFH reconstruction. The assumed SFH prior influences the SFR at early times, where spectral diagnostic power is limited. Simulated z ~ 3 quiescent galaxies from IllustrisTNG, SHARK, and Magneticum broadly match the average SFHs of the observed sample but struggle to capture the full diversity, particularly at early stages. Our results emphasize the need for mechanisms that rapidly build stellar mass and quench star formation within the first billion years of the Universe.
Purpose To investigate the relationship between training data volume and performance of a deep learning artificial intelligence (AI) algorithm developed to assess the malignancy risk of pulmonary nodules detected on low-dose CT scans in lung cancer screening. Materials and Methods This retrospective study used a dataset of 16 077 annotated nodules (1249 malignant, 14 828 benign) from the National Lung Screening Trial (NLST) to systematically train an AI algorithm for pulmonary nodule malignancy risk prediction across various stratified subsets ranging from 1.25% to the full dataset. External testing was conducted using data from the Danish Lung Cancer Screening Trial (DLCST) to determine the amount of training data at which the performance of the AI was statistically noninferior to the AI trained on the full NLST cohort. A size-matched cancer-enriched subset of DLCST, in which each malignant nodule had been paired in diameter with the closest two benign nodules, was used to investigate the amount of training data at which the performance of the AI algorithm was statistically noninferior to the average performance of 11 clinicians. Results The external testing set included 599 participants (mean age ± SD, 57.65 years ± 4.84 for female participants and 59.03 years ± 4.94 for male participants) with 883 nodules (65 malignant, 818 benign). The AI achieved a mean area under the receiver operating characteristic curve (AUC) of 0.92 (95% CI: 0.88, 0.96) on the DLCST cohort when trained on the full NLST dataset. Training with 80% of the NLST data resulted in noninferior performance (mean AUC, 0.92; 95% CI: 0.89, 0.96; P = .005). On the size-matched DLCST subset (59 malignant, 118 benign), the AI reached noninferior clinician-level performance (mean AUC, 0.82; 95% CI: 0.77, 0.86) with 20% of the training data (P = .02). Conclusion The deep learning AI algorithm demonstrated excellent performance in assessing pulmonary nodule malignancy risk, achieving clinical level performance with a fraction of the training data and reaching peak performance before using the full dataset. Keywords: Convolutional Neural Network (CNN), CT, Lung, Screening, Diagnosis, Supervised Learning, Lung Cancer Screening, Pulmonary Nodule Malignancy Risk, Deep Learning, Pulmonary Nodule Management Supplemental material is available for this article. © RSNA, 2025 See also commentary by Archer in this issue.
The European Society of Thoracic Imaging (ESTI) nodule management recommendation for lung cancer screening with low-dose CT builds on existing nodule management guidelines but puts a stronger focus on lesion aggressiveness and measurement error. Key objectives included finding a compromise between the overall number of follow-up examinations, avoiding a major stage shift, and reducing the risk for overtreatment. Nodule management categories at baseline are chosen depending on the size of a solid nodule or the solid component of a subsolid or cystic nodule, with suspicious morphology upgrading risk to the next higher category. Higher risk categories mandate shorter follow-up times or diagnostic workup. Volume is the preferred size measure, with diameter measurements as a fallback if segmentation for volumetry is inaccurate at visual control. Nodule aggressiveness at follow-up is estimated from growth rate, calculated as volume doubling time (VDT), or yearly diameter change. Calculation of growth rate, however, is strongly affected by measurement variability, with large error margins for short follow-up and slower growing lesions. Growth thresholds were therefore set so that rapidly growing lesions can be identified while still small, while unnecessary workups for benign or slow-growing lesions could be kept low. New lesions that are retrospectively visible on earlier scans are managed according to their growth rate. New nodules not visible on earlier scans are followed after 3 months if they have a volume of ≥ 30 mm3. Question This work strives to reduce follow-up examinations while preventing major stage shift and overtreatment. It provides nodule management based on estimated nodule aggressiveness. Findings Calculation of the growth rate of pulmonary nodules is strongly affected by measurement variability, with large error margins for short follow-up and slower growing lesions. Clinical relevance Growth thresholds that trigger management are adjusted to the follow-up time so that rapidly growing lesions can be identified while still being small while unnecessary workups for benign or slow-growing lesions can be reduced.