Neonatal hyperbilirubinemia is common in early neonatal life, and delayed recognition of high-risk infants may result in bilirubin-induced neurological injury. Prediction models based on routinely available maternal, perinatal, neonatal, and standardized laboratory variables may support early risk stratification as an adjunct to bilirubin measurement and clinical assessment. This single-center retrospective cohort included 986 neonates admitted between January 1, 2016, and March 1, 2026. The hemolytic-disease variable was excluded before preprocessing and model fitting. Admissions from 2016 through 2023 formed the development cohort (n = 776), and those from 2024 through 2026 formed a held-out temporal validation cohort (n = 210). Stability was assessed using repeated stratified 5-fold cross-validation repeated 5 times. The primary outcome was peak total bilirubin of at least 230 µmol/L during hospitalization. Only predictors documented by the predefined early-admission prediction time were used. Tabular prior-data fitted network (TabPFN) was compared with 5 conventional machine learning models. Performance included discrimination, calibration, Brier score, decision-curve analysis, bootstrap confidence intervals (CIs), and grouped permutation importance. The development and temporal validation cohorts included 468 and 54 events, respectively. Across 25 held-out development folds, TabPFN achieved a mean area under the receiver operating characteristic curve of 0.997 (standard deviation [SD], 0.002), mean area under the precision-recall curve of 0.998 (SD, 0.001), and mean Brier score of 0.023 (SD, 0.009). In temporal validation, TabPFN achieved an area under the receiver operating characteristic curve of 0.999 (95% CI, 0.997-1.000), area under the precision-recall curve of 0.997 (95% CI, 0.991-1.000), and Brier score of 0.013. At the development-derived threshold of 0.538, sensitivity was 1.000, specificity 0.981, positive predictive value 0.947, negative predictive value 1.000, and F1 score 0.973. Random forest had marginally higher temporal discrimination, whereas TabPFN had the lowest Brier score. Leading predictors were maternal disease count, abortion or preterm history, postnatal abnormality, and birth-weight category. After exclusion of the hemolytic-disease variable, TabPFN showed stable repeated-cross-validation performance and strong temporal discrimination, calibration, and clinical net benefit. These single-center findings support TabPFN as a candidate decision-support approach requiring independent multicenter validation and prospective benchmarking against standard-of-care bilirubin risk assessment.
更多