Background Early identification of children at high risk of obesity can provide clinicians with the information needed to provide targeted lifestyle counseling to high-risk children at a critical time to change the disease course. Objectives This study aimed to develop predictive models of childhood obesity, applying advanced machine learning methods to a large unaugmented electronic health record (EHR) dataset. This work improves on other studies that have (i) relied on data not routinely available in EHRs (like prenatal data), (ii) focused on single-age predictions, or (iii) not been rigorously validated. Methods A customized sequential deep-learning model to predict the development of obesity was built, using EHR data from 36,191 diverse children aged 0–10 years. The model was evaluated using extensive discrimination, calibration, and utility analysis; and was validated temporally, geographically, and across various subgroups. Results Our results are mostly better or comparable to similar studies. Specifically, the model achieved an AUROC above 0.8 in all cases (with most cases around 0.9) for predicting obesity within the next 3 years for children 2–7 years of age. Validation results show the model’s robustness and top predictors match important risk factors of obesity. Conclusions Our model can predict the risk of obesity for young children at multiple time points using only routinely collected EHR data, greatly facilitating its integration into clinical care. Our model can be used as an objective screening tool to provide clinicians with insights into a patient’s risk for developing obesity so that early lifestyle counseling can be provided to prevent future obesity in young children.
Amid the ongoing global repercussions of SARS-CoV-2, it is crucial to comprehend its potential long-term psychiatric effects. Several recent studies have suggested a link between COVID-19 and subsequent mental health disorders. Our investigation joins this exploration, concentrating on Schizophrenia Spectrum and Psychotic Disorders (SSPD). Different from other studies, we took acute respiratory distress syndrome (ARDS) and COVID-19 lab-negative cohorts as control groups to accurately gauge the impact of COVID-19 on SSPD. Data from 19,344,698 patients, sourced from the N3C Data Enclave platform, were methodically filtered to create propensity matched cohorts: ARDS (n = 222,337), COVID-19 positive (n = 219,264), and COVID-19 negative (n = 213,183). We systematically analyzed the hazard rate of new-onset SSPD across three distinct time intervals: 0-21 days, 22-90 days, and beyond 90 days post-infection. COVID-19 positive patients consistently exhibited a heightened hazard ratio (HR) across all intervals [0-21 days (HR: 4.6; CI: 3.7-5.7), 22-90 days (HR: 2.9; CI: 2.3 -3.8), beyond 90 days (HR: 1.7; CI: 1.5-1.)]. These are notably higher than both ARDS and COVID-19 lab-negative patients. Validations using various tests, including the Cochran Mantel Haenszel Test, Wald Test, and Log-rank Test confirmed these associations. Intriguingly, our data indicated that younger individuals face a heightened risk of SSPD after contracting COVID-19, a trend not observed in the ARDS and COVID-19 negative groups. These results, aligned with the known neurotropism of SARS-CoV-2 and earlier studies, accentuate the need for vigilant psychiatric assessment and support in the era of Long-COVID, especially among younger populations.
Background: Understanding social determinants of health (SDOH) that may be risk factors for childhood obesity is important to developing targeted interventions to prevent obesity. Prior studies have examined these risk factors, mostly examining obesity as a static outcome variable. Methods: We extracted electronic health record data from 2012 to 2019 for a children's health system that includes two hospitals and wide network of outpatient clinics spanning five East Coast states in the United States. Using data-driven and algorithmic clustering, we have identified distinct BMI-percentile classification groups in children from 0 to 7 years of age. We used two separate algorithmic clustering methods to confirm the robustness of the identified clusters. We used multinomial logistic regression to examine the associations between clusters and 27 neighborhood SDOHs and compared positive and negative SDOH characteristics separately. Results: From the cohort of 36,910 children, five BMI-percentile classification groups emerged: always having obesity (n = 429; 1.16%), overweight most of the time (n = 15,006; 40.65%), increasing BMI percentile (n = 9,060; 24.54%), decreasing BMI percentile (n = 5,058; 13.70%), and always normal weight (n = 7,357; 19.89%). Compared to children in the decreasing BMI percentile and always normal weight groups, children in the other three groups were more likely to live in neighborhoods with higher poverty, unemployment, crowded households, single-parent households, and lower preschool enrollment. Conclusions: Neighborhood-level SDOH factors have significant associations with children's BMI-percentile classification and changes in classification. This highlights the need to develop tailored obesity interventions for different groups to address the barriers faced by communities that can impact the weight and health of children living within them.
As clinical understanding of pediatric Post-Acute Sequelae of SARS CoV-2 (PASC) develops, and hence the clinical definition evolves, it is desirable to have a method to reliably identify patients who are likely to have post-acute sequelae of SARS CoV-2 (PASC) in health systems data. In this study, we developed and validated a machine learning algorithm to classify which patients have PASC (distinguishing between Multisystem Inflammatory Syndrome in Children (MIS-C) and non-MIS-C variants) from a cohort of patients with positive SARS- CoV-2 test results in pediatric health systems within the PEDSnet EHR network. Patient features included in the model were selected from conditions, procedures, performance of diagnostic testing, and medications using a tree-based scan statistic approach. We used an XGboost model, with hyperparameters selected through cross-validated grid search, and model performance was assessed using 5-fold cross-validation. Model predictions and feature importance were evaluated using Shapley Additive exPlanation (SHAP) values. The model provides a tool for identifying patients with PASC and an approach to characterizing PASC using diagnosis, medication, laboratory, and procedure features in health systems data. Using appropriate threshold settings, the model can be used to identify PASC patients in health systems data at higher precision for inclusion in studies or at higher recall in screening for clinical trials, especially in settings where PASC diagnosis codes are used less frequently or less reliably. Analysis of how specific features contribute to the classification process may assist in gaining a better understanding of features that are associated with PASC diagnoses.
Introduction: Preservation of cardiovascular health (CVH) across the lifespan is essential to reducing cardiovascular disease burden. Greater knowledge of the relationship between neighborhood level deprivation and CVH in youth is needed. Hypothesis: We hypothesized that worse socio/environmental deprivation is associated with poor/intermediate CVH status in adolescents. Methods: Data from 2009-2019 were extracted from PEDSnet (a PCORI-funded network of pediatric health data). We modeled the relationship between CVH and neighborhood deprivation. National Area Deprivation Index values (ADI; scaled 1-100) were coded into PEDSnet and divided into tertiles (ADI<25=best, 25-52, > 52=worst) for analysis. CVH status was scored from a subset of available AHA Life’s Essential 8 (LE8) scores including blood pressure, blood glucose, blood cholesterol, body mass index, smoking/tobacco exposure, and sleep. Overall CVH scores were derived as the average of the sum of the individual scores. Univariate and multivariate analyses were performed. SAS 9.4 was used ( p<0.05) . Results: Data from 122,177 youth, 13-17 years old were analyzed. 51% were female, 24% lived in neighborhoods with ADI <25; 35% lived in neighborhoods with ADI > 52. 58% were non-Hispanic white and 22% non-Hispanic Black. According to a multivariate logistic model, ADI > 52 was associated with a 48% greater odds of poor/intermediate CVH status compared to ADI <25 (Figure). Female sex, ethnic-race categories of Non-Hispanic White and Other, as well private or other forms of non-public insurance coverage were associated with a lower odds of intermediate/poor CVH (Table). Interaction terms for age, ADI, race and ethnicity were not significant. Conclusions: Higher neighborhood deprivation is associated with poor/intermediate CVH status. We recommend greater attention to preserving CVH status within high deprivation communities to preserve and promote CVH status across the lifespan.
Introduction: Efficient methods to obtain and benchmark national data are needed to improve comparative quality assessment for children with type 1 diabetes (T1D). PCORnet is a network of clinical data research networks whose infrastructure includes standardization to a Common Data Model (CDM) incorporating electronic health record (EHR)-derived data across multiple clinical institutions. The study aimed to determine the feasibility of the automated use of EHR data to assess comparative quality for T1D. Methods: In two PCORnet networks, PEDSnet and OneFlorida, the study assessed measures of glycemic control, diabetic ketoacidosis admissions, and clinic visits in 2016–2018 among youth 0–20 years of age. The study team developed measure EHR-based specifications, identified institution-specific rates using data stored in the CDM, and assessed agreement with manual chart review. Results: Among 9,740 youth with T1D across 12 institutions, one quarter (26%) had two or more measures of A1c greater than 9% annually (min 5%, max 47%). The median A1c was 8.5% (min site 7.9, max site 10.2). Overall, 4% were hospitalized for diabetic ketoacidosis (min 2%, max 8%). The predictive value of the PCORnet CDM was >75% for all measures and >90% for three measures. Conclusions: Using EHR-derived data to assess comparative quality for T1D is a valid, efficient, and reliable data collection tool for measuring T1D care and outcomes. Wide variations across institutions were observed, and even the best-performing institutions often failed to achieve the American Diabetes Association HbA1C goals (<7.5%).
Introduction: Youth with sickle cell anemia (SCA) are at risk for significant morbidity and mortality and impaired quality of life. Hydroxyurea (HU) is one of few disease-modifying therapies for SCA, with strong evidence of efficacy and safety. In 2014, the National Heart, Lung, and Blood Institute (NHLBI) published guidelines advocating that HU be offered to all youth with SCA (HbSS and HbSβ0-thalassemia genotypes) ages >9 months, regardless of clinical severity. We sought to examine hydroxyurea use among youth with SCA in the five years following the release of this clinical practice guideline across eight pediatric sickle cell care centers nationally. Methods: Retrospective data from 2014-2019 were obtained from PEDSnet, a network of pediatric health systems within the United States. A base cohort of sickle cell disease patients who received care at any time across the eight health systems was identified using a validated computable phenotype (Michalik et al., 2017). Next, patients with HbSS genotype were identified via diagnosis codes assigned at Hematology/Oncology outpatient visits. Finally, the proportion of youth with SCA (HbSS genotype) ages 9 months to 21 years who received >1 HU prescription was calculated for each sickle cell care center for each year. We selected >1 HU prescriptions within a study year as an indicator that the patient/family agreed to initiate or continue this therapy. Patients who had <1 Hematology/Oncology outpatient visit during the study year and those likely receiving chronic transfusion therapy (defined as >7 blood transfusion procedures during the study year) were excluded. Results: Between 1340 and 1601 patients per year met selection criteria and were included in this analysis; data were missing from one site for the final study year (2018-19). Most centers demonstrated a trend of increasing HU prescription rates over the study period. There was significant variability between sites in the proportion of youth with SCA who received >1 HU prescription within each study year: 31% - 64% (2014-15), 41% - 73% (2015-16), 42% - 74% (2016-17), 42% - 69% (2017-18), and 47% - 77% (2018-19; see Figure 1). Conclusions: Despite national guidelines and decades of strong evidence of efficacy and safety, use of HU is suboptimal among patients with SCA. Varying HU prescribing practices across clinicians and sites (e.g., length of supply provided) and lack of pharmacy dispensing data limited our ability to examine maintenance of HU therapy over time within this analysis. Additional research is needed to examine clinical implementation processes that may contribute to high variability in HU uptake across sickle cell care centers. Evidence-based interventions to improve scale up of HU are critical for this high-risk population. Acknowledgements: The research reported in this abstract was conducted using PEDSnet, A National Pediatric Learning Health System, and includes data from the following PEDSnet institutions: Children's Hospital Colorado, Children's Hospital of Philadelphia, Cincinnati Children's Hospital Medical Center, Nationwide Children's Hospital, Nemours Children's Health, Seattle Children's Hospital, and St. Louis Children's Hospital. Figure 1View largeDownload PPTFigure 1View largeDownload PPT Close modal
BACKGROUND:Children with glomerular disease have unique risk factors for compromised bone health. Studies addressing skeletal complications in this population are lacking.METHODS:This retrospective cohort study utilized data from PEDSnet, a national network of pediatric health systems with standardized electronic health record data for more than 6.5 million patients from 2009 to 2021. Incidence rates (per 10,000 person-years) of fracture, slipped capital femoral epiphysis (SCFE), and avascular necrosis/osteonecrosis (AVN) in 4598 children and young adults with glomerular disease were compared with those among 553,624 general pediatric patients using Poisson regression analysis. The glomerular disease cohort was identified using a published computable phenotype. Inclusion criteria for the general pediatric cohort were two or more primary care visits 1 year or more apart between 1 and 21 years of age, one visit or more every 18 months if followed >3 years, and no chronic progressive conditions defined by the Pediatric Medical Complexity Algorithm. Fracture, SCFE, and AVN were identified using SNOMED-CT diagnosis codes; fracture required an associated x-ray or splinting/casting procedure within 48 hours.RESULTS:We found a higher risk of fracture for the glomerular disease cohort compared with the general pediatric cohort in girls only (incidence rate ratio [IRR], 1.6; 95% CI, 1.3 to 1.9). Hip/femur and vertebral fracture risk were increased in the glomerular disease cohort: adjusted IRR was 2.2 (95% CI, 1.3 to 3.7) and 5 (95% CI, 3.2 to 7.6), respectively. For SCFE, the adjusted IRR was 3.4 (95% CI, 1.9 to 5.9). For AVN, the adjusted IRR was 56.2 (95% CI, 40.7 to 77.5).CONCLUSIONS:Children and young adults with glomerular disease have significantly higher burden of skeletal complications than the general pediatric population.
Importance There is limited information on severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) testing and infection among pediatric patients across the United States. Objective To describe testing for SARS-CoV-2 and the epidemiology of infected patients. Design, Setting, and Participants A retrospective cohort study was conducted using electronic health record data from 135 794 patients younger than 25 years who were tested for SARS-CoV-2 from January 1 through September 8, 2020. Data were from PEDSnet, a network of 7 US pediatric health systems, comprising 6.5 million patients primarily from 11 states. Data analysis was performed from September 8 to 24, 2020. Exposure Testing for SARS-CoV-2. Main Outcomes and Measures SARS-CoV-2 infection and coronavirus disease 2019 (COVID-19) illness. Results A total of 135 794 pediatric patients (53% male; mean [SD] age, 8.8 [6.7] years; 3% Asian patients, 15% Black patients, 11% Hispanic patients, and 59% White patients; 290 per 10 000 population [range, 155-395 per 10 000 population across health systems]) were tested for SARS-CoV-2, and 5374 (4%) were infected with the virus (12 per 10 000 population [range, 7-16 per 10 000 population]). Compared with White patients, those of Black, Hispanic, and Asian race/ethnicity had lower rates of testing (Black: odds ratio [OR], 0.70 [95% CI, 0.68-0.72]; Hispanic: OR, 0.65 [95% CI, 0.63-0.67]; Asian: OR, 0.60 [95% CI, 0.57-0.63]); however, they were significantly more likely to have positive test results (Black: OR, 2.66 [95% CI, 2.43-2.90]; Hispanic: OR, 3.75 [95% CI, 3.39-4.15]; Asian: OR, 2.04 [95% CI, 1.69-2.48]). Older age (5-11 years: OR, 1.25 [95% CI, 1.13-1.38]; 12-17 years: OR, 1.92 [95% CI, 1.73-2.12]; 18-24 years: OR, 3.51 [95% CI, 3.11-3.97]), public payer (OR, 1.43 [95% CI, 1.31-1.57]), outpatient testing (OR, 2.13 [1.86-2.44]), and emergency department testing (OR, 3.16 [95% CI, 2.72-3.67]) were also associated with increased risk of infection. In univariate analyses, nonmalignant chronic disease was associated with lower likelihood of testing, and preexisting respiratory conditions were associated with lower risk of positive test results (standardized ratio [SR], 0.78 [95% CI, 0.73-0.84]). However, several other diagnosis groups were associated with a higher risk of positive test results: malignant disorders (SR, 1.54 [95% CI, 1.19-1.93]), cardiac disorders (SR, 1.18 [95% CI, 1.05-1.32]), endocrinologic disorders (SR, 1.52 [95% CI, 1.31-1.75]), gastrointestinal disorders (SR, 2.00 [95% CI, 1.04-1.38]), genetic disorders (SR, 1.19 [95% CI, 1.00-1.40]), hematologic disorders (SR, 1.26 [95% CI, 1.06-1.47]), musculoskeletal disorders (SR, 1.18 [95% CI, 1.07-1.30]), mental health disorders (SR, 1.20 [95% CI, 1.10-1.30]), and metabolic disorders (SR, 1.42 [95% CI, 1.24-1.61]). Among the 5374 patients with positive test results, 359 (7%) were hospitalized for respiratory, hypotensive, or COVID-19-specific illness. Of these, 99 (28%) required intensive care unit services, and 33 (9%) required mechanical ventilation. The case fatality rate was 0.2% (8 of 5374). The number of patients with a diagnosis of Kawasaki disease in early 2020 was 40% lower (259 vs 433 and 430) than in 2018 or 2019. Conclusions and Relevance In this large cohort study of US pediatric patients, SARS-CoV-2 infection rates were low, and clinical manifestations were typically mild. Black, Hispanic, and Asian race/ethnicity; adolescence and young adulthood; and nonrespiratory chronic medical conditions were associated with identified infection. Kawasaki disease diagnosis is not an effective proxy for multisystem inflammatory syndrome of childhood.
BACKGROUND:Clinical data research networks (CDRNs) aggregate electronic health record data from multiple hospitals to enable large-scale research. A critical operation toward building a CDRN is conducting continual evaluations to optimize data quality. The key challenges include determining the assessment coverage on big datasets, handling data variability over time, and facilitating communication with data teams. This study presents the evolution of a systematic workflow for data quality assessment in CDRNs.IMPLEMENTATION:Using a specific CDRN as use case, the workflow was iteratively developed and packaged into a toolkit. The resultant toolkit comprises 685 data quality checks to identify any data quality issues, procedures to reconciliate with a history of known issues, and a contemporary GitHub-based reporting mechanism for organized tracking.RESULTS:During the first two years of network development, the toolkit assisted in discovering over 800 data characteristics and resolving over 1400 programming errors. Longitudinal analysis indicated that the variability in time to resolution (15day mean, 24day IQR) is due to the underlying cause of the issue, perceived importance of the domain, and the complexity of assessment.CONCLUSIONS:In the absence of a formalized data quality framework, CDRNs continue to face challenges in data management and query fulfillment. The proposed data quality toolkit was empirically validated on a particular network, and is publicly available for other networks. While the toolkit is user-friendly and effective, the usage statistics indicated that the data quality process is very time-intensive and sufficient resources should be dedicated for investigating problems and optimizing data for research.
Objective PEDSnet is a clinical data research network (CDRN) that aggregates electronic health record data from multiple children's hospitals to enable large-scale research. Assessing data quality to ensure suitability for conducting research is a key requirement in PEDSnet. This study presents a range of data quality issues identified over a period of 18 months and interprets them to evaluate the research capacity of PEDSnet. Materials and Methods Results were generated by a semiautomated data quality assessment workflow. Two investigators reviewed programmatic data quality issues and conducted discussions with the data partners' extract-transform-load analysts to determine the cause for each issue. Results The results include a longitudinal summary of 2182 data quality issues identified across 9 data submission cycles. The metadata from the most recent cycle includes annotations for 850 issues: most frequent types, including missing data (>300) and outliers (>100); most complex domains, including medications (>160) and lab measurements (>140); and primary causes, including source data characteristics (83%) and extract-transform-load errors (9%). Discussion The longitudinal findings demonstrate the network's evolution from identifying difficulties with aligning the data to a common data model to learning norms in clinical pediatrics and determining research capability. Conclusion While data quality is recognized as a critical aspect in establishing and utilizing a CDRN, the findings from data quality assessments are largely unpublished. This paper presents a real-world account of studying and interpreting data quality findings in a pediatric CDRN, and the lessons learned could be used by other CDRNs.