This study investigates the role of cumulative adverse and positive childhood experiences (ACEs and PCEs) in the development of youth violence using longitudinal data from the UK millennium cohort study (N = 18,282). Findings show a strong association between higher ACE exposure and increased likelihood of assault, weapon involvement, and gang affiliation in adolescence. Conversely, greater exposure to PCEs is associated with reduced youth violence, with evidence that PCEs attenuate the harmful association between ACEs and violence. This study contributes UK-based evidence and highlights the need for early, multi-domain interventions. Targeting both familial risks and extra-familial protective factors could significantly reduce youth violence and potentially improve broader developmental outcomes across adolescence and beyond.
Background Enhancing longitudinal cohort studies by linking routine external data to them is increasingly used to evaluate how local environments impact participants' outcomes (e.g. crime on adolescents' perception of security and victimisation). Objective To describe the geographical linkage between the UK Millennium Cohort Study (MCS) and street-level crime incidents reported to the Police in England and Wales, and to estimate crime count and rates around MCS participants' residences. Methods Eight years of monthly street-level police data were linked to the residential postcodes of MCS participants living in England and Wales in surveys 5, 6 and 7 to create individual-level variables of neighbourhood crime counts and rates (28,724 surveys and 11,365 individuals). Radial buffers around participants' residences were created at ages 11, 14 and 17. Crime counts and rates were created prior to the month of interview (at 1, 3, 6, 9, and 12 months prior). A homogenisation of crime categories reported in the police data was conducted to evaluate changes over time and areas. Multivariate models were used to study the association between MCS participants' demographic characteristics and derived measures of neighbourhood crime. Results While total crime rates and counts around MCS participants remain stable over the period, they hide heterogeneous upward and downward trends in specific sub-categories, with violence and sexual offences showing a larger increase. We observe a negative socioeconomic gradient between household income deciles, recorded at age 11, and subsequent exposure to neighbourhood crime. Conclusion Linking routine crime data to longitudinal studies, such as the MCS, which follow children and their families through a critical period of development, can provide a new resource to understand how local crime impacts child and adolescent outcomes.
Objectives Children with chronic liver disease (CLD) are at increased risk of neurodevelopmental difficulties compared with their healthy peers. However, current evidence is limited to studies of small, single centres focusing predominantly on children after liver transplantation. To address this evidence gap, we evaluated educational outcomes of children with CLD. Methods We used ECHILD, which contains linked, longitudinal administrative health and education records for a national cohort of children born in England. Our study population included pupils born between September 2002 and August 2012 with developmental assessment recorded at age 5. CLD by age 5 was identified from ICD-10 diagnosis codes and OPCS procedure codes recorded since birth. Outcomes, including total point scores, up to seven main areas of child development, and achieving a ‘Good level of development’, were modelled using multivariable regression, adjusting for a range of confounders. Results Of 5,084,585 pupils, 0.2% (N=9,937) developed CLD by age 5, ans 51.2% of children with CLD by age 5 did not achieve a good level of development, compared with 35.4% of those without CLD (risk difference: 15.8, 95% CI [14.7,17.0]; adjusted relative risk: 1.19; 95% CI [1.16,1.22]). Standardised total point scores were 0.35 lower (95% CI [0.41, 0.28]) in 2007/08–2010/11, and 0.19 lower (95% CI [0.21, 0.16]) in 2011/12–2016/17. Children with CLD performed worse than those without CLD in all areas of development. Effect sizes were strongest for the Physical Development domain. Conclusions Children with CLD are at higher risk of poor cognitive development compared to the general population. Our data should be used to inform educational and health services policies of the urgent need for neuro-developmental assessment to be included in the routine care to ensure early detection and referral to specialist services. Further exploration of longitudinal educational outcomes and its relationship with specific liver disease related factors is needed.
Introduction We aimed to generate evidence about child development measured through school attainment and provision of special educational needs (SEN) across the spectrum of gestational age, including for children born early term and >41 weeks of gestation, with and without chronic health conditions. Methods We used a national linked dataset of hospital and education records of children born in England between 1 September 2004 and 31 August 2005. We evaluated school attainment at Key Stage 1 (KS1; age 7) and Key Stage 2 (KS2; age 11) and any SEN by age 11. We stratified analyses by chronic health conditions up to age 2, and size-for-gestation, and calculated population attributable fractions (PAF). Results Of 306 717 children, 5.8% were born <37 weeks gestation and 7.0% had a chronic condition. The percentage of children not achieving the expected level at KS1 increased from 7.6% at 41 weeks, to 50.0% at 24 weeks of gestation. A similar pattern was seen at KS2. SEN ranged from 29.0% at 41 weeks to 82.6% at 24 weeks. Children born early term (37-38 weeks of gestation) had poorer outcomes than those born at 40 weeks; 3.2% of children with SEN were attributable to having a chronic condition compared with 2.0% attributable to preterm birth. Conclusions Children born with early identified chronic conditions contribute more to the burden of poor school outcomes than preterm birth. Evaluation is needed of how early health characteristics can be used to improve preparation for education, before and at entry to school.
We study the relationship between proximity to fast food restaurants and weight gain from late childhood to early adolescence. We use the Millennium Cohort Study, a UK-wide nationally representative longitudinal study, linked with granular geocoded food outlet data to measure the presence of fast food outlets around children's homes and schools from ages 7 to 14. We find that proximity to fast food outlets is associated with increased weight (body mass index, overweight, obese, body fat, weight), but only among those with maternal education below degree level. Within this sample, those with lower levels of emotional regulation are at heightened risk of weight gain.
ObjectivesGovernments often struggle to accurately estimate the number of migrants using public services due to the lack of a unique national ID. We aim to study this in the context of migrant access to immunization programs in Chile and estimate vaccine coverage in school-age children. MethodsTo estimate vaccine coverage for migrant school-age children, we combined data from two databases: the Chilean National Immunization Register (which contained 77.9 million records) and the School Enrollment database (which contained around 68 million records, representing about 3.6 pupils per year). Using Splink, a Python package developed by the UK Ministry of Justice, we created a probability linkage model to link and deduplicate records of migrants who lack a unique national ID. The following linkage keys were considered in the model: first and second name, first and last name and date of birth. Linkage quality was evaluated using ‘gold standard data. ResultsIn 2022, we find that out of 3,644,467 students enrolled in school, 140,317 of them were migrants who didn't have a Chilean national ID. Additionally, in the NIR database, 5.2 out of 77.9 million records belonged to migrants without a national ID. After removing duplicates from both databases, our linkage model determined that 52,524 of the 140,317 students without a national ID in SE were linked to NIR (37.4%). We find that excluding migrants without national IDs when estimating national vaccine coverage for school-aged children leads to an underestimation of 2%, from 86% to 88%. ConclusionOur findings emphasize the significance of utilizing linkage techniques in order to accurately estimate access to public services for migrant populations who typically lack a national ID. By linking their records across public institutions, more reliable data can be obtained.
Data Resource Profile: The Education and Child Health Insights from Linked Data (ECHILD) Database Louise Mc Grath-Lone ,* Nicolás Libuy, Katie Harron , Matthew A Jay , Linda Wijlaars, David Etoori, Matthew Lilliman, Ruth Gilbert, and Ruth Blackburn University College London, Institute of Health Informatics, London, UK, Centre for Longitudinal Studies, University College London, Institute of Education, London, UK and University College London, Great Ormond Street Institute of Child Health, London, UK
Background The COVID-19 pandemic has raised concerns about long-term harms to children’s health and education and re-emphasised how strongly interconnected these domains are in childhood and adolescence. It has also highlighted the need to maximise the utility of administrative datasets (which reflect service provision for the whole population) as an evidence base for policy and practice. To date, technical and governance barriers have limited the potential for wide-scale analyses across health and education. Here, we report linking Hospital Episode Statistics (HES) to the National Pupil Database (NPD), which includes important information on children’s functional health and wellbeing, such as attainment in national exams, special educational needs (SEN) support and absence rates. This newly linked health-education database can generate evidence for paediatricians, policymakers and the public on, for example, educational outcomes for children with rare or common health conditions or how SEN support in schools might improve health outcomes. Objectives Create a de-identified, linked HES-NPD database for all children and young people in England aged 0–24 years who were born on or after 01/09/1995 (the Education and Child Health Insights from Linked Data (ECHILD) Database) Assess linkage quality in the ECHILD Database Methods To create the ECHILD Database, NHS Digital applied multi-step rules-based algorithms to longitudinal records of names, date of birth, gender and postcodes extracted from HES and NPD (to separate them from health- and education-related information). This produced a bridging file of pseudonymised IDs to link extracts of de-identified NPD and HES data (the ECHILD Database). If data linkage is biased (for example, less accurate for ethnic minority groups), then subsequent analyses could underestimate health needs and further entrench disadvantage. We evaluated linkage quality for three academic cohorts born 1st September to 31st August in 1996/7, 1999/00 and 2004/5. Permissions to create the ECHILD Database are described at: https://www.ucl.ac.uk/child-health/echild Results In total, the newly-created ECHILD Database includes de-identified, linked HES-NPD records for approximately 14.7 million individuals. It currently covers a 25-year period (01/09/1995 to 31/03/2020) and will be updated with more recent data as it is available. Our initial assessments indicate high linkage rates, particularly for more recent cohorts. Of pupils born in 2004/05, 99% linked to a HES record and, overall, 96% of pupils linked (1,609,670/1,674,899). Ethnic minority pupils and those living in more deprived areas were less likely to link; however, differences in linked and unlinked pupil characteristics were moderate to small. Throughout childhood, two-thirds of children had at least one admission to hospital (excluding being born in hospital). Conclusions The ECHILD Database enables large-scale, longitudinal research exploring interrelationships between health and education. For example, we are exploring how gestational age at birth relates to attainment and SEN. These results will be useful for policymakers and service providers for estimating future need for SEN support in schools based on the population’s birth characteristics. As more recent data becomes available, the ECHILD Database represents a unique opportunity to explore the impact of recent disruptions to health services on health and educational outcomes for children and young people during and after the COVID-19 pandemic.
Introduction Linkage of administrative data for universal state education and National Health Service (NHS) hospital care would enable research into the inter-relationships between education and health for all children in England. Objectives We aim to describe the linkage process and evaluate the uality of linkage of four one-year birth cohorts within the National Pupil Database (NPD) and Hospital Episode Statistics (HES). Methods We used multi-step deterministic linkage algorithms to link longitudinal records from state schools to the chronology of records in the NHS Personal Demographics Service (PDS; linkage stage 1), and HES (linkage stage 2). We calculated linkage rates and compared pupil characteristics in linked and unlinked samples for each stage of linkage and each cohort (1990/91, 1996/97, 1999/00, and 2004/05). Results Of the 2,287,671 pupil records, 2,174,601 (95%) linked to HES. Linkage rates improved over time (92% in 1990/91 to 99% in 2004/05). Ethnic minority pupils and those living in more deprived areas were less likely to be matched to hospital records, but differences in pupil characteristics between linked and unlinked samples were moderate to small. Conclusion We linked nearly all pupils to at least one hospital record. The high coverage of the linkage represents a unique opportunity for wide-scale analyses across the domains of health and education. However, missed links disproportionately affected ethnic minorities or those living in the poorest neighbourhoods: selection bias could be mitigated by increasing the quality and completeness of identifiers recorded in administrative data or the application of statistical methods that account for missed links. Highlights • Longitudinal administrative records for all children attending state school and acute hospital services in England have been used for research for more than two decades, but lack of a shared unique identifier has limited scope for linkage between these databases. • We applied multi-step deterministic linkage algorithms to 4 one-year cohorts of children born 1 September-31 August in 1990/91, 1996/97, 1999/00 and 2004/05. In stage 1, full names, date of birth, and postcode histories from education data in the National Pupil Database were linked to the NHS Personal Demographic Service. In stage 2, NHS number, postcode, date of birth and sex were linked to hospital records in Hospital Episode Statistics. • Between 92% and 99% of school pupils linked to at least one hospital record. Ethnic minority pupils and pupils who were living in the most deprived areas were least likely to link. Ethnic minority pupils were less likely than white children to link at the first step in both algorithms. • Bias due to linkage errors could lead to an underestimate of the health needs in disadvantaged groups. Improved data quality, more sensitive linkage algorithms, and/or statistical methods that account for missed links in analyses, should be considered to reduce linkage bias.
In The Lancet Digital Health, Hannah Knight and colleagues1Knight HE Deeny SR Dreyer K et al.Challenging racism in the use of health data.Lancet Digit Health. 2021; 3: e144-e146Summary Full Text Full Text PDF PubMed Scopus (14) Google Scholar highlight stages in the data science pipeline that are affected by and lead to racism. Data linkage is a further stage in which ethnic bias can be encoded into datasets. Ethnic bias occurs when linkage error (false or missed matches) is more likely to occur for particular ethnic groups. The problem of ethnic bias in health data linkage is well described in the literature2Bohensky MA Jolley D Sundararajan V et al.Data linkage: a powerful research tool with potential problems.BMC Health Serv Res. 2010; 10: 346Crossref PubMed Scopus (131) Google Scholar and is concerning because health data are widely used for monitoring, service planning, research, evaluation, and policy. Systematic biases in data linkage misestimate health needs for ethnic minorities and further entrench existing disadvantages. Accurate data linkage relies on accurately recorded identifying information and well designed linkage algorithms. However, ethnic minorities are more likely to have missing or incorrect information in their health records,3Hagger-Johnson G Harron K Fleming T et al.Data linkage errors in hospital administrative data when applying a pseudonymisation algorithm to paediatric intensive care records.BMJ Open. 2015; 5e008118Crossref PubMed Scopus (23) Google Scholar which might reflect structural biases in health systems (eg, ethnic minorities are more likely to be treated at health facilities with poorer overall data quality).2Bohensky MA Jolley D Sundararajan V et al.Data linkage: a powerful research tool with potential problems.BMC Health Serv Res. 2010; 10: 346Crossref PubMed Scopus (131) Google Scholar Data capture systems are also typically designed around Western name standards (ie, a first, middle, and last name) and do not account for cultural differences in name structures (eg, Hispanic groups can have multiple first or middle names, and often two surnames, and Asian names can follow different ordering norms). Linkage methods that require exact agreement on names can therefore contribute to ethnic bias. Requiring consent for linkage can also exacerbate bias, since ethnic minorities have higher rates of non-consent for linkage,2Bohensky MA Jolley D Sundararajan V et al.Data linkage: a powerful research tool with potential problems.BMC Health Serv Res. 2010; 10: 346Crossref PubMed Scopus (131) Google Scholar perhaps reflecting lower levels of trust in health systems and how their data are used. Data providers and users should routinely explore ethnic bias by assessing data quality and linkage error.4Gilbert R Lafferty R Hagger-Johnson G et al.GUILD: guidance for information about linking data sets.J Public Health (Oxf). 2018; 40: 191-198Crossref PubMed Scopus (60) Google Scholar Greater transparency of linkage processes, including routine reporting by disaggregated ethnic subgroups, would allow ethnic biases to be accounted for by statistical methods, and considered when assessing the validity of analyses and interpreting results. Data providers need to continually improve data quality and linkage methods (eg, through training of patient-facing staff in recording data for ethnic minorities, more inclusive data capture systems, and more flexible linkage algorithms). For example, we recently showed that, when linking administrative health and education records, relaxing requirements for exact matching on name improved linkage rates for ethnic minorities, although they remained disproportionately low.5Mc Grath-Lone L Blackburn R Gilbert R The Education and Child Health Insights from Linked Data (ECHILD) database: an introductory guide for researchers. University College London, London2021Google Scholar Crucially, echoing Knight and colleagues,1Knight HE Deeny SR Dreyer K et al.Challenging racism in the use of health data.Lancet Digit Health. 2021; 3: e144-e146Summary Full Text Full Text PDF PubMed Scopus (14) Google Scholar we must all strive for greater diversity in the data linkage community, and more meaningful engagement with ethnic minorities to increase understanding of data linkage and address their concerns. We declare no competing interests. Challenging racism in the use of health dataData and data-driven technologies are playing an increasingly influential role in health care, helping to detect disease earlier, move care closer to home, encourage health-promoting behaviours, and improve the efficiency of service delivery. Although data-driven technologies have potential for good, they can also exacerbate existing health inequalities, which are deep-rooted and have been laid bare during the COVID-19 pandemic. In this Comment, we examine how structural inequalities, biases, and racism in society are easily encoded in datasets and in the application of data science, and how this practice can reinforce existing social injustices and health inequalities. Full-Text PDF Open Access
We use longitudinal data across a key developmental period, spanning much of childhood and adolescence (age 5 to 17, years 2006–2018) from the UK Millennium Cohort Study, a nationally representative study with an initial sample of just over 19,000. We first examine the extent to which inequalities in overweight, obesity, BMI and body fat over this period are consistent with the evolution of inequalities in health behaviours, including exercise and healthy diet markers (i.e., skipping breakfast) (n = 7,220). We next study the links between SES, health behaviours and adiposity (BMI, body fat), using rich models that account for the influence of a range of unobserved factors that are fixed over time. In this way, we improve on existing estimates measuring the relationship between SES and health behaviours on the one hand and adiposity on the other. The advantage of the individual fixed effects models is that they exploit within-individual changes over time to help mitigate biases due to unobserved fixed characteristics (n = 6,883). We observe stark income inequalities in BMI and body fat in childhood (age 5), which have further widened by age 17. Inequalities in obesity, physical activity, and skipping breakfast are observed to widen from age 7 onwards. Ordinary Least Square estimates reveal the previously documented SES gradient in adiposity, which is reduced slightly once health behaviours including breakfast consumption and physical activity are accounted for. The main substantive change in estimates comes from the fixed effects specification. Here we observe mixed findings on the SES associations, with a positive association between income and adiposity and a negative association with wealth. The role of health behaviours is attenuated but they remain important, particularly for body fat.