Missing data arising from sweep non-response is a major challenge in longitudinal cohort studies, threatening statistical power and the validity of inferences. In the UK Millennium Cohort Study (MCS), non-response has increased substantially from sweep 1 (9 months old) to sweep 7 (17 years old), underscoring the need for robust strategies to handle non-response. We applied a systematic, data-driven approach to identify predictors of non-response at each sweep of the MCS, drawing on all available survey data at the time of analysis. The strongest and most consistent predictor of non-response was prior sweep non-response. Additional robust predictors included lower parental occupational social class, parental non-participation in the latest general elections, parent not being in paid work, higher cohort member’s age and lower cognitive test scores. We then evaluated whether incorporating the identified predictors of non-response as auxiliary variables in multiple imputation (MI) or as covariates in inverse probability weighting (IPW) improved sample representativeness. Validation analyses, using both external benchmarks (2021 Census) and internal comparisons to known early-life distributions, showed that MI and IPW models including the identified predictors substantially reduced or eliminated bias in key variables such as housing tenure and parental social class. Our findings demonstrate that the use of systematically identified auxiliary variables can improve the validity of inferences drawn from the MCS. The resulting predictor set offers a practical resource for applied researchers using MCS data and provides a replicable framework for addressing sweep non-response in other longitudinal studies.
BACKGROUND:One in seven households in England live in accommodation not meeting housing quality standards. Low-quality housing is linked to adverse child health, but less is known about the relationship with educational outcomes. This study evaluated the relationship between housing quality, school absences and educational attainment. METHODS:Data were drawn from the Millennium Cohort Study, a nationally representative cohort of children born in 2000/2002. Housing quality at age 7 years was computed from six indicators: accommodation type, floor level, access to a garden, damp, heating and overcrowding. Percentage of missed school sessions and standardised test scores in Maths and English at age 7, 11 and 16 were linked from the National Pupil Database. Confounder-adjusted linear regressions with survey weights were fitted. RESULTS:Approximately 16% of children lived in lower quality housing (ie, disadvantage in ≥2 conditions); after confounder adjustment, these children had 0.74% (or 1.4 days) more absences per year than those living in higher quality housing (n=7272, 95% CI 0.34% to 1.13%). Damp, overcrowding and accommodation type were the strongest predictors of absence. Test scores in Maths and English across compulsory schooling were between 0.07 and 0.13 SD lower for children living in lower versus higher quality housing (n=6741), mainly driven by overcrowding and lack of central heating. CONCLUSION:Children living in homes with lower quality housing conditions missed 15.5 days more of school throughout compulsory schooling and performed worse on national tests than those in higher quality housing. Targeting specific housing conditions, such as damp and overcrowding, could be beneficial for children's school outcomes.
This study investigates the role of cumulative adverse and positive childhood experiences (ACEs and PCEs) in the development of youth violence using longitudinal data from the UK millennium cohort study (N = 18,282). Findings show a strong association between higher ACE exposure and increased likelihood of assault, weapon involvement, and gang affiliation in adolescence. Conversely, greater exposure to PCEs is associated with reduced youth violence, with evidence that PCEs attenuate the harmful association between ACEs and violence. This study contributes UK-based evidence and highlights the need for early, multi-domain interventions. Targeting both familial risks and extra-familial protective factors could significantly reduce youth violence and potentially improve broader developmental outcomes across adolescence and beyond.
Mental ill-health often emerges during adolescence, and the environments in which children grow up may shape this risk. Still, evidence is limited to single environmental exposures, urban samples, and short follow-ups. We investigated how built, social and chemical environments across childhood relate to common mental illness in adolescence, and whether associations differ by population density. Data were drawn from the Millennium Cohort Study, a nationally representative cohort born in 2000/2002 in the UK. High-resolution environmental data was linked to home addresses at birth and ages 3, 5, 7, 11, 14 and 17. Psychological distress (Kessler-6) and doctor diagnosed depression or serious anxiety were assessed at age 17. Confounder-adjusted and weighted generalized linear mixed models were fitted for single and domain-specific exposures. At age 17, ∼16% of participants had high psychological distress and ∼ 10.5% had been diagnosed with depression or anxiety (n = 7769-8374). Built and social environment from early childhood onwards (5 year: OR = 1.22 [95% CI: 1.07-1.39]; 7 year: OR = 1.16 [1.02, 1.33]; 11 year: OR = 1.15 [1.01-1.30]; 17 year: OR = 1.22 [1.05-1.39]; accumulation: OR = 1.21 [1.05-1.39])-especially lower greenness, greater distance to green space, more grey space, higher area deprivation and crime-were associated with high psychological distress, and, to a lesser extent, with diagnosis. Living close to the sea was associated with higher likelihood of common mental disorder, particular their diagnosis. Findings on air pollution were inconclusive. Associations with built and social environments were stronger in rural areas. Built and social environments in childhood and adolescence were significant correlates of adolescent common mental disorders. Future studies and interventions should consider urban/rural differences and timing of exposure.
Investigating the relationship between self-reported mental health and secondary care utilisation can provide evidence on the link between population-level common mental conditions and clinical care; however, cohort studies with linked administrative data are rare. We explored the link between self-reported mental health in adolescence and mental health-related hospital attendance in young adulthood. Data from a nationally representative English cohort (Next Steps) were linked to NHS Hospital Episode Statistics. GHQ-12 assessed psychological distress in Next Steps at age 15; participants were followed up until their first mental health-related hospital presentations and outpatient treatments or were censored at the end of the study (age 27). Cox proportionate hazard models with survey weights estimated associations. Out of 4058 young people, 19
Polygenic indexes (PGIs) - DNA-based predictors of individual phenotypes - have become essential tools across biomedical and social sciences. We introduce Version 2 of the Polygenic Index Repository, which expands phenotype coverage from 47 to 61, increases the number of participating datasets from 11 to 20, and adopts a more consistent and improved methodology for PGI construction. For 16 phenotypes, we leverage summary statistics from an updated GWAS meta-analysis with greater statistical power compared to the original release, thereby improving the PGI's predictive power. To improve power for family-based analyses, we provide imputed parental PGIs in all datasets with first-degree relatives and offer a framework for interpreting results from analyses that control for parental PGIs. We illustrate the utility of parental PGIs using two applications: (1) comparing PGI associations with and without parental PGI controls for all phenotypes in two Repository datasets with family data, and (2) for BMI and diastolic blood pressure, exploring the contribution of causal versus non-causal components of PGI associations to the imperfect portability of PGIs across subgroups within a genetic ancestry. Collectively, the updates enhance predictive performance, broaden the Repository's scope, and introduce novel resources that reduce confounding bias and improve interpretability.
OBJECTIVE:To explore how birth weight and size-for-gestation may contribute to school absences and educational attainment and whether there are different associations across sex and income groups. DESIGN:Longitudinal linked cohort study. METHODS:Data were drawn from the Millennium Cohort Study, a nationally representative cohort of children born in 2000-2001; percentage of authorised and unauthorised absences from Year 1 to Year 11, and Key Stage test scores at ages 7, 11 and 16 in English and Maths were linked from the National Pupil Database. Birth outcomes and covariates were derived from the 9-month survey, and linear regressions with complex survey weights were fitted. RESULTS:Being born small-for-gestational-age (vs average-for-gestational-age) was associated with an increase of 0.47%, 0.55% and 0.40% in authorised absences in Years 1, 3 and 4 (n=6659) and with a reduction of 0.16-0.26 SD in all English and Maths test scores (n=6204). Similar associations were found for birth weight. After adjusting for prior test scores, English (b=0.07) and Maths (b=0.05) performance at age 11 remained associated with birth weight. Socioeconomic status modified the associations: there were larger disparities in test scores among higher-income families, suggesting that higher income did not compensate for being born small-for-gestational-age. CONCLUSION:Children born smaller missed slightly more classes (~1 day per year) during primary school and had lower English and Maths performance across compulsory education. Exploring specific health conditions and understanding how education and health systems can work together to support children may help to reduce the burden.
Objectives Our study has three objectives: (a) to document the educational achievement gap between children from richer and poorer backgrounds; (b) to show how much of the achievement gap is explained by school quality, parental investment and children’s own investment; and (c) to estimate the effect of parental investments, and children’s activities on achievement. Methods Using linked data from MCS and NPD at age 7, 11 and 16, we estimate the test score gap between children from advantaged and disadvantaged backgrounds. We use socio-economic status measured by eligibility to free school meals and family income. First, we estimate the test score gap between the two groups controlling for predetermined characteristics. Then, we progressively add the measures of school quality; parental time and material factors; and children’s time investments. We compare the coefficients that include school, parent and child factors with the ones estimated without to understand the role each played in explaining the achievement gap. Results At age 7 children from poorer families score 0.58 of a standard deviation (SD) less in maths than their richer peers. When we control for school quality and parental investments, the gap declined to 0.35 SD- implying these factors account for 40% of the gap. Similarly, by age 16, poorer children score 1.33 of a grade less in their average attainment 8 GCSEs. Controlling for school quality, and parental and children investments reduced the gap to 0.82 of a grade. These factors explain about 38% of the attainment gap. We also find that school quality, educational activities by parents and children significantly and positively related with achievements. Children’s activities in unorganised leisure and parent’s leisure intensive material investments are negatively correlated with achievements. Conclusion Understanding the determinants of human capital inequality is very important to identify the role of different stakeholders and to design relevant policies. In this study, we document the role of schools, parents and children themselves in driving achievements through primary school outcomes to high-stakes exams.
Despite policy efforts, early childhood inequalities have remained relatively stable in England over recent decades. In this study, we aimed to understand the determinants of early childhood inequality and how these have changed in their composition and impact on early child development, using two nationally representative cohorts of pre-school children collected in 2004 (N = 8,990) and 2014 (N = 3,877). Child development was assessed in both cohorts at age three, using the British Ability Scales – Naming Vocabulary test, and developmental inequality was depicted by concentration indices across quintiles of the English Index of Multiple Deprivation. A between-sample test of the concentration index revealed there was no real change to developmental inequality over time. However, decomposition of the concentration index revealed there were some changes to the determinants of inequality. Most notably, higher rates of maternal education were associated with reduced inequality, while living in the most deprived areas was associated with increased inequality, helping to explain the overall net zero change over time. Other factors, such as attending early childhood education and care by age three, made little contribution to developmental inequality, despite much higher attendance in 2014. This study highlights the complexity of factors contributing to early childhood inequality and includes a discussion on where policy efforts may be useful going forward.
Objectives One in seven households in England lives in homes that do not meet housing standards, and poor-quality housing is associated with worse health outcomes among children. This study aimed to explore the relationship between housing quality, school absences, and educational attainment during compulsory schooling. Methods The Millennium Cohort Study is a nationally representative cohort of children born in 2000/2001 in the UK. Housing quality was assessed using six indicators—accommodation type, floor level, garden access, damp and mould, heating, and overcrowding—reported by the main caregiver at age 7. Educational administrative data were linked from the National Pupil Database, providing information on the percentage of missed sessions from Year 1 to Year 11, as well as Maths and English exam scores in Key Stage (KS) 1, KS2, and KS4. Confounder-adjusted linear regression models with complex survey weights were conducted. Results Poorer housing quality was associated with higher school absences across the 11 years of compulsory schooling (β = 0.24, 95% CI: 0.11, 0.38; n = 7272), with associations observed for both authorized and unauthorized absences. Similarly, poorer housing quality was linked to lower Maths (KS1: -0.03 [95% CI: -0.05, 0.00]; KS2: -0.04 [95% CI: -0.07, -0.01]; KS4: -0.04 [95% CI: -0.06, -0.01]) and English (KS1: -0.03 [95% CI: -0.05, 0.00]; KS4: -0.04 [95% CI: -0.06, -0.01]) exam scores (n = 6741). Indicator-specific analyses suggested that damp and condensation, accommodation type, and overcrowding contributed to higher school absences, while overcrowding was associated with lower test performance. Conclusions Children living in overcrowded homes, in flats or flat shares, and in accommodations with damp and condensation, missed school more often and performed worse in high stake exams. Future research should explore how specific housing policies could benefit child health and educational attainment.
Air pollution is associated with health in childhood. However, there is limited evidence on sensitive periods during the first 18 years of life. Data were drawn from the Millennium Cohort Study, a large and nationally representative cohort born in 2000/2002. Self-reported general health was assessed at age 17; number of hospital records were derived from linked health data (Hospital Episode Statistics) for consented participants. Residential history was linked to 25 × 25 m grid resolution annual PM2.5, PM10 and NO2 maps between 2000 and 2019; year-specific air pollution exposure in 200-m buffers around postcode centroids were computed. After adjusting for individual and time-variant area-level confounders, children exposed to higher air pollution in early (2-4 y) (n = 9137; PM2.5: OR = 1.06, 95% CI: 1.01-1.11; PM10: OR = 1.05, 95% CI: 1.01-1.09; NO2: OR = 1.01, 95% CI: 1.00-1.02) and middle childhood (5-7) (n = 9171; PM2.5: OR = 1.04, 95% CI: 1.00-1.07; PM10: OR = 1.03, 95% CI: 1.01-1.06) reported worse general health at age 17. Higher PM2.5 and NO2 exposure in adolescence increased the number of hospital episodes in young adulthood. Individuals from non-White and disadvantaged backgrounds were exposed to higher levels of air pollution. Air pollution in early and middle childhood might contribute to worse general health, with ethnic minority and disadvantaged children being more exposed.
Background:The study examines the adolescent developmental outcomes in education, mental health, and physical health of children born to teenage mothers at the start of the millennium. Objective:It aims to understand the extent to which long-term developmental outcomes of children born to adolescent mothers are due to selection effects versus other factors. Methods:It uses longitudinal data from the UK Millennium Cohort Study. Multivariate regressions examine the extent to which the association between maternal age at birth and adolescent outcomes is explained by selection into teenage motherhood, and how the relationship is mediated by the early environment and maternal behaviours. Results:Teenage mothers are disadvantaged in terms of their backgrounds, and their children faced more adversity in their early environment. An unadjusted comparison shows that their adolescent offspring have lower academic achievement, and are more likely to be overweight or obese, but there are no differences in their socio-emotional adjustment. The 'penalty' from teenage motherhood in excess weight is due to negative selection into teenage motherhood. However, the differences in educational attainment of adolescents born to teenage and older mothers reflect both pre-childbearing selection and differences in the child's early environment. A decomposition analysis shows that maternal age accounts for only a low proportion of the variance in adolescent development. Contribution:The study provides the first evidence on long-term outcomes of children born to teenage mothers for the UK. It studies the entire range of key developmental outcomes. It uses a novel decomposition to examine the relative importance of different variables for explaining variation in the outcomes of interest.
Abstract Background When collecting data from human participants, it is often important to minimise the length of questionnaire-based measures. This makes it possible to ensure that the data collection is as engaging as possible, while it also reduces response burden, which may protect data quality. Brevity is especially important when assessing eating disorders and related phenomena, as minimising questions pertaining to shame-ridden, unpleasant experiences may in turn minimise any negative affect experienced whilst responding. Methods We relied on item response theory to shorten three eating disorder and body dysmorphia measures, while aiming to ensure that the information assessed by the scales remained as close to that assessed by the original scales as possible. We further tested measurement invariance, correlations among different versions of the same scales as well as different measures, and explored additional properties of each scale, including their internal consistency. Additionally, we explored the performance of the 3-item version of the modified Weight Bias Internalisation Scale and compared it to that of the 11-item version of the scale. Results We introduce a 5-item version of the Eating Disorder Examination Questionnaire, a 3-item version of the SCOFF questionnaire, and a 3-item version of the Dysmorphic Concern Questionnaire. The results revealed that, across a sample of UK adults (N = 987, ages 18–86, M = 45.21), the short scales had a reasonably good fit. Significant positive correlations between the longer and shorter versions of the scales and their significant positive, albeit somewhat weaker correlations to other, related measures support their convergent and discriminant validity. The results followed a similar pattern across the young adult subsample (N = 375, ages 18–39, M = 28.56). Conclusions These results indicate that the short forms of the tested scales may perform similarly to the full versions.
Birth cohort studies involve repeated surveys of large numbers of individuals from birth and throughout their lives. They collect information useful for a wide range of life course research domains, and biological samples which can be used to derive data from an increasing collection of omic technologies. This rich source of longitudinal data, when combined with genomic data, offers the scientific community valuable insights ranging from population genetics to applications across the social sciences. Here we present quality-controlled whole exome sequencing data from three UK birth cohorts: the Avon Longitudinal Study of Parents and Children (8,436 children and 3,215 parents), the Millenium Cohort Study (7,667 children and 6,925 parents) and Born in Bradford (8,784 children and 2,875 parents). The overall objective of this coordinated effort is to make the resulting high-quality data widely accessible to the global research community in a timely manner. We describe how the datasets were generated and subjected to quality control at the sample, variant and genotype level. We then present some preliminary analyses to illustrate the quality of the datasets and probe potential sources of bias. We introduce measures of ultra-rare variant burden to the variables available for researchers working on these cohorts, and show that the exome-wide burden of deleterious protein-truncating variants, S het burden, is associated with educational attainment and cognitive test scores. The whole exome sequence data from these birth cohorts (CRAM & VCF files) are available through the European Genome-Phenome Archive, and here we provide guidance for their use.
Birth cohort studies involve repeated surveys of large numbers of individuals from birth and throughout their lives. They collect information useful for a wide range of life course research domains, and biological samples which can be used to derive data from an increasing collection of omic technologies. This rich source of longitudinal data, when combined with genomic data, offers the scientific community valuable insights ranging from population genetics to applications across the social sciences. Here we present quality-controlled whole exome sequencing data from three UK birth cohorts: the Avon Longitudinal Study of Parents and Children (8,436 children and 3,215 parents), the Millenium Cohort Study (7,667 children and 6,925 parents) and Born in Bradford (8,784 children and 2,875 parents). The overall objective of this coordinated effort is to make the resulting high-quality data widely accessible to the global research community in a timely manner. We describe how the datasets were generated and subjected to quality control at the sample, variant and genotype level. We then present some preliminary analyses to illustrate the quality of the datasets and probe potential sources of bias. We introduce measures of ultra-rare variant burden to the variables available for researchers working on these cohorts, and show that the exome-wide burden of deleterious protein-truncating variants, S het burden, is associated with educational attainment and cognitive test scores. The whole exome sequence data from these birth cohorts (CRAM & VCF files) are available through the European Genome-Phenome Archive, and here we provide guidance for their use.
Decades of research shows that sexual minority youth (SMY) display heightened risk for mental health problems, although the onset of such disparities remains unclear. The Millennium Cohort Study is the largest nationally representative longitudinal study of adolescents in the United Kingdom. In this study, participants (N = 10,047, 50% female) self-reported their sexual identity at age 17 and had parent-reported mental health data, from the Strengths and Difficulties Questionnaire, reported across five waves at ages 5, 7, 11, 14, and 17. Multilevel linear spline models, stratified by sex, were used to examine mental health trajectories between sexual identity groups (completely heterosexual, mostly heterosexual, SMY). SMY showed heightened peer problems from the baseline assessment at age five, increasing over time, and heightened emotional problems from age 11, increasing over time. Mostly heterosexual youth showed heightened emotional problems at age 11 in males, and at age 17 in females. Findings are discussed in light of the literature on minority stress and gender conformity in youth. The use of parent-reported mental health data means that estimates are likely to be conservative. We conclude that interventions supporting SMY should start early and be available throughout adolescence.
Immunotherapy has revolutionised cancer therapy but current immune checkpoint inhibitors (ICI) produce low response rates in most cancers, indicating that new therapeutic options are needed. Conventional immune-oncology (IO) discovery uses preclinical models with limited translation capacity as they do not fully recapitulate human tumour complexity. We use multimodal patient molecular data with modern machine learning (ML) methods to identify new IO targets with improved clinical potential.
Despite the pervasiveness of cyber crime victimisation, knowledge is limited regarding the prevalence, characteristics and pathways of offenders. The present study examines predictors of self-reported engagement in cyber crime in middle adolescence in a large (N=13,277) longitudinal dataset from the UK Millennium Cohort Study. We adopted an ecological systems approach to examine a range of multicausal, intersecting factors across individual, familial, psychosocial and environmental systems. The overall prevalence of self-reported cyber offending (account hacking or the deployment of viruses) was 5.6% at age 14 and 3.8% at age 17, although persistence over time by the same individuals was relatively low (1.1%). Significant predictors of cyber offending at age 17 were being male, domestic violence between parents, low parental monitoring, low wellbeing, self-harm, exclusion from school, spending more time online gaming, participating in offline leisure activities, and engaging in serious violence (weapon carrying or use), assault, and cyber crime at age 14. Findings indicate that young cyber offenders are often males and those who have experienced a range of risk factors that are connected to poorer wellbeing and engaging in multiple risky/offending behaviours. Implications for theory, policy and practice are discussed.
Background Enhancing longitudinal cohort studies by linking routine external data to them is increasingly used to evaluate how local environments impact participants' outcomes (e.g. crime on adolescents' perception of security and victimisation). Objective To describe the geographical linkage between the UK Millennium Cohort Study (MCS) and street-level crime incidents reported to the Police in England and Wales, and to estimate crime count and rates around MCS participants' residences. Methods Eight years of monthly street-level police data were linked to the residential postcodes of MCS participants living in England and Wales in surveys 5, 6 and 7 to create individual-level variables of neighbourhood crime counts and rates (28,724 surveys and 11,365 individuals). Radial buffers around participants' residences were created at ages 11, 14 and 17. Crime counts and rates were created prior to the month of interview (at 1, 3, 6, 9, and 12 months prior). A homogenisation of crime categories reported in the police data was conducted to evaluate changes over time and areas. Multivariate models were used to study the association between MCS participants' demographic characteristics and derived measures of neighbourhood crime. Results While total crime rates and counts around MCS participants remain stable over the period, they hide heterogeneous upward and downward trends in specific sub-categories, with violence and sexual offences showing a larger increase. We observe a negative socioeconomic gradient between household income deciles, recorded at age 11, and subsequent exposure to neighbourhood crime. Conclusion Linking routine crime data to longitudinal studies, such as the MCS, which follow children and their families through a critical period of development, can provide a new resource to understand how local crime impacts child and adolescent outcomes.