Survey data are self-reported data collected directly from respondents by a questionnaire or an interview and are commonly used in epidemiology. Such data are traditionally collected via a single mode (eg, face-to-face interview alone), but use of mixed-mode designs (eg, offering face-to-face interview or online survey) has become more common. This introduces two key challenges. First, individuals may respond differently to the same question depending on the mode; these differences due to measurement are known as "mode effects." Second, different individuals may participate via different modes; these differences in sample composition between modes are known as "mode selection." Where recognized, mode effects are often handled by straightforward approaches, such as conditioning on survey mode. However, while reducing mode effects, this and other equivalent approaches may introduce collider bias in the presence of mode selection. The existence of mode effects and the consequences of naïve conditioning may be underappreciated in epidemiology. This paper offers a simple introduction to these challenges using directed acyclic graphs by exploring a range of possible data structures. We discuss the potential implications of using conditioning- or imputation-based approaches and outline the advantages of quantitative bias analysIs for dealing with mode effects.
Recent advances in artificial intelligence (AI) - particularly generative AI - present new opportunities to accelerate, or even automate, epidemiological research. Unlike disciplines based on physical experimentation, a sizable fraction of Epidemiology relies on secondary data analysis and thus is well-suited for such augmentation. Yet, it remains unclear which specific tasks can benefit from AI interventions or where roadblocks exist. Awareness of current AI capabilities is also mixed. Here, we map the landscape of epidemiological tasks using existing datasets - from literature review to data access, analysis, writing up, and dissemination - and identify where existing AI tools offer efficiency gains. While AI can increase productivity in some areas such as coding and administrative tasks, its utility is constrained by limitations of existing AI models (e.g. hallucinations in literature reviews) and human systems (e.g. barriers to accessing datasets). Through examples of AI-generated epidemiological outputs, including fully AI-generated papers, we demonstrate that recently developed agentic systems can now design and execute epidemiological analysis, albeit to varied quality (see https://github.com/edlowther/automated-epidemiology). Epidemiologists have new opportunities to empirically test and benchmark AI systems; realising the potential of AI will require two-way engagement between epidemiologists and engineers.
Large language models (LLMs) and agents are now widely used tools in code development, with data typically sent to third-party cloud-based models. Their adoption in research using personal data is constrained by governance requirements that typically prohibit data transmission to external services. Locally deployable open-weight models offer an alternative since sensitive data never leave the local environment. We introduce an open-source framework for evaluating the efficacy of AI agents powered by open-weight LLMs on one of the most persistent bottlenecks in research on longitudinal population studies: data preparation. The framework comprises: a curated ground-truth dataset (cleaning scripts preparing six sweeps of data from a British cohort study), task definitions encompassing tasks such as category harmonization and multi-wave merging, and automated routines for evaluating the LLM-produced R code and outputted data. We benchmark LLMs across the (consumer grade) deployment spectrum to assess their efficacy in 20 data preparation tasks (creation of 102 variables). Current state-of-the-art, 31-35B parameter models almost saturated our benchmark ('average task completion' up to 87.9
Survey data are increasingly collected using mixed-mode designs. However, the measurement of survey items may differ across modes, introducing ‘mode effects’, a type of systematic measurement error which can bias analyses of mixed-mode data. While the theoretical mechanisms giving rise to mode effects have been discussed in detail, the empirical evidence on their occurrence and size is fragmented. In addition, while many existing statistical approaches for handling mode effects require unrealistic assumptions, other more suitable approaches remain underutilised due to the need for external evidence on the magnitude of mode effects. To address this, we conducted a systematic review of the experimental literature on mode effects. We searched multiple bibliographic databases, grey literature sources, and implemented backwards and forwards citation screening. Studies eligible for inclusion were (quasi-)experimental, sampled from the general population (or age-, sex-, region-specific strata), and reported mode effect estimates on item measurement. We extracted comprehensive information relating to the study design, sampling, mode effect estimates, and reporting. Ninety experimental studies published between 1967 and 2024 met the inclusion criteria, which included 4,113 mode effect estimates for 3,545 unique variables in total. Mode effects were generally small, typically below 0.2 SD. However, larger mode effects were more commonly observed when modes differed by interviewer involvement or by question delivery (visual vs aural), as well as for sensitive items (e.g., sexual behaviour, social life), which aligns with pre-existing theory on the causes of mode effects. Generally, where mode effects occur, they are item-, mode-, and population-specific. Reporting quality varied substantially and insufficient details regarding randomisation compliance, non- response, and uncertainty of estimates were common. We collated all mode effect estimates into a free online database and provide a set of recommendations to improve the reporting of future studies.
Should original research articles routinely contain prominent policy claims or broad calls to action? Growing emphasis on research impact might be welcome yet have unintended consequences (eg, incentivising overextrapolation, and undermining the perceived objectivity of scientists). We examined 45 807 abstracts from ten leading Epidemiology and Public Health journals (1990-2024). Using a large language model with human validation, we classified policy claims and mapped trends. Claims markedly increased from 17.6% to 35.8%, with wide variation across countries and journals (>60% vs < 4%). Keywords linked to higher claim rates differed by topic and time: some corresponded to topics with clear causal evidence of harm as well as topics with notable advocacy. Claims were most common in qualitative or cross-sectional studies, and less common in cohort, quasi-experimental, or experimental studies. We argue that these patterns reflect a research culture increasingly oriented toward claiming policy relevance-and incentives that encourage attaching claims to single studies. Our findings raise questions about how scientists and journals balance evidence, advocacy, and scientific credibility. Ensuring that policy claims and calls to action remain commensurate with evidence will be central to building trust as policy impact continues to be incentivised.
Background: Mixed-mode designs can generate bias in all types of analysis of survey data due to systematic differences in how survey participants may respond according to mode – a phenomenon referred to as a mode effect. There is widespread evidence of mode effects upon item measurement (e.g., item means and proportions), but less on associations between survey items, despite estimating associations being a primary use of survey data. Methods: We estimated average differences (means and proportions) in item responses according to mode using data from an experiment in which participants were randomized to respond via web questionnaire, video interview, or in-person interview. We then used these estimates to simulate (counterfactual) data from two web-then-in-person sequential mixed-mode surveys of young adults in the United Kingdom (Next Steps and the Millennium Cohort Study, MCS) as if they were collected using a single mode, web. We estimated associations between variables using this simulated data and compared these against associations obtained using the observed (i.e., factual) data; the difference between the two provides a measure of variation in variable correlations reflecting the mixed-mode survey design. Results: Mode effects were observed for an array of items comparing web with in-person and video modes; differences between video and in-person were more muted. Participants in web were more likely to report negatively valent values for sensitive items, in particular (e.g., poorer physical and mental health, lower life satisfaction, and greater financial difficulty), though effect sizes were generally small with few exceeding 0.2 SD and all below 0.3 SD. The impact of mode effects on variable associations in the mixed-mode surveys was also generally small: a 0.004 SD change in absolute effect sizes on average in both Next Steps and the MCS. Conclusion: Items related to sensitive characteristics showed small mode effects when comparing web and in-person or video interview modes. Mode effects did not appear to translate into substantial changes to associations between variables as assessed in a sequential web-then-in-person mixed-mode survey compared with a counterfactual web-only survey.
Dimensional models of adversity posit partially distinct mechanisms linking dimensions of early-life adversity (ELA) to mental health. Although evidence supports some of these mechanistic differences, there has yet to be a review that centers on development of attention biases to emotional stimuli. This systematic review examined associations between the dimensions of early-life threat and deprivation with attention biases in childhood and adolescence across nine different task types. Articles were identified using six online databases organized using Covidence with our search spanning from inception to June 17, 2025. Forward and backward snowball searches were conducted to identify additional articles. Articles were selected for inclusion if: (1) participants were aged 0–18 years; (2) participants had experienced at least one form of ELA with a non-exposed comparison group (or ELA measured continuously); (3) at least one measure of attention bias indexed; (4) they were written in the English language; and (5) they were published in a peer-reviewed journal. PRISMA guidelines were used to inform data extraction, which was conducted by multiple independent reviewers. A total of 46 articles were identified for extraction. Evidence was heterogeneous and revealed that, across different dimensions of ELA, significant associations tended to be conditional on specific study characteristics such as participant age, stimulus presentations time, and emotion type. Overall, partial support for the dimensional model was found, though reliable patterns were hampered by inconsistencies in study design and sampling. Recommendations for future investigation are provided.
Internalizing symptoms such as depression and anxiety rise dramatically during emerging adulthood. Although both childhood and adulthood adversity are associated with internalizing symptoms during this period, the underlying transdiagnostic processes connecting adversity to internalizing symptoms are unclear. To investigate this, we examined how childhood and adulthood adversity are related to internalizing symptoms during emerging adulthood via individual differences in executive functioning. In a cross-sectional study of 203 participants aged 18-24 years (Mage=20.36, SDage=1.73, 66.01% women/transwomen, 28.08% White-European/North American), lifetime stressor exposure was indexed using the Stress and Adversity Inventory, internalizing symptoms were measured using the Kessler Psychological Distress Scale, and executive functioning was measured as a latent factor indicated by multiple performance-based measures. Analyses were conducted using regression and path analyses. In separate models, both childhood and adulthood adversity exposure predicted greater internalizing symptoms, but when entered as competing predictors in the same model, only adulthood adversity continued to predict psychopathology. No evidence for an indirect association involving executive functioning was detected, although higher childhood adversity was associated with better executive functioning. These results highlight differential associations between adversity timing, internalizing symptoms, and executive functioning during emerging adulthood, with minimal evidence of executive functioning as an indirect path linking adversity to internalizing symptoms.
Survey data are increasingly collected using mixed-mode designs (e.g. personal interview and web questionnaire). Little is known, however, about the extent to which this introduces bias in subsequent analyses, or whether simply including mode as a covariate addresses it. Using data simulations, we identified the conditions under which mode effects (mode measurement differences) and mode selection (respondent differences) introduce bias, complementing this with empirical illustrations using mixed-mode data from the 1958 National Child Development Study. In simulations, absent mode selection, substantial bias arose only from unusually large mode effects, but was amplified when mode selection was present. Controlling for mode under strong mode selection introduced substantial bias, including sign-reversal of the estimate. Bias was more pronounced when the sample was split equally between modes. The direction and size of bias will depend on the direction of all effects. In our empirical illustration, the results were largely unchanged after controlling for mode, possibly reflecting weaker mode selection and mode effects in these data. These findings highlight that, while possible, substantial bias from mixed-mode designs may be unlikely in many practical settings. However, the risk of bias, and the appropriate strategy to address it, should be considered on an analysis-specific basis.
Obesity is a highly heritable trait, but rising obesity rates suggest environmental change is also of profound importance. We conducted a cross-cohort analysis to examine how associations between genetic risk for high BMI and observed BMI differed in four British birth cohorts born before and amidst the obesity epidemic (1946, 1958, 1970 and ~2001; N = 19,379). BMI (kg/m 2 ) was measured at multiple time points between ages 3 and 69 years. We used polygenic indices (PGI) derived from GWAS of adulthood and childhood BMI, respectively, with mixed effects models used to estimate associations with mean BMI and quantile regression used to assess associations across the distribution of BMI. We further used linear regression to estimate PGI-heritability (PGI-h 2 ; incremental variance explained by the PGI) and Genomic Relatedness Restricted Maximum Likelihood (GREML) to calculate SNP-heritability (SNP-h 2 ) by cohort and age. Adulthood BMI PGI was associated with BMI in all cohorts and ages but was more strongly associated with BMI in more recently born generations, e.g., at age 16y, a 1 SD increase in the adulthood PGI was associated with 0.46 kg/m 2 (0.37, 0.55) higher BMI in the 1946c and 0.90 kg/m 2 (0.83, 0.97) higher BMI in the 2001c. Cross-cohort differences widened with age and were larger at the upper end of the BMI distribution, indicating disproportionate increases in obesity in more recent generations for those with higher PGIs. Differences were also observed when using the childhood PGI. There were no clear, consistent differences in PGI-h 2 or SNP-h 2 , possibly due to limited statistical power, except that PGI-h 2 was highest in the most recently born cohort (2001c) when using the most predictive PGI for adulthood BMI. Findings highlight how the environment can modify genetic associations; genetic associations with BMI differed by birth cohort, age, and outcome centile.
With the rise of web-based data collection, researchers are increasingly exploring how to measure cognition online, including in web surveys and video interviews. This paper provides empirical evidence on mode effects (web, in-person, video interviewing) in the measurement of cognitive abilities using the TestMyBrain Backward Digit Span test (an immediate serial recall task) among adults age 20-40 years in England. Participants completed surveys including the same cognitive test at two time points two weeks apart and were randomly allocated to mode at each wave. The 1,692 wave 1 respondents and 1,510 wave 2 respondents were analysed using mixed effects models to estimate mode effects relating to several assessment outcomes. We found that participation by web was associated with a higher likelihood of an incomplete assessment relative to both in-person and video interviewing, suggesting that the presence of an interviewer – even if not physically present as in video interviewing – enhances the likelihood of a successfully completed test. For participants who completed the cognitive assessment, we found no evidence that the observed mean scores differed by mode, though this appeared to mask some heterogeneity. Although based on small numbers, there was a trend for both web and video respondents to be more likely to obtain the highest test scores than those participating in-person. This is consistent with prior suggestions that individuals in modes where there is limited or no interviewer presence are more likely to cheat. Mean test scores were consistently higher at wave 2 than at wave 1, with this ‘retest gain’ differing little by mode. These findings contribute to the limited literature on mode effects in cognitive assessment and offer valuable insights for researchers designing and analysing mixed-mode studies.
Social scientists have long sought to investigate whether the predictors of educational attainment (EA) have changed across time. Here, we provide insights by incorporating genetic predictors of education in three nationally representative British birth cohorts born in 1946, 1958, and 1970. We investigated whether individual characteristics as proxied by polygenic indexes (PGIs) for EA and cognition have become more relevant to educational success over time and whether returns to genetic predisposition were moderated by early life socioeconomic background. We present three findings. First, associations between the EA PGI and attainment increased over time, with increasing incremental variance explained by the EA PGI. Second, associations between the cognition PGI and attainment were broadly consistent across cohorts, and there was no clear change in explained variance. Since the EA PGI captures multiple traits related to educational success, factors other than those related to cognition may have become more relevant over time. Third, we observed strong evidence of interaction: Associations between the EA PGI and EA were disproportionately larger among those from more advantaged socioeconomic backgrounds. The strength and pattern of associations varied when using EA PGIs that were less conservatively filtered for SNPs. Our findings suggest EA is influenced by social and genetic factors both independently and jointly. Genetic liability and social background could be considered as two forms of inherited advantage which synergistically influence education attainment.
Birth cohort studies have a rich history of contributing to science across disciplinary fields, notably health and social sciences. Here, we introduce a curated resource comprising genomic data from five British birth cohort studies: longitudinal studies with extensive data collected prospectively across life, each deliberately sampled to be nationally representative (born 1946 to 2001). These contain health and social data from birth to older age, enabling longitudinal and cross-cohort genetically informed research. The Millennium Cohort Study additionally includes data on parents and offspring, enabling within-family analyses. Across five cohorts born in 1946, 1958, 1970, 1989/90, and 2000/2002, 27,432 participants have harmonized, imputed, and quality-controlled genetic data from genotyping arrays covering 6.7 million common SNPs. The Millennium Cohort Study contains over 6,000 mother-offspring pairs and over 3,000 mother-father-offspring trios. Pseudonymized data are freely available to the global research community upon approval of a data access request (https://cls.ucl.ac.uk/data-access-training). ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement The Centre for Longitudinal Studies is funded by the Economic and Social Research Council (grant numbers ES/M001660/1 and ES/W013142/1). DB, LW and NMD are supported by the Medical Research Council (MR/V002147/1). NMD is supported via a Norwegian Research Council Grant (295989) and the UCL Division of Psychiatry (https://www.ucl.ac.uk/psychiatry/division-psychiatry). NC and AW are supported by the Medical Research Council (MR/Y014022/1). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Ethical approval was obtained in each study: 1946c (North Thames Multicentre Research Ethics Committee: reference 98/2/121 and 07/H1008/168), 1958c (South East Multi-centre Research Ethics Committee: ref 01/1/44), 1970c (South East Coast Brighton & Sussex: ref 15/LO/1446), 1989c (East of England Cambridge Central Research Ethics Committee: ref 22/EE/0052), 2001c (London-Central REC: 13/LO/1786). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Pseudonymised data are freely available to the research community upon approval of a data application request.
Should original research routinely contain prominent policy claims, such as recommendations for policymakers or broad calls to action? Growing emphasis on research impact might be welcome but also have unintended consequences that include risks of overextrapolation, the blurring of roles between scientists and advocates, and potential erosion of scientific credibility. To inform this debate, we examined 45,807 abstracts from ten leading epidemiology and public health journals (1990-2024). Using a large language model with human validation, we classified policy claims and mapped their prevalence by time, country, journal, field of study, and study design. Policy claims markedly increased in frequency from 17.6% in 1990-1999 to 35.8% in 2020-2024, with wide variation across countries (>40% in Italy and Australia vs <19% in Norway and Japan, 2020-2024) and journals (>60% in some vs <6% in others). Keywords linked to higher claim rates differed by topic and time: some corresponded to topics with clear causal evidence (tobacco), others to topics with more complex causal evidence and notable researcher advocacy (health inequalities and COVID-19). Claims were most common in qualitative or cross-sectional studies, and less common in cohort, quasiexperimental, or experimental studies. We argue that these patterns reflect a research culture increasingly oriented toward claiming policy relevance, and incentives that encourage attaching claims to single studies. Our findings raise questions about how scientists and journals balance evidence, advocacy, and credibility. Ensuring that policy claims remain commensurate with evidence will be central to build trust as policy impact continues to be incentivised. We discuss alternative ways for researchers and publishers to engage meaningfully with policy beyond attaching claims to individual studies, and share our data and scripts to catalyse further work in this area. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement DB and LW are funded by the Economic and Social Research Council (ES/W013142/1) and the UKRI Digital Research Infrastructure (DRI) Programme (UKRI/ST/B000295/1). NMD is supported by the Norwegian Research Council 295989. EC and MW are supported by a European Research Council Starting Grant (UKRI guarantee, EP/Y010345/1). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Data are available at https://github.com/dbann/policyclaims
BACKGROUND AND AIMS:Tobacco smoking has declined dramatically in many high-income countries over the past seventy years. Studies that have mapped this trend have relied on repeat cross-sectional or retrospectively measured smoking data, which have limitations regarding accurate measurement, inclusion of early smokers, and capturing of within-person change over time. Here, we introduce a new resource detailing harmonisable smoking data in four British birth cohort studies spanning 1946-2018 and use these data to document age and cohort changes in smoking. DESIGN:Longitudinal data from four nationally representative British Birth Cohort Studies, born 1946, 1958, 1970 and 2000/02, respectively. SETTING:Great Britain. PARTICIPANTS:50 942 participants were eligible for inclusion in this study (5362 in the 1946c, 16 178 in the 1958c, 16 036 in the 1970c, and 13 366 in the 2001c). Data collection spanned the years 1946-2018. MEASUREMENTS:Prevalence of daily smoking and cigarettes smoked per day were measured prospectively at various points across the life course via self-report. FINDINGS:The prevalence of smoking and the average number of cigarettes smoked by daily smokers declined between each successive cohort. At age 42/43y, prevalence of daily smoking was 33.6% (95% confidence interval [CI] = 31.8%, 35.5%) in the 1946c, 27.3% (95% CI = 26.5%, 28.2%) in the 1958c, and 22.1% (95% CI = 21.3%, 22.8%) in the 1970c. Males smoked more and with greater intensity than females, on average, though sex differences were smaller in latter cohorts. Within a cohort, the prevalence and intensity of smoking peaked in early adulthood (< age 30y) and declined thereafter; participants who continued to smoke daily smoked fewer cigarettes as they grew older. CONCLUSIONS:In Great Britain, smoking prevalence and cigarette consumption appear to have declined substantially between cohorts born across the latter half of the twentieth and early twenty-first centuries. The British Birth Cohorts represent a unique and largely underutilized resource for investigating trends in smoking across life (prenatal to old age) and by year of birth (1946-2001), including changes in the determinants, correlates, and consequences of smoking. We provide syntax and information on items on smoking in these cohorts to catalyse future research, also available at: https://cls-data.github.io/smoking-in-the-cohorts/.
Objectives:To investigate whether adolescent social connections influence body mass index (BMI) trajectories into adulthood and explore whether associations are moderated by gender, ethnicity or age. Methods:Data came from 17,719 American adolescents in grades 7-12 at baseline (1994-95) from the National Longitudinal Study of Adolescent to Adult Health. Growth curve models tested associations between baseline social connections and BMI trajectories from waves II-V including interactions for gender, ethnicity and age. Results:Stronger peer connections were associated with flatter BMI trajectories. For example, BMI for those with high peer contact was 0.79 kg/m2 lower [95% CI -1.20, -0.38] 22 years after baseline, compared to those with low contact. Stronger family connections were associated with steeper trajectories. For example, BMI for those with high family contact was 0.52 kg/m2 higher [95% CI 0.01, 1.02] 22 years after baseline, compared to those with low contact. Discussion:Among adolescents, stronger peer connections were associated with flatter BMI trajectories and stronger family connections with steeper trajectories. Promotion of peer-based interventions could be explored as a strategy to promote healthy weight trajectories.
BACKGROUND:Children with obesity are more likely to have parents with obesity than those without. Several environmental explanations have been proposed for this correlation, including foetal programming and parenting practices. However, body mass index (BMI) is a heritable trait; child-parent correlations may reflect direct inheritance of adiposity-related genes. There is some evidence that mothers' BMI associates with offspring BMI net of direct genetic inheritance, consistent with both intrauterine and parenting effects, but this requires replication. Here, we also investigate the role of fathers' BMI as well as offsprings' diet as a mediating factor. METHODS:We used Mendelian Randomization (MR) with genetic trio (mother-father-offspring) data from 2,630 families in the Millennium Cohort Study, a UK birth cohort study of individuals born in 2000/02, to examine the association between parental BMI (kg/m2) and offspring birthweight and BMI and diet measured at six-time points between ages 3y and 17y. Paternal and maternal BMI were instrumented with polygenic indices (PGI) for BMI conditioning upon offspring PGI. This allowed us to separate direct and indirect ("genetic nurture") genetic effects. We compared these results with associations obtained using standard multivariable regression techniques using phenotypic BMI data only. RESULTS:Mothers' and fathers' BMI were positively associated with offspring BMI to similar degrees. However, in MR analysis, associations between father's BMI and offspring BMI were close to the null. In contrast, mother's BMI was consistent in MR analysis with phenotypic associations. Maternal indirect genetic effects were between 25-50% the size of direct genetic effects. There was limited and inconsistent evidence of associations with offspring diet and some evidence that mothers', but not fathers', BMI was related to birthweight in both MR and multivariable regression models. CONCLUSIONS:Results suggest maternal BMI may be particularly important for offspring BMI: associations may arise due to both direct transmission of genetic effects and indirect (genetic nurture) effects. Associations of father's and offspring adiposity that do not account for direct genetic inheritance may yield biased estimates of paternal influence. Larger studies are required to confirm these findings.
Background: Analyses of sequential mixed-mode survey data can be biased due to non-random selection of participants into mode. Understanding the consequence of this bias requires knowledge of the predictors of mode selection. However, research on these predictors is sparse, in contrast with research on the predictors of survey non-response. The ‘continuum of resistance’ model of survey response predicts that delayed responders – who, in sequential mixed-mode surveys, appear in later offered modes – and non-responders share similar characteristics. If the model, which is testable in longitudinal data, is correct, this would suggest that research on non-response could be generalized to understand mode selection.Methods: We used data from a major UK birth cohort study (the 1958 National Child Development Study) which embedded a sequential web-then-telephone mixed-mode survey at the age 55y sweep to assess whether (a) descriptively, late (i.e. telephone) and non-responders share similar characteristics and (b) whether predictions from models of non-response are accurate when used to predict telephone response. For (a), we calculated univariate descriptive statistics and performed cluster analysis to compare an array of participant characteristics across response groups (web, telephone and non-response). For (b), we estimated random forest models for non-response and telephone response (conditional on response) and compared two metrics of predictive accuracy (Area Under the Receiver Operating Characteristic curve [AUC ROC] and Brier scores) when using the non-response model to instead predict telephone response, against predictions from the model generated for telephone response, specifically.Results: Telephone and non-respondents were similar on almost all (measured) characteristics, and dissimilar in most regards to web respondents. Predictions from non-response models had similar predictive accuracy to predictions from models trained on telephone response, specifically – for instance, AUC ROC values in hold-out samples not used to train models of 0.72 (95% CI = 0.70, 0.74) and 0.74 (95% CI = 0.72, 0.75), respectively.Conclusions: The characteristics of late- and non-responders in a sequential (web-then-telephone) mixed-mode survey were very similar, consistent with the ‘continuum of resistance’ model of survey response. This suggests that research on non-response could transport to understanding mode selection in sequential mixed-mode surveys, though replications in other surveys with different mixed-mode designs (e.g., modes adopted and their order) is required.
Socioeconomic inequalities in cardiovascular disease (CVD) persist in high-income countries despite marked overall declines in CVD-related morbidity and mortality. After decades of research, the field has struggled to unequivocally answer a crucial question: is the association between low socioeconomic position (SEP) and the development of CVD causal? We review relevant evidence from various study designs and disciplinary perspectives. Traditional observational, family-based and Mendelian randomization studies support the widely accepted view that low SEP causally influences CVD. However, results from quasi-experimental and experimental studies are both limited and equivocal. While more experimental and quasi-experimental studies are needed to aid causal understanding and inform policy, high-quality descriptive studies are also required to document inequalities, investigate their contextual dependence and consider SEP throughout the lifespan; no simple hierarchy of evidence exists for an exposure as complex as SEP. The COVID-19 pandemic illustrates the context-dependent nature of CVD inequalities, with the generation of potentially new causal pathways linking SEP and CVD. The linked goals of understanding the causal nature of SEP and CVD associations, their contextual dependence, and their remediation by policy interventions necessitate a detailed understanding of society, its change over time and the phenotypes of CVD. Interdisciplinary research is therefore key to advancing both causal understanding and policy translation.