Clustering of common causes by geography (confounding via geography) can induce bias if ignored. Analyses using polygenic indicators (PGI) for mental health traits are traditionally adjusted for genetic principal components (PCs), but not local area. In this study, we investigate whether accounting for geography more explicitly helps ameliorate such confounding. Using UK Biobank data (N=209,391 to 293,851), we construct PGIs across a range of p-value thresholds and contrast single-level regression with multilevel Mundlak models for the relationship between genetically predicted mental health and greenspace, whilst adjusting for 25 genetic PCs. Our formulation decomposes single level estimates into within-area, between-area and contextual estimates of mental health PGIs on two greenspace measures. Single-level models for the least filtered PGIs find genetically predicted depression and schizophrenia associate with lower greenspace, while genetically predicted wellbeing associates with higher greenspace. However stricter p-value selection on PGI exposures results in estimates largely attenuating and, for schizophrenia, changing sign entirely. Within-area estimates of less-filtered PGIs align more closely with strictly filtered PGIs. Differences appear driven by between-area effects, with single-level models underperforming particularly in the most urbanised areas, and National Parks. We advocate for investigating the inclusion of local context as a routine sensitivity analysis for genetic epidemiologists working with spatially structured populations. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement This work was supported in part by the UK Medical Research Council Integrative Epidemiology Unit at the University of Bristol (Grant ref: MC\_UU\_00032/1 and MC\_UU\_00032/7). GJG was supported by an ESRC Postdoctoral Fellowship and MQ Fellows Award (Grant ref: ES/T009101/1 and MQF22\22). MRM is supported by the National Institute for Health Research Bristol Biomedical Research Centre. OSPD is funded by the Alan Turing Institute under the EPSRC grant EP/N510129/1. MRM and OSPD were also supported by the National Institute for Health Research (NIHR) Biomedical Research Centre at the University Hospitals Bristol NHS Foundation Trust and the University of Bristol. TTM is funded by the ESRC (ES/W013142/1). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The UK Biobank has approval from the North West Multi-centre Research Ethics Committee as a research tissue bank. This means that researchers can operate under this approval with no need for further ethical approval, other than exceptions such as re-contact applications. This RTB approval was granted initially in 2011 and it is renewal on a five-yearly cycle: we successfully applied to renew it in 2016 and 2021. UK Biobank will apply for renewal effective in 2026. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes UK Biobank data are available through a procedure described at: https://www.ukbiobank.ac.uk/enable-your-research. [https://github.com/ZoeReed/Greenspace\_MH\_PGS_analysis][1] MQ: Transforming Mental Health, https://ror.org/04g6r1b21, MQF22/22 Medical Research Council, https://ror.org/03x94j517, MC\_UU\_00032/1, MC\_UU\_00032/7 Economic and Social Research Council, https://ror.org/03n0ht308, ES/W013142/1 The Alan Turing Institute, EP/N510129/1 NIHR Bristol Biomedical Research Centre, https://ror.org/02mtt1z51 [1]: https://github.com/ZoeReed/Greenspace_MH_PGS_analysis
It has been suggested that active use of social media (i.e. direct interactions with other users) is associated with improved mental health. Active use could help build relationships, and lead to feelings of social and emotional support. However, evidence is inconsistent, potentially because active use comprises many different behaviours. The impact of active use may depend on whether individuals are targeting specific users and the feedback they receive from others.Our study explored the relationship between the public active use Twitter users performed and received, and mental health outcomes. We linked Twitter data from 2020 to 2022 to self-reported measures of mental health from the Avon Longitudinal Study of Parents and Children. We generated variables capturing the type of active use each participant performed, and the feedback they received from other users. We produced models predicting mental health from these variables, using data from 310 adults.We found little to no evidence that engaging in more targeted interactions or receiving greater amounts of feedback on Twitter was linked to better mental health outcomes. Likewise, there was limited evidence that the amount of feedback received influenced the relationship between targeted interactions and mental health. We also found no evidence that posting more nontargeted Tweets (i.e. those not directed at any specific user) was related to mental health. However, this relationship was influenced by the amount of feedback received. As feedback increased, the relationship between nontargeted Tweets and mental wellbeing became more negative.Given that Twitter is primarily designed for information sharing purposes, social interactions on this platform may often result in weak relationships that do not provide meaningful social or emotional benefit.
Mental health can influence both the intensity and dynamics of emotion expression. For example, persistent and intense negative emotions are symptoms of depression and anxiety. Such patterns could reflect maladaptive or impaired emotion regulation. Researching the relationships between the dynamics of emotion expression and mental health can improve our understanding of the experiences that characterize these conditions and help inform their prevention and treatment. Many previous studies rely on self-reports of emotion, a limitation that could be addressed by assessing emotion expression in social media posts. We used Twitter (now "X") data, from 2020 to 2022, and gold-standard questionnaire measures of mental health from 230 adult participants in a U.K. longitudinal study to explore the relationships between the dynamics of emotion expression in tweets and mental health. We compared results generated using three different sentiment analysis methods. We found evidence that posting tweets that expressed more positive and less negative emotions was associated with reduced symptoms of anxiety. Expressing positive emotions at a greater variability was also associated with reduced anxiety symptoms in our participants. There was much less evidence that variability in negative emotions and instability in any emotion were associated with mental health. Mood disorders, such as anxiety, may be characterized by more negative emotions and a reduced ability to respond to positive internal and external stimuli. (PsycInfo Database Record (c) 2026 APA, all rights reserved).
The publisher regrets that this article has been temporarily removed. A replacement will appear as soon as possible in which the reason for the removal of the article will be specified, or the article will be reinstated.The full Elsevier Policy on Article Withdrawal can be found at https://www.elsevier.com/about/policies-and-standards/article-withdrawal.
It has been suggested that use of social media late at night could lead to worse mental health outcomes. We linked Twitter (‘X’) data to self-reported measures of mental health from the Avon Longitudinal Study of Parents and Children. We aimed to predict these measures from the average hour a participant posted Tweets. We used data from 310 adult participants and 18,288 Tweets. We found strong evidence the average hour participants posted Tweets was associated with depressive symptoms, anxiety symptoms, and mental wellbeing. Average hour explained almost 2% of the variation in mental wellbeing, comparable to reports of the impact of binge drinking and exercise. Participants who, on average, Tweeted through the night (23:00 to 05:00) showed meaningfully worse mental wellbeing than those who Tweeted during the daytime. Although the average hour that a participant posted their Tweets explained less of the variation in their depressive (0.2%) and anxiety (0.7%) symptoms, after stratification by age and sex these relationships became stronger. Our results could inform behavioural interventions for improving the safety of social media platforms.
Background:Linking digital footprint data into longitudinal population studies (LPS) presents an opportunity to enrich our understanding of how digitally captured behaviours relate to health traits and disease. However, this linkage introduces significant methodological challenges that require systematic exploration. Objectives:To develop a robust framework for successful digital footprint linkage into LPS, informed by discussions from a workshop from the Digital Footprints Conference 2024. Methods:We propose a structured, four-stage framework to facilitate successful linkage of digital footprint data into LPS: (1) understand participant expectations and acceptability; (2) collect and link the data; (3) evaluate properties of the data; and (4) ensure secure and ethical access for research. This framework addresses the key methodological challenges identified at each stage, discussed through the lens of two LPS case studies: the Avon Longitudinal Study of Parents and Children and Generation Scotland. Results:Key methodological challenges identified include privacy and confidentiality concerns, reliance on third-party platforms, data quality issues like missing data and measurement error. We also emphasize the role of trusted research environments and synthetic datasets in enabling secure, privacy-sensitive data sharing for research. Conclusions:While the linkage digital footprint data to LPS remains in early stages, our framework provides a methodological foundation for overcoming current challenges. Through iterative refinement of these methods there is significant potential to advance population-level insights into health and wellbeing.
Associations between mental health and residential characteristics of a person’s local area are well established in epidemiological literature. However, the clustering of common causes by geographical context (hereafter referred to as confounding via geography), may induce bias into associations for these relationships. Such confounding via geography is also likely to influence genetic exposures of interest, such as polygenic indices (PGI) for mental health. In this study, we investigate the extent to which this may have an impact on PGI analyses by considering within-area and contextual effects. Data from UK Biobank was used to examine this (N=209,391 to 293,851). We conducted analyses contrasting single-level linear models with multilevel Mundlak models to estimate single-level, and geographically decomposed within-area and contextual effects between PGI for depression, wellbeing and schizophrenia and two greenspace outcomes (one measuring greenspace via land use and one measuring greenness via satellite spectroscopy). We used UK Census geography membership as our area level specification. Analyses were conducted with PGI derived at different value thresholds to assess effects for more and less strict p-value thresholds. Single-level and within-area estimates were more similar for the most strictly filtered PGI exposures, with single-level effects between genetically predicted depression (at p<5e-8) and lower percentage greenspace (-0.004, 95%CI -0.007 to -0.00023), genetically predicted wellbeing (at p<0.1) and higher percentage greenspace 0.007, 95%CI 0.003 to 0.01) and genetically predicted schizophrenia (p<0.05) and lower percentage greenspace (-0.005, 95%CI -0.008 to -0.0006). Similar, but weaker effects were observed with the greenness outcome. At more relaxed p-value thresholds discordance between the within-area estimate and the single-level estimate is larger, with the within-area estimate being consistent with the more strictly selected PGI. Contextual effects also differed and, although have wide confidence intervals, they appear to in-part drive the single-level effects. We see that the areas in which single-level models perform worst are heavily urbanised areas (e.g., central London) and areas in National Parks (e.g. Peak District and North Yorkshire Moors). This study demonstrates that accounting for confounding via geography meaningfully alters associations between mental health PGI and greenspace outcomes. Given these results, we propose that genetic epidemiologists examining relationships in a spatially clustered population should consider routinely adopting a sensitivity analysis accounting for local context, and consider exploring within- and between-area decomposition of effects.
Introduction & Background To use digital footprint data for mental health and well-being research we often need to collect concurrent, high-quality measures of ground truth. Delivering frequent surveys to participants using an ecological momentary assessment (EMA) methodology is one way to collect such data. However, existing surveys tend to be long, not focused on momentary states or rely on rating images which are not platform agnostic. Here we present a five-item test-based survey designed with participants and validated for use in EMA studies to collect data about momentary changes in mood. We describe its methodological development and how it has been used to investigate music listening on Spotify as a digital footprint of mood. Objectives & Approach The survey is based on the circumplex model of affect. It was co-produced with a participant advisory group (N=5), who gave feedback on the length, content and delivery of the survey. It was then piloted in a group of N=98 participants to assess statistical validity, and congruence with the 20-item Positive and Negative Affect Schedule (PANAS). Following this it was delivered in a wider sample (N=150) four times a day over a two-week period using an EMA app on participant’s phones. Relevance to Digital Footprints EMA is an increasingly popular method for collecting ground truth to support the interpretation of digital footprint data. This newly developed and tested mood survey offers an opportunity to reduce participant burden for collecting mood data in EMA studies which will support the collection of high quality and high time-resolution ground truth for digital footprints research. Results Together with participants we selected four emotions across the axes of arousal and valence, as well as rumination which participants considered important in their music listening behaviors. Factor analysis of pilot data showed that the questions represented two factors of positive and negative affect. The ratings on a 0-10 scale of the emotions ‘cheerful’ and ‘relaxed’ explained 44% of the variance in positive affect, and ratings of ‘worried’, ‘sad’ and ‘frustrated’ explained 40% of the variance in negative affect. Delivery of the questionnaire in a wider student sample (N=150) four times per day for two weeks allowed for the opportunity to assess typical response rates in a realistic EMA setting. On average participants completed 3 out of the 4 surveys a day. Conclusions & Implications The co-created, short mood survey for the collection of ground truth in digital footprint studies was validated across two independent samples, and shown to allow for good response rates in a two week study. Future testing on wider samples will provide opportunities to validate the survey and assess its effectiveness across demographic groups and different sample types.
The way in which socioeconomic status (SES) moderates the etiology of reading attainment has been explored many times, with past work often finding that genetic influences are suppressed under conditions of socioeconomic deprivation and more fully realized under conditions of socioeconomic advantage: a gene–SES interaction. Additionally, past work has pointed toward the presence of gene–location interactions, with the relative influence of genes and environment varying across geographic regions of the same country/state. This study investigates the extent to which SES and geographical location interact to moderate the genetic and environmental components of reading attainment. Utilizing data from 2,135 twin pairs in Florida (mean age 13.82 years, range 10.71–17.77), the study operationalized reading attainment as reading comprehension scores from a statewide test and SES as household income. We applied a spatial twin analysis procedure to investigate how twin genetic and environmental estimates vary by geographic location. We then expanded this analysis to explore how the moderating role of SES on said genetic and environmental influences also varied by geographic location. A gene–SES interaction was found, with heritability of reading being suppressed in lower- (23%) versus higher-SES homes (78%). The magnitude of the moderating parameters were not consistent by location, however, and ranged from −0.10 to 0.10 for the moderating effect on genetic influences, and from −0.30 to 0.05 for the moderating effect on environmental influences. For smaller areas and those with less socioeconomic variability, the magnitude of the genetic moderating parameter was high, giving rise to more fully realized genetic influences on reading there. SES significantly influences reading variability. However, a child's home location matters in both the overall etiology and how strongly SES moderates said etiologies. These results point toward the presence of multiple significant environmental factors that simultaneously, and inseparably, influence the underlying etiology of reading attainment.
Educational attainment is associated with a range of positive outcomes, yet its impact on wellbeing is unclear, and complicated by high correlations with intelligence. We use genetic and observational data to investigate for the first time, whether educational attainment and intelligence are causally and independently related to wellbeing. Results from our multivariable Mendelian randomisation demonstrated a positive causal impact of a genetic predisposition to higher educational attainment on wellbeing that remained after accounting for intelligence, and a negative impact of intelligence that was independent of educational attainment. Observational analyses suggested that these associations may be subject to sex differences, with benefits to wellbeing greater for females who attend higher education compared to males. For intelligence, males scoring more highly on measures related to happiness were those with lower intelligence. Our findings demonstrate a unique benefit for wellbeing of staying in school, over and above improving cognitive abilities, with benefits likely to be greater for females compared to males.
Introduction & Background An estimated 4.95 billion people used social media in 2023, with the average user active on around seven platforms for over two hours per day. This widespread use leads to abundant digital footprint data around interactions with social media. These data can be collected continuously and reflect real behaviour of users in naturalistic settings. These strengths have led researchers to propose the use of social media data in digital phenotyping, where digital footprints can be used to quantify and predict health conditions. Mental health assessment in particular could benefit, as existing approaches, such as self-report questionnaires and inpatient assessment, are unable to perform the real-time monitoring that digital phenotyping could potentially achieve. Digital phenotyping models for mental health require careful consideration of what aspects of social media data to include. Including all data users generate could result in models that are overfitted and difficult to explain. Studies are required that explore the relationship between specific aspects of social media data, such as the time course of expressed emotion, and gold-standard measures of mental health. Objectives & Approach With participants’ consent, we linked Twitter data to self-reported measures of mental health from the Avon Longitudinal Study of Parents and Children. We performed sentiment analysis using three different approaches—LIWC, VADER and RoBERTa—to estimate the amount, variability and instability of positive and negative emotional content in each participant’s Tweets over a one-year period. We explored the association between these measures of emotion expression and self-reported scores of depressive symptoms, anxiety symptoms and wellbeing. These mental health measures are the Short Mood and Feelings Questionnaire, the Generalized Anxiety 7 and the Warwick Edinburgh Mental Wellbeing Scale. Relevance to Digital Footprints Our research is highly relevant to digital footprint research, as it involves the use of digital footprint data (i.e. Twitter data) to predict mental health outcomes. Conclusions & Implications The results of our analysis will inform the development of digital footprint based phenotyping for mental health that could one day provide information to supplement clinical assessments.
BackgroundThe genetic and environmental aetiology of autistic and Attention Deficit Hyperactivity Disorder (ADHD) traits is known to vary spatially, but does this translate into variation in the association of specific common genetic variants?MethodsWe mapped associations between polygenic scores for autism and ADHD and their respective traits in the Avon Longitudinal Study of Parents and Children (N = 4,255–6,165) across the area surrounding Bristol, UK, and compared them to maps of environments associated with the prevalence of autism and ADHD.ResultsOur results suggest genetic associations vary spatially, with consistent patterns for autistic traits across polygenic scores constructed at different p‐value thresholds. Patterns for ADHD traits were more variable across thresholds. We found that the spatial distributions often correlated with known environmental influences.ConclusionsThese findings shed light on the factors that contribute to the complex interplay between the environment and genetic influences in autistic and ADHD traits.
This data note describes the collection and linkage of participants' Twitter data as a digital phenotype in the Avon Longitudinal Study of Parents and Children (ALSPAC) multi-generational birth cohort study. Twitter (renamed X in 2023) is a social media platform based around a micro-blog format. Digital phenotyping represents a novel opportunity for cohort studies to collect data with a low participant burden, and outside of discrete measurement periods. The ALSPAC governance framework supports the ethical consenting, storage and sharing of social media data, and linking Twitter data with wider cohort data provides opportunities to assess Twitter data quality concerns in a research context. All adults currently participating in ALSPAC (N=26,205) were invited to take part, which included the index cohort and their parents. N=3,247 indicated that they were Twitter users, 26% of these (N=835) consented and 19% (N=623) had their data successfully linked. Data were collected using our open-source software, Epicosm in February 2023. Approximately two thirds of the linked Twitter cohort are from the index cohort generation, and the remainder from the parent generation. In general, linked participants are representative of the general ALSPAC cohort, with the exception of having slightly higher educational attainment. This is consistent with previous research into the demographics of Twitter users. Overall the linked dataset contains 1,488,517 posts (tweets) from between 2008 and 2023, with 27% of these being 'retweets'. The available data includes information derived from a range of commonly used sentiment scoring algorithms, type of tweet, public metrics such as likes and retweets, and the time and date of the tweet. Controls are in place to maintain the anonymity of cohort participants, and data linkage is managed by ALSPAC’s data linkage team to reduce disclosure risk. This ensures high standards of data security and ethical use of social media data.
How DNA is folded and packaged in nucleosomes is an essential regulator of gene expression. Abnormal patterns of chromatin folding are implicated in a wide range of diseases and disorders, including epilepsy and autism spectrum disorder (ASD). These disorders are thought to have a shared pathogenesis involving an imbalance in the number of excitatory-inhibitory neurons formed during neurodevelopment; however, the underlying pathological mechanism behind this imbalance is poorly understood. Studies are increasingly implicating abnormal chromatin folding in neural stem cells as one of the candidate pathological mechanisms, but no review has yet attempted to summarise the knowledge in this field. This meta-synthesis is a systematic search of all the articles on epilepsy, ASD, and chromatin folding. Its two main objectives were to determine to what extent abnormal chromatin folding is implicated in the pathogenesis of epilepsy and ASD, and secondly how abnormal chromatin folding leads to pathological disease processes. This search produced 22 relevant articles, which together strongly implicate abnormal chromatin folding in the pathogenesis of epilepsy and ASD. A range of mutations and chromosomal structural abnormalities lead to this effect, including single nucleotide polymorphisms, copy number variants, translocations and mutations in chromatin modifying. However, knowledge is much more limited into how abnormal chromatin organisation subsequently causes pathological disease processes, not yet showing, for example, whether it leads to abnormal excitation-inhibitory neuron imbalance in human brain organoids.
Introduction & Background Social media use has been proposed as a cause of worsening mental health and wellbeing over the last decade, but its role in mitigating some of the effects of social distancing during the pandemic showed that it also has the potential to improve these outcomes. Whilst existing research disagrees on the degree to which social media use harms or helps, there is growing consensus around the need to move from global measures of social media use to specific measures of types of social media use. These new measures can enable an exploration of proposed mechanisms and causal pathways linking social media use and mental health and wellbeing. A commonly proposed mechanism is nighttime social media use reducing sleep quality, and consequently harming mental health and wellbeing. Objectives & Approach We aimed to investigate the relationships between the time Twitter users post content and their mental health, wellbeing and sleep quality using direct measurements of Twitter use linked to standardised mental health measures in a well-characterized cohort. This study uses approximately 1.5 million Tweets harvested between January 2008 and March 2023 from 622 participants in the Avon Longitudinal Study of Parents and Children (ALSPAC). These Tweets have been linked to questionnaire data collected on six occasions spanning April 2019 to May 2021. These questionnaires included standard measures of depressive symptoms, anxiety symptoms, mental wellbeing and difficulty sleeping. We have taken two approaches to explore these relationships, using circular statistical methods novel to social media data analysis to account for day/night cycles. The first approach used mixed effect models to investigate the association between the time a Tweet was posted and the mental health, mental wellbeing and sleep quality of the poster. The second approach explored the relationships between the mean hour participants post Tweets in a given time period, and their mental health, mental wellbeing and sleep quality. Relevance to Digital Footprints This research is highly relevant to Digital Footprints, due to its use of data directly extracted from a social media site. The methodologies employed in analysing this alongside more traditional epidemiological survey data provides an example of how digital footprint data can complemented by high quality ground truths. Results There was evidence that the timing of Twitter activity was predictive of the mental wellbeing and sleep quality of participants, even after adjustment for demographic, educational and socio-economic covariates. However, the hour a Tweet was posted at explained very little of the variation in the mental wellbeing or sleep quality of the participant who posted it (0.1% and less than 0.1% respectively). There was weak to no evidence that the timing of Twitter activity was predictive of the depressive and anxiety symptoms of participants. Conclusions & Implications Whilst this study found evidence that the hour participants post on Twitter is predictive of their mental wellbeing and sleep quality, the amount of variation explained by these models suggests that this is not a clinically relevant risk factor. This study supports arguments in the literature that the use of social media has a very small and insignificant effect on mental health, wellbeing and sleep quality.
Abstract Motivation Social media represent an unrivalled opportunity for epidemiological cohorts to collect large amounts of high-resolution time course data on mental health. Equally, the high-quality data held by epidemiological cohorts could greatly benefit social media research as a source of ground truth for validating digital phenotyping algorithms. However, there is currently a lack of software for doing this in a secure and acceptable manner. We worked with cohort leaders and participants to co-design an open-source, robust and expandable software framework for gathering social media data in epidemiological cohorts. Implementation Epicosm is implemented as a Python framework that is straightforward to deploy and run inside a cohort’s data safe haven. General features The software regularly gathers Tweets from a list of accounts and stores them in a database for linking to existing cohort data. Availability This open-source software is freely available at [https://dynamicgenetics.github.io/Epicosm/].
Introduction & BackgroundHealth research using digital footprint data often involves the collection and use of large datasets that contain deeply personal information to make inferences about the course and onset of illness. In this context, innovating responsibly is essential for the field to develop safe, trustworthy and, ultimately, ethical research. The inherent interdisciplinarity of digital footprints research can be a challenge to this aim, with different fields having different ethical norms and standards. As well as this, there has been a strong focus to date on traditional ethical issues such as privacy, which do not necessarily account for the breadth of issues that arise in data science and internet-based work. Objectives & ApproachData Hazards is an open-source project that aims to provide a controlled vocabulary of ethical risks (Data Hazards) that can arise from data science research and its implementation. This vocabulary is presented as a set of 11 Hazard labels (v1.0) each with a visual icon and a set of safety precautions. Over three events in 2021-2022 we invited feedback from researchers who volunteered to take part in a Data Hazards workshop (N=15). They varied from PhD students to professors and worked across a range of disciplines, and were asked to discuss the case of mental health prediction from Twitter. Relevance to Digital FootprintsSince digital footprint technologies have great potential to pave the way for earlier and more personal medical treatment, it is important for researchers to be able to innovate whilst considering and communicating risk. We can then collaborate to establish effective safety precautions that allow us to maintain research momentum, without compromising safety or trust. ResultsBased on discussion at the workshops and surveys completed by participants, four main Data Hazards were raised for consideration by the digital footprint research community. These were: 'Lack of Community Involvement' relating to the need to further involve those with lived experience in the development of new technologies; 'Reinforces Existing Bias' due to the potential for automated labelling of ground-truth data to bias training datasets; 'Privacy' given the potential disclosure of sensitive information without consent; and 'Danger of Misuse' due to strong potential for malicious use of such technologies. Other considerations included the potential psychological risk to those labelling suicide and self-harm content with limited support. Conclusions & ImplicationsThe Data Hazards identified provide a means of communicating and clarifying ethical concerns so that they can be more easily addressed in this complex and multidisciplinary field. Further collaboration by the research community to develop and agree appropriate safety precautions would help to build trust in these new technologies before they are deployed in practice.
BACKGROUND:The way in which socioeconomic status (SES) moderates the etiology of reading attainment has been explored many times, with past work often finding that genetic influences are suppressed under conditions of socioeconomic deprivation and more fully realized under conditions of socioeconomic advantage: a gene-SES interaction. Additionally, past work has pointed toward the presence of gene-location interactions, with the relative influence of genes and environment varying across geographic regions of the same country/state. METHOD:This study investigates the extent to which SES and geographical location interact to moderate the genetic and environmental components of reading attainment. Utilizing data from 2,135 twin pairs in Florida (mean age 13.82 years, range 10.71-17.77), the study operationalized reading attainment as reading comprehension scores from a statewide test and SES as household income. We applied a spatial twin analysis procedure to investigate how twin genetic and environmental estimates vary by geographic location. We then expanded this analysis to explore how the moderating role of SES on said genetic and environmental influences also varied by geographic location. RESULTS:A gene-SES interaction was found, with heritability of reading being suppressed in lower- (23%) versus higher-SES homes (78%). The magnitude of the moderating parameters were not consistent by location, however, and ranged from -0.10 to 0.10 for the moderating effect on genetic influences, and from -0.30 to 0.05 for the moderating effect on environmental influences. For smaller areas and those with less socioeconomic variability, the magnitude of the genetic moderating parameter was high, giving rise to more fully realized genetic influences on reading there. CONCLUSIONS:SES significantly influences reading variability. However, a child's home location matters in both the overall etiology and how strongly SES moderates said etiologies. These results point toward the presence of multiple significant environmental factors that simultaneously, and inseparably, influence the underlying etiology of reading attainment.
IntroductionDigital footprint records -- the tracks and traces amassed by individuals as a result of their interactions with the internet, digital devices and services -- can provide ecologically valid data on individual behaviours. These could enhance longitudinal population study databanks; but few UK longitudinal studies are attempting this. When using novel sources of data, study managers must engage with participants in order to develop ethical data processing frameworks that facilitate data sharing whilst safeguarding participant interests. ObjectivesThis paper aims to summarise the participant involvement approach used by the ALSPAC birth cohort study to inform the development of a framework for using linked participant digital footprint data, and provide an exemplar for other data linkage infrastructures. MethodsThe paper synthesises five qualitative forms of inquiry. Thematic analysis was used to code transcripts for common themes in relation to conditions associated with the acceptability of sharing digital footprint data for longitudinal research. ResultsWe identified six themes: participant understanding; sensitivity of location data; concerns for third parties; clarity on data granularity; mechanisms of data sharing and consent; and trustworthiness of the organisation. For cohort members to consider the sharing of digital footprint data acceptable, they require information about the value, validity and risks; control over sharing elements of the data they consider sensitive; appropriate mechanisms to authorise or object to their records being used; and trust in the organisation. ConclusionRealising the potential for using digital footprint records within longitudinal research will be subject to ensuring that this use of personal data is acceptable; and that rigorously controlled population data science benefiting the public good is distinguishable from the misuse and lack of personal control of similar data within other settings. Participant co-development informs the ethical-governance framework for these novel linkages in a manner which is acceptable and does not undermine the role of the trusted data custodian.