Abstract Background There is currently no proven treatment to prevent the growth or rupture of abdominal aortic aneurysm (AAA). Metformin, a commonly prescribed drug for type 2 diabetes, has been linked in observational studies to a reduced risk of AAA, although causality has not been established. We investigated this association in an observational analysis using the UK Biobank cohort, and tested for a causal relationship using Mendelian randomisation (MR), a genomic approach that infers causality from genetic variation. Methods Logistic regression was used to test the association between AAA and self-reported metformin use in 2972 AAA cases and 89 160 propensity matched controls from UK Biobank. MR analyses were conducted using genetic variants at putative metformin targets to proxy drug effects. Results In the observational analysis, metformin was significantly associated with lower AAA risk (OR 0.49, 95% c.i.: 0.41–0.59, P = 3.15x10−14). MR analysis supported a causal effect, estimating a 43% risk reduction per standard deviation (OR = 0.57, 95% c.i.: 0.38–0.88 P = 0.010) per standard deviation decrease in HbA1c via metformin gene targets, equivalent to the effect of a prescribed dose of metformin. Conclusions Both observational and MR analyses found strong evidence that metformin use reduces the risk of developing AAA, supporting our findings that there is evidence of a potential causal association. Clinical trials are warranted to assess the efficacy of metformin to reduce the risk of aneurysm growth and rupture in people with AAA.
Early temperament, such as socio-emotional development and activity level, varies widely, yet its underlying biological associations are not understood. We identified genetic variation associated with infant and toddler temperament using genome-wide association meta-analyses. We studied parent-rated emotionality, activity, shyness and sociability (n = 43,963-72,663) in the second and third postnatal years and a cross-age average. Cross-age single nucleotide polymorphism heritabilities for emotionality, activity, shyness and sociability were 6.79% (95% confidence interval (CI), (4.71%, 8.87%)), 9.55% (95% CI, (7.04%, 12.06%)), 15.26% (95% CI, (12.24%, 18.28%)) and 3.42% (95% CI, (1.30%, 5.54%)), respectively. Ten genome-wide significant loci were discovered. Two loci colocalized with expression quantitative trait loci in the adult cortex: RHEBL1 (posterior probability, 0.93; associated with activity) and MR1 (posterior probability, 0.99; with emotionality). Genetic correlations were observed between early temperament and later outcomes, such as emotionality and adult neuroticism, activity and attention deficit/hyperactivity disorder (ADHD), sociability and autism, and shyness and adult extraversion. Multi-ancestry (n = 56,083-78,894) and European-ancestry analyses gave similar results. Infant and toddler temperament is associated with genetic variation and shows genetic continuity with later outcomes.
AIMS:There is no proven treatment to prevent the growth or rupture of abdominal aortic aneurysm (AAA). As an aneurysm enlarges over time, the risk of fatal aortic rupture increases. Metformin, a drug usually prescribed to treat Type 2 diabetes, has previously been associated with reduced AAA risk in observational studies. Our aim was to assess whether there was a causal association between metformin treatment and AAA using Mendelian randomisation (MR). METHODS:Logistic regression analysis was conducted in UK Biobank with 2972 AAA cases and 89,160 propensity-matched controls. In addition, two-sample MR analysis was performed, using a genetic proxy for metformin consisting of variants associated with both gene expression of seven metformin drug targets and decreased glycated haemoglobin (HbA1c) levels. Effect sizes for HbA1c were obtained from within UK Biobank, and for AAA risk from AAAgen, a multi-ancestry meta-GWAS analysis of 39,221 cases and 1,086,107 controls. RESULTS:We found evidence of a protective association between self-reported metformin treatment and reduced AAA risk in the observational analysis, OR 0.49 (95% CI: 0.41-0.59, p = 1.5 × 10-14). MR results support this finding, with an estimated decrease in AAA risk of 43%, OR = 0.57 (95% CI: 0.38-0.88, p = 0.010) per one standard deviation (sd) decrease in HbA1c via metformin gene targets, equivalent to the effect of a prescribed dose of metformin. This effect is specific to metformin target genes and was not seen using a general untargeted instrument. CONCLUSION:Our observational study found evidence that metformin use reduces the risk of developing AAA. The MR results support this finding, providing evidence that the association may be causal. Clinical trials are warranted to assess the efficacy of metformin to reduce the risk of aneurysm growth and rupture in people with AAA.
Abstract Polygenic scores are imperfect measures of the additive genetic effects of common genetic variants. The resulting measurement error biases estimates of quantities of interest in epidemiological analyses integrating polygenic scores. For example, how much of an exposure-outcome association is genetically confounded can be substantially underestimated when using polygenic scores alone. Here we present extensions to Gsens , a genetic sensitivity analysis, which aims to correct for such measurement error using both polygenic scores and heritability estimates. Gsens now allows for multiple exposures and estimates several quantities of interest, i.e. genetic confounding, adjusted residual association (net of genetic confounding), genetic overlap and environmentally mediated genetic effects. We present derivations and simulations showing how Gsens accounts for measurement error in the polygenic score; we also show how estimation may be affected by misspecifications of the causal structure between exposures. Applying Gsens in the Norwegian Mother, Father and Child Cohort Study (MoBa), we uncover, among other results, substantial genetic confounding in the associations between multiple known risk factors for attention deficit hyperactivity disorder (ADHD), such as low birth weight and temperament, and measures of ADHD in childhood. The updated Gsens R package offers multiple options, including for missing data handling and customisable syntax. Our extended version of Gsens is applicable to a broad range of substantive questions in multiple disciplines.
BACKGROUND:Air pollution and diet both affect lung health and may interact. We investigated whether a healthy diet may modify associations between air pollution and lung function in adults. METHODS:Modelled annual-average concentrations of nitrogen dioxide (NO2) and particulate matter with aerodynamic diameters ≤10 μm (PM10) and ≤2.5 μm (PM2.5) were linked to residential address points of 260,982 individuals in the UK Biobank cohort. Averaged air pollution concentrations in the year of spirometry and two years prior to spirometry measurements were used. The healthy diet score (HDS) was calculated based on dietary data collected at baseline. Effect modifications by HDS and individual food components (fruit and vegetables) on the associations of air pollution and lung function were investigated. RESULTS:Participants in the highest HDS group had higher forced expiratory volume in 1-s(FEV1) and forced vital capacity(FVC) than those in the lowest group in both males and females. Interactions between HDS and air pollution were not seen. Suggestive evidence of total fruit intake effect modification on PM2.5-FEV1 was observed in females. Exposure to PM2.5 per 5 μg/m3 increment was associated with reduced FEV1 in the low of -14.4 mL(95%CI: -26.8, -2.2) but not in medium and high fruit intake groups (+2.9 mL(95%CI: -13.8,19.7) and +9.7 mL(95%CI: -4.1,23.6), respectively). A similar pattern was observed for FVC. CONCLUSIONS:We found suggestive evidence that higher consumption of fruit may partially reduce the adverse effects of air pollution on lung function in females. These findings merit investigation to see if they replicate in other cohorts.
Background:Despite multiple clinical trials, disease-modifying treatments for COPD are currently limited. Since many drugs target proteins, identifying causality between proteins and lung function informs understanding of COPD pathophysiology and may suggest novel targets. We used Mendelian randomisation (MR) to prioritise proteins as potentially causal for imparied lung function. For prioritised proteins, we explored their potential suitability as drug targets by predicting their effects on a range of clinical outcomes. Methods:We used genome-wide association study (GWAS) data on 2923 proteins (n=48 195, UK Biobank) to identify single genetic variants (protein quantitative trait loci (cis-pQTLs)) associated with protein levels (p≤5×10-9, variant ≤100 kb of a transcription start site). We performed cis-pQTL-MR analyses of four spirometric traits (n=149 166, 36 independent cohorts). Sensitivity analyses included colocalisation and reverse direction MR. We report associations between cis-pQTLs for prioritised proteins and multiple clinical respiratory outcomes, and use phenome-wide analysis to explore potential adverse effects or drug repurposing opportunities. Findings:1841 proteins had a suitable cis-pQTL. We implicated 16 proteins as potentially causal for lung function (p<1.71×10-5): seven proteins have not been implicated by previous lung function GWAS or MR (CCND2, DTD1, PILRA, PTPRK, TDRKH, GRHPR, NUDT5), and we provide corroborative evidence for 10 proteins. We add to the literature identifying surfactant protein D (SFTPD) as a candidate, yet predict that integrin subunit alpha V (ITGAV) inhibition could impair some lung function measures, mimicking adverse results from a recent trial. Interpretation:Our approach identifies proteins (some novel) that are potentially therapeutic targets for respiratory disease, and which warrant follow-up for utility and safety.
RATIONALE:Idiopathic Pulmonary Fibrosis (IPF) is characterized by chronic progressive pulmonary fibrosis and high mortality. Genetic markers, summarized into a polygenic risk score (PRS), associate with IPF in well-phenotyped research cohorts. OBJECTIVES:To evaluate the performance of the PRS using real-world data from routinely captured electronic healthcare records. METHODS:We conducted an observational study evaluating the association of a PRS for IPF with electronic healthcare record diagnosis of IPF as well as lung transplant-free survival in four independent cohorts; the Mass General Brigham Biobank (MGBB), Mayo Clinic Biobank (MCBB), Mayo Clinic Tapestry Cohort (Tapestry), and U.K. Biobank (UKBB). We used multivariable logistic regression and multivariable Cox proportional hazards models adjusting for age, gender, and principal components of ancestry. The cohorts then underwent fixed and random effects meta-analysis. MEASUREMENTS AND MAIN RESULTS:Of 37,709; 44,195; 43,202; and 447,422 participants from MGBB, MCBB, Tapestry and UKBB respectively, 1,015 (2.7%) 2,879 (6.5%), 1,310 (3.0%), and 2,742 (0.6%) participants had an IPF diagnosis. Meta-analysis demonstrated a high-risk PRS associated with IPF diagnosis, OR 2.88 (95%CI 2.41-3.44) compared to all other individuals. A high-risk PRS also associated with the composite endpoint of mortality or lung transplant among those with an IPF diagnosis, HR 1.23(95%CI 1.11-1.35) compared to all other individuals. CONCLUSIONS:A PRS can identify those at risk for an IPF diagnosis and mortality in biobank-scale data, which may have implications for clinical decisions. Further work is necessary to evaluate the utility of adding genetics in clinical settings.
BACKGROUND AND AIMS:Risk factor associations for subsequent events among coronary heart disease (CHD) survivors often appear weakened or paradoxical compared with first CHD event associations, confounding clinical interpretation and risk prediction. This study sought to test whether these observations are a systemic feature of studying disease survivors or unique to specific risk factors requiring biological explanation. METHODS:A retrospective longitudinal cohort study of 3 275 736 cardiovascular disease (CVD)-free adults and 57 165 CHD survivors was conducted using linked electronic health records from England (Clinical Practice Research Datalink, CPRD). The association of 15 CHD risk factors, documented at cohort entry and again at the time of the first CHD, was compared with both first and subsequent CHD, respectively. The performance of a 10-year primary prevention risk score was also assessed in both settings. RESULTS:Absolute risk over 5 years was 32.4% in those with CHD compared with 1.33% in the CVD-free cohort. There was attenuation of all subsequent CHD event risk factor associations, proportional to first CHD event effect size. Nine risk factors retained a concordant association, while others showed neutral or paradoxical association with subsequent CHD events compared with first CHD events. A 10-year risk model for primary prevention performed well for first events [area under the curve (AUC) .84-.85] but poorly for subsequent events (AUC ∼.55). CONCLUSIONS:Risk factor 'weakness' after CHD appears to be a systemic phenomenon and consistent with the operation of index event bias. Association findings in this context should not be used to infer or dismiss causality nor to deprioritize proven therapies such as low-density lipoprotein cholesterol lowering. Broader progress in understanding subsequent event risk associations will require new methods that address index event bias.
We previously identified genetic correlation between pairs of musculoskeletal (MSK) and respiratory conditions. Strategies to prevent or delay their onset remain underexplored in the context of multimorbidity. This study investigated whether MSK–respiratory disease pairs show evidence of potential causal relationships, identified modifiable risk factors, and quantified intervention windows to prevent progression to multimorbidity. We examined combinations of one respiratory condition (asthma, COPD) and one MSK condition [rheumatoid arthritis (RA), osteoarthritis (OA), polymyalgia rheumatica (PMR), psoriasis]. Two-sample Mendelian randomisation (MR) evaluated potential causal relationships in both directions. Linked electronic health records from CPRD (N = 11,042,985; age ≥ 40 years) were used to assess longitudinal disease trajectories, prognostic consequences, and mediation by potentially modifiable or treatable factors. We found evidence for bidirectional relationships between COPD and RA/OA (ORs 1.10–1.19) and between asthma and RA/OA (ORs 1.03–1.14). COPD genetic liability also increased PMR risk (OR 1.14, 95
BACKGROUND:Multimorbidity, the co-occurrence of multiple long-term conditions (LTCs), is an increasingly important clinical problem, but little is known about the underlying causes. We investigate the role of a critical multimorbidity risk factor, obesity, as measured by body mass index (BMI), in explaining shared genetics amongst 71 common LTCs. METHODS:In a population of northern Europeans, we estimated genetic correlation, between LTCs and partial genetic correlations after adjustment for the genetics of BMI. We used multiple causal inference methods to confirm that BMI causally affects individual LTCs, and their co-occurrence. Finally, we quantified the population-level impact of intervening and lowering BMI on the prevalence of 15 key common multimorbid LTC pairs. RESULTS:BMI partially explains some of the shared genetics for 740 LTC pairs (30% of all pairs considered). For a further 161 LTC pairs, the genetic similarity between the LTCs was entirely accounted for by BMI genetics. This list included diabetes and osteoarthritis and gout and osteoarthritis: Causal inference methods confirmed that higher BMI acts as a common risk factor for a subset of these pairs, and therefore BMI-lowering interventions would likely reduce their prevalence. For example, we estimated that a 1 standard deviation or 4.5 unit decrease in BMI would result in 17 fewer people with both chronic kidney disease and osteoarthritis per 1000 who currently have both LTCs. CONCLUSIONS:Our genetics-centred approach quantifies the contribution of obesity to multi-morbidity. Our method for calculating full and partial genetic correlations is published as an R package {partialLDSC}.
Introduction Clinical guidelines may reduce statistical power in epidemiological studies by discarding informative measures. Epidemiological studies of lung function may discard one-third to one-half of participants due to spirometry measures deemed "low quality" using criteria adapted from clinical practice. Objectives To optimise the signal-to-noise ratio in epidemiological studies of lung function, we aimed to develop a data-driven method to refine spirometry quality control (QC) criteria. Methods We proposed a genetic risk score (GRS) informed strategy to categorise spirometer blows by quality criteria. GRS was built using SNPs associated with lung function traits in non-UK Biobank cohorts. In the UK Biobank, we applied a step-wise testing of the GRS association across groups of spirometry blows stratified by acceptability flags to rank the blow quality. We reassessed QC criteria by comparing the genetic associations under different acceptability flags and repeatability thresholds to determine the trade-off between sample size and measurement error. Results We found that including blows previously excluded by strict QC criteria would maximise the statistical power for genome-wide association study and retain acceptable precision in the UK Biobank. This approach allowed the inclusion of 29% more participants compared to the strictest clinical guidelines and demonstrated genetic signals could be identified earlier. Conclusions Our GRS-based method offers an important framework to challenge prevailing practices that exclude informative measures and limit power in epidemiological studies.
Background/Objectives: Autism spectrum disorder (ASD), a neurodevelopmental condition characterised by social and communication differences, is complex and aetiologically heterogeneous. Untargeted metabolomics is emerging as a tool in screening for biochemical abnormalities. This research was conducted using the Australian Autism Biobank resource and involved analysis of plasma metabolites to characterise metabolite differences between autistic children and controls. Methods: We sought to identify molecular signatures in the plasma of study subjects using mass-spectrometry methods. We included 955 untargeted plasma metabolites from autistic children (n = 491; 2–18 years; 78% male) and control subjects (n = 97; 2–17 years of age; 51% male). Statistical analyses were performed using questionnaire data for both groups, including standardised scores from the Autism Diagnostic Observation Schedule—Second Edition (ADOS-2), which measures the severity of autism-related behaviours. We also evaluated intellectual disability by examining the relationships between metabolites and clinical phenotypes. Results: After controlling the false discovery rate at 5%, we identified significant negative associations between the uncharacterised metabolites X-21383 and X-24970 and ASD status (p = 1.85 × 10−6 and p = 1.92 × 10−5 respectively). X-21383 was also found to be significantly reduced in autistic children with coexisting intellectual disability when compared with controls (p = 6.06 × 10−6). No significant associations were identified between the metabolite data and ADOS-2 scores. However, greater levels of X-16938, N1-methyladenosine, and 2-oxoarginine were found to be suggestively associated with higher ADOS-2 scores (p = 2.95 × 10−4–9.6 × 10−5). Conclusion: This metabolomics study in the Australian Autism Biobank has identified several novel metabolites associated with core autism diagnostic behaviours.
Neuropathic pain is a common and debilitating symptom with limited treatment options. Genetic studies, which can provide vital evidence for drug development, have identified only 3 genome-wide significant signals for neuropathic pain traits. To address this, we performed the largest genome-wide association study (GWAS) to date of all-cause neuropathic pain and neuropathic pain subtypes. We defined all-cause neuropathic pain and 33 neuropathic pain subtypes using DeepPheWAS software in the UK Biobank, taking advantage of the longitudinal drug prescription data alongside clinical and self-reported records. We performed a GWAS of all-cause neuropathic pain (33,278 cases, 140,134 controls) as our primary analysis and GWASs of neuropathic pain subtypes as secondary analyses. We used 8 variant-to-gene criteria to identify putative causal genes. We identified 7 independent novel genome-wide associations for neuropathic pain phenotypes, which mapped to 22 novel putative causal genes. NCAM1 was the only gene identified from the primary analysis of all-cause neuropathic pain and met the most variant-to-gene criteria (4) of any identified gene. Of the 21 other genes, ASCC1, CHST3, C4A/C4B, and KCNN2 had the most compelling evidence for mechanistic involvement in neuropathic pain. We have performed the largest GWAS to date of all-cause neuropathic pain and more than doubled the number of genome-wide significant associations for neuropathic pain traits, identifying putative causal genes. There is strong evidence for the involvement of NCAM1 in neuropathic pain, which merits for further study for drug development.
RATIONALE: Impaired lung function predicts mortality and is a diagnostic criterion for chronic obstructive pulmonary disease (COPD). Proteins are often the target of pharmacological interventions, therefore identifying causal links between proteins and lung function could inform understanding of COPD pathophysiology and suggest therapeutic targets. We aim to infer the potential impact of circulating protein levels on lung function, using strictly defined cis protein quantitative trait loci (cis-pQTLs) as genetic instrumental variables for Mendelian randomisation (MR). METHODS: We applied two-sample MR by integrating protein GWAS data (2,923 proteins, 48,195 UK Biobank European participants) with lung function GWAS data (four lung function traits, 149,166 European participants from 36 non-UK Biobank cohorts). We selected strictly defined cis-pQTLs, within 100 kilobase pairs of a transcription start site and strongly associated (P≤5×10-9) with protein levels, and applied single-cis-MR analysis (Wald ratio method). Sensitivity analyses included colocalization analysis (to distinguish causal effects from genomic confounding by linkage disequilibrium), and bidirectional MR to explore possible reverse causation. Replication analysis was conducted where possible. We used the Drug-Gene Interaction Database and phenome-wide association studies (PheWAS) to inform biological and clinical interpretation of identified proteins. RESULTS: We curated 1,841 proteins with a suitable cis-pQTL instrument, and evaluated evidence for causal effects of these proteins on four lung function traits. The single-cis MR analysis implicated 18 proteins for lung function at a Bonferroni-corrected threshold (Wald ratio estimator P<2.72×10-5). Of 10 proteins previously implicated by reported lung function signals, surfactant protein D (SFTPD) has been highlighted in previous respiratory MR analyses and variants in SFTPD have been previously reported to be associated with emphysema; our PheWAS suggested that this variant has a relatively specific effect on lung function as it was associated with no non-respiratory traits at a FDR<1%. In contrast to previous expression QTL evidence, our study suggested that ITGAV inhibition could reduce FEV1/FVC; we note that reduced lung function was also seen in a recent trial of an ITGAV inhibitor (NCT01371305). Our MR analysis implicated 8 novel proteins not implicated by previous GWAS (CCND2, DTD1, PILRA, PTPRK, TDRKH, GRHPR, NUDT5, SLITRK6); in our PheWAS the variants instrumenting these protein levels were associated with a wide range of traits. CONCLUSIONS: Our protein-based approach identified proteins that may be causally related for lung function variability. We highlight known protein drug targets, and identify several new proteins which are potentially therapeutic targets but warrant further follow up for potential utility and safety.
Background: Hypertension and type 2 diabetes (T2D) are two of the most frequently co-occurring long-term conditions, but their shared mechanisms are not fully understood, often being attributed to adiposity pathways. Here, we aimed to identify shared genetic mechanisms independent of adiposity. Methods: We performed genome-wide association study meta-analyses of T2D and, separately, hypertension. We investigated the bidirectional causal relationship using Mendelian randomisation and quantified genetic correlation before and after accounting for common modifiable risk factors. We then applied a Bayesian GWAS approach to re-estimate SNP-disease effects after accounting for the causal genetic effects of adiposity-related traits. Colocalisation analysis identified shared causal genetic variants, and we investigated the biological pathways involved. Results: We observed a bidirectional causal relationship, and substantial genetic correlation between the two traits (rg = 0.48, 95%CI 0.45-0.52), which persisted after accounting for the genetic contributions of BMI, waist-hip ratio (WHR), and triglycerides (rg = 0.29, 95%CI 0.24-0.34). This indicated shared mechanisms beyond those captured by standard measures of adiposity. We found 98 genetic loci containing variants significantly associated with both hypertension and T2D; colocalisation analysis identified 37 that contained specific shared causal variants. Of these, eight remained statistically significant after adjusting for genetic measures of adiposity, and four were identified only after removing the causal effect of adiposity measures. Shared variants include an allele within PCSK7 associated with risk of both T2D and hypertension and with circulating PCSK7 protein levels, as well as a variant in the 3′ untranslated region of ZNF101 , within the TM6SF2 locus, likely reflecting regulatory variation affecting hepatic lipid metabolism and cardiometabolic traits.
BACKGROUND:Multimorbidity, the presence of two or more conditions in one person, is common but studies are often limited to observational data and single datasets. We address this gap by integrating large-scale primary-care and genetic data from multiple studies to interrogate multimorbidity patterns and producing digital resources to support future research. METHODS:We defined chronic, common, and heritable conditions in individuals aged ≥65 years, using two large primary-care databases [CPRD (UK) N = 2,425,014 and SIDIAP (Spain) N = 1,053,640], and estimated heritability using the same definitions in UK Biobank (N = 451,197). We used logistic regression to estimate the co-occurrence of pairs of conditions in the primary care data. Linkage disequilibrium score regression was used to estimate genetic similarity between pairs of conditions. Meta-analyses were conducted across databases, and up to three sources of genetic data, for each pair of conditions. We classified pairs of conditions as across or within-domain based on the international classification of disease. FINDINGS:We identified 72 chronic conditions, with 43.6% of 2546 pairs showing higher co-occurrence than chance in primary care and evidence of shared genetics. Many across-domain pairs exhibited substantial shared genetics (e.g., iron deficiency anaemia and peripheral arterial disease: genetic correlation Rg = 0.45 [95% Confidence Intervals 0.27:0.64]). 33 pairs displayed negative genetic correlations, such as skin cancer and rheumatoid arthritis (Rg = -0.14 [-0.21:-0.06]), due to potential adverse drug effects. Discordance between genetic and primary care data was also observed, e.g., abdominal aortic aneurysm and bladder cancer co-occurred in primary care but were not genetically correlated (Odds-Ratio = 2.23 [2.09:2.37], Rg = 0.04 [-0.20:0.28]) and schizophrenia and fibromyalgia were less likely to co-occur together in primary care but were positively genetically correlated (OR = 0.84 [0.75:0.94], Rg = 0.20 [0.11:0.29]). INTERPRETATION:Most pairs of chronic conditions show evidence of shared genetics, and co-occurrence in primary care, suggesting shared mechanisms. The identified patterns of shared genetics, negative correlations and discordance between genetic and observational data provide a foundation for future multimorbidity research. FUNDING:UK Medical Research Council [MR/W014548/1].
Age at onset of walking is an important early childhood milestone which is used clinically and in public health screening. In this genome-wide association study meta-analysis of age at onset of walking (N = 70,560 European-ancestry infants), we identified 11 independent genome-wide significant loci. SNP-based heritability was 24.13% (95% confidence intervals = 21.86-26.40) with ~11,900 variants accounting for about 90% of it, suggesting high polygenicity. One of these loci, in gene RBL2, co-localized with an expression quantitative trait locus (eQTL) in the brain. Age at onset of walking (in months) was negatively genetically correlated with ADHD and body-mass index, and positively genetically correlated with brain gyrification in both infant and adult brains. The polygenic score showed out-of-sample prediction of 3-5.6%, confirmed as largely due to direct effects in sib-pair analyses, and was separately associated with volume of neonatal brain structures involved in motor control. This study offers biological insights into a key behavioural marker of neurodevelopment.
RATIONALE: Impaired lung function predicts mortality and is a diagnostic criterion for chronic obstructive pulmonary disease (COPD). Proteins are often the target of pharmacological interventions, therefore identifying causal links between proteins and lung function could inform understanding of COPD pathophysiology and suggest therapeutic targets. We aim to infer the potential impact of circulating protein levels on lung function, using strictly defined cis protein quantitative trait loci (cis-pQTLs) as genetic instrumental variables for Mendelian randomisation (MR). METHODS: We applied two-sample MR by integrating protein GWAS data (2,923 proteins, 48,195 UK Biobank European participants) with lung function GWAS data (four lung function traits, 149,166 European participants from 36 non-UK Biobank cohorts). We selected strictly defined cis-pQTLs, within 100 kilobase pairs of a transcription start site and strongly associated (P<5E-9) with protein levels, and applied single-cis-MR analysis (Wald ratio method). Sensitivity analyses included colocalization analysis (to distinguish causal effects from genomic confounding by linkage disequilibrium), and bidirectional MR to explore possible reverse causation. Replication analysis was conducted where possible. We used the Drug-Gene Interaction Database and phenome-wide association studies (PheWAS) to inform biological and clinical interpretation of identified proteins. RESULTS: We curated 1,841 proteins with a suitable cis-pQTL instrument, and evaluated evidence for causal effects of these proteins on four lung function traits. The single-cis MR analysis implicated 16 proteins for lung function at a Bonferroni-corrected threshold (Wald ratio estimator P<1.71E-5), with evidence from colocalization. Of these, 10 proteins have been previously implicated either by lung function GWAS, or from other MR analyses with colocalization. Surfactant protein D (SFTPD) has been highlighted in previous respiratory MR analyses and variants in SFTPD have been previously reported to be associated with emphysema; our PheWAS suggested that this variant has a relatively specific effect on lung function as it was associated with no non-respiratory traits at a FDR<1%. In contrast to previous expression QTL evidence, our study suggested that ITGAV inhibition could reduce FEV1/FVC; we note that reduced lung function was also seen in a recent trial of an ITGAV inhibitor ([NCT01371305][1]). Our MR analysis implicated six proteins not implicated by previous lung function GWAS or MR (DTD1, PILRA, PTPRK, TDPRK, GRHPR, NUDT5). CONCLUSIONS: Our protein-based approach identified proteins that may be causally related for lung function variability. We highlight known protein drug targets, and identify several new proteins which are potentially therapeutic targets but warrant further follow up for potential utility and safety. ### Competing Interest Statement Richard J. Packer, Martin D Tobin and Anna L Guyatt receive collaborative funding from Orion Pharma, unrelated to the submitted work. ### Funding Statement This research was supported by a Wellcome Discovery Award (WT 225221/Z/22/Z). The research was partially supported by the NIHR Leicester Biomedical Research Centre and through an NIHR Senior Investigator Award to M.D.T. and I.P.H.; views expressed are those of the author(s) and not necessarily those of the NHS, the NIHR or the Department of Health. The funders had no role in the design of the study. For the purpose of open access, the author has applied a CC BY public copyright licence to any Author Accepted Manuscript version arising from this submission. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The study is carried out in UK Biobank, which has approval from the North West Multi-centre Research Ethics Committee (MREC) as a Research Tissue Bank (RTB) approval (21/NW/0157). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present study are available upon reasonable request to the authors [1]: /lookup/external-ref?link_type=CLINTRIALGOV&access_num=NCT01371305&atom=%2Fmedrxiv%2Fearly%2F2025%2F02%2F10%2F2025.02.07.25321860.atom