The Women's Health Study (WENDY) was conducted to improve insights into women's health and health burden. It provides a unique, comprehensive data source that can be broadly utilized to understand gynecological symptoms, diseases, and their relation to metabolic and overall health more deeply in a population-based setting. The study was conducted in Finland from May 2020 to October 2022. It included 1918 women (33-37 years old) who were born in northern Finland between July 1985 and December 1987. Data collection comprised one 3- to 4-hour study visit that included clinical measurements, biological samples, ultrasound examinations and an extensive questionnaire on gynecological and reproductive history, physical and mental health, quality of life, lifestyles, current life situations, health awareness, and opinions. The study also included a menstrual cycle follow-up and cognitive testing up to 3 months via a mobile application. Given that all participants' data can be linked to all Finnish national registers, and the Northern Finland Birth Cohort participants' data can be linked to the birth cohort data set collected from gestational week 24 onward, WENDY study forms one of the largest data sets worldwide to investigate gynecological and metabolic health burden in women.
Introduction A better understanding of the earliest stages of Alzheimer’s disease (AD) could expedite the development or administration of treatments. Large population biobanks hold the promise to identify individuals at an elevated risk of AD and related dementias based on health registry information. Here, we establish the protocol for an observational clinical recall and biomarker study called TWINGEN with the aim to identify individuals at high risk of AD by assessing cognition, health and AD-related biomarkers. Suitable candidates were identified and invited to participate in the new study among THL Biobank donors according to TWINGEN study criteria.Methods and analysis A multi-centre study (n=800) to obtain blood-based biomarkers, telephone-administered and web-based memory and cognitive parameters, questionnaire information on lifestyle, health and psychological factors, and accelerometer data for measures of physical activity, sedentary behaviour and sleep. A subcohort is being asked to participate in an in-person neuropsychological assessment (n=200) and wear an Oura ring (n=50). All participants in the TWINGEN study have genome-wide genotyping data and up to 48 years of follow-up data from the population-based older Finnish Twin Cohort (FTC) study of the University of Helsinki. The data collected in TWINGEN will be returned to THL Biobank from where it can later be requested for other biobank studies such as FinnGen that supported TWINGEN.Ethics and dissemination This recall study consists of FTC/THL Biobank/FinnGen participants whose data were acquired in accordance with the Finnish Biobank Act. The recruitment protocols followed the biobank protocols approved by Finnish Medicines Agency. The TWINGEN study plan was approved by the Ethics Committee of Hospital District of Helsinki and Uusimaa (number 16831/2022). THL Biobank approved the research plan with the permission no: THLBB2022_83.
Additional file 5: Table S4. Frequency of lipid-related publications for the PoPS+ prioritized genes.
Abstract Shared genetic factors may contribute to the associations between higher levels of physical activity (PA) and lower risk for cardiometabolic diseases (CMDs), and may partially explain these associations observed in cohort studies. To explore this, we used novel methodology to calculate PA genotypes (polygenic risk score, PRS) and validated them against measured or reported PA in three independent cohorts. We then investigated the associations between polygenic inheritance of PA and cardiometabolic risk factors and diseases in two large population-based biobank datasets, and examined whether selected associations were independent of self-reported PA. Our study utilized the UK Biobank as a base dataset (N = 400,124) and constructed genomewide PRSs for both self-reported and device-measured PA using single nucleotide polymorphism (SNP)-specific weights and SBayesR methodology. Both PRSs for PA included over one million SNPs. PRSs were constructed in the Finnish Twin cohort (N = 759–11,528), the Northern Finland Birth Cohort 1966 (N = 3,263–4,061), the Trøndelag Health Study cohort (HUNT, N = 47,148), and the FinnGen (N = 218,792). Cardiometabolic risk factors were measured in laboratory conditions, and CMD outcomes were derived from national health registers (ICD codes). We utilized linear, logistic, and cox regression methods for analysis. Our results showed that genotypes predisposing to higher PA were associated with higher levels of PA in independent datasets, but PRSs accounted for only a limited amount of variation (0.13-1.44%). Genotypes supporting higher PA were associated with lower body mass index [B=-0.002 in HUNT and B=-0.025 in FinnGen] and favorable cardiometabolic health in HUNT (waist circumference [B=-0.003] and HDL cholesterol [B = 0.004]). Genotypes supporting higher PA volumes were associated with lower incidence of CMDs in both HUNT and FinnGen. The strongest associations were found in hypertensive diseases and Type 2 Diabetes. In HUNT, the observed associations were not materially changed after accounting for self-reported PA. Higher PRS for PA was also associated with lower risk of mortality in FinnGen. Our findings suggest small pleiotropic effects between PA and CMDs. This means that same genetic variation may explain both physical activity behaviour and risk of diseases. PRSs provide new tools for genetic studies in sport science, but they currently have substantial practical limitations.
Additional file 17: Table S9. PheWAS UKB-MVP meta-analysis results for each index lipid variant at Bonferroni threshold for multiple testing p<=3.5e-8)
Increased blood lipid levels are heritable risk factors of cardiovascular disease with varied prevalence worldwide owing to different dietary patterns and medication use1. Despite advances in prevention and treatment, in particular through reducing low-density lipoprotein cholesterol levels2, heart disease remains the leading cause of death worldwide3. Genome-wideassociation studies (GWAS) of blood lipid levels have led to important biological and clinical insights, as well as new drug targets, for cardiovascular disease. However, most previous GWAS4–23 have been conducted in European ancestry populations and may have missed genetic variants that contribute to lipid-level variation in other ancestry groups. These include differences in allele frequencies, effect sizes and linkage-disequilibrium patterns24. Here we conduct a multi-ancestry, genome-wide genetic discovery meta-analysis of lipid levels in approximately 1.65 million individuals, including 350,000 of non-European ancestries. We quantify the gain in studying non-European ancestries and provide evidence to support the expansion of recruitment of additional ancestries, even with relatively small sample sizes. We find that increasing diversity rather than studying additional individuals of European ancestry results in substantial improvements in fine-mapping functional variants and portability of polygenic prediction (evaluated in approximately 295,000 individuals from 7 ancestry groupings). Modest gains in the number of discovered loci and ancestry-specific variants were also achieved. As GWAS expand emphasis beyond the identification of genes and fundamental biology towards the use of genetic variants for preventive and precision medicine25, we anticipate that increased diversity of participants will lead to more accurate and equitable26 application of polygenic scores in clinical practice. A genome-wide association meta-analysis study of blood lipid levels in roughly 1.6 million individuals demonstrates the gain of power attained when diverse ancestries are included to improve fine-mapping and polygenic score generation, with gains in locus discovery related to sample size.
Additional file 18: Table S10. Lambda GC values across minor allele frequency bins for sex-specific meta-analyses.
Additional file 23: Table S15. Comparison of the sex-specific effects.
Background Epidemiological and experimental evidence has linked chronic inflammation to cancer aetiology. It is unclear whether associations for specific inflammatory biomarkers are causal or due to bias. In order to examine whether altered genetically predicted concentration of circulating cytokines are associated with cancer development, we performed a two-sample Mendelian randomisation (MR) analysis. Methods Up to 31,112 individuals of European descent were included in genome-wide association study (GWAS) meta-analyses of 47 circulating cytokines. Single nucleotide polymorphisms (SNPs) robustly associated with the cytokines, located in or close to their coding gene (c is ), were used as instrumental variables. Inverse-variance weighted MR was used as the primary analysis, and the MR assumptions were evaluated in sensitivity and colocalization analyses and a false discovery rate (FDR) correction for multiple comparisons was applied. Corresponding germline GWAS summary data for five cancer outcomes (breast, endometrial, lung, ovarian, and prostate), and their subtypes were selected from the largest cancer-specific GWASs available (cases ranging from 12,906 for endometrial to 133,384 for breast cancer). Results There was evidence of inverse associations of macrophage migration inhibitory factor with breast cancer (OR per SD = 0.88, 95% CI 0.83 to 0.94), interleukin-1 receptor antagonist with endometrial cancer (0.86, 0.80 to 0.93), interleukin-18 with lung cancer (0.87, 0.81 to 0.93), and beta-chemokine-RANTES with ovarian cancer (0.70, 0.57 to 0.85) and positive associations of monokine induced by gamma interferon with endometrial cancer (3.73, 1.86 to 7.47) and cutaneous T-cell attracting chemokine with lung cancer (1.51, 1.22 to 1.87). These associations were similar in sensitivity analyses and supported in colocalization analyses. Conclusions Our study adds to current knowledge on the role of specific inflammatory biomarker pathways in cancer aetiology. Further validation is needed to assess the potential of these cytokines as pharmacological or lifestyle targets for cancer prevention.
The objective was to study the genetic etiology of Ménière's disease (MD) using next-generation sequencing in three families with three cases of MD. Whole exome sequencing was used to identify rare genetic variants co-segregating with MD in Finnish families. In silico estimations and population databases were used to estimate the frequency and pathogenicity of the variants. Variants were validated and genotyped from additional family members using capillary sequencing. A geneMANIA analysis was conducted to investigate the functional pathways and protein interactions of candidate genes. Seven rare variants were identified to co-segregate with MD in the three families: one variant in the CYP2B6 gene in family I, one variant in GUSB and EPB42 in family II, and one variant in each of the SLC6A, ASPM, KNTC1, and OVCH1 genes in family III. Four of these genes were linked to the same co-expression network with previous familial MD candidate genes. Dysfunction of CYP2B6 and SLC6A could predispose to MD via the oxidative stress pathway. Identification of ASPM and KNTC1 as candidate genes for MD suggests dysregulation of mitotic spindle formation in familial MD. The genetic etiology of familial MD is heterogenic. Our findings suggest a role for genes acting on oxidative stress and mitotic spindle formation in MD but also highlight the genetic complexity of MD.
A major challenge of genome-wide association studies (GWASs) is to translate phenotypic associations into biological insights. Here, we integrate a large GWAS on blood lipids involving 1.6 million individuals from five ancestries with a wide array of functional genomic datasets to discover regulatory mechanisms underlying lipid associations. We first prioritize lipid-associated genes with expression quantitative trait locus (eQTL) colocalizations and then add chromatin interaction data to narrow the search for functional genes. Polygenic enrichment analysis across 697 annotations from a host of tissues and cell types confirms the central role of the liver in lipid levels and highlights the selective enrichment of adipose-specific chromatin marks in high-density lipoprotein cholesterol and triglycerides. Overlapping transcription factor (TF) binding sites with lipid-associated loci identifies TFs relevant in lipid biology. In addition, we present an integrative framework to prioritize causal variants at GWAS loci, producing a comprehensive list of candidate causal genes and variants with multiple layers of functional evidence. We highlight two of the prioritized genes, CREBRF and RRBP1, which show convergent evidence across functional datasets supporting their roles in lipid biology.
Copy number variants (CNVs) are associated with syndromic and severe neurological and psychiatric disorders (SNPDs), such as intellectual disability, epilepsy, schizophrenia, and bipolar disorder. Although considered high-impact, CNVs are also observed in the general population. This presents a diagnostic challenge in evaluating their clinical significance. To estimate the phenotypic differences between CNV carriers and non-carriers regarding general health and well-being, we compared the impact of SNPD-associated CNVs on health, cognition, and socioeconomic phenotypes to the impact of three genome-wide polygenic risk score (PRS) in two Finnish cohorts (FINRISK, n = 23,053 and NFBC1966, n = 4895). The focus was on CNV carriers and PRS extremes who do not have an SNPD diagnosis. We identified high-risk CNVs (DECIPHER CNVs, risk gene deletions, or large [>1 Mb] CNVs) in 744 study participants (2.66%), 36 (4.8%) of whom had a diagnosed SNPD. In the remaining 708 unaffected carriers, we observed lower educational attainment (EA; OR = 0.77 [95% CI 0.66–0.89]) and lower household income (OR = 0.77 [0.66–0.89]). Income-associated CNVs also lowered household income (OR = 0.50 [0.38–0.66]), and CNVs with medical consequences lowered subjective health (OR = 0.48 [0.32–0.72]). The impact of PRSs was broader. At the lowest extreme of PRS for EA, we observed lower EA (OR = 0.31 [0.26–0.37]), lower-income (OR = 0.66 [0.57–0.77]), lower subjective health (OR = 0.72 [0.61–0.83]), and increased mortality (Cox’s HR = 1.55 [1.21–1.98]). PRS for intelligence had a similar impact, whereas PRS for schizophrenia did not affect these traits. We conclude that the majority of working-age individuals carrying high-risk CNVs without SNPD diagnosis have a modest impact on morbidity and mortality, as well as the limited impact on income and educational attainment, compared to individuals at the extreme end of common genetic variation. Our findings highlight that the contribution of traditional high-risk variants such as CNVs should be analyzed in a broader genetic context, rather than evaluated in isolation.
Cohort Profile: 46 years of follow-up of the Northern Finland Birth Cohort 1966 (NFBC1966) Tanja Nordström, Jouko Miettunen, Juha Auvinen, Leena Ala-Mursula, Sirkka Keinänen-Kiukaanniemi, Juha Veijola, Marjo-Riitta Järvelin, Sylvain Sebert and Minna Männikkö* Northern Finland Birth Cohorts, Infrastructure for Population Studies, Faculty of Medicine, University of Oulu, Oulu, Finland, Center for Life Course Health Research, Faculty of Medicine, University of Oulu, Oulu, Finland, Medical Research Center Oulu, Oulu University Hospital and University of Oulu, Oulu, Finland, Unit of Primary Care, Oulu University Hospital, Oulu, Finland, Oulunkaari Health Center, Ii, Finland, Healthcare and Social Services of Selänne, Pyhäjärvi, Finland, Healthcare and Social Services of City of Oulu, Oulu, Finland, Department of Psychiatry, Research Unit of Clinical Neuroscience, University of Oulu, Oulu, Finland, Department of Psychiatry, University Hospital of Oulu, Oulu, Finland, Department of Epidemiology and Biostatistics, School of Public Health, Imperial College London, London, UK, MRC-PHE Centre for Environment and Health, School of Public Health, Imperial College London, London, UK and Department of Life Sciences, College of Health and Life Sciences, Brunel University London, London, UK
The understanding of the biological and environmental risk factors of fractures in pediatrics is limited. Previous studies have reported that fractures involve heritable traits, but the genetic factors contributing to the risk of fractures remain elusive. Furthermore, genetic influences specific to immature bone have not been thoroughly studied. Therefore, the aim of the present study was to identify genetic variations that are associated with fractures in early childhood. The present study used a prospective Northern Finland Birth Cohort (year 1986; n=9,432). The study population was comprised of 3,230 cohort members with available genotype data. A total of 48 members of the cohort (1.5%) had in-hospital treated bone fractures during their first 6 years of life. Furthermore, individuals without fracture (n=3,182) were used as controls. A genome-wide association study (GWAS) was performed using a frequentist association test. In the GWAS analysis, a linear regression model was fitted to test for additive effects of single-nucleotide polymorphisms (SNPs; genotype dosage) adjusting for sex and performing population stratification using genotypic principal components. Using the GWAS analysis, the present study identified one locus with a significant association with fractures during childhood on chromosome 10 (rs112635931) and six loci with a suggested implication. The lead SNP rs112635931 was located near proline- and serine-rich 2 (PROSER2) antisense RNA 1 (PROSER2-AS1) and PROSER2, thus suggesting that these may be novel candidate genes associated with the risk of pediatric fractures.
Cytokines are the signalling molecules that underlie inflammatory processes. Here, we performed genome-wide association study (GWAS) analyses of 47 circulating cytokines in up to 13,365 individuals to identify protein quantitative trait loci (pQTL). Applying a novel approach, we incorporated pQTL and expression quantitative trait loci (eQTL) data of 10,361 tissue samples in 635 individuals to identify biologically plausible genetic instruments to proxy the effect of cytokines. Using Mendelian randomization analysis, we explored the causal determinants of inflammatory cytokines, investigated inflammatory cascades and evaluated their effects on 20 diseases. We show evidence of body mass index (BMI), smoking and systolic blood pressure (SBP) being associated with inflammation, and specifically BMI affecting levels of active PAI-1, HGF, MCP1, sE-Selectin, sICAM1, TRAIL, IL6 and CRP. Our analysis highlights a key role of VEGF in influencing the levels of eight other inflammatory cytokines. Finally, we report evidence of sICAM affecting waist circumference and risk of major depressive disorder, evidence for TRAIL affecting the risk of cardiovascular diseases, breast and prostate cancer, and evidence for MIG affecting the risk of stroke. Overall, our results offer insight into inflammatory mediators of BMI, smoking and SBP, pleiotropic effects of VEGF, and circulating cytokines that increase the risk of cancer, cardiovascular, metabolic and neuropsychiatric diseases. All the studied cytokines represent pharmacological targets and therefore offer opportunities for clinical translation in diseases with inflammatory components.
Background/objective Children BMI is a longitudinal phenotype, developing through interplays between genetic and environmental factors. Whilst childhood obesity is escalating, we require a better understanding of its early origins and variation across generations to prevent it. Subjects/methods We designed a cross-cohort study including 12,040 Finnish children from the Northern Finland Birth Cohorts 1966 and 1986 (NFBC1966 and NFBC1986) born before or at the start of the obesity epidemic. We used group-based trajectory modelling to identify BMI trajectories from 2 to 20 years. We subsequently tested their associations with early determinants (mother and child) and the possible difference between generations, adjusted for relevant biological and socioeconomic confounders. Results We identified four BMI trajectories, 'stable-low' (34.8%), 'normal' (44.0%), 'stable-high' (17.5%) and 'early-increase' (3.7%). The 'early-increase' trajectory represented the highest risk for obesity. We analysed a dose-response association of maternal pre-pregnancy BMI and smoking with BMI trajectories. The directions of effect were consistent across generations and the effect sizes tended to increase from earlier generation to later. Respectively for NFBC1966 and NFBC1986, the adjusted risk ratios of being in the early-increase group were 1.08 (1.06-1.10) and 1.12 (1.09-1.15) per unit of pre-pregnancy BMI and 1.44 (1.05-1.96) and 1.48 (1.17-1.87) in offspring of smoking mothers compared to non-smokers. We observed similar relations with infant factors including birthweight for gestational age and peak weight velocity. In contrast, the age at adiposity peak in infancy was associated with the BMI trajectories in NFBC1966 but did not replicate in NFBC1986. Conclusions Exposures to adverse maternal predictors were associated with a higher risk obesity trajectory and were consistent across generations. However, we found a discordant association for the timing of adiposity peak over a 20-year period. This suggests the role of residual environmental factors, such as nutrition, and warrants additional research to understand the underlying gene-environment interplay.
We studied a family with severe primary osteoporosis carrying a heterozygous p.Arg8Phefs*14 deletion in COL1A2, leading to haploinsufficiency. Three affected individuals carried the mutation and presented nearly identical spinal fractures but lacked other typical features of either osteogenesis imperfecta or Ehlers-Danlos syndrome. Although mutations leading to haploinsufficiency in COL1A2 are rare, mutations in COL1A1 that lead to less protein typically result in a milder phenotype. We hypothesized that other genetic factors may contribute to the severe phenotype in this family. We performed whole-exome sequencing in five family members and identified in all three affected individuals a rare nonsense variant (c.1282C > T/p.Arg428*, rs150257846) in ZNF528. We studied the effect of the variant using qPCR and Western blot and its subcellular localization with immunofluorescence. Our results indicate production of a truncated ZNF528 protein that locates in the cell nucleus as per the wild-type protein. ChIP and RNA sequencing analyses on ZNF528 and ZNF528-c.1282C > T indicated that ZNF528 binding sites are linked to pathways and genes regulating bone morphology. Compared with the wild type, ZNF528-c.1282C > T showed a global shift in genomic binding profile and pathway enrichment, possibly contributing to the pathophysiology of primary osteoporosis. We identified five putative target genes for ZNF528 and showed that the expression of these genes is altered in patient cells. In conclusion, the variant leads to expression of truncated ZNF528 and a global change of its genomic occupancy, which in turn may lead to altered expression of target genes. ZNF528 is a novel candidate gene for bone disorders and may function as a transcriptional regulator in pathways affecting bone morphology and contribute to the phenotype of primary osteoporosis in this family together with the COL1A2 deletion. (c) 2020 The Authors.Journal of Bone and Mineral Researchpublished by Wiley Periodicals LLC on behalf of American Society for Bone and Mineral Research (ASBMR).
ABSTRACT Purpose Polygenic risk scores (PRS) summarize genome-wide genotype data into a single variable that produces an individual-level risk score for genetic liability. PRS has been used for prediction of chronic diseases and some risk factors. As PRS has been studied less for physical activity (PA), we constructed PRS for PA and studied how much variation in PA can be explained by this PRS in independent population samples. Methods We calculated PRS for self-reported and objectively measured PA using UK Biobank genome-wide association study summary statistics, and analyzed how much of the variation in self-reported (MET-hours per day) and measured (steps and moderate-to-vigorous PA minutes per day) PA could be accounted for by the PRS in the Finnish Twin Cohorts (FTC; N = 759–11,528) and the Northern Finland Birth Cohort 1966 (NFBC1966; N = 3263–4061). Objective measurement of PA was done with wrist-worn accelerometer in UK Biobank and NFBC1966 studies, and with hip-worn accelerometer in the FTC. Results The PRS accounted from 0.07% to 1.44% of the variation ( R 2 ) in the self-reported and objectively measured PA volumes ( P value range = 0.023 to <0.0001) in the FTC and NFBC1966. For both self-reported and objectively measured PA, individuals in the highest PRS deciles had significantly (11%–28%) higher PA volumes compared with the lowest PRS deciles ( P value range = 0.017 to <0.0001). Conclusions PA is a multifactorial phenotype, and the PRS constructed based on UK Biobank results accounted for statistically significant but overall small proportion of the variation in PA in the Finnish cohorts. Using identical methods to assess PA and including less common and rare variants in the construction of PRS may increase the proportion of PA explained by the PRS.
Background: Obesity is an established risk factor for multiple cancer types. Lower microbial richness has been linked to obesity, but human studies are inconsistent, and associations of early-life body mass index (BMI) with the fecal microbiome and metabolome are unknown. Methods: We characterized the fecal microbiome (n = 563) and metabolome (n = 340) in the Northern Finland Birth Cohort 1966 using 16S rRNA gene sequencing and untargeted metabolomics. We estimated associations of adult BMI and BMI history with microbial features and metabolites using linear regression and Spearman correlations (rs) and computed correlations between bacterial sequence variants and metabolites overall and by BMI category. Results: Microbial richness, including the number of sequence variants (rs = −0.21, P < 0.0001), decreased with increasing adult BMI but was not independently associated with BMI history. Adult BMI was associated with 56 metabolites but no bacterial genera. Significant correlations were observed between microbes in 5 bacterial phyla, including 18 bacterial genera, and metabolites in 49 of the 62 metabolic pathways evaluated. The genera with the strongest correlations with relative metabolite levels (positively and negatively) were Blautia, Oscillospira, and Ruminococcus in the Firmicutes phylum, but associations varied by adult BMI category. Conclusions: BMI is strongly related to fecal metabolite levels, and numerous associations between fecal microbial features and metabolite levels underscore the dynamic role of the gut microbiota in metabolism. Impact: Characterizing the associations between the fecal microbiome, the fecal metabolome, and BMI, both recent and early-life exposures, provides critical background information for future research on cancer prevention and etiology.