Substantial shifts in reproductive behaviors have recently taken place in many high-income countries including earlier age at menarche, advanced age at childbearing, rising childlessness and a lower number of children. As reproduction shifts to later ages, genetic factors may become increasingly important. Although monogenic genetic effects are known, the genetics underlying human reproductive traits are complex, with both causal effects and statistical bias often confounded by socioeconomic factors. Here, we review genome-wide association studies (GWASs) of 44 reproductive traits of both female and male individuals from 2007 to early 2024, examining reproductive behavior, reproductive lifespan and aging, infertility and hormonal concentration. Using the GWAS Catalog as a basis, from 159 relevant studies, we isolate 37 genes that harbor association signals for four or more reproductive traits, more than half of which are linked to rare Mendelian disorders, including ten genes linked to reproductive-related disorders: FSHB, MCM8, DNAH2, WNT4, ESR1, IGSF1, THRB, BRWD1, CYP19A1 and PTPRF. We also review the relationship of reproductive genetics to related health and behavioral traits, aging and longevity and the effect of parental age on offspring outcomes as well as reflecting on limitations, open questions and challenges in this fast-moving field.
Genome-wide association studies (GWAS) have discovered thousands of replicable genetic associations, guiding drug target discovery and powering genetic prediction of human phenotypes and diseases. However, genetic associations can be affected by gene-environment correlations and non-random mating, which can lead to biased inferences in downstream analyses. Family-based GWAS (FGWAS) uses the natural experiment of random assignment of genotype within families to separate out the contribution of direct genetic effects (DGEs) - causal effects of alleles in an individual on an individual - from other factors contributing to genetic associations. Here, we report results from an FGWAS meta-analysis of 34 phenotypes from 17 cohorts. We found evidence that factors uncorrelated with DGEs make substantial contributions to genetic associations for 27 phenotypes, with population stratification confounding - a form of gene-environment correlation - likely the major cause. By estimating SNP heritability and genetic correlations using DGEs, we found evidence that assortative mating has led to overestimation of SNP heritability for 5 phenotypes and overestimation of the degree of shared genetic effects (pleiotropy) between 22 pairs of phenotypes. Polygenic predictors constructed from DGEs are particularly useful for studying natural selection, assortative mating, and indirect genetic effects (effects of relatives' genes mediated through the family environment). We validate our meta-analysis results by predicting phenotypes in hold-out samples using polygenic predictors constructed from DGEs, achieving statistically significant out-of-sample prediction for 24 phenotypes with little attenuation of predictive power within-families. We provide FGWAS summary statistics for 34 phenotypes that can be used for downstream analyses. Our study provides both a template for performing FGWAS and an argument for its value for debiasing inferences and understanding the impact of environment and mating patterns.
The scale and population coverage of Our Future Health, alongside other next-generation biobanks, offers unique opportunities to advance genomic medicine. Focusing on the UK context, we provide a researcher’s perspective of how this new resource could reach its full potential in a way that is impactful, user-friendly and informs related global efforts.
It is increasingly recognized that participation bias can pose problems for genetic studies. Recently, to overcome the challenge that genetic information of nonparticipants is unavailable, it is shown that by comparing the IBD (identity by descent) shared and not-shared segments between participating relative pairs, one can estimate the genetic component underlying participation. That, however, does not directly address how to adjust estimates of heritability and genetic correlation for phenotypes correlated with participation. Here, we demonstrate a way to do so by adopting a statistical framework that separates the genetic and nongenetic correlations between participation and these phenotypes. Crucially, our method avoids making the assumption that the effect of the genetic component underlying participation is manifested entirely through these other phenotypes. Applying the method to 12 UK Biobank phenotypes, we found eight that have significant genetic correlations with participation, including body mass index, educational attainment, and smoking status. For most of these phenotypes, without adjustments, estimates of heritability and the absolute value of genetic correlation would have underestimation biases.
Supplementary Tables 1-7 from Genome-Wide Significant Association Between a Sequence Variant at 15q15.2 and Lung Cancer Risk
The trait of participating in a genetic study probably has a genetic component. Identifying this component is difficult as we cannot compare genetic information of participants with nonparticipants directly, the latter being unavailable. Here, we show that alleles that are more common in participants than nonparticipants would be further enriched in genetic segments shared by two related participants. Genome-wide analysis was performed by comparing allele frequencies in shared and not-shared genetic segments of first-degree relative pairs of the UK Biobank. In nonoverlapping samples, a polygenic score constructed from that analysis is significantly associated with educational attainment, body mass index and being invited to a dietary study. The estimated correlation between the genetic components underlying participation in UK Biobank and educational attainment is estimated to be 36.6%-substantial but far from total. Taking participation behaviour into account would improve the analyses of the study data, including those of health traits.
We conduct a genome-wide association study (GWAS) of educational attainment (EA) in a sample of ~3 million individuals and identify 3,952 approximately uncorrelated genome-wide-significant single-nucleotide polymorphisms (SNPs). A genome-wide polygenic predictor, or polygenic index (PGI), explains 12-16% of EA variance and contributes to risk prediction for ten diseases. Direct effects (i.e., controlling for parental PGIs) explain roughly half the PGI's magnitude of association with EA and other phenotypes. The correlation between mate-pair PGIs is far too large to be consistent with phenotypic assortment alone, implying additional assortment on PGI-associated factors. In an additional GWAS of dominance deviations from the additive model, we identify no genome-wide-significant SNPs, and a separate X-chromosome additive GWAS identifies 57.
Participation in a genetic study likely has a genetic component. Identifying such component is difficult as we cannot compare genetic information of participants with non-participants directly, the latter being unavailable. Here, we show that alleles that are more common in participants than non-participants would be further enriched in genetic segments shared by two related participants. Genome-wide analysis was performed by comparing allele frequencies in shared and not-shared genetic segments of first-degree relative pairs of the UK Biobank. A polygenic score constructed from that analysis, in non-overlapping samples, is associated with educational attainment ( P = 2.1 × 10 −52 ), body mass index ( P = 1.5 × 10 −19 ), and participation in a dietary study ( P = 6.9 × 10 −21 ). Further analysis shows that inclination to participate is a behavioural trait in its own right, and not simply a consequence of other established phenotypes. Understanding the basis of this trait is important for data analyses and the design of future surveys, genetic or otherwise.
Effects estimated by genome-wide association studies (GWASs) include effects of alleles in an individual on that individual (direct genetic effects), indirect genetic effects (for example, effects of alleles in parents on offspring through the environment) and bias from confounding. Within-family genetic variation is random, enabling unbiased estimation of direct genetic effects when parents are genotyped. However, parental genotypes are often missing. We introduce a method that imputes missing parental genotypes and estimates direct genetic effects. Our method, implemented in the software package snipar (single-nucleotide imputation of parents), gives more precise estimates of direct genetic effects than existing approaches. Using 39,614 individuals from the UK Biobank with at least one genotyped sibling/parent, we estimate the correlation between direct genetic effects and effects from standard GWASs for nine phenotypes, including educational attainment ( r = 0.739, standard error (s.e.) = 0.086) and cognitive ability ( r = 0.490, s.e. = 0.086). Our results demonstrate substantial confounding bias in standard GWASs for some phenotypes.
Associations between genotype and phenotype derive from four sources: direct genetic effects, indirect genetic effects from relatives, population stratification, and correlations with other variants affecting the phenotype through assortative mating. Genome-wide association studies (GWAS) of unrelated individuals have limited ability to distinguish the different sources of genotype-phenotype association, confusing interpretation of results and potentially leading to bias when those results are applied – in genetic prediction of traits, for example. With genetic data on families, the randomisation of genetic material during meiosis can be used to distinguish direct genetic effects from other sources of genotype-phenotype association. Genetic data on siblings is the most common form of genetic data on close relatives. We develop a method that takes advantage of identity-by-descent sharing between siblings to impute missing parental genotypes. Compared to no imputation, this increases the effective sample size for estimation of direct genetic effects and indirect parental effects by up to one third and one half respectively. We develop a related method for imputing missing parental genotypes when a parent-offspring pair is observed. We provide the imputation methods in a software package, SNIPar (single nucleotide imputation of parents), that also estimates genome-wide direct and indirect effects of SNPs. We apply this to a sample of 45,826 White British individuals in the UK Biobank who have at least one genotyped first degree relative. We estimate direct and indirect genetic effects for ∼5 million genome-wide SNPs for five traits. We estimate the correlation between direct genetic effects and effects estimated by standard GWAS to be 0.61 (S.E. 0.09) for years of education, 0.68 (S.E. 0.10) for neuroticism, 0.72 (S.E. 0.09) for smoking initiation, 0.87 (S.E. 0.04) for BMI, and 0.96 (S.E. 0.01) for height. These results suggest that GWAS based on unrelated individuals provides an inaccurate picture of direct genetic effects for certain human traits.### Competing Interest StatementThe authors have declared no competing interest.
Genotype-phenotype associations can be results of direct effects, genetic nurturing effects and population stratification confounding. Genotypes from parents and siblings of the proband can be used to statistically disentangle these effects. To maximize power, a comprehensive framework for utilizing various combinations of parents’ and siblings’ genotypes is introduced. Central to the approach is mendelian imputation , a method that utilizes identity by descent (IBD) information to non-linearly impute genotypes into untyped relatives using genotypes of typed individuals. Applying the method to UK Biobank probands with at least one parent or sibling genotyped, for an educational attainment (EA) polygenic score that has an R 2 of 5.7% with EA, its predictive power based on direct genetic effect alone is demonstrated to be only about 1.4%. For women, the EA polygenic score has a bigger estimated direct effect on age-at-first-birth than EA itself.
Efforts to link variation in the human genome to phenotypes have progressed at a tremendous pace in recent decades. Most human traits have been shown to be affected by a large number of genetic variants across the genome. To interpret these associations and to use them reliably—in particular for phenotypic prediction—a better understanding of the many sources of genotype-phenotype associations is necessary. We summarize the progress that has been made in this direction in humans, notably in decomposing direct and indirect genetic effects as well as population structure confounding. We discuss the natural next steps in data collection and methodology development, with a focus on what can be gained by analyzing genotype and phenotype data from close relatives.
In the version of this article published, statements about the impact of insertions and deletions on gene conversions were incorrect. We reported a bias toward deletions, whereas in fact the bias was toward insertions. We are deeply indebted to Laurent Duret and Brice Letcher for noticing this mistake in our manuscript. The following statements are incorrect in the published manuscript.
De novo mutations (DNMs) cause a large proportion of severe rare diseases of childhood. DNMs that occur early may result in mosaicism of both somatic and germ cells. Such early mutations can cause recurrence of disease. We scanned 1,007 sibling pairs from 251 families and identified 878 DNMs shared by siblings (ssDNMs) at 448 genomic sites. We estimated DNM recurrence probability based on parental mosaicism, sharing of DNMs among siblings, parent-of-origin, mutation type and genomic position. We detected 57.2% of ssDNMs in the parental blood. The recurrence probability of a DNM decreases by 2.27% per year for paternal DNMs and 1.78% per year for maternal DNMs. Maternal ssDNMs are more likely to be T>C mutations than paternal ssDNMs, and less likely to be C>T mutations. Depending on the properties of the DNM, the recurrence probability ranges from 0.011% to 28.5%. We have launched an online calculator to allow estimation of DNM recurrence probability for research purposes.
A genome is a mosaic of chromosome fragments from ancestors who existed some arbitrary number of generations earlier. Here, we reconstruct the genome of Hans Jonatan (HJ), born in the Caribbean in 1784 to an enslaved African mother and European father. HJ migrated to Iceland in 1802, married and had two children. We genotyped 182 of his 788 descendants using single-nucleotide polymorphism (SNP) chips and whole-genome sequenced (WGS) 20 of them. Using these data, we reconstructed 38% of HJ's maternal genome and inferred that his mother was from the region spanned by Benin, Nigeria and Cameroon.
Heritability measures the proportion of trait variation that is due to genetic inheritance. Measurement of heritability is important in the nature-versus-nurture debate. However, existing estimates of heritability may be biased by environmental effects. Here, we introduce relatedness disequilibrium regression (RDR), a novel method for estimating heritability. RDR avoids most sources of environmental bias by exploiting variation in relatedness due to random Mendelian segregation. We used a sample of 54,888 Icelanders who had both parents genotyped to estimate the heritability of 14 traits, including height (55.4%, s.e. 4.4%) and educational attainment (17.0%, s.e. 9.4%). Our results suggest that some other estimates of heritability may be inflated by environmental effects.
Heritability measures the proportion of trait variation that is due to genetic inheritance. Measurement of heritability is of importance to the nature-versus-nurture debate. However, existing estimates of heritability could be biased by environmental effects. Here we introduce relatedness disequilibrium regression (RDR), a novel method for estimating heritability. RDR removes environmental bias by exploiting variation in relatedness due to random segregation. We use a sample of 54,888 Icelanders with both parents genotyped to estimate the heritability of 14 traits, including height (55.4%, S.E. 4.4%) and educational attainment (17.0%, S.E. 9.4%). Our results suggest that some other estimates of heritability could be inflated by environmental effects.
Epidemiological and genetic association studies show that genetics play an important role in the attainment of education. Here, we investigate the effect of this genetic component on the reproductive history of 109,120 Icelanders and the consequent impact on the gene pool over time. We show that an educational attainment polygenic score, POLYEDU, constructed from results of a recent study is associated with delayed reproduction (P < 10-100) and fewer children overall. The effect is stronger for women and remains highly significant after adjusting for educational attainment. Based on 129,808 Icelanders born between 1910 and 1990, we find that the average POLYEDU has been declining at a rate of ∼0.010 standard units per decade, which is substantial on an evolutionary timescale. Most importantly, because POLYEDU only captures a fraction of the overall underlying genetic component the latter could be declining at a rate that is two to three times faster.