Detecting microevolutionary responses to natural selection by observing temporal changes in individual breeding values is challenging. The collection of suitable datasets can take many years and disentangling the contributions of the environment and genetics to phenotypic change is not trivial. Furthermore, pedigree-based methods of obtaining individual breeding values have known biases. Here, we apply a genomic prediction approach to estimate breeding values of adult weight in a 35-year dataset of Soay sheep (Ovis aries). Comparisons are made with a traditional pedigree-based approach. During the study period, adult body weight decreased, but the underlying genetic component of body weight increased, at a rate that is unlikely to be attributable to genetic drift. Thus cryptic microevolution of greater adult body weight has probably occurred. Genomic and pedigree-based approaches gave largely consistent results. Thus, using genomic prediction to study microevolution in wild populations can remove the requirement for pedigree data, potentially opening up new study systems for similar research.
Genomic prediction, the technique whereby an individual’s genetic component of their phenotype is estimated from its genome, has revolutionised animal and plant breeding and medical genetics. However, despite being first introduced nearly two decades ago, it has hardly been adopted by the evolutionary genetics community studying wild organisms. Here, genomic prediction is performed on eight traits in a wild population of Soay sheep. The population has been the focus of a >30 year evolutionary ecology study and there is already considerable understanding of the genetic architecture of the focal Mendelian and quantitative traits. We show that the accuracy of genomic prediction is high for all traits, but especially those with loci of large effect segregating. Five different methods are compared, and the two methods that can accommodate zero-effect and large-effect loci in the same model tend to perform best. If the accuracy of genomic prediction is similar in other wild populations, then there is a real opportunity for pedigree-free molecular quantitative genetics research to be enabled in many more wild populations; currently the literature is dominated by studies that have required decades of field data collection to generate sufficiently deep pedigrees. Finally, some of the potential applications of genomic prediction in wild populations are discussed.
Meta‐analysis is the synthesis of findings from research projects, which enables an estimate of the average or pooled effect across various studies. This study presents findings from the intention to treat analysis for a series of educational evaluations in England using a two‐stage meta‐analysis with standardised outcome data and individual participant data meta‐analyses. The research estimates the overall impact of educational trials on pupils eligible for Free School Meals (FSM) and the attainment gap in literacy and mathematics performance between FSM and non‐FSM pupils based on analysis of 88 trials and data from over half a million pupils. For the meta‐analyses, frequentist and Bayesian multilevel models were used to estimate the individual and pooled effect size across categories of explanatory variables such as age groups (key stages in England) and aspects of the type of interventions (one‐to‐one, small group, whole class). Results indicated that the overall impact of interventions on the literacy outcomes of FSM pupils was positive, with a pooled effect size of 0.06 (0.03, 0.08). However, for mathematics, no overall effect on FSM pupils was observed. Analysis of the attainment gap indicated that literacy outcomes for FSM pupils were improved by interventions marginally more than for non‐FSM pupils (pooled attainment gap 0.01 (−0.01, 0.04)). The risk of bias assessment showed that estimates were consistent across different methodological approaches. Overall, evidence from this study can be used to identify, test and scale educational interventions in schools to improve educational outcomes for disadvantaged pupils.
Most complex traits evolved in the ancestors of all modern humans and have been under negative or balancing selection to maintain the distribution of phenotypes observed today. Yet all large studies mapping genomes to complex traits occur in populations that have experienced the Out-of-Africa bottleneck. Does this bottleneck affect the way we characterise complex traits? We demonstrate using the 1000 Genomes dataset and hypothetical complex traits that genetic drift can strongly affect the joint distribution of effect size and SNP frequency, and that the bias can be positive or negative depending on subtle details. Characterisations that rely on this distribution therefore conflate genetic drift and selection. We provide a model to identify the underlying selection parameter in the presence of drift, and demonstrate that a simple sensitivity analysis may be enough to validate existing characterisations. We conclude that biobanks characterising more worldwide diversity would benefit studies of complex traits.
In the original article publication, there is an incorrect impression that Fig. 1 formed a formal Directed Acyclic Graph (DAG) by describing it as a causal model. However, it was not correct if interpreted in this way.
Replicable genetic association signals have consistently been found through genome-wide association studies in recent years. The recent dramatic expansion of study sizes improves power of estimation of effect sizes, genomic prediction, causal inference, and polygenic selection, but it simultaneously increases susceptibility of these methods to bias due to subtle population structure. Standard methods using genetic principal components to correct for structure might not always be appropriate and we use a simulation study to illustrate when correction might be ineffective for avoiding biases. New methods such as trans-ethnic modeling and chromosome painting allow for a richer understanding of the relationship between traits and population structure. We illustrate the arguments using real examples (stroke and educational attainment) and provide a more nuanced understanding of population structure, which is set to be revisited as a critical aspect of future analyses in genetic epidemiology. We also make simple recommendations for how problems can be avoided in the future. Our results have particular importance for the implementation of GWAS meta-analysis, for prediction of traits, and for causal inference.
By using the genotyping-by-sequencing method, it is feasible to characterize genomic relationships directly at the level of family pools and to estimate genomic heritabilities from phenotypes scored on family-pools in outbreeding species.
Until now, genomic prediction (GP) in plant breeding has only used information from individuals that have been genotyped. Information from nongenotyped relatives of genotyped individuals can also be used. Single‐step GP combines marker and pedigree information into a single relationship matrix to perform GP. The objective of this study was to evaluate single‐step GP in a wheat breeding program. We compared the performance of pedigree‐based, marker‐based, and single‐step models (ABLUP, GBLUP, and HBLUP, respectively). Data consisted of 1176 genotyped (via genotyping‐by‐sequencing) and 11,131 nongenotyped wheat lines replicated in five management environments at the CIMMYT experiment station in Obregon, Mexico. Analyses involved three scenarios: (i) all lines had pedigree information but only some were genotyped, with phenotypes from one or two environments in the 2011–2012 season, (ii) all lines had genotype and pedigree information and phenotypes from four or five environments in the 2012–2013 season, and (iii) the combination of Scenarios 1 and 2. Prediction accuracies were calculated by five‐fold cross validation on plant height, maturity, heading date, and grain yield. The single‐step HBLUP outperformed GBLUP and pedigree‐based ABLUP in all cases. We conclude that the single‐step procedure combining pedigree and genomic marker data should be favored where appropriate data is available for GP in wheat breeding programs.
The implementation of genomic selection (GS) in plant breeding, so far, has been mainly evaluated in crops farmed as homogeneous varieties, and the results have been generally positive. Fewer results are available for species, such as forage grasses, that are grown as heterogenous families (developed from multiparent crosses) in which the control of the genetic variation is far more complex. Here we test the potential for implementing GS in the breeding of perennial ryegrass ( L.) using empirical data from a commercial forage breeding program. Biparental F and multiparental synthetic (SYN) families of diploid perennial ryegrass were genotyped using genotyping-by-sequencing, and phenotypes for five different traits were analyzed. Genotypes were expressed as family allele frequencies, and phenotypes were recorded as family means. Different models for genomic prediction were compared by using practically relevant cross-validation strategies. All traits showed a highly significant level of genetic variance, which could be traced using the genotyping assay. While there was significant genotype × environment (G × E) interaction for some traits, accuracies were high among F families and between biparental F and multiparental SYN families. We have demonstrated that the implementation of GS in grass breeding is now possible and presents an opportunity to make significant gains for various traits.
Background Genomic selection (GS) has become a commonly used technology in animal breeding. In crops, it is expected to significantly improve the genetic gains per unit of time. So far, its implementation in plant breeding has been mainly investigated in species farmed as homogeneous varieties. Concerning crops farmed in family pools, only a few theoretical studies are currently available. Here, we test the opportunity to implement GS in breeding of perennial ryegrass, using real data from a forage breeding program. Heading date was chosen as a model trait, due to its high heritability and ease of assessment. Genome Wide Association analysis was performed to uncover the genetic architecture of the trait. Then, Genomic Prediction (GP) models were tested and prediction accuracy was compared to the one obtained in traditional Marker Assisted Selection (MAS) methods. Results Several markers were significantly associated with heading date, some locating within or proximal to genes with a well-established role in floral regulation. GP models gave very high accuracies, which were significantly better than those obtained through traditional MAS. Accuracies were higher when predictions were made from related families and from larger training populations, whereas predicting from unrelated families caused the variance of the estimated breeding values to be biased downwards. Conclusions We have demonstrated that there are good perspectives for GS implementation in perennial ryegrass breeding, and that problems resulting from low linkage disequilibrium (LD) can be reduced by the presence of structure and related families in the breeding population. While comprehensive Genome Wide Association analysis is difficult in species with extremely low LD, we did identify variants proximal to genes with a known role in flowering time (e.g. CONSTANS and Phytochrome C).
We propose a method in which GBS data can be conveniently analyzed without calling genotypes. F2 families are frequently used in breeding of outcrossing species, for instance to obtain trait measurements on plots. We propose to perform association studies by obtaining a matching "family genotype" from sequencing a pooled sample of the family, and to directly use allele frequencies computed from sequence read-counts for mapping. We show that, under additivity assumptions, there is a linear relationship between the family phenotype and family allele frequency, and that a regression of family phenotype on family allele frequency will estimate twice the allele substitution effect at a locus. However, medium-to-low sequencing depth causes underestimation of the true allele substitution effect. An expression for this underestimation is derived for the case that parents are diploid, such that F2 families have up to four dosages of every allele. Using simulation studies, estimation of the allele effect from F2-family pools was verified and it was shown that the underestimation of the allele effect is correctly described. The optimal design for an association study when sequencing budget would be fixed is obtained using large sample size and lower sequence depth, and using higher SNP density (resulting in higher LD with causative mutations) and lower sequencing depth. Therefore, association studies using genotyping by sequencing are optimal and use low sequencing depth per sample. The developed framework for association studies using allele frequencies from sequencing can be modified for other types of family pools and is also directly applicable for association studies in polyploids.