La selection genetique est organisee de differentes facons selon les especes animales ou vegetales et le contexte economique. La selection genomique a seulement ete mise en oeuvre pour la filiere des bovins laitiers. Il est attendu que les manieres de mettre en oeuvre la selection genomique pour les autres especes soient assez differentes. L'objectif du projet R2D2 est d'etudier l'efficacite statistique et economique de la selection genomique dans differents contextes, selon les ressources disponibles (populations de reference, capacite de genotypage) et des specificites de gestion des populations (croisements, multiraciale, pratiques de selection, caracteres choisis, etc.). La premiere tâche consiste a proposer des methodes d'evaluation des principaux parametres genetiques (progres genetique, precision des estimations, etc.) requis pour les methodes de selection genomique. La deuxieme tâche consiste a developper une methode de comparaison de l'efficacite technico-economique des differents schemas de selection, basee ou non sur la selection genomique.
BACKGROUND:Integration of multiple results from Quantitative Trait Loci (QTL) studies is a key point to understand the genetic determinism of complex traits. Up to now many efforts have been made by public database developers to facilitate the storage, compilation and visualization of multiple QTL mapping experiment results. However, studying the congruency between these results still remains a complex task. Presently, the few computational and statistical frameworks to do so are mainly based on empirical methods (e.g. consensus genetic maps are generally built by iterative projection).RESULTS:In this article, we present a new computational and statistical package, called MetaQTL, for carrying out whole-genome meta-analysis of QTL mapping experiments. Contrary to existing methods, MetaQTL offers a complete statistical process to establish a consensus model for both the marker and the QTL positions on the whole genome. First, MetaQTL implements a new statistical approach to merge multiple distinct genetic maps into a single consensus map which is optimal in terms of weighted least squares and can be used to investigate recombination rate heterogeneity between studies. Secondly, assuming that QTL can be projected on the consensus map, MetaQTL offers a new clustering approach based on a Gaussian mixture model to decide how many QTL underly the distribution of the observed QTL.CONCLUSION:We demonstrate using simulations that the usual model choice criteria from mixture model literature perform relatively well in this context. As expected, simulations also show that this new clustering algorithm leads to a reduction in the length of the confidence interval of QTL location provided that across studies there are enough observed QTL for each underlying true QTL location. The usefulness of our approach is illustrated on published QTL detection results of flowering time in maize. Finally, MetaQTL is freely available at http://bioinformatics.org/mqtl.
We present a maximum likelihood method for mapping quantitative trait loci that uses linkage disequilibrium information from single and multiple markers. We made paired comparisons between analyses using a single marker, two markers and six markers. We also compared the method to single marker regression analysis under several scenarios using simulated data. In general, our method outperformed regression (smaller mean square error and confidence intervals of location estimate) for quantitative trait loci with dominance effects. In addition, the method provides estimates of the frequency and additive and dominance effects of the quantitative trait locus.
Recently, the use of linkage disequilibrium (LD) to locate genes which affect quantitative traits (QTL) has received an increasing interest, but the plausibility of fine mapping using linkage disequilibrium techniques for QTL has not been well studied. The main objectives of this work were to (1) measure the extent and pattern of LD between a putative QTL and nearby markers in finite populations and (2) investigate the usefulness of LD in fine mapping QTL in simulated populations using a dense map of multiallelic or biallelic marker loci. The test of association between a marker and QTL and the power of the test were calculated based on single-marker regression analysis. The results show the presence of substantial linkage disequilibrium with closely linked marker loci after 100 to 200 generations of random mating. Although the power to test the association with a frequent QTL of large effect was satisfactory, the power was low for the QTL with a small effect and/or low frequency. More powerful, multi-locus methods may be required to map low frequent QTL with small genetic effects, as well as combining both linkage and linkage disequilibrium information. The results also showed that multiallelic markers are more useful than biallelic markers to detect linkage disequilibrium and association at an equal distance.
Crop models have a large number of parameters compared to the amount of field data used for adjustment. It is not possible in general to adjust all the parameters to data, so one uses in addition prior information about the parameter values, but this prior information is imperfect. It is then important to evaluate the effect of the uncertainty in this prior information on the final model. Here we evaluate the effect of such uncertainty on a model of corn growth and development, when one uses a recently proposed algorithm for parameter estimation. This is done by applying the algorithm with different sets of initial parameter values. The algorithm automatically determines which parameters to adjust. The choice of parameters is found not to be identical for different sets of initial values. Nevertheless, the first two parameters chosen for adjustment are always the same or (for the one exception) play essentially the same role in the model. Also, the mean squared error of prediction, and the individual predictions, of the adjusted model are quite stable for different initial values.
Parameter estimation for mechanistic crop models is an important but not well-studied subject. We explore here the possibility of fitting a small number of linear combinations of the original parameters. The hope is that this approach will avoid overparameterization while still providing realistic parameter values. We first study a very simple linear model, and show the advantages of fitting a linear combination of parameters in this simple case. We then propose a method of fitting linear combinations of parameters that can be applied to mechanistic crop models. First, a linearized version of the crop model is calculated and used to generate linear combinations of the original parameters according to the method of continuum regression. Then these linear combinations of parameters are fitted to the data using the real crop model. We apply this procedure to an example. The conclusion is that the overall approach seems promising but needs further study to become operational.
The adjustment of the parameters in mechanistic crop models to field data, using an automatic procedure, is essential to ensure efficient and objective use of measured data. However, it is in general numerically impossible, and in any case undoubtedly unwise, to adjust all the model parameters to the measured data. There is currently no widely accepted solution to this problem. This paper proposes a new approach to parameter adjustment, and applies it to a model of corn growth and development. One begins by defining a criterion of model goodness-of-fit, which should be adapted to the goal of the modeling exercise, and a corresponding criterion of model prediction error. For the latter we propose a cross validation version of the goodness-of-fit criterion. In Step 1 of the algorithm, one orders the parameters according to how much each improves the goodness-of-fit of the model. In the second step, the number of parameters actually adjusted is chosen to minimize the prediction error criterion. This approach has the advantage of explicitly using prediction quality as a criterion. As a by-product, it leads to adjusting relatively few parameters (in our example, 3 out of the 26 potentially adjustable parameters), which considerably reduces the numerical problems. The procedure is quite straightforward to apply, although it does require substantial computing time.
We evaluated concordance of AFLP and RAPD markers for estimating genetic distances of 47 pepper inbred lines belonging to five varietal types. It enabled us to see the efficiency of these markers for identification, estimation of distances between varieties and variety discrimination. Genetic distance and multidimensional scaling results showed a general agreement between AFLP and RAPD markers. Based on pattern scores, dendrograms were produced by the UPGMA method. Phenetic trees based on molecular data were consistent with the classification of variety group. The precision of the estimation of the genetic distance was given. The molecular genetic distances were correlated with distances based on a set of discriminating agronomic traits measured for identification and distinctiveness tests. The relationship between molecular and morphological distances appeared to be triangular. These results and their implications in the cultivar protection purposes of pepper hybrids are discussed.
The genetic diversity among strains in a worldwide collection of Ralstonia solanacearum, causal agent of bacterial wilt, was assessed by using three different molecular methods. PCR-RFLP analysis of the hrp gene region was extended from previous studies to include additional strains and showed that five amplicons were produced not only with all R. solanacearum strains but also with strains of the closely related bacteria Pseudomonas syzygii and the blood disease bacterium (BDB). However, the three bacterial taxa could be discriminated by specific restriction profiles. The PCR-RFLP clustering, which agreed with the biovar classification and the geographical origin of strains, was confirmed by AFLP. Moreover, AFLP permitted very fine discrimination between different isolates and was able to differentiate strains that were not distinguishable by PCR-RFLP. AFLP and PCR-RFLP analyses confirmed the results of previous investigations which split the species into two divisions, but revealed a further subdivision. This observation was further supported by 16S rRNA sequence data, which grouped biovar 1 strains originating from the southern part of Africa.
We present a general regression-based method for mapping quantitative trait loci (QTL) by combining different populations derived from diallel designs. The model expresses, at any map position, the phenotypic value of each individual as a function of the specific-mean of the population to which the individual belongs, the additive and dominance effects of the alleles carried by the parents of that population and the probabilities of QTL genotypes conditional on those of neighbouring markers. Standard linear model procedures (ordinary or iteratively reweighted least-squares) are used for estimation and test of the parameters.
This article presents a method to combine QTL results from different independent analyses. This method provides a modified Akaike criterion that can be used to decide how many QTL are actually represented by the QTL detected in different experiments. This criterion is computed to choose between models with one, two, three, etc., QTL. Simulations are carried out to investigate the quality of the model obtained with this method in various situations. It appears that the method allows the length of the confidence interval of QTL location to be consistently reduced when there are only very few “actual” QTL locations. An application of the method is given using data from the maize database available online at http://www.agron.missouri.edu/.
Whereas resistance genes (R-genes) governing qualitative resistance have been isolated and characterized, the biological roles of genes governing quantitative resistance (quantitative trait loci, QTLs) are still unknown. We hypothesized that genes at QTLs could share homologies with cloned R-genes. We used a PCR-based approach to isolate R-gene analogs (RGAs) with consensus primers corresponding with conserved domains of cloned R-genes: (i) the nucleotide binding site (NBS) and hydrophobic domain, and (ii) the kinase domain. PCR-amplified fragments were sequenced and mapped on a pepper intraspecific map. NBS-containing sequences of pepper, most similar to the N gene of tobacco, were classified into seven families and all mapped in a unique region covering 64 cM on the Noir chromosome. Kinase domain containing sequences and cloned R-gene homologs (Pto, Fen, Cf-2) were mapped on four different linkage groups. A QTL involved in partial resistance to cucumber mosaic virus (CMV) with an additive effect was closely linked or allelic to one NBS-type family. QTLs with epistatic effects were also detected at several RGA loci. The colocalizations between NBS-containing sequences and resistance QTLs suggest that the mechanisms of qualitative and quantitative resistance may be similar in some cases.
In a series of papers, alternative models for QTL detection in livestock are proposed and their properties evaluated using simulations. This first paper describes the basic model used, applied to independent half-sib families, with marker phenotypes measured for a two or three generation pedigree and quantitative trait phenotypes measured only for the last generation. Hypotheses are given and the formulae for calculating the likelihood are fully described. Different alternatives to this basic model were studied, including variation in the performance modelling and consideration of full-sib families. Their main features are discussed here and their influence on the result illustrated by means of a numerical example. (C) Inra/Elsevier, Paris.
This paper describes two kinds of alternative models for QTL detection in livestock: an heteroskedastic model, and models corresponding to several hypotheses concerning the distribution of the QTL substitution effect among the sires: a fixed and limited number of alleles or an infinite number of alleles.The power of different tests built with these hypotheses were computed under different situations.The genetic variance associated with the QTL was shown in some situations.The results showed small power differences between the different models, but important differences in the quality of the estimations.In addition, a model was built in a simplified situation to investigate the gain in using possible linkage disequilibrium.© Inra/Elsevier, Paris half-sib families / heteroskedastic model / linkage disequilibrium / QTL detection Résumé -Modèles alternatifs pour la détection de QTL dans les populations animales.III.Modèle hétéroscédastique et modèles correspondant à différentes distributions de l'effet du QTL.Ce papier décrit deux types de modèles alternatifs pour la détection de QTL dans les populations animales : un modèle hétéroscédastique
In this paper, we compare four different methods of dealing with the unknown linkage phase of sire markers which occurs in the detection of quantitative trait loci (QTL) in a half-sib family structure when no information is available on grandparents.The methods are compared by considering a Gaussian approximation of the progeny likelihood instead of the mixture likelihood.In the first simulation study, the properties of the Gaussian model and of the mixture model were investigated, using the simplest method for sire gamete reconstruction.Both models lead to comparable results as regards the test power but the mean square error of sib QTL effect estimates was larger for the Gaussian likelihood than for the mixture likelihood, especially for maps with widely spaced markers.The second simulation study revealed that the simplest method for sire marker genotype estimation was as powerful as complicated methods and that the method including all the possible sire marker genotypes was never the most powerful.© Inra/Elsevier, Paris half-sib family / QTL detection / unknown linkage phase / Gaussian approxi- mation / log-likelihood ratio test Résumé -Modèles alternatifs pour la détection de QTL dans les populations animales.II.Approximations de la vraisemblance et estimations du génotype des mâles aux marqueurs.Dans ce papier, nous comparons quatre méthodes, qui permettent de résoudre le problème relatif à la phase inconnue des mâles
The aim of this paper was to compare different methods for testing the presence of one versus multiple QTLs on a same chromosome. We describe different methods that have partially been taken out of the literature. We perform simulations covering different situations to compare the power of these methods for detecting more than one QTL. None of the tests considered appear to be similar; that is, the first-type error depends on the value of the parameters concerning the first QTL. The method starting with a two-QTL model is the most powerful in many situations.
SummaryA Bayesian approach to map Quantitative Trait Loci (QTL) is compared to the flanking markers regression method (asymptotically equivalent to the traditional Interval Mapping method) using simulated backcross data. The main part consists of a comparison of the properties for a one QTL model. The Bayesian approach gives less biases and more accurate estimates of the genetic effect and significance thresholds close to χ2 quantiles when the number of individuals and markers become large. The behaviour of both approaches in the case of two linked QTL is mentioned.ZusammenfassungEigenschaften eines BAYES Ansatz zur QTL Identifikation im Vergleich zur flankierende Marker Regressions‐Methode.Der Bayes Ansatz zur Kartierung von Quantitativen Merkmals Loci (QTL) wurde mittels simulierter Rückkreuzungsdaten mit der Flankierenden Marker Regressionsmethode (asymptotisch äquivalent mit der traditionellen Intervall Markierung) verglichen. Es werden hauptsächlich die Eigenschaften für ein ein‐QTL Modell geprüft. Der Bayes Ansatz ergibt weniger verzerrte und genauere Schätz‐werte der genetischen Wirkung und Signifikanz Schwellen nahe den χ2 Quantilen bei großen Individuen und Marker Zahlen. Die Eigenschaften beider Ansätze im Falle von zwei gekoppelten QTL werden erwähnt.
Alfalfa (Medicago sativa L.) is a forage legume of world-wide importance whose both allogamous and autotetraploid nature maximizes the genetic diversity within natural and cultivated populations. This genetic diversity makes difficult the discrimination between two related populations. We analyzed this genetic diversity by screening DNA from individual plants of eight cultivated and natural populations of M. sativa and M. falcata using the RAPD method. A high level of genetic variation was found within and between populations. Using five primers, 64 intense bands were scored as present or absent across all populations. Most of the loci were revealed to be highly polymorphic whereas very few population-specific polymorphisms were identified. From these observations, we adopted a method based on the Roger’s genetic distance between populations using the observed frequency of bands to discriminate populations pairwise. Except for one case, the between-population distances were all significantly different from zero. We have also determined the minimal number of bands and individuals required to test for the significance of between-population distances.
Selective genotyping, i.e. increasing the size of the population phenotyped and genotyping only individuals from the high and low tails of the population, can considerably improve the efficiency of experiments aimed at detecting and locating quantitative trait loci (QTLs) affecting a single trait. In this paper we study how selective genotyping can increase the efficiency of multitrait QTL experiments. By selecting on an index combining the variables of interest and having the maximum correlation with each variable, the efficiency of QTL detection is increased for each trait. The efficiency of selective genotyping relative to random selection strongly depends on the correlation between the index and each variable. The optimum selection rate that minimizes costs for a given experimental power depends also on this correlation and on the genotyping costs relative to phenotyping costs. When the population segregating for the quantitative traits and the markers is not as simple as a backcross or an F 2 population, but is composed of several connected or unconnected families, selective genotyping can be used to improve the efficiency of the QTL study. In this case, the extreme individuals should be selected within each family. A method is provided to choose the selection rates within each family in order to optimize the global power of the experiment when the family sizes are unequal.
Usually, experiments designed for mapping quantitative trait loci (QTL) involve the genotyping of all individuals of a segregating population but are not very efficient with populations of feasible sizes. The use of selective genotyping, i.e., genotyping only a selected sample of the population (the high and low phenotypic tails), can considerably increase the efficiency of an experiment, provided the size of the population phenotyped is increased. In this paper, Ne examine the consequences of the use of this method for backcross and similar populations. New formulas, derived from likelihood functions, are proposed to estimate easily, without numerical maximization of the likelihood function, the effect of a QTL on the selected trait and on other traits of interest. This estimation procedure is shown to be reliable and quite robust to nonnormality, at least when the selection rate is not too low. The effect of selection on the efficiency of QTL detection for a trait other than the trait selected on is also considered: we show that selective genotyping never results in a loss of accuracy compared to random selection. Formulas to estimate QTL effects in more complex segregating populations could be easily derived in the same way as those proposed here.