Despite the availability of whole genome sequences of apple and peach, there has been a considerable gap between genomics and breeding. To bridge the gap, the European Union funded the FruitBreedomics project (March 2011 to August 2015) involving 28 research institutes and private companies. Three complementary approaches were pursued: (i) tool and software development, (ii) deciphering genetic control of main horticultural traits taking into account allelic diversity and (iii) developing plant materials, tools and methodologies for breeders. Decisive breakthroughs were made including the making available of ready-to-go DNA diagnostic tests for Marker Assisted Breeding, development of new, dense SNP arrays in apple and peach, new phenotypic methods for some complex traits, software for gene/QTL discovery on breeding germplasm via Pedigree Based Analysis (PBA). This resulted in the discovery of highly predictive molecular markers for traits of horticultural interest via PBA and via Genome Wide Association Studies (GWAS) on several European genebank collections. FruitBreedomics also developed pre-breeding plant materials in which multiple sources of resistance were pyramided and software that can support breeders in their selection activities. Through FruitBreedomics, significant progresses were made in the field of apple and peach breeding, genetics, genomics and bioinformatics of which advantage will be made by breeders, germplasm curators and scientists. A major part of the data collected during the project has been stored in the FruitBreedomics database and has been made available to the public. This review covers the scientific discoveries made in this major endeavour, and perspective in the apple and peach breeding and genomics in Europe and beyond.
Background Peach ( Prunus persica (L.) Batsch) is a major temperate fruit crop with an intense breeding activity. Breeding is facilitated by knowledge of the inheritance of the key traits that are often of a quantitative nature. QTLs have traditionally been studied using the phenotype of a single progeny (usually a full-sib progeny) and the correlation with a set of markers covering its genome. This approach has allowed the identification of various genes and QTLs but is limited by the small numbers of individuals used and by the narrow transect of the variability analyzed. In this article we propose the use of a multi-progeny mapping strategy that used pedigree information and Bayesian approaches that supports a more precise and complete survey of the available genetic variability. Results Seven key agronomic characters (data from 1 to 3 years) were analyzed in 18 progenies from crosses between occidental commercial genotypes and various exotic lines including accessions of other Prunus species. A total of 1467 plants from these progenies were genotyped with a 9 k SNP array. Forty-seven QTLs were identified, 22 coinciding with major genes and QTLs that have been consistently found in the same populations when studied individually and 25 were new. A substantial part of the QTLs observed (47%) would not have been detected in crosses between only commercial materials, showing the high value of exotic lines as a source of novel alleles for the commercial gene pool. Our strategy also provided estimations on the narrow sense heritability of each character, and the estimation of the QTL genotypes of each parent for the different QTLs and their breeding value. Conclusions The integrated strategy used provides a broader and more accurate picture of the variability available for peach breeding with the identification of many new QTLs, information on the sources of the alleles of interest and the breeding values of the potential donors of such valuable alleles. These results are first-hand information for breeders and a step forward towards the implementation of DNA-informed strategies to facilitate selection of new cultivars with improved productivity and quality.
Genetic variability is a key requirement for breeding. Although new peach cultivars are released yearly to the market, the genetic pool of cultivated peaches is very limited. To evaluate the variability available in commercial but also in old local peach accessions we selected a panel of 1,580 accessions maintained and evaluated in four European and one Chinese germplasm collections. Phenotypic data collected over years following common protocols have been integrated in a database, generating a useful tool for breeders and researchers. These accessions were genotyped with the peach 9K SNP array v1. T. Genotypic data distributed the accessions in three main subpopulations (Occidental obtained in breeding programs, Occidental old local varieties and Chinese cultivars). Linkage disequilibrium (LD) was in agreement with previous studies reporting long extension. Phenotypic and genotypic data have been combined in a GWAS study allowing the design of markers for marker assisted selection (MAS). Preliminary analyses on quantitative traits are promising, while further analysis will be required to integrate all data in a single genome-wide association analysis.
Although many peach QTLs responsible for the variability of traits have been identified and published so far, the number of molecular markers currently used in breeding is still limited. One of the reasons is the large QTL intervals produced, in part, by the limited progeny size. Here we report a QTL mapping approach that enlarges the progeny size by analyzing jointly multiple progenies. Our analysis included 1467 individuals from eighteen Prunus progenies, 13 from intra-specific crosses and five inter-specific between peach and closely related species. The progenies were grown in five locations, with no duplication between orchards. Data from phenology, tree, flower and fruit traits, fruit quality and yield measured in different locations and years were subjected to different standardization methods and integrated in a single data file. The populations were genotyped with the 9K SNP Illumina array, which increased considerably the marker density compared to previous studies. The QTL analysis was conducted with FlexQTLTM software. Here we describe and discuss the preliminary QTLs obtained for some of the traits analyzed (maturity date, percentage of red skin color and soluble solid content). The identification of donors of favorable alleles will represent an important tool for marker-assisted breeding. This study has been conducted in the frame of the Fruit Breedomics European project.
The distribution of the N-glycoproteome in integral membrane proteins of the vacuolar membrane (tonoplast) or the plasma membrane of Arabidopsis thaliana and, for further comparison, of the Rattus norvegicus lysosomal and plasma membranes, was analyzed. In silico analysis showed that potential N-glycosylation sites are much less frequent in tonoplast proteins. Biochemical analysis of Arabidopsis subcellular fractions with the lectin concanavalin A, which recognizes mainly unmodified N-glycans, or with antiserum against Golgi-modified N-glycans confirmed the in silico results and showed that, unlike the plant plasma membrane, the tonoplast is almost or totally devoid of N-glycoproteins with Golgi-modified glycans. Lysosomes share with vacuoles the hydrolytic functions and the position along the secretory pathway; however, our results indicate that their membranes had a divergent evolution. We propose that protection against the luminal hydrolases that are abundant in inner hydrolytic compartments, which seems to have been achieved in many lysosomal membrane proteins by extensive N-glycosylation of the luminal domains, has instead been obtained in the vast majority of tonoplast proteins by limiting the length of such domains.
Background In recent years, the use of genomic information in livestock species for genetic improvement, association studies and many other fields has become routine. In order to accommodate different market requirements in terms of genotyping cost, manufacturers of single nucleotide polymorphism (SNP) arrays, private companies and international consortia have developed a large number of arrays with different content and different SNP density. The number of currently available SNP arrays differs among species: ranging from one for goats to more than ten for cattle, and the number of arrays available is increasing rapidly. However, there is limited or no effort to standardize and integrate array- specific (e.g. SNP IDs, allele coding) and species-specific (i.e. past and current assemblies) SNP information. Results Here we present SNPchiMp v.3, a solution to these issues for the six major livestock species (cow, pig, horse, sheep, goat and chicken). Original data was collected directly from SNP array producers and specific international genome consortia, and stored in a MySQL database. The database was then linked to an open-access web tool and to public databases. SNPchiMp v.3 ensures fast access to the database (retrieving within/across SNP array data) and the possibility of annotating SNP array data in a user-friendly fashion. Conclusions This platform allows easy integration and standardization, and it is aimed at both industry and research. It also enables users to easily link the information available from the array producer with data in public databases, without the need of additional bioinformatics tools or pipelines. In recognition of the open-access use of Ensembl resources, SNPchiMp v.3 was officially credited as an Ensembl E!mpowered tool. Availability at http://bioinformatics.tecnoparco.org/SNPchimp.
Since the beginning of the genomic era, the number of available single nucleotide polymorphism (SNP) arrays has grown considerably. In the bovine species alone, 11 SNP chips not completely covered by intellectual property are currently available, and the number is growing. Genomic/genotype data are not standardized, and this hampers its exchange and integration. In addition, software used for the analyses of these data usually requires not standard (i.e. case specific) input files which, considering the large amount of data to be handled, require at least some programming skills in their production. In this work, we describe a software toolkit for SNP array data management, imputation, genome-wide association studies, population genetics and genomic selection. However, this toolkit does not solve the critical need for standardization of the genotypic data and software input files. It only highlights the chaotic situation each researcher has to face on a daily basis and gives some helpful advice on the currently available tools in order to navigate the SNP array data complexity.
Since the beginning of the genomic era, the SNP chip market in livestock species has grown almost exponentially. Today, researchers are asked to deal with many SNP chips on daily basis, and this requires having the (general and specific) information on the SNPs available and at hand. However, the information is often difficult to obtain (e.g. data on chips no longer on the market), integrate and standardise. Here we present SNPchiMp v.2, a multi-species database linked to an open-access web-based interface that solves many of these problems. This second version of the tool includes 9 bovine, 2 porcine and 1 equine SNP chips, marketed by Illumina, Affymetrix and GeneSeek, and genomic exchange indexes from Interbull (for bovine data). The interest of the animal genetics community on the first version of this tool has rapidly increased as have the number of chips and species included.
The expression profile of flavour-related genes during ripening was investigated in two peach genotypes, Bolero and OroA, which have been selected for their contrasting aroma/ripening behaviour. A new peach microarray containing 4776 oligonucleotide probes corresponding to a set of ESTs specifically enriched in secondary metabolism (μPEACH2.0) was designed to investigate transcriptome changes during three fruit ripening stages, revealing 1807 transcripts differentially expressed within and between the two genotypes. Differences in the expression of genes involved in the biosynthesis of aroma compounds were detected during the ripening process within and between the two genotypes. In particular, a subset of 12 transcripts involved in metabolism of esters, norisoprenoids, phenylpropanoids and lactones, varied in expression during ripening and between Bolero and OroA.
Porcine reproductive and respiratory syndrome (PRRS) is one of the most significant swine diseases worldwide. Despite its relevance, serum biomarkers associated with early-onset viral infection, when clinical signs are not detectable and the disease is characterized by a weak anti-viral response and persistent infection, have not yet been identified. Surface-enhanced laser desorption ionization time of flight mass spectrometry (SELDI-TOF MS) is a reproducible, accurate, and simple method for the identification of biomarker proteins related to disease in serum. This work describes the SELDI-TOF MS analyses of sera of 60 PRRSV-positive and 60 PRRSV-negative, as measured by PCR, asymptomatic Large White piglets at weaning. Sera with comparable and low content of hemoglobin (< 4.52 μg/mL) were fractionated in 6 different fractions by anion-exchange chromatography and protein profiles in the mass range 1–200 kDa were obtained with the CM10, IMAC30, and H50 surfaces.
In the intensive pig industry, control of infectious diseases is a major production challenge. Not only infectious diseases cause great losses to the producer, but they are important also from the animal welfare perspective. Approaches applied to achieve disease eradication include various control measures designed to reduce infection pressure within the herd, e.g. management changes, medication and vaccination (Christensen and Mousing, (1992)). Additionally, drugs used to treat infectious diseases in pigs, if not properly applied, have the potential to remain as residues in meat destined for human consumption. For these reasons the prevention of infectious diseases is of great interest not only to the pig producer but also to the consumer. Evidence for genetic variation in pigs in response to different pathogens has already been reported, such as breed differences in incidence of respiratory and enteric diseases (Van Diemen et al., (2002)) and in immune response (Henryon et al., (2002), Petry et al., (2007)). The identification of genetic variation might allow to use it in selective breeding program (Lewis et al., (2009a)). Moreover, advances in genomics of main livestock species, including pigs, will provide a powerful set of tools for understanding the genetic variation underlying economically important and complex phenotypes, such as susceptibility to metabolic and infectious diseases (Green et al., (2007), Tuggle et al., (2007), Chen et al., (2007) ). In this new scenario Genome-Wide Association (GWA) studies are being used in livestock, as in humans, to map genes affecting complex traits (Goddard and Hayes, 2009) but prior to them heritability, defined as the proportion of variation in a particular trait that is attributable to genetic factors (Visscher et al., (2008)), needs to be estimated in order not only to assess the genetic contribution to the disease outcome, but even to obtain a proper evaluation of SNPs across chromosomes and lines or breeds (Hassen et al., (2008)) and hence accurately predicting the response to artificial and natural selection. Porcine Reproductive and Respiratory Syndrome (PRRS) represents one of the most economically important disease in pig populations worldwide and causes reproductive failure, abortions, stillbirths, interstitial pneumonia and decreased growth rate (Neumann et al., 2005). The causative agent is a small envelope RNA virus (Arteriviridae family) that infects alveolar macrophages (Murtaugh et al., (2002)) and induces apoptosis and virus persistence for several weeks (Mateu et al., (2007)). The goal of the present study was to estimate heritability of susceptibility to PRRS virus in commercial pigs in Italy.
Expressed sequence tag (EST) represents a resource for gene discovery, genome annotation and comparative genomics in plants. ESTs were derived by sequencing clones from five libraries created from two different fruit tissues (skin and mesocarp), at four ripening stages (from post-allegation to post-climacteric) in three different genotypes of peach (OroA, Bolero and Suncrest). A total of 10,847 EST sequences were produced (dataset A); in addition, 21,857 peach ESTs (dataset B) were obtained from public databases. Clustering and assembly of both datasets gave 17,858 unigenes. Analysis of the sequences allowed the assignment of a putative function to 70.8% of the ESTs. In order to define the relationship among fruit tissues transcriptome, a gene ontology analysis was performed. Differences among organs and among different maturation stages of the same organs were identified in organelle, signal transducer and antioxidant activity. A distance matrix of pairwise correlation coefficients analysis was applied between the libraries. Shoot appeared to outgroup and our analysis proved to be an efficient tool to parallel and complement gene expression studies (for example, based on microarray analysis). We conducted an analysis of the frequency of genes putatively involved in the metabolism of some volatiles, which pointed to a predominant presence of those transcripts in the skin. The metabolic pathways of esters and lactones were selected for further isolation and cloning of key genes. The EST database is available at the web site www.itb.cnr.it/estree.
The identification of the genetic variations controlling phenotypes, including immune function, can be achieved by a combination of linkage mapping, association studies or candidate gene approaches. The availability of the draft bovine genome sequence together with annotation information provides considerable new information to identify positional candidate genes. There are now over 2.2 million putative single nucleotide polymorphisms for cattle in DBSNP identified from the genome sequencing project. However, up to now, few have been confirmed. Therefore, the identification and validation of polymorphisms in the functional or regulatory regions of the genes remains an important task.
BACKGROUND:With the rapid growth in the availability of genome sequence data, the automated identification of orthologous genes between species (orthologs) is of fundamental importance to facilitate functional annotation and studies on comparative and evolutionary genomics. Genes with no apparent orthologs between the bovine and human genome may be responsible for major differences between the species, however, such genes are often neglected in functional genomics studies.RESULTS:A BLAST-based method was exploited to explore the current annotation and orthology predictions in Ensembl. Genes with no orthologs between the two genomes were classified into groups based on alignments, ontology, manual curation and publicly available information. Starting from a high quality and specific set of orthology predictions, as provided by Ensembl, hidden relationship between genes and genomes of different mammalian species were unveiled using a highly sensitive approach, based on sequence similarity and genomic comparison.CONCLUSIONS:The analysis identified 3,801 bovine genes with no orthologs in human and 1010 human genes with no orthologs in cow, among which 411 and 43 genes, respectively, had no match at all in the other species. Most of the apparently non-orthologous genes may potentially have orthologs which were missed in the annotation process, despite having a high percentage of identity, because of differences in gene length and structure. The comparative analysis reported here identified gene variants, new genes and species-specific features and gave an overview of the other side of orthology which may help to improve the annotation of the bovine genome and the knowledge of structural differences between species.
BACKGROUND:Two complete genome sequences are available for Vitis vinifera Pinot noir. Based on the sequence and gene predictions produced by the IASMA, we performed an in silico detection of putative microRNA genes and of their targets, and collected the most reliable microRNA predictions in a web database. The application is available at http://www.itb.cnr.it/ptp/grapemirna/.DESCRIPTION:The program FindMiRNA was used to detect putative microRNA genes in the grape genome. A very high number of predictions was retrieved, calling for validation. Nine parameters were calculated and, based on the grape microRNAs dataset available at miRBase, thresholds were defined and applied to FindMiRNA predictions having targets in gene exons. In the resulting subset, predictions were ranked according to precursor positions and sequence similarity, and to target identity. To further validate FindMiRNA predictions, comparisons to the Arabidopsis genome, to the grape Genoscope genome, and to the grape EST collection were performed. Results were stored in a MySQL database and a web interface was prepared to query the database and retrieve predictions of interest.CONCLUSION:The GrapeMiRNA database encompasses 5,778 microRNA predictions spanning the whole grape genome. Predictions are integrated with information that can be of use in selection procedures. Tools added in the web interface also allow to inspect predictions according to gene ontology classes and metabolic pathways of targets. The GrapeMiRNA database can be of help in selecting candidate microRNA genes to be validated.
BACKGROUND:The NCBI dbEST currently contains more than eight million human Expressed Sequenced Tags (ESTs). This wide collection represents an important source of information for gene expression studies, provided it can be inspected according to biologically relevant criteria. EST data can be browsed using different dedicated web resources, which allow to investigate library specific gene expression levels and to make comparisons among libraries, highlighting significant differences in gene expression. Nonetheless, no tool is available to examine distributions of quantitative EST collections in Gene Ontology (GO) categories, nor to retrieve information concerning library-dependent EST involvement in metabolic pathways. In this work we present the Human EST Ontology Explorer (HEOE) http://www.itb.cnr.it/ptp/human_est_explorer, a web facility for comparison of expression levels among libraries from several healthy and diseased tissues.RESULTS:The HEOE provides library-dependent statistics on the distribution of sequences in the GO Direct Acyclic Graph (DAG) that can be browsed at each GO hierarchical level. The tool is based on large-scale BLAST annotation of EST sequences. Due to the huge number of input sequences, this BLAST analysis was performed with the aid of grid computing technology, which is particularly suitable to address data parallel task. Relying on the achieved annotation, library-specific distributions of ESTs in the GO Graph were inferred. A pathway-based search interface was also implemented, for a quick evaluation of the representation of libraries in metabolic pathways. EST processing steps were integrated in a semi-automatic procedure that relies on Perl scripts and stores results in a MySQL database. A PHP-based web interface offers the possibility to simultaneously visualize, retrieve and compare data from the different libraries. Statistically significant differences in GO categories among user selected libraries can also be computed.CONCLUSION:The HEOE provides an alternative and complementary way to inspect EST expression levels with respect to approaches currently offered by other resources. Furthermore, BLAST computation on the whole human EST dataset was a suitable test of grid scalability in the context of large-scale bioinformatics analysis. The HEOE currently comprises sequence analysis from 70 non-normalized libraries, representing a comprehensive overview on healthy and unhealthy tissues. As the analysis procedure can be easily applied to other libraries, the number of represented tissues is intended to increase.
BACKGROUND:The ESTree database (db) is a collection of Prunus persica and Prunus dulcis EST sequences that in its current version encompasses 75,404 sequences from 3 almond and 19 peach libraries. Nine peach genotypes and four peach tissues are represented, from four fruit developmental stages. The aim of this work was to implement the already existing ESTree db by adding new sequences and analysis programs. Particular care was given to the implementation of the web interface, that allows querying each of the database features.RESULTS:A Perl modular pipeline is the backbone of sequence analysis in the ESTree db project. Outputs obtained during the pipeline steps are automatically arrayed into the fields of a MySQL database. Apart from standard clustering and annotation analyses, version VI of the ESTree db encompasses new tools for tandem repeat identification, annotation against genomic Rosaceae sequences, and positioning on the database of oligomer sequences that were used in a peach microarray study. Furthermore, known protein patterns and motifs were identified by comparison to PROSITE. Based on data retrieved from sequence annotation against the UniProtKB database, a script was prepared to track positions of homologous hits on the GO tree and build statistics on the ontologies distribution in GO functional categories. EST mapping data were also integrated in the database. The PHP-based web interface was upgraded and extended. The aim of the authors was to enable querying the database according to all the biological aspects that can be investigated from the analysis of data available in the ESTree db. This is achieved by allowing multiple searches on logical subsets of sequences that represent different biological situations or features.CONCLUSIONS:The version VI of ESTree db offers a broad overview on peach gene expression. Sequence analyses results contained in the database, extensively linked to external related resources, represent a large amount of information that can be queried via the tools offered in the web interface. Flexibility and modularity of the ESTree analysis pipeline and of the web interface allowed the authors to set up similar structures for different datasets, with limited manual intervention.