DNA barcoding using the nuclear internal transcribed spacer (ITS) has become prevalent in surveys of fungal diversity. This approach is, however, associated with numerous caveats, including the desire for speed, rather than accuracy, through the use of automated analytical pipelines, and the shortcomings of reference sequence repositories. Here we use the case of a specimen of the bracket fungus Trametes s.lat. (which includes the common and widespread turkey tail, T. versicolor) to illustrate these problems. The material was collected in Vietnam as part of a biodiversity inventory including DNA barcoding approaches for arthropods, plants and fungi. The ITS barcoding sequence of the query taxon was compared against reference sequences in GenBank and the curated fungal ITS database UNITE, using BLASTn and MegaBLAST, and was subsequently analysed in a multiple alignment-based phylogenetic context through a maximum likelihood tree including related sequences. Our results initially indicated issues with BLAST searches, including the use of Frairwise local alignments and sorting through Total score and E value, rather than Percentage identity, as major shortcomings of the DNA barcoding approach. However, after thorough analysis of the results, we concluded that the single most important problem of this approach was incorrect sequence labelling, calling for the implementation of third-party annotations or analogous approaches in primary sequence repositories. In addition, this particular example revealed problems of improper fungal nomenclature, which required reinstatement of the genus name Cubamyces (= Leioirametes), with three new combinations: C. flavidus, C lactineus and C. menziesii. The latter was revealed as the correct identification of the query taxon, although the name did not appear among the best BLAST hits. While the best BLAST hits did correspond to the target taxon in terms of sequence data, their label names were misleading or unresolved, including [Fungal endophyte], [Uncultured fungus], Basidiomycota, Trametes cf. cubensis, Lenzites elegans and Geotrichum candidum (an unrelated ascomycetous contaminant). Our study demonstrates that accurate identification of fungi through molecular barcoding is currently not a fast-track approach that can be achieved through automated pipelines.
Forests of SW Ethiopia constitute the native habitat of Coffea arabica and also the place where domestication of Arabica coffee started. Selection from wild populations has led to numerous landraces (farmer’s varieties) and cultivars. Inter-simple sequence repeats (ISSRs) were generated from a representative set of forest coffee populations and landraces across Ethiopia. For the broad diversity assessment, nine di- and tri-nucleotide ISSR primers were applied, as chosen from a total of 102 primers tested initially. Tetranucleotide ISSR primers differed in amplifying fingerprints that could hardly be analysed due to excessive variation. Tree building analysis (NJ, UPGMA) of 84 polymorphic loci amplified for 125 C. arabica individuals provided evidence for several groups of related genotypes occurring in certain geographical areas of Ethiopia and underscored the existence of wild coffee distinct from landraces. Landraces seem to have originated in different geographical areas of Ethiopia in a stepwise domestication process. While the overall geographical signal in the dataset was weak, analysis in a Bayesian framework using the admixture model with geographical priors in STRUCTURE recovered some genetic clustering. Based on Shannon’s diversity index, populations from Yayu (0.47) and Bonga (0.46) showed highest diversity, followed by individuals from Berhane Kontir (0.41). A likely scenario for the differentiation of C. arabica after an allopolyploidization event is that the hierarchical-geographical patterning of wild Coffea genotypes expected from stepwise range extension was obscured by recent or ancient gene flow. The diversity and geographical distribution of autochthonous C. arabica genotypes indicates the need for a multi-site in situ conservation approach.
Field observations of morphologically intermediate water lilies in central Canada suggested a hybrid origin involving the parents Nymphaea odorata Aiton and Nymphaea leibergii Morong despite the fertile nature of these plants. Sequencing of the nrITS and the plastid rps4–trnT–trnF regions further including all members of the north temperate Nymphaea subg. Nymphaea clade, and samples from other hybrids occurring in North America in the wild, allowed us to determine that individuals of N. leibergii and of N. odorata were the maternal and paternal parents, respectively. Hybrids of New England have all proven to be sterile, are genetically variable, and probably are F1. By comparison, the plants of east-central Saskatchewan and west-central Manitoba are fully fertile and genetically uniform based on ISSR and sequence data. On the basis of this evidence, the latter are here described as a new species, Nymphaea loriana sp. nov., which may have originated during the Holocene climatic optimum about 6000 years ago in a past contact zone of the parents. Further hybrids detected between N. odorata and Nymphaea tetragona Georgi, as well as between N. leibergii and N. tetragona, were always sterile. Gene trees of the temperate clade of Nymphaea converge on a clade of small-flowered water lilies (sect. Chamaenymphaea), including N. leibergii, N. tetragona, and Nymphaea pygmaea. Nuclear ITS further resolves an American clade (Nymphaea mexicana Zucc. – N. odorata) sister to all remaining species. This split into two major subclades also appears in the otherwise less resolved rps4–trnT–trnF tree. Thus the origin of N. loriana is a reticulation between long-separated parental lineages.
Comparative sequencing of >7 kb of highly variable chloroplast genome regions (atpB-rbcL, trnS-trnG, rpl22-rps19, and rps19-rpl2 spacers; introns in atpF, trnG, trnK, and rpl16) with microsatellites known from other angiosperms was carried out in Coffea. Samples comprised 8 diploid species of Coffea, 5 individuals of tetraploid C. arabica representing geographically distant wild populations from Ethiopia, 2 commercial cultivars of C. arabica, and Psilanthus leroyi and Ixora coccinea as outgroups. Phylogeny reconstruction using maximum parsimony and Bayesian inference resulted in congruent topologies with high support for C. arabica and C. eugenioides being sisters. Partitioned analyses showed that all regions except the atpB-rbcL spacer resolved this sister-group, although this was often unsupported. The large sequence data set further shows that chloroplast genomes of C. arabica and C. eugenioides each possess apomorphies, indicating that not C. eugenioides but an ancestor or close relative of C. eugenioides is the maternal parent of C. arabica. Seven variable chloroplast microsatellites were characterized in Coffea. Most microsatellites are poly(A/T) stretches, whereas one in the trnS-trnG spacer has an (AT)n motif. Most strikingly, all individuals of C. arabica possess identical sequences, suggesting a single chloroplast haplotype. This can be explained by a recent origin of C. arabica in a unique allopolyploidization event, or by severe bottleneck effects in the evolutionary history of the species. Reconstruction of the evolution of microstructural mutations shows much higher levels of homoplasy in microsatellite loci than in other parts of spacers and introns. Microsatellites are inferred to evolve by insertion and deletion of 1 to 3 motif copies in one step.