Quantifying dissimilarity between ecological communities is fundamental to functional community ecology. In this study, we develop a conceptual and analytical framework that integrates species-based and trait-based dissimilarity measures. The core of the proposed approach involves two steps: the first computes the products of species abundances and trait values, while the second combines products into aggregated trait abundances (ATAs) for use in both species- and trait-based analyses. Building upon the additive decomposition of the Marczewski-Steinhaus and Bray-Curtis indices into difference and replacement components, we introduce a suite of novel methods that allow for the independent weighting of species abundances and trait values. We first detail the methodology and elucidate its conceptual foundations. Subsequently, we assess its performance using both illustrative toy examples and ecologically realistic simulated datasets. To demonstrate its practical utility, we apply the method to compare macroinvertebrate assemblages from natural and anthropogenically impacted stream sections. Our findings indicate that the proposed framework provides a continuum between traditional species-based and trait-based approaches. Finally, we offer practical guidance for ecologists on selecting the most appropriate dissimilarity measure based on specific research objectives and data characteristics.
Among the many diversity indices in the ecologist toolbox, measures that can be partitioned into additive terms are particularly useful as the different components can be related to different ecological processes shaping community structure. In this paper, an additive diversity decomposition is proposed to partition the diversity structure of a given community into three complementary fractions: functional diversity, functional redundancy and species dominance. These three components sum up to one. Therefore, they can be used to portray the community structure in a ternary diagram. Since the identification of community‐level patterns is an essential step to investigate the main drivers of species coexistence, the ternary diagram of functional diversity can be used to relate different facets of diversity to community assembly processes more exhaustively than looking only at one index at a time. The value of the proposed diversity decomposition is demonstrated by the analysis of actual abundance data on plant assemblages sampled in grazed and ungrazed grasslands in Tuscany (Central Italy).
We document the makeup of four understory epiphyllous (leaf-inhabiting) bryophyte assemblages along an elevational gradient on the eastern slopes of the Cordillera El Sira in Ucayali, Peru. Epiphylls from two lowland rainforest (250 and 350 m) and two upland cloud forest (1550 and 1800 m) sites were sampled along a transect spanning 6.8 km horizontal and 1.6 km vertical distance, with 74 epiphyllous taxa (69 liverworts and 5 mosses) identified. Change in community composition along the elevational gradient and factors affecting the diversity and distribution of understory epiphyllous bryophytes are explored using various approaches to diversity measurement and multivariate analysis.
Lice were collected from 579 hummingbirds, representing 49 species, in 19 locations in Brazil, Costa Rica, Honduras, Paraguay and Peru, at elevations 0–3000 m above sea level. The following variables were included in an ecological analysis (1) host species' mean body mass, sexual size dimorphism, sexual dichromatism, migratory behaviour and dominance behaviour; (2) mean elevation, mean and predictability of temperature, mean and predictability of precipitation of the host species' geographic area; (3) prevalence and mean abundance of species of lice as measures of infestation. Ordination methods were applied to evaluate data structure. Since the traits are expressed at different scales (nominal, interval and ratio), a principal component analysis based on d-correlations for the traits and a principal coordinates analysis based on the Gower index for species were applied. Lice or louse eggs were found on 80 (13.8%) birds of 22 species. A total of 267 lice of 4 genera, Trochiloecetes, Trochiliphagus, Myrsidea and Leremenopon, were collected, with a total mean intensity of 4.6. There were positive interactions between migration behaviour and infestation indices, with elevational migrants having a higher prevalence and abundance of lice than resident birds. Further, we found weak negative correlations between host body mass and infestation indices and positive correlations between mean elevation and prevalence and abundance of Trochiliphagus. Thus, formerly unknown differences in the ecological characteristics and infestation measures of Trochiliphagus and Trochiloecetes lice were revealed, which allows a better understanding of these associations and their potential impacts on hummingbirds.
Ordinations are compared most commonly by Procrustes methods applicable to points belonging to the same domain, either the objects or the variables describing them. However, no published approach applies to biplots which visualize principal component ordinations of objects and variables simultaneously.To fill this gap, this paper provides a new procedure based on two fundamental, scale-invariant properties of biplots: the cosine of the angle between the vectors pointing to every pair of objects and variables and the mean of the relative length of the two vectors. The method applies to situations in which the number of objects is the same in all ordinations, the number of variables is also the same and, additionally, there is one-to-one correspondence between the points in the different ordinations. The new method is illustrated with two artificial examples and is also applied to the comparison of biplots obtained from real field data representing samples repeated in the same locations over time.It is demonstrated that the new method reveals the similarities and differences between alternative biplots obtained for a given set of objects and variables.We expect that, thanks to the widespread use of principal component analysis in science, the method will receive applications in any studies in which interest lies in the comparison of simultaneous ordinations of objects and variables.
Functional diversity is regarded as a key concept for understanding the link between ecosystem function and biodiversity. The different and ecologically well-defined aspects of the concept are reflected by the so-called functional components, for example, functional richness and divergence. Many authors proposed that components be distinguished according to the multivariate technique on which they rely, but more recent studies suggest that several multivariate techniques, providing different functional representations (such as dendrograms and ordinations) of the community can in fact express the same functional component. Here, we review the relevant literature and find that (1) general ecological acceptance of the field is hampered by ambiguous terminology and (2) our understanding of the role of multivariate techniques in defining components is unclear. To address these issues, we provide new definitions for the three basic functional diversity components namely functional richness, functional divergence and functional regularity. In addition, we present a classification of presence-/absence-based approaches suitable for quantifying these components. We focus exclusively on the binary case for its relative simplicity. We find illogical, as well as logical but unused combinations of components and representations; and reveal that components can be quantified almost independently from the functional representation of the community. Finally, theoretical and practical implications of the new classification are discussed.
Ecological variables may be expressed on four basic measurement scales (nominal, ordinal, interval or ratio), whereas circular variables and those combining a nominal state with other scale types are also common. However, existing methods are not suited to calculate correlations between all pairwise combinations of such variables, preventing the application of standard multivariate techniques. The essence of the new approach is to derive a so‐called difference semimatrix for all pairs of observations for each variable, and then to calculate the matrix correlation based on two such semimatrices. The advantage of this function, termed d ‐correlation, is that comparisons are made on the same logical basis regardless of the measurement scale, allowing for the use of principal components analysis to visualize interrelationships among many variables simultaneously. Further advantages are that missing values in the data are tolerated and that the Gower index of dissimilarity between objects may also be computed. The use of the method is demonstrated on a small toy matrix, an artificial plant trait matrix and a large dataset summarizing ecological features of all vascular plant species of Sardinia, Italy. The source code in R and FORTRAN, and applications for three different operation systems, are provided for computations with results serving as input for other statistical software. The new computational framework will allow the comparison of any types of ecological traits in a mathematically meaningful manner. This option was not available earlier in the field of multivariate statistics, and the method is expected to receive applications in other subject areas as well in which many objects are described in terms of variables expressed on different measurement scales.
Variation in community composition and species turnover are different types of beta diversity, expressing non-directional and directional changes, respectively. While directional changes (e.g. turnover) along geographic gradients can be studied in any direction depending on the hypothesis of interest to researchers, temporal changes can only be meaningfully studied from past to present. Although a wide variety of methods exist for partitioning variation and related community-level phenomena such as similarity, richness difference and nestedness, approaches evaluating species turnover along geographic or temporal gradients, based on an analogous conceptual framework, are rare. We therefore look into the possibilities for examining different aspects of directional changes along a gradient when presence-absence community data are available. Measures of community overlap, as well as species loss and gain from one sampling unit to another along a gradient are combined to define a variety of turnover and nestedness concepts and to derive functions for their quantification. Each concept represents an ecological phenomenon to be indicated (indicandum), whereas measures (indicators) quantify relevant properties of these concepts. The measures use the raw number of species as well as relativized forms in accordance with the well-known Jaccard and Sørensen indices. The main innovation is the development of new measures of directional community change. We demonstrate differences between traditional non-directional and the new directional measures and use several examples to show that actual communities display directional responses to a particular ecological gradient. The new measures therefore reveal an uncovered aspect of community ecology.
A central issue of ecological data analysis is the pairwise comparison of variables describing biological entities and the environment. Difficulties arise with calculations if the measurement scales of the variables differ. In particular, no method is available for measuring the association between a nominal and a fully ranked ordinal variable. Here two coefficients are suggested by reducing this problem to the evaluation of pattern in string representations. The first one is a topological measure that counts the number of other types of elements occurring between pairs of elements of a given state along the entire length of the string, thus providing a global coefficient of aggregation/segregation. The second coefficient is based on counting the number of different elements within substrings generated from the complete string with the moving window technique. Thus, it is a local measure. There is no compact and general formula for calculating these measures, and heuristics are involved for finding the possible minimum and maximum values by algorithmic approximation and Markov Chain Monte Carlo simulation. An R function is provided for computations. The methods are applied to the comparison of nominal variables (biological traits) categorizing marine food web nodes with fully ranked variables describing major graph theory properties of the same nodes in the network. The most descriptive traits (mobility, major functional group) significantly associated with network metrics (weighted indices) were identified from a variety of combinations across three marine ecosystems. These coefficients thus provide an objective, statistically-sound method for identifying ecologically meaningful traits.
We examined the functional strategies and the trait space of 596 European taxa of freshwater macroinvertebrates characterized by 63 fuzzy coded traits belonging to 11 trait groups. Principal component analysis was used to reduce trait dimensionality, to explain ecological strategies, and to quantify the trait space occupied by taxa. Null models were used to compare observed occupancy with theoretical models, and randomization-based analyses were performed to test whether taxonomic relatedness, a proxy of phylogenetic signal, constrains the functional trait space of freshwater macroinvertebrates. We identified four major strategies along which functional traits of the taxa examined show trade-offs. In agreement with expectations and in contrast to existing evidence we found that life cycles and aquatic strategies are important in shaping functional structure of freshwater macroinvertebrates. Our results showed that the taxonomic groups examined fill remarkably different niches in the functional trait space. We found that the functional trait space of freshwater macroinvertebrates is reduced compared to the range of possibilities that would exist if traits varied independently. The observed decrease was between 23.44 and 44.61% depending on the formulation of the null expectations. We demonstrated also that taxonomic relatedness constrains the functional trait space of macroinvertebrates.
Networks of trophic interactions provide a lot of information on the functioning of marine ecosystems. Beyond feeding habits, three additional traits (mobility, size, and habitat) of various organisms can complement this trophic view. The combination of traits and food web positions are studied here on a large food web database. The aim is a better description and understanding of ecological roles of organisms and the identification of the most important keystone species. This may contribute to develop better ecological indicators (e.g., keystoneness) and help in the interpretation of food web models. We use food web data from the Ecopath with Ecosim (EwE) database for 92 aquatic ecosystems. We quantify the network position of organisms by 18 topological indices (measuring centrality, hierarchy, and redundancy) and consider their three, categorical traits (e.g., for mobility: sessile, drifter, limited mobility, and mobile). Relationships are revealed by multivariate analysis. We found that topological indices belong to six different categories and some of them nicely separate various trait categories. For example, benthic organisms are richly connected and mobile organisms occupy higher food web positions.
Sister groups at the root of large plant clades (with ≥1000 extant species) are described in terms of variables expressing species richness, geographical distribution, age, diversification rate and speed of molecular evolution. Standard statistical tools and principal component analysis are used to determine and visualize their relationships. In the majority of the 167 groups examined, one clade (the minor) has many fewer species than its sister (the major clade). A striking phenomenon is that many minor clades (17%) contain four or fewer species only. Asymmetry is largely, but not generally reflected by proportions in geographical range and relative speed of molecular evolution: species in the minor clade are more widely distributed than the major clade in 8 groups and – based on available information – we detected longer maximum branch lengths on the minor lineage in 14 cases. The basal lineages considerably differ in diversification rate; the differences are of the same magnitude in many groups as the diversification rate itself calculated for families and orders by other authors. For the major clade, we found a positive relationship between the species richness of clades and geographical range, and negative correlation between the logarithm of species richness and clade age.
Thinking about the dynamics of populations of plants and animals goes back to Linnaeus. He used at least three examples to show what happens when the population of a species grows without limitations and to illustrate the potential reproductive capacity of organisms. We examined the mathematical precision of calculations Linnaeus used in presenting these examples and reviewed the assumptions under which Linnaeus's conclusions are valid. In the case of a slowly reproducing annual plant, additionally cited by Darwin, the final result was incorrect, although little different from the true value. In the example of a pair of pigeons, the calculations were accurate, although the well-known fact that pigeons breed several times throughout their lifetime was ignored. Though the input parameters must have been unknown to Linnaeus, a short statement in Systema naturae regarding the population increase and feeding capacity of bluebottle flies was found fairly correct and robust enough to withstand minor changes in input parameters.
A long-standing problem in biological data analysis is the unintentional absence of values for some observations or variables, preventing the use of standard multivariate exploratory methods, such as principal component analysis (PCA). Solutions include deleting parts of the data by which information is lost, data imputation, which is always arbitrary, and restriction of the analysis to either the variables or observations, thereby losing the advantages of biplot diagrams. We describe a minor modification of eigenanalysis-based PCA in which correlations or covariances are calculated using different numbers of observations for each pair of variables, and the resulting eigenvalues and eigenvectors are used to calculate component scores such that missing values are skipped. This procedure avoids artificial data imputation, exhausts all information from the data and allows the preparation of biplots for the simultaneous display of the ordination of variables and observations. The use of the modified PCA, called InDaPCA (PCA of Incomplete Data) is demonstrated on actual biological examples: leaf functional traits of plants, functional traits of invertebrates, cranial morphometry of crocodiles and fish hybridization data – with biologically meaningful results. Our study suggests that it is not the percentage of missing entries in the data matrix that matters; the success of InDaPCA is mostly affected by the minimum number of observations available for comparing a given pair of variables. In the present study, interpretation of results in the space of the first two components was not hindered, however.
Large ecological data matrices may be incomplete for various reasons, preventing the use of standard multidimensional scaling (ordination) and cluster analysis packages. Although there exist a few resemblance functions that allow missing scores, there is no theoretical background and software support for most distance and similarity coefficients potentially applied in multivariate data analysis. We provide a general framework for a precise mathematical redefinition of a large set of resemblance functions originally developed for complete data sets with presence-absence (binary) or ratio-scale variables. Included are coefficients which consider double absences in abundance data. Potential problems with the use of these functions are discussed, with the conclusion that incompleteness of data would rarely if ever influence greatly the interpretability of ordinations and classifications. An R function described in the Appendix represents a link to R. We also provide a stand-alone WINDOWS application for users of other computer programs. The new software will allow users of standard data analysis packages to perform multivariate analysis using a wide variety of resemblance coefficients even if the data are incomplete for whatever reason.
BACKGROUND:Hawthorn species (Crataegus L.; Rosaceae tribe Maleae) form a well-defined clade comprising five subgeneric groups readily distinguished using either molecular or morphological data. While multiple subsidiary groups (taxonomic sections, series) are recognized within some subgenera, the number of and relationships among species in these groups are subject to disagreement. Gametophytic apomixis and polyploidy are prevalent in the genus, and disagreement concerns whether and how apomictic genotypes should be recognized taxonomically. Recent studies suggest that many polyploids arise from hybridization between members of different infrageneric groups.METHODS:We used target capture and high throughput sequencing to obtain nucleotide sequences for 257 nuclear loci and nearly complete chloroplast genomes from a sample of hawthorns representing all five currently recognized subgenera. Our sample is structured to include two examples of intersubgeneric hybrids and their putative diploid and tetraploid parents. We queried the alignment of nuclear loci directly for evidence of hybridization, and compared individual gene trees with each other, and with both the maximum likelihood plastome tree and the nuclear concatenated and multilocus coalescent-based trees. Tree comparisons provided a promising, if challenging (because of the number of comparisons involved) method for visualizing variation in tree topology. We found it useful to deploy comparisons based not only on tree-tree distances but also on a metric of tree-tree concordance that uses extrinsic information about the relatedness of the terminals in comparing tree topologies.RESULTS:We obtained well-supported phylogenies from plastome sequences and from a minimum of 244 low copy-number nuclear loci. These are consistent with a previous morphology-based subgeneric classification of the genus. Despite the high heterogeneity of individual gene trees, we corroborate earlier evidence for the importance of hybridization in the evolution of Crataegus. Hybridization between subgenus Americanae and subgenus Sanguineae was documented for the origin of Sanguineae tetraploids, but not for a tetraploid Americanae species. This is also the first application of target capture probes designed with apple genome sequence. We successfully assembled 95% of 257 loci in Crataegus, indicating their potential utility across the genera of the apple tribe.
The present article has two primary objectives. First, the article provides a historical overview of graphical tools used in the past centuries for summarizing the classification and phylogeny of plants. It is emphasized that each published diagram focuses on only a single or a few aspects of the present and past of plant life on Earth. Therefore, these diagrams are less useful for communicating general knowledge in botanical research and education. Second, the article offers a solution by describing the principles and methods of constructing a lesserknown image type, the coral, whose potential usefulness in phylogenetics was first raised by Charles Darwin. Cladogram topology, phylogenetic classification and nomenclature, diversity of taxonomic groups, geological timescale, paleontological records, and other relevant information on the evolution of Archaeplastida are simultaneously condensed for the first time into the same figure – the Coral of Plants. This image is shown in two differently scaled parts to efficiently visualize as many details as possible, because the evolutionary timescale is much longer, and the extant diversity is much lower for red and green algae than for embryophytes. A fundamental property of coral diagrams, that is their self-similarity, allows for the redrawing of any part of the diagram at smaller scales.