DNA sequences were obtained from 32 blade-forming Ulva specimens collected in 2018 and 2019 from four islands in the Galapagos Archipelago: Fernandina, Floreana, Isabela and San Cristobal. The loci sequenced were nuclear encoded ITS and plastid encoded rbcL and tufA, all recognized as barcode markers for green algae. Four species were found, Ulva adhaerens, U. lactuca, U. ohnoi and U. tanneri, all of which have had their type specimens sequenced, ensuring the correct application of these names. Only one of these, U. lactuca, was reported historically from the archipelago. Ulva adhaerens was the species most commonly collected and widely distributed, occurring on all four islands. Previously known only from Japan and Korea, this is the first report of U. adhaerens from the southeast Pacific Ocean. Ulva ohnoi was collected on three islands, Isabela, Floreana, and San Cristobal, and U. lactuca only on the last two. Ulva tanneri is a diminutive, 1-2 cm tall, high intertidal species that is easily overlooked, but likely far more common than the one specimen that was collected. This study of blade-forming Ulva species confirms that a concerted effort, using DNA sequencing, is needed to document the seaweed flora of the Galapagos Archipelago.
AimHybridization is thought to have played an important role in shaping the evolutionary history of diverse island taxa. Here, we propose an ecological and evolutionary framework for understanding the causes and consequences of heterospecific mating on islands-with and without introgressive hybridization. We use this framework to support our main contention that cases of secondary contact among endemic species should commonly result in introgressive hybridization whereas cases of contact between endemic and introduced species should commonly result in reproductive interference (RIN)-resulting in two qualitatively different faces of secondary contact on islands.LocationCanary Islands, Galapagos, New Zealand, Caribbean and Hawaii.Taxa705 vertebrate, invertebrate and plant species spanning 167 genera and 99 families.MethodsUsing a quantitative analysis of empirical research on secondary contact on islands, we weigh evidence for the causes and consequences of secondary contact and heterospecific mating on islands. In particular, we compare cases of secondary contact between endemic species versus secondary contact between endemic and introduced species.ResultsWe find that the drivers of secondary contact and heterospecific mating on islands most frequently reported in the literature are disturbance, long-distance (e.g. inter-island) dispersal and compromised assortative mating. We find support for the hypothesis that extensive introgression is a more common outcome between endemic species while RIN is a more common outcome between endemic and introduced species.Main ConclusionsWe conclude that there are biological reasons to predict secondary contact and heterospecific mating to be common on islands for all taxa, but that the consequence of secondary contact is categorically different for contact between endemic species and contact between endemic and introduced species. We conclude that the former likely explains the apparent frequency of hybridization on islands, while the latter presents a cryptic and underappreciated conservation threat.
Populations suffer two types of stochasticity: demographic stochasticity, from sampling error in offspring number, and environmental stochasticity, from temporal variation in the growth rate. By modelling evolution through phenotypic selection following an abrupt environmental change, we investigate how genetic and demographic dynamics, as well as effects on population survival of the genetic variance and of the strength of stabilizing selection, differ under the two types of stochasticity. We show that population survival probability declines sharply with stronger stabilizing selection under demographic stochasticity, but declines more continuously when environmental stochasticity is strengthened. However, the genetic variance that confers the highest population survival probability differs little under demographic and environmental stochasticity. Since the influence of demographic stochasticity is stronger when population size is smaller, a slow initial decline of genetic variance, which allows quicker evolution, is important for population persistence. In contrast, the influence of environmental stochasticity is population-size-independent, so higher initial fitness becomes important for survival under strong environmental stochasticity. The two types of stochasticity interact in a more than multiplicative way in reducing the population survival probability. Our work suggests the importance of explicitly distinguishing and measuring the forms of stochasticity during evolutionary rescue.
Organismal anatomy is a hierarchical system of anatomical entities often imposing dependencies among multiple morphological characters. Ontologies provide a formal and computable framework for incorporating prior biological knowledge about anatomical dependencies in models of trait evolution. They also offer new opportunities for working with semantic representations of morphological data. In this work, we present a new R package— rphenoscate —that enables incorporating ontological knowledge in evolutionary analyses and exploring semantic patterns of morphological data. In conjunction with rphenoscape , it allows for assembling synthetic phylogenetic character matrices from semantic phenotypes of morphological data. We showcase the package functionality with data sets from bees and fishes. We demonstrate that ontologies can be employed to automatically set up evolutionary models accounting for trait dependencies in stochastic character mapping. We also demonstrate how ontology annotations can be explored to interrogate patterns of morphological evolution. Finally, we demonstrate that synthetic character matrices assembled from semantic phenotypes retain most of the phylogenetic information from their original data sets. Ontologies will become important tools for integrating anatomical knowledge into phylogenetic methods and making morphological data FAIR compliant—a critical step of the ongoing ‘phenomics’ revolution. Our new package offers key advancements towards this goal.
Invasive species are a major threat to Earth's biodiversity, particularly in unique ecosystems such as the Galapagos Islands. Research on the ecology and genetics of these invasive species is essential to understand their interactions with native and endemic flora, and to alleviate the negative effects of these invasions. In one of our studies, it was found that that the most likely origin of the invasive tomatillo in Galapagos is the central region of mainland Ecuador. Hybridization between the invasive and two endemic tomato species was observed. This could imply an imminent fast extinction risk for the endemic tomatoes, as several of the populations reported decades ago couldn't be found anymore. Moreover, genetic hijacking by the invasive species could lead to an even more aggressive invasive tomatillo in the Galapagos Islands.We also study the Guava, which is one of the most aggressive invasive plants in the Galapagos Islands, displacing and outcompeting its endemic relative, the guayabillo, for resources and space. There is also the possibility of hybridization of guava with its endemic relative, which could lead to the fast extinction of the latter. However, this hybridization is probably not occurring, yet guava could still interfere with the successful reproduction of guayabillo, decreasing its populations. The most likely origin of the guava in Galapagos would be the Central Highlands of mainland Ecuador. A better understanding of the interactions between invasive and endemic plants can contribute to the conservation of the endemic species and a better management of invasive species.
Invasive species can interact with native relatives in a variety of ways which may jeopardize their long-term coexistence. Here we show that interactions with an invasive species of guava ( Psidium guajava ) appear to be driving the local exclusion and regional decline of guayabillo ( Psidium galapageium ), a tree species endemic to the Galápagos archipelago. We find evidence consistent with recent historic exclusion of guayabillo from the highlands of San Cristóbal Island, signatures of ongoing demographic decline in sympatric populations at lower elevations, and evidence suggesting that the four coinhabited islands represent points along a time series of regional decline, with the extent of guayabillo decline depending on the date that guava was introduced to each island. Based on these results, we then use the percentage of guava cover surrounding guayabillo populations to target populations that are at imminent risk of exclusion to aid in prioritizing management targets.
Morphology remains a primary source of phylogenetic information for many groups of organisms, and the only one for most fossil taxa. Organismal anatomy is not a collection of randomly assembled and independent "parts", but instead a set of dependent and hierarchically nested entities resulting from ontogeny and phylogeny. How do we make sense of these dependent and at times redundant characters? One promising approach is using ontologies-structured controlled vocabularies that summarize knowledge about different properties of anatomical entities, including developmental and structural dependencies. Here, we assess whether evolutionary patterns can explain the proximity of ontology-annotated characters within an ontology. To do so, we measure phylogenetic information across characters and evaluate if it matches the hierarchical structure given by ontological knowledge-in much the same way as across-species diversity structure is given by phylogeny. We implement an approach to evaluate the Bayesian phylogenetic information (BPI) content and phylogenetic dissonance among ontology-annotated anatomical data subsets. We applied this to data sets representing two disparate animal groups: bees (Hexapoda: Hymenoptera: Apoidea, 209 chars) and characiform fishes (Actinopterygii: Ostariophysi: Characiformes, 463 chars). For bees, we find that BPI is not substantially explained by anatomy since dissonance is often high among morphologically related anatomical entities. For fishes, we find substantial information for two clusters of anatomical entities instantiating concepts from the jaws and branchial arch bones, but among-subset information decreases and dissonance increases substantially moving to higher-level subsets in the ontology. We further applied our approach to address particular evolutionary hypotheses with an example of morphological evolution in miniature fishes. While we show that phylogenetic information does match ontology structure for some anatomical entities, additional relationships and processes, such as convergence, likely play a substantial role in explaining BPI and dissonance, and merit future investigation. Our work demonstrates how complex morphological data sets can be interrogated with ontologies by allowing one to access how information is spread hierarchically across anatomical concepts, how congruent this information is, and what sorts of processes may play a role in explaining it: phylogeny, development, or convergence. [Apidae; Bayesian phylogenetic information; Ostariophysi; Phenoscape; phylogenetic dissonance; semantic similarity.]
Synthesis centers are a recently-developed form of scientific organization that catalyzes and supports a form of interdisciplinary research that integrates diverse theories, methods and data across spatial or temporal scales, scientific phenomena, and forms of expertise to increase the generality, parsimony, applicability, or empirical soundness of scientific explanations. Research has shown the synthesis working group to be a distinctive form of scientific collaboration that reliably produces consequential, high-impact publications, but no one has asked: do synthesis working groups produce publications that are substantially more diverse than those produced outside of synthesis centers, and if so, how and with what effects? We have investigated these questions through a novel textual analysis. We found that if diversity is measured solely by mean difference in the Rao-Stirling (aggregate) measure of diversity, then the answer is no. But synthesis center papers have significantly greater variety and balance, but significantly lower disparity, than papers in the reference corpus. Synthesis center influence is mediated by the greater size of synthesis center collaborations (numbers of authors, distinct institutions, and references) but even when taking size into account, there is a persistent direct effect: synthesis center papers have significantly greater variety and balance, but less disparity, than papers in the reference corpus. We conclude by inviting further exploration of what this novel textual analysis approach might reveal about interdisciplinary research and by offering some practical implications of our results.
Despite the increase in the number of journals issuing data policies requiring authors to make data underlying reporting findings publicly available, authors do not always do so, and when they do, the data do not always meet standards of quality that allow others to verify or extend published results. This phenomenon suggests the need to consider the effectiveness of journal data policies to present and articulate transparency requirements, and how well they facilitate (or hinder) authors' ability to produce and provide access to data, code, and associated materials that meet quality standards for computational reproducibility. This article describes the results of a research study that examined the ability of journal-based data policies to: 1) effectively communicate transparency requirements to authors, and 2) enable authors to successfully meet policy requirements. To do this, we conducted a mixed-methods study that examined individual data policies alongside editors' and authors' interpretation of policy requirements to answer the following research questions. Survey responses from authors and editors along with results from a content analysis of data policies found discrepancies among editors' assertion of data policy requirements, authors' understanding of policy requirements, and the requirements stated in the policy language as written. We offer explanations for these discrepancies and offer recommendations for improving authors' understanding of policies and increasing the likelihood of policy compliance.
The role of genetic architecture in adaptation to novel environments has received considerable attention when the source of adaptive variation is de novo mutation. Relatively less is known when the source of adaptive variation is inter- or intraspecific hybridization. We model hybridization between divergent source populations and subsequent colonization of an unoccupied novel environment using individual-based simulations to understand the influence of genetic architecture on the timing of colonization and the mode of adaptation. We find that two distinct categories of genetic architecture facilitate rapid colonization but that they do so in qualitatively different ways. For few and/or tightly linked loci, the mode of adaptation is via the recovery of adaptive parental genotypes. With many unlinked loci, the mode of adaptation is via the generation of novel hybrid genotypes. The first category results in the shortest colonization lag phases across the widest range of parameter space, but further adaptation is mutation limited. The second category takes longer and is more sensitive to genetic variance and dispersal rate, but can facilitate adaptation to environmental conditions that exceed the tolerance of parental populations. These findings have implications for understanding the origins of biological invasions and the success of hybrid populations.
Conferences with contributed talks grouped into multiple concurrent sessions pose an interesting scheduling problem. From an attendee's perspective, choosing which talks to visit when there are many concurrent sessions is challenging since an individual may be interested in topics that are discussed in different sessions simultaneously. The frequency of topically similar talks in different concurrent sessions is, in fact, a common cause for complaint in post-conference surveys. Here, we introduce a practical solution to the conference scheduling problem by heuristic optimization of an objective function that weighs the occurrence of both topically similar talks in one session and topically different talks in concurrent sessions. Rather than clustering talks based on a limited number of preconceived topics, we employ a topic model to allow the topics to naturally emerge from the corpus of contributed talk titles and abstracts. We then measure the topical distance between all pairs of talks. Heuristic optimization of preliminary schedules seeks to balance the topical similarity of talks within a session and the dissimilarity between concurrent sessions. Using an ecology conference as a test case, we find that stochastic optimization dramatically improves the objective function relative to the schedule manually produced by the program committee. Approximate Integer Linear Programming can be used to provide a partially-optimized starting schedule, but the final value of the discrimination ratio (an objective function used to estimate coherence within a session and disparity between concurrent sessions) is surprisingly insensitive to the starting schedule. Furthermore, we show that, in contrast to the manual process, arbitrary scheduling constraints are straightforward to include. We applied our method to a second biology conference with over 1,000 contributed talks plus scheduling constraints. In a randomized experiment, biologists responded similarly to a machine-optimized schedule and a highly modified schedule produced by domain experts on the conference program committee.
There is a growing body of research on the evolution of anatomy in a wide variety of organisms. Discoveries in this field could be greatly accelerated by computational methods and resources that enable these findings to be compared across different studies and different organisms and linked with the genes responsible for anatomical modifications. Homology is a key concept in comparative anatomy; two important types are historical homology (the similarity of organisms due to common ancestry) and serial homology (the similarity of repeated structures within an organism). We explored how to most effectively represent historical and serial homology across anatomical structures to facilitate computational reasoning. We assembled a collection of homology assertions from the literature with a set of taxon phenotypes for vertebrate fins and limbs from the Phenoscape Knowledgebase (KB). Using six competency questions, we evaluated the reasoning ramifications of two logical models: the Reciprocal Existential Axioms Homology Model (REA) and the Ancestral Value Axioms Homology Model (AVA). Both models returned the user-expected results for all but one historical homology query and all serial homology queries. Additionally, for each competency question, the AVA model returns the search term and any subtypes. We identify some challenges of implementing complete homology queries due to limitations of OWL reasoning. This work lays the foundation for homology reasoning to be incorporated into other ontology-based tools, such as those that enable synthetic supermatrix construction and candidate gene discovery.
Semantic similarity has been used for comparing genes, proteins, phenotypes, diseases, etc. for various biological applications. The rise of ontology-based data representation in biology has also led to the development of several semantic similarity metrics that use different statistics to estimate similarity. Although semantic similarity has become a crucial computational tool in several applications, there has not been a formal evaluation of the statistical sensitivity of these metrics and their ability to recognize similarity between distantly related biological objects. Here, we present a statistical sensitivity comparison of five semantic similarity metrics (Jaccard, Resnik, Lin, Jiang u0026 Conrath, and Hybrid Relative Specificity Similarity) representing three different kinds of metrics (Edge based, Node based, and Hybrid) and explore key parameter choices that can impact sensitivity. Furthermore, we compare four methods of aggregating individual annotation similarities to estimate similarity between two biological objects - All Pairs, Best Pairs, Best Pairs Symmetric, and Groupwise. To evaluate sensitivity in a controlled fashion, we explore two different models for simulating data with varying levels of similarity and compare to the noise distribution using resampling. Source data are derived from the Phenoscape Knowledgebase of evolutionary phenotypes. Our results indicate that the choice of similarity metric along with different parameter choices can substantially affect sensitivity. Among the five metrics evaluated, we find that Resnik similarity shows the greatest sensitivity to weak semantic similarity. Among the ways to combine pairwise statistics, the Groupwise approach provides the greatest discrimination among values above the sensitivity threshold, while the Best Pairs statistic can be parametrically tuned to provide the highest sensitivity. Our findings serve as a guideline for an appropriate choice and parameterization of semantic similarity metrics, and point to the need for improved reporting of the statistical significance of semantic similarity matches in cases where weak similarity is of interest.
Natural language descriptions of organismal phenotypes - a principal object of study in biology, are abundant in biological literature. Expressing these phenotypes as logical statements using formal ontologies would enable large-scale analysis on phenotypic information from diverse systems. However, considerable human effort is required to make the semantics of phenotype descriptions amenable to machine reasoning by (a) recognizing appropriate on-tological terms for entities in text and (b) stringing these terms into logical statements. Most existing Natural Language Processing tools stop at entity recognition, leaving a need for tools that can assist with both aspects of the task. The recently described Semantic CharaParser aims to meet this need. We describe the first expert-curated Gold Standard corpus for ontology-based annotation of phenotypes from the systematics literature. We use it to evaluate Semantic CharaParser’s annotations and explore differences in performance between humans and machine. We use four annotation accuracy metrics that can account for both semantically identical and similar matches. We found that machine-human consistency was significantly lower than inter-curator (human–human) consistency. Surprisingly, allowing curators access to external information that was not available to Semantic CharaParser did not significantly increase the similarity of their annotations to the Gold Standard nor have a significant effect on inter-curator consistency. We found that the similarity of machine annotations to the Gold Standard increased after new ontology terms relevant to the input text had been added. Evaluation by the original authors of the character descriptions indicated that the Gold Standard annotations came closer to representing their intended meaning than did either the curator or machine annotations. These findings point toward ways to better design of software to augment human curators, and the Gold Standard corpus will allow training and assessment of new tools to improve phenotype annotation accuracy at scale.
We report the complete plastome sequences of an endemic and an unidentified species from the genus Psidium in the Galápagos Islands ( P. galapageium and Psidium sp. respectively).
The study of how the observable features of organisms, i.e., their phenotypes, result from the complex interplay between genetics, development, and the environment, is central to much research in biology. The varied language used in the description of phenotypes, however, impedes the large scale and interdisciplinary analysis of phenotypes by computational methods. The Phenoscape project (www.phenoscape.org) has developed semantic annotation tools and a gene–phenotype knowledgebase, the Phenoscape KB, that uses machine reasoning to connect evolutionary phenotypes from the comparative literature to mutant phenotypes from model organisms. The semantically annotated data enables the linking of novel species phenotypes with candidate genes that may underlie them. Semantic annotation of evolutionary phenotypes further enables previously difficult or novel analyses of comparative anatomy and evolution. These include generating large, synthetic character matrices of presence/absence phenotypes based on inference, and searching for taxa and genes with similar variation profiles using semantic similarity. Phenoscape is further extending these tools to enable users to automatically generate synthetic supermatrices for diverse character types, and use the domain knowledge encoded in ontologies for evolutionary trait analysis. Curating the annotated phenotypes necessary for this research requires significant human curator effort, although semi-automated natural language processing tools promise to expedite the curation of free text. As semantic tools and methods are developed for the biodiversity sciences, new insights from the increasingly connected stores of interoperable phenotypic and genetic data are anticipated.
Scientists are adept at comparing genomic sequences. The collection of more such data promises to increase our ability to determine gene function, discover and describe biological processes, and prioritize causative variants of interest that underlie disease response. Yet the question remains: Can we compare phenotypes or traits of interest across disciplines in a manner similar to how we compare genomic sequences? Here we present examples of 'semantic reasoning' - computational methodologies that enable computation across organized formal phenotypic representations. These methods facilitate the analysis of phenotype information across species, domains of knowledge, people, and computers. We review representative examples of successful semantic reasoning to recover known biological phenomena in medical and agricultural applications. Necessary changes in how we collect, analyze, and share data to enable such computations are presented, and database and analytic tool suites for these sorts of analyses are described.