The Poritiinae are a diverse subfamily of lycaenid butterflies with about 700 species divided into two major groups: the Asian endemic tribe Poritiini, and the African endemic tribe Liptenini. Among these, the Liptenini are notable for their lichenivorous diet and the strong but apparently non‐mutualistic ant associations of many species. We present the first molecular phylogeny for this subfamily, based on data from 14 gene regions, and including 218 representatives from 177 taxa (approximately 25% of species) in 50 of the 58 (86%) recognized genera. From this analysis, we confirm the division of the subfamily into two tribes, and we rearrange the Liptenini tribe into six subtribes, Durbaniina, Pentilina, Liptenina, Iridanina and Epitolina, plus a new tribe, Cooksoniina subtrib. n., to fill a gap in the nomenclature revealed by the phylogenetic analysis. We also point to several genera in need of further taxonomic revision. Ancestral range reconstruction could not infer the range of the common ancestor of the Poritiinae; however, the common ancestor of the Poritiini was likely Asian, while that of the Liptenini was likely African, with subsequent narrowing of ranges in several lineages.
In intimate ecological interactions, the interdependency of species may result in correlated demographic histories. For species of conservation concern, understanding the long-term dynamics of such interactions may shed light on the drivers of population decline. Here, we address the demographic history of the monarch butterfly, Danaus plexippus, and its dominant host plant, the common milkweed Asclepias syriaca (A. syriaca), using broad-scale sampling and genomic inference. Because genetic resources for milkweed have lagged behind those for monarchs, we first release a chromosome-level genome assembly and annotation for common milkweed. Next, we show that despite its enormous geographic range across eastern North America, A. syriaca is best characterized as a single, roughly panmictic population. Using approximate Bayesian computation with random forests (ABC-RF), a machine learning method for reconstructing demo-graphic histories, we show that both monarchs and milkweed experienced population expansion during the most recent recession of North American glaciers 10,000-20,000 years ago. Our data also identify concurrent population expansions in both species during the large-scale clearing of eastern forests (-200 years ago). Finally, we find no evidence that either species experienced a reduction in effective population size over the past 75 years. Thus, the well-documented decline of monarch abundance over the past 40 years is not visible in our genomic dataset, reflecting a possible mismatch of the overwintering census population to effective population size in this species.
In intimate ecological interactions, the interdependency of species may result in correlated demographic histories. For species of conservation concern, understanding the long-term dynamics of such interactions may shed light on the drivers of population decline. Here we address the demographic history of the monarch butterfly, Danaus plexippus , and its dominant host plant, the common milkweed Asclepias syriaca , using broad-scale sampling and genomic inference. Because genetic resources for milkweed have lagged behind those for monarchs, we first release a chromosome-level genome assembly and annotation for common milkweed. Next, we show that despite its enormous geographic range across eastern North America, A. syriaca is best characterized as a single, roughly panmictic population. Using Approximate Bayesian Computation via Random Forests (ABC-RF), a machine learning method for reconstructing demographic histories, we show that both monarchs and milkweed experienced concurrent range expansion during the most recent recession of North American glaciers ∼12,000 years ago. Our data identify an expansion of milkweed during the large-scale clearing of eastern forests (∼200 years ago) but was inconclusive as to expansion or contraction of the monarch butterfly population during this time. Finally, our results indicate that neither species experienced a population contraction over the past 75 years. Thus, the well-documented decline of monarch abundance over the past 40 years is not visible in our genomic dataset, reflecting a possible mismatch of the overwintering census population to effective population size in this species. ### Competing Interest Statement The authors have declared no competing interest.
The Asian tiger mosquito, Aedes albopictus, is an invasive vector mosquito of substantial public health concern. The large genome size (~1.19–1.28 Gb by cytofluorometric estimates), comprised of ~68% repetitive DNA sequences, has made it difficult to produce a high-quality genome assembly for this species. We constructed a high-density linkage map for Ae. albopictus based on 111,328 informative SNPs obtained by RNAseq. We then performed a linkage-map anchored reassembly of AalbF2, the genome assembly produced by Palatini et al. (2020). Our reassembled genome sequence, AalbF3, represents several improvements relative to AalbF2. First, the size of the AalbF3 assembly is 1.45 Gb, almost half the size of AalbF2. Furthermore, relative to AalbF2, AalbF3 contains a higher proportion of complete and single-copy BUSCO genes (84.3%) and a higher proportion of aligned RNAseq reads that map concordantly to a single location of the genome (46%). We demonstrate the utility of AalbF3 by using it as a reference for a bulk-segregant-based comparative genomics analysis that identifies chromosomal regions with clusters of candidate SNPs putatively associated with photoperiodic diapause, a crucial ecological adaptation underpinning the rapid range expansion and climatic adaptation of A. albopictus.
The association between the African ant plant, Vachellia drepanolobium, and the ants that inhabit it has provided insight into the boundaries between mutualism and parasitism, the response of symbioses to environmental perturbations, and the ecology of species coexistence. We use a landscape genomics approach at sites sampled throughout the range of this system in Kenya to investigate the demographics and genetic structure of the different partners in the association. We find that different species of ant associates of V. drepanolobium show striking differences in their spatial distribution throughout Kenya, and these differences are only partly correlated with abiotic factors. A comparison of the population structure of the host plant and its three obligately arboreal ant symbionts, Crematogaster mimosae, Crematogaster nigriceps, and Tetraponera penzigi, shows that the ants exhibit somewhat similar patterns of structure throughout each of their respective ranges, but that this does not correlate in any clear way with the respective genetic structure of the populations of their host plants. A lack of evidence for local coadaptation in this system suggests that all partners have evolved to cope with a wide variety of biotic and abiotic conditions.
The role of specialization in diversification can be explored along two geological axes in the butterfly family Lycaenidae. In addition to variation in host-plant specialization normally exhibited by butterflies, the caterpillars of most Lycaenidae have symbioses with ants ranging from no interactions through to obligate and specific associations, increasing niche dimensionality in ant-associated taxa. Based on mitochondrial sequences from 8282 specimens from 967 species and 249 genera, we show that the degree of ecological specialization of lycaenid species is positively correlated with genetic divergence, haplotype diversity and an increase in isolation by distance. Nucleotide substitution rate is higher in carnivorous than phytophagous lycaenids. The effects documented here for both micro- and macroevolutionary processes could result from increased spatial segregation as a consequence of reduced connectivity in specialists, niche-based divergence or a combination of both. They could also provide an explanation for the extraordinary diversity of the Lycaenidae and, more generally, for diversity in groups of organisms with similar multi-dimensional ecological specialization.
The Acacia drepanolobium (also known as Vachellia drepanolobium) ant-plant symbiosis is considered a classic case of species coexistence, in which four species of tree-defending ants compete for nesting space in a single host tree species. Coexistence in this system has been explained by trade-offs in the ability of the ant associates to compete with each other for occupied trees versus the ability to colonize unoccupied trees. We seek to understand the proximal reasons for how and why the ant species vary in competitive or colonizing abilities, which are largely unknown. In this study, we use RADseq-derived SNPs to identify relatedness of workers in colonies to test the hypothesis that competitively dominant ants reach large colony sizes due to polygyny, that is, the presence of multiple egg-laying queens in a single colony. We find that variation in polygyny is not associated with competitive ability; in fact, the most dominant species, unexpectedly, showed little evidence of polygyny. We also use these markers to investigate variation in mating behavior among the ant species and find that different species vary in the number of males fathering the offspring of each colony. Finally, we show that the nature of polygyny varies between the two commonly polygynous species, Crematogaster mimosae and Tetraponera penzigi: in C.mimosae, queens in the same colony are often related, while this is not the case for T.penzigi. These results shed light on factors influencing the evolution of species coexistence in an ant-plant mutualism, as well as demonstrating the effectiveness of RADseq-derived SNPs for parentage analysis.
The Vachellia drepanolobium ant-plant symbiosis is a classic case of species coexistence, in which four species of tree-defending ants compete for nesting space in host trees. Species coexistence in this system has been explained by trade-offs in the ability of the ant associates to compete with their neighbors for occupied trees and establish new colonies on unoccupied trees. Proximal reasons for how and why the ant species vary in competitive or colonizing abilities are largely unknown. In this paper, we use RADseq derived SNPs to identify relatedness of workers in colonies to test the hypothesis that competitively dominant ants reach large colony sizes due to polygyny, i.e., the presence of multiple egg-laying queens in a single colony. We found that variation in queen number is not associated with competitive ability; in fact, the most dominant species, unexpectedly, was usually monogynous. We also used these markers to investigate variation in mating behavior among the ant species, and found that different species varied in the number of males fathering the offspring of each queen. Finally, we show that among Crematogaster mimosae and Tetraponera penzigi, the two commonly polygynous species, the manner of polygyny varies: in C. mimosae, queens in the same colony are often related, while this is not the case for T. penzigi. These results demonstrate the effectiveness of RADseq-derived SNPs for parentage analysis, and shed light on factors influencing the evolution of species coexistence in an ant-plant mutualism.
The Aphnaeinae (Lepidoptera: Lycaenidae) are a largely African subfamily of 278 described species that exhibit extraordinary life-history variation. The larvae of these butterflies typically form mutualistic associations with ants, and feed on a wide variety of plants, including 23 families in 19 orders. However, at least one species in each of 9 of the 17 genera is aphytophagous, parasitically feeding on the eggs, brood or regurgitations of ants. This diversity in diet and type of symbiotic association makes the phylogenetic relations of the Aphnaeinae of particular interest. A phylogenetic hypothesis for the Aphnaeinae was inferred from 4.4kb covering the mitochondrial marker COI and five nuclear markers (wg, H3, CAD, GAPDH and EF1) for each of 79 ingroup taxa representing 15 of the 17 currently recognized genera, as well as three outgroup taxa. Maximum Parsimony, Maximum Likelihood and Bayesian Inference analyses all support Heath's systematic revision of the clade based on morphological characters. Ancestral range inference suggests an African origin for the subfamily with a single dispersal into Asia. The common ancestor of the aphnaeines likely associated with myrmicine ants in the genus Crematogaster and plants of the order Fabales.
Too many data-management projects fail because they ignore the changing nature of life-sciences data, argues John Boyle.
Vacating industry and joining the academic ranks means leaving the world of order for the satisfying chaos of research.
BACKGROUND:As the volume, complexity and diversity of the information that scientists work with on a daily basis continues to rise, so too does the requirement for new analytic software. The analytic software must solve the dichotomy that exists between the need to allow for a high level of scientific reasoning, and the requirement to have an intuitive and easy to use tool which does not require specialist, and often arduous, training to use. Information visualization provides a solution to this problem, as it allows for direct manipulation and interaction with diverse and complex data. The challenge addressing bioinformatics researches is how to apply this knowledge to data sets that are continually growing in a field that is rapidly changing.RESULTS:This paper discusses an approach to the development of visual mining tools capable of supporting the mining of massive data collections used in systems biology research, and also discusses lessons that have been learned providing tools for both local researchers and the wider community. Example tools were developed which are designed to enable the exploration and analyses of both proteomics and genomics based atlases. These atlases represent large repositories of raw and processed experiment data generated to support the identification of biomarkers through mass spectrometry (the PeptideAtlas) and the genomic characterization of cancer (The Cancer Genome Atlas). Specifically the tools are designed to allow for: the visual mining of thousands of mass spectrometry experiments, to assist in designing informed targeted protein assays; and the interactive analysis of hundreds of genomes, to explore the variations across different cancer genomes and cancer types.CONCLUSIONS:The mining of massive repositories of biological data requires the development of new tools and techniques. Visual exploration of the large-scale atlas data sets allows researchers to mine data to find new meaning and make sense at scales from single samples to entire populations. Providing linked task specific views that allow a user to start from points of interest (from diseases to single genes) enables targeted exploration of thousands of spectra and genomes. As the composition of the atlases changes, and our understanding of the biology increase, new tasks will continually arise. It is therefore important to provide the means to make the data available in a suitable manner in as short a time as possible. We have done this through the use of common visualization workflows, into which we rapidly deploy visual tools. These visualizations follow common metaphors where possible to assist users in understanding the displayed data. Rapid development of tools and task specific views allows researchers to mine large-scale data almost as quickly as it is produced. Ultimately these visual tools enable new inferences, new analyses and further refinement of the large scale data being provided in atlases such as PeptideAtlas and The Cancer Genome Atlas.
BACKGROUND:For shotgun mass spectrometry based proteomics the most computationally expensive step is in matching the spectra against an increasingly large database of sequences and their post-translational modifications with known masses. Each mass spectrometer can generate data at an astonishingly high rate, and the scope of what is searched for is continually increasing. Therefore solutions for improving our ability to perform these searches are needed.RESULTS:We present a sequence database search engine that is specifically designed to run efficiently on the Hadoop MapReduce distributed computing framework. The search engine implements the K-score algorithm, generating comparable output for the same input files as the original implementation. The scalability of the system is shown, and the architecture required for the development of such distributed processing is discussed.CONCLUSION:The software is scalable in its ability to handle a large peptide database, numerous modifications and large numbers of spectra. Performance scales with the number of processors in the cluster, allowing throughput to expand with the available resources.
Genomic studies are now being undertaken on thousands of samples requiring new computational tools that can rapidly analyze data to identify clinically important features. Inferring structural variations in cancer genomes from mate-paired reads is a combinatorially difficult problem. We introduce Fastbreak, a fast and scalable toolkit that enables the analysis and visualization of large amounts of data from projects such as The Cancer Genome Atlas.
In computational biology, permutation tests have become a widely used tool to assess the statistical significance of an event under investigation. However, the common way of computing the P-value, which expresses the statistical significance, requires a very large number of permutations when small (and thus interesting) P-values are to be accurately estimated. This is computationally expensive and often infeasible. Recently, we proposed an alternative estimator, which requires far fewer permutations compared to the standard empirical approach while still reliably estimating small P-values [1].
Access to public data sets is important to the scientific community as a resource to develop new experiments or validate new data. Projects such as the PeptideAtlas, Ensembl and The Cancer Genome Atlas (TCGA) offer both access to public data and a repository to share their own data. Access to these data sets is often provided through a web page form and a web service API. Access technologies based on web protocols (e.g. http) have been in use for over a decade and are widely adopted across the industry for a variety of functions (e.g. search, commercial transactions, and social media). Each architecture adapts these technologies to provide users with tools to access and share data. Both commonly used web service technologies (e.g. REST and SOAP), and custom-built solutions over HTTP are utilized in providing access to research data. Providing multiple access points ensures that the community can access the data in the simplest and most effective manner for their particular needs. This article examines three common access mechanisms for web accessible data: BioMart, caBIG, and Google Data Sources. These are illustrated by implementing each over the PeptideAtlas repository and reviewed for their suitability based on specific usages common to research. BioMart, Google Data Sources, and caBIG are each suitable for certain uses. The tradeoffs made in the development of the technology are dependent on the uses each was designed for (e.g. security versus speed). This means that an understanding of specific requirements and tradeoffs is necessary before selecting the access technology.
This paper discusses a general purpose software architecture, called Addama, which is used to support the rapid integration and analysis of high volumes of complex biological data. It does this by providing: adaptable software which enables interoperable data access; a step-wise and flexible integration strategy, allowing new information to be overlaid on top of existing annotations and context graphs; and through the provision of asynchronous messaging to support rapid integration of new analysis mechanisms. This work is illustrated through the Cancer Genome Atlas (TCGA) study. Addama is being used within a TCGA analysis center to identity new therapeutic intervention approaches by equating clinical outcomes with underlying genomic effects across heterogeneous data from approximately 20,000 patient samples. Addama supports projects like the TCGA through accepting that biological understanding continually changes, and that the rapid integration of new information and analyses is an essential requirement when supporting research.
Stable incorporation of labeled amino acids in cell culture is a simple approach to label proteins in vivo for mass spectrometric quantification. Full incorporation of isotopically heavy amino acids facilitates accurate quantification of proteins from different cultures, yet analysis methods for determination of incorporation are cumbersome and time-consuming. We present QTIPS, Quantification by Total Identified Peptides for SILAC, a straightforward, accurate method to determine the level of heavy amino acid incorporation throughout a population of peptides detected by mass spectrometry. Using QTIPS, we show that the incorporation of heavy amino acids in baker's yeast is unaffected by the use of prototrophic strains, indicating that auxotrophy is not a requirement for SILAC experiments in this organism. This method has general utility for multiple applications where isotopic labeling is used for quantification in mass spectrometry.
Background Public proteomics databases such as PeptideAtlas contain peptides and proteins identified in mass spectrometry experiments. However, these databases lack information about human disease for researchers studying disease-related proteins. We have developed mspecLINE, a tool that combines knowledge about human disease in MEDLINE with empirical data about the detectable human proteome in PeptideAtlas. mspecLINE associates diseases with proteins by calculating the semantic distance between annotated terms from a controlled biomedical vocabulary. We used an established semantic distance measure that is based on the co-occurrence of disease and protein terms in the MEDLINE bibliographic database. Results The mspecLINE web application allows researchers to explore relationships between human diseases and parts of the proteome that are detectable using a mass spectrometer. Given a disease, the tool will display proteins and peptides from PeptideAtlas that may be associated with the disease. It will also display relevant literature from MEDLINE. Furthermore, mspecLINE allows researchers to select proteotypic peptides for specific protein targets in a mass spectrometry assay. Conclusions Although mspecLINE applies an information retrieval technique to the MEDLINE database, it is distinct from previous MEDLINE query tools in that it combines the knowledge expressed in scientific literature with empirical proteomics data. The tool provides valuable information about candidate protein targets to researchers studying human disease and is freely available on a public web server.