The health, growth, and fitness of boreal forest trees are impacted and improved by their associated microbiomes. Microbial gene expression and functional activity can be assayed with RNA sequencing (RNA-Seq) data from host samples. In contrast, phylogenetic marker gene amplicon sequencing data are used to assess taxonomic composition and community structure of the microbiome. Few studies have considered how much of this structural and taxonomic information is included in transcriptomic data from matched samples. Here, we described fungal communities using both host-derived RNA-Seq and fungal ITS1 DNA amplicon sequencing to compare the outcomes between the methods. We used a panel of root and needle samples from the coniferous tree species Picea abies (Norway spruce) growing in untreated (nutrient-deficient) and nutrient-enriched plots at the Flakaliden forest research site in boreal northern Sweden. We show that the relationship between samples and alpha and beta diversity indicated by the fungal transcriptome is in agreement with that generated by the ITS data, while also identifying a lack of taxonomic overlap due to limitations imposed by current database coverage. Furthermore, we demonstrate how metatranscriptomics data additionally provide biologically informative functional insights. At the community level, there were changes in starch and sucrose metabolism, biosynthesis of amino acids, and pentose and glucuronate interconversions, while processing of organic macromolecules, including aromatic and heterocyclic compounds, was enriched in transcripts assigned to the genus Cortinarius. IMPORTANCE A deeper understanding of microbial communities associated with plants is revealing their importance for plant health and productivity. RNA extracted from plant field samples represents the host and other organisms present. Typically, gene expression studies focus on the plant component or, in a limited number of studies, expression in one or more associated organisms. However, metatranscriptomic data are rarely used for taxonomic profiling, which is currently performed using amplicon approaches. We created an assembly-based, reproducible, and hardware-agnostic workflow to taxonomically and functionally annotate fungal RNA-Seq data obtained from Norway spruce roots, which we compared to matching ITS amplicon sequencing data. While we identified some limitations and caveats, we show that functional, taxonomic, and compositional insights can all be obtained from RNA-Seq data. These findings highlight the potential of metatranscriptomics to advance our understanding of interaction, response, and effect between host plants and their associated microbial communities.
Microbial communities are major players in carbon and nitrogen cycling globally and are of particular importance for plant communities in the nutrient poor soils of boreal forests. Especially relevant are the fungal communities in the soil that interact with the plants in multiple ways, indirectly through their pivotal role in the breakdown of organic matter and, more directly, through mycorrhizal symbiosis with plant roots. Large-scale disturbances of these complex microbial communities can lead to shifts in soil carbon storage with unknown and global-scale long-term consequences. To understand the dynamics of these communities and their relationship to associated plants in response to climate change and anthropogenic influence, we need a better understanding of how modern “omics” methods can help us to understand compositional and functional shifts of these microbiomes. Microbial gene expression and functional activity can be assayed with RNA sequencing (RNA-Seq) data from environmental samples. In contrast, currently phylogenetic marker gene amplicon sequencing data is generally used to assess taxonomic composition and community structure of the microbiome. Few studies have considered how much of this structural and taxonomic information is included in RNA-Seq transcriptomic data from matched samples. Here we describe fungal communities using both RNA-Seq and fungal ITS1 DNA amplicon sequencing to compare the outcomes between the methods. We used a panel of root and needle samples from mature stands of the coniferous tree species Picea abies (Norway spruce) growing in untreated (nutrient deficient) and nutrient enriched plots at the Flakaliden forest research site in boreal northern Sweden. We created an assembly-based, reproducible and hardware agnostic workflow to taxonomically and functionally annotate fungal RNA-Seq data obtained from Norway spruce roots, which we compared to matching ITS amplicon sequencing data. We show that the community structure indicated by the fungal transcriptome is in agreement with that generated by the ITS data, while also identifying limitations imposed by current database coverage. Furthermore, we show examples to demonstrate how metatranscriptomics data additionally provides biologically informative functional insight at the community and individual species level. These findings highlight the potential of metatranscriptomics to advance our understanding of interaction, response and effect both between host plants and their associated microbial communities, and among the members of microbial communities in environmental samples in general.
Enrichment of forest soils with inorganic nitrogen (N) tends to inhibit oxidative enzyme expression by microbes and reduces plant litter and soil organic matter decomposition rates. Without further explanation than is currently presented in the scientific literature, we argue that upregulation of oxidative enzymes seems a more competitive response to prolonged N enrichment at high rates than the observed downregulation. Thus, as it stands, observed responses are inconsistent with predicted responses. In this article, we present a hypothesis that resolves this conflict. We suggest that high rates of N addition alter the competitive balance between enzymatic lignin mineralisation and non-enzymatic lignin oxidation. Using metatransciptomics and chemical assays to examine boreal forest soils, we found that N addition suppressed peroxidase activity, but not iron reduction activity (involved in non-enzymatic lignin oxidation). Our hypothesis seems positioned as a parsimonious and empirically consistent working model that warrants further testing.
Background Studies that aim at explaining phenotypes or disease susceptibility by genetic or epigenetic variants often rely on clustering methods to stratify individuals or samples. While statistical associations may point at increased risk for certain parts of the population, the ultimate goal is to make precise predictions for each individual. This necessitates tools that allow for the rapid inspection of each data point, in particular to find explanations for outliers. Results ACES is an integrative cluster- and phenotype-browser, which implements standard clustering methods, as well as multiple visualization methods in which all sample information can be displayed quickly. In addition, ACES can automatically mine a list of phenotypes for cluster enrichment, whereby the number of clusters and their boundaries are estimated by a novel method. For visual data browsing, ACES provides a 2D or 3D PCA or Heat Map view. ACES is implemented in Java, with a focus on a user-friendly, interactive, graphical interface. Conclusions ACES has been proven an invaluable tool for analyzing large, pre-filtered DNA methylation data sets and RNA-Sequencing data, due to its ease to link molecular markers to complex phenotypes. The source code is available from https://github.com/GrabherrGroup/ACES .
The Populus genus is one of the major plant model systems, but genomic resources have thus far primarily been available for poplar species, and primarily Populus trichocarpa (Torr. & Gray), which was the first tree with a whole-genome assembly. To further advance evolutionary and functional genomic analyses in Populus, we produced genome assemblies and population genetics resources of two aspen species, Populus tremula L. and Populus tremuloides Michx. The two aspen species have distributions spanning the Northern Hemisphere, where they are keystone species supporting a wide variety of dependent communities and produce a diverse array of secondary metabolites. Our analyses show that the two aspens share a similar genome structure and a highly conserved gene content with P. trichocarpa but display substantially higher levels of heterozygosity. Based on population resequencing data, we observed widespread positive and negative selection acting on both coding and noncoding regions. Furthermore, patterns of genetic diversity and molecular evolution in aspen are influenced by a number of features, such as expression level, coexpression network connectivity, and regulatory variation. To maximize the community utility of these resources, we have integrated all presented data within the PopGenIE web resource (PopGenIE.org).
Cancer is one of the most common causes of death in humans. It can arise from many different cell types, and even cancers originating from the same tissue can constitute a heterogeneous group of di ...
The understanding of the etiology of type 1 diabetes (T1D) remains limited. One objective of the Diabetes Virus Detection (DiViD) study was to collect pancreatic tissue from living subjects shortly after the diagnosis of T1D. Here we report the insulin secretion ability by in vitro glucose perifusion and explore the expression of insulin pathway genes in isolated islets of Langerhans from these patients. Whole-genome RNA sequencing was performed on islets from six DiViD study patients and two organ donors who died at the onset of T1D, and the findings were compared with those from three nondiabetic organ donors. All human transcripts involved in the insulin pathway were present in the islets at the onset of T1D. Glucose-induced insulin secretion was present in some patients at the onset of T1D, and a perfectly normalized biphasic insulin release was obtained after some days in a nondiabetogenic environment in vitro. This indicates that the potential for endogenous insulin production is good, which could be taken advantage of if the disease process was reversed at diagnosis.
The evolution of the opioid peptides and nociceptin/orphanin as well as their receptors has been difficult to resolve due to variable evolutionary rates. By combining sequence comparisons with information on the chromosomal locations of the genes, we have deduced the following evolutionary scenario: The vertebrate predecessor had one opioid precursor gene and one receptor gene. The two genome doublings before the vertebrate radiation resulted in three peptide precursor genes whereupon a fourth copy arose by a local gene duplication. These four precursors diverged to become the pre-propeptides for endorphin (POMC), enkephalins, dynorphins, and nociceptin, respectively. The ancestral receptor gene was quadrupled in the genome doublings leading to delta, kappa, and mu and the nociceptin/orphanin receptor. This scenario is corroborated by new data presented here for coelacanth and spotted gar, representing two basal branches in the vertebrate tree. A third genome doubling in the ancestor of teleost fishes generated additional gene copies. These results show that the opioid system was quite complex already in the first vertebrates and that it has more components in teleost fishes than in mammals. From an evolutionary point of view, nociceptin and its receptor can be considered full-fledged members of the opioid system.
UNLABELLEDWhiteboard is a class library implemented in C++ that enables visualization to be tightly coupled with computation when analyzing large and complex datasets.AVAILABILITY AND IMPLEMENTATIONthe C++ source code, coding samples and documentation are freely available under the Lesser General Public License from http://whiteboard-class.sourceforge.net/.
The vertebrate gene family for neuropeptide Y (NPY) receptors expanded by duplication of the chromosome carrying the ancestral Y1-Y2-Y5 gene triplet. After loss of some duplicates, the ancestral jawed vertebrate had seven receptor subtypes forming the Y1 (including Y1, Y4, Y6, Y8), Y2 (including Y2, Y7) and Y5 (only Y5) subfamilies. Lampreys are considered to have experienced the same chromosome duplications as gnathostomes and should also be expected to have multiple receptor genes. However, previously only a Y4-like and a Y5 receptor have been cloned and characterized. Here we report the cloning and characterization of two additional receptors from the sea lamprey Petromyzon marinus. Sequence phylogeny alone could not with certainty assign their identity, but based on synteny comparisons of P. marinus and the Arctic lamprey, Lethenteron camtschaticum, with jawed vertebrates, the two receptors most likely are Y1 and Y2. Both receptors were expressed in human HEK293 cells and inositol phosphate assays were performed to determine the response to the three native lamprey peptides NPY, PYY and PMY. The three peptides have similar potencies in the nanomolar range for Y1. No obvious response to the three peptides was detected for Y2. Synteny analysis supports identification of the previously cloned receptor as Y4. No additional NPY receptor genes could be identified in the presently available lamprey genome assemblies. Thus, four NPY-family receptors have been identified in lampreys, orthologs of the same subtypes as in humans (Y1, Y2, Y4 and Y5), whereas many other vertebrate lineages have retained additional ancestral subtypes.
After performing de novo transcript assembly of >1 billion RNA-Sequencing reads obtained from 22 samples of different Norway spruce (Picea abies) tissues that were not surface sterilized, we found that assembled sequences captured a mix of plant, lichen, and fungal transcripts. The latter were likely expressed by endophytic and epiphytic symbionts, indicating that these organisms were present, alive, and metabolically active. Here, we show that these serendipitously sequenced transcripts need not be considered merely as contamination, as is common, but that they provide insight into the plant's phyllosphere. Notably, we could classify these transcripts as originating predominantly from Dothideomycetes and Leotiomycetes species, with functional annotation of gene families indicating active growth and metabolism, with particular regards to glucose intake and processing, as well as gene regulation.
Background Genomic duplications constitute major events in the evolution of species, allowing paralogous copies of genes to take on fine-tuned biological roles. Unambiguously identifying the orthology relationship between copies across multiple genomes can be resolved by synteny, i.e. the conserved order of genomic sequences. However, a comprehensive analysis of duplication events and their contributions to evolution would require all-to-all genome alignments, which increases at N 2 with the number of available genomes, N. Results Here, we introduce Kraken, software that omits the all-to-all requirement by recursively traversing a graph of pairwise alignments and dynamically re-computing orthology. Kraken scales linearly with the number of targeted genomes, N, which allows for including large numbers of genomes in analyses. We first evaluated the method on the set of 12 Drosophila genomes, finding that orthologous correspondence computed indirectly through a graph of multiple synteny maps comes at minimal cost in terms of sensitivity, but reduces overall computational runtime by an order of magnitude. We then used the method on three well-annotated mammalian genomes, human, mouse, and rat, and show that up to 93% of protein coding transcripts have unambiguous pairwise orthologous relationships across the genomes. On a nucleotide level, 70 to 83% of exons match exactly at both splice junctions, and up to 97% on at least one junction. We last applied Kraken to an RNA-sequencing dataset from multiple vertebrates and diverse tissues, where we confirmed that brain-specific gene family members, i.e. one-to-many or many-to-many homologs, are more highly correlated across species than single-copy (i.e. one-to-one homologous) genes. Not limited to protein coding genes, Kraken also identifies thousands of newly identified transcribed loci, likely non-coding RNAs that are consistently transcribed in human, chimpanzee and gorilla, and maintain significant correlation of expression levels across species. Conclusions Kraken is a computational genome coordinate translator that facilitates cross-species comparisons, distinguishes orthologs from paralogs, and does not require costly all-to-all whole genome mappings. Kraken is freely available under LPGL from http://github.com/nedaz/kraken .
The domestic dog, Canis familiaris, is a well-established model system for mapping trait and disease loci. While the original draft sequence was of good quality, gaps were abundant particularly in promoter regions of the genome, negatively impacting the annotation and study of candidate genes. Here, we present an improved genome build, canFam3.1, which includes 85 MB of novel sequence and now covers 99.8% of the euchromatic portion of the genome. We also present multiple RNA-Sequencing data sets from 10 different canine tissues to catalog ∼175,000 expressed loci. While about 90% of the coding genes previously annotated by EnsEMBL have measurable expression in at least one sample, the number of transcript isoforms detected by our data expands the EnsEMBL annotations by a factor of four. Syntenic comparison with the human genome revealed an additional ∼3,000 loci that are characterized as protein coding in human and were also expressed in the dog, suggesting that those were previously not annotated in the EnsEMBL canine gene set. In addition to ∼20,700 high-confidence protein coding loci, we found ∼4,600 antisense transcripts overlapping exons of protein coding genes, ∼7,200 intergenic multi-exon transcripts without coding potential, likely candidates for long intergenic non-coding RNAs (lincRNAs) and ∼11,000 transcripts were reported by two different library construction methods but did not fit any of the above categories. Of the lincRNAs, about 6,000 have no annotated orthologs in human or mouse. Functional analysis of two novel transcripts with shRNA in a mouse kidney cell line altered cell morphology and motility. All in all, we provide a much-improved annotation of the canine genome and suggest regulatory functions for several of the novel non-coding transcripts.
Background The fundamental challenge in optimally aligning homologous sequences is to define a scoring scheme that best reflects the underlying biological processes. Maximising the overall number of matches in the alignment does not always reflect the patterns by which nucleotides mutate. Efficiently implemented algorithms that can be parameterised to accommodate more complex non-linear scoring schemes are thus desirable. Results We present Cola , alignment software that implements different optimal alignment algorithms, also allowing for scoring contiguous matches of nucleotides in a nonlinear manner. The latter places more emphasis on short, highly conserved motifs, and less on the surrounding nucleotides, which can be more diverged. To illustrate the differences, we report results from aligning 14,100 sequences from 3' untranslated regions of human genes to 25 of their mammalian counterparts, where we found that a nonlinear scoring scheme is more consistent than a linear scheme in detecting short, conserved motifs. Conclusions Cola is freely available under LPGL from https://github.com/nedaz/cola .
A peptide ending with RFamide (Arg-Phe-amide) was discovered independently by three different laboratories in 2003 and named 26RFa or QRFP. In mammals, a longer version of the peptide, 43 amino acids, was identified and found to bind to the orphan G protein-coupled receptor GPR103. We searched the genome database of Branchiostoma floridae (Bfl) for receptor sequences related to those that bind peptides ending with RFa or RYa (including receptors for NPFF, PRLH, GnIH, and NPY). One receptor clustered in phylogenetic analyses with mammalian QRFP receptors. The gene has 3 introns in Bfl and 5 in human, but all intron positions differ, implying that the introns were inserted independently. A QRFP-like peptide consisting of 25 amino acids and ending with RFa was identified in the amphioxus genome. Eight of the ten last amino acids are identical between Bfl and human. The prepro-QRFP gene in Bfl has one intron in the propeptide whereas the human gene lacks introns. The Bfl QRFP peptide was synthesized and the receptor was functionally expressed in human cells. The response was measured as inositol phosphate (IP) turnover. The Bfl QRFP peptide was found to potently stimulate the receptor's ability to induce IP turnover with an EC50 of 0.28nM. Also the human QRFP peptides with 26 and 43 amino acids were found to stimulate the receptor (1.9 and 5.1nM, respectively). Human QRFP with 26 amino acids without the carboxyterminal amide had dramatically lower potency at 1.3μM. Thus, we have identified an amphioxus QRFP-related peptide and a corresponding receptor and shown that they interact to give a functional response.
The neuropeptide Y (NPY) system influences numerous physiological functions including feeding behavior, endocrine regulation, and cardiovascular regulation. In jawed vertebrates it consists of 3-4 peptides and 4-7 receptors. Teleost fishes have unique duplicates of NPY and PYY as well as the Y8 receptor. In the zebrafish, the NPY system consists of the peptides NPYa, PYYa, and PYYb (NPYb appears to have been lost) and at least seven NPY receptors: Y1, Y2, Y2-2, Y4, Y7, Y8a, and Y8b. Previously PYYb binding has been reported for Y2 and Y2-2. To search for peptide-receptor preferences, we have investigated PYYb binding to four of the remaining receptors and compared with NPYa and PYYa. Taken together, the most striking observations are that PYYa displays reduced affinity for Y2 (3 nM) compared to the other peptides and receptors and that all three peptides have higher affinity for Y4 (0.028-0.034 nM) than for the other five receptors. The strongest peptide preference by any receptor selectivity is the one previously reported for PYYb by the Y2 receptor, as compared to NPY and PYYa. These affinity differences may be helpful to elucidate specific details of peptide-receptor interactions. Also, we have investigated the level of mRNA expression in different organs using qPCR. All peptides and receptors have higher expression in heart, kidney, and brain. These quantitative aspects on receptor affinities and mRNA distribution help provide a more complete picture of the NPY system.
Background: Vertebrate color vision is dependent on four major color opsin subtypes: RH2 (green opsin), SWS1 (ultraviolet opsin), SWS2 (blue opsin), and LWS (red opsin). Together with the dim-light receptor rhodopsin (RH1), these form the family of vertebrate visual opsins. Vertebrate genomes contain many multi-membered gene families that can largely be explained by the two rounds of whole genome duplication (WGD) in the vertebrate ancestor (2R) followed by a third round in the teleost ancestor (3R). Related chromosome regions resulting from WGD or block duplications are said to form a paralogon. We describe here a paralogon containing the genes for visual opsins, the G-protein alpha subunit families for transducin (GNAT) and adenylyl cyclase inhibition (GNAI), the oxytocin and vasopressin receptors (OT/VP-R), and the L-type voltage-gated calcium channels (CACNA1-L).Results: Sequence-based phylogenies and analyses of conserved synteny show that the above-mentioned gene families, and many neighboring gene families, expanded in the early vertebrate WGDs. This allows us to deduce the following evolutionary scenario: The vertebrate ancestor had a chromosome containing the genes for two visual opsins, one GNAT, one GNAI, two OT/VP-Rs and one CACNA1-L gene. This chromosome was quadrupled in 2R. Subsequent gene losses resulted in a set of five visual opsin genes, three GNAT and GNAI genes, six OT/VP-R genes and four CACNA1-L genes. These regions were duplicated again in 3R resulting in additional teleost genes for some of the families. Major chromosomal rearrangements have taken place in the teleost genomes. By comparison with the corresponding chromosomal regions in the spotted gar, which diverged prior to 3R, we could time these rearrangements to post-3R.Conclusions: We present an extensive analysis of the paralogon housing the visual opsin, GNAT and GNAI, OT/VP-R, and CACNA1-L gene families. The combined data imply that the early vertebrate WGD events contributed to the evolution of vision and the other neuronal and neuroendocrine functions exerted by the proteins encoded by these gene families. In pouched lamprey all five visual opsin genes have previously been identified, suggesting that lampreys diverged from the jawed vertebrates after 2R.