Background Microbiome analysis is becoming a standard component in many scientific studies, but also requires extensive quality control of the 16S rRNA gene sequencing data prior to analysis. In particular, when investigating low-biomass microbial environments such as human skin, contaminants distort the true microbiome sample composition and need to be removed bioinformatically. We introduce MicrobIEM, a novel tool to bioinformatically remove contaminants using negative controls. Results We benchmarked MicrobIEM against five established decontamination approaches in four 16S rRNA amplicon sequencing datasets: three serially diluted mock communities (10 8 –10 3 cells, 0.4–80% contamination) with even or staggered taxon compositions and a skin microbiome dataset. Results depended strongly on user-selected algorithm parameters. Overall, sample-based algorithms separated mock and contaminant sequences best in the even mock, whereas control-based algorithms performed better in the two staggered mocks, particularly in low-biomass samples (≤ 10 6 cells). We show that a correct decontamination benchmarking requires realistic staggered mock communities and unbiased evaluation measures such as Youden’s index. In the skin dataset, the Decontam prevalence filter and MicrobIEM’s ratio filter effectively reduced common contaminants while keeping skin-associated genera. Conclusions MicrobIEM’s ratio filter for decontamination performs better or as good as established bioinformatic decontamination tools. In contrast to established tools, MicrobIEM additionally provides interactive plots and supports selecting appropriate filtering parameters via a user-friendly graphical user interface. Therefore, MicrobIEM is the first quality control tool for microbiome experts without coding experience.
Plants are resistant to most microbial species due to nonhost resistance (NHR), providing broad-spectrum and durable immunity. However, the molecular components contributing to NHR are poorly characterised. We address the question of whether failure of pathogen effectors to manipulate nonhost plants plays a critical role in NHR. RxLR (Arg-any amino acid-Leu-Arg) effectors from two oomycete pathogens, Phytophthora infestans and Hyaloperonospora arabidopsidis, enhanced pathogen infection when expressed in host plants (Nicotiana benthamiana and Arabidopsis, respectively) but the same effectors performed poorly in distantly related nonhost pathosystems. Putative target proteins in the host plant potato were identified for 64 P. infestans RxLR effectors using yeast 2-hybrid (Y2H) screens. Candidate orthologues of these target proteins in the distantly related non-host plant Arabidopsis were identified and screened using matrix Y2H for interaction with RxLR effectors from both P. infestans and H. arabidopsidis. Few P. infestans effector-target protein interactions were conserved from potato to candidate Arabidopsis target orthologues (cAtOrths). However, there was an enrichment of H. arabidopsidis RxLR effectors interacting with cAtOrths. We expressed the cAtOrth AtPUB33, which unlike its potato orthologue did not interact with P. infestans effector PiSFI3, in potato and Nicotiana benthamiana. Expression of AtPUB33 significantly reduced P. infestans colonization in both host plants. Our results provide evidence that failure of pathogen effectors to interact with and/or correctly manipulate target proteins in distantly related non-host plants contributes to NHR. Moreover, exploiting this breakdown in effector-nonhost target interaction, transferring effector target orthologues from non-host to host plants is a strategy to reduce disease.
S-Adenosyl-L-methionine (SAM) is a key enzyme involved in many important biological processes, such as ethylene and polyamine biosynthesis, transmethylation, and transsulfuration. Here, the SAM synthetase (SAMS) gene family was studied in ten different plants (Arabidopsis, tomato, eggplant, sunflower, Medicago truncatula, soybean, rice, barley, Triticum urartu and sorghum) with respect to its physical structure, physicochemical characteristics, and post-transcriptional and post-translational modifications. Additionally, the expression patterns of SAMS genes in tomato were analyzed based on a real-time quantitative PCR assay and an analysis of a public expression dataset. SAMS genes of monocots were more conserved according to the results of a phylogenetic analysis and the prediction of phosphorylation and glycosylation patterns. SAMS genes showed differential expression in response to abiotic stresses and exogenous hormone treatments. Solyc01g101060 was especially expressed in fruit and root tissues, while Solyc09g008280 was expressed in leaves. Additionally, our results revealed that exogenous BR and ABA treatments strongly reduced the expression of tomato SAMS genes. Our research provides new insights and clues about the role of SAMS genes. In particular, these results can inform future functional analyses aimed at revealing the molecular mechanisms underlying the functions of SAMS genes in plants.
Background: Pollen exposure induces local and systemic allergic immune responses in sensitized individuals, but nonsensitized individuals also are exposed to pollen. The kinetics of symptom expression under natural pollen exposure have never been systematically studied, especially in subjects without allergy. Objective: We monitored the humoral immune response under natural pollen exposure to potentially uncover nasal biomarkers for in-season symptom severity and identify protective factors. Methods: We compared humoral immune response kinetics in a panel study of subjects with seasonal allergic rhinitis (SAR) and subjects without allergy and tested for cross-sectional and interseasonal differences in levels of serum and nasal, total, and Betula verrucosa 1-specific immunoglobulin isotypes; immunoglobulin free light chains; cytokines; and chemokines. Nonsupervised principal component analysis was performed for all nasal immune variables, and single immune variables were correlated with in-season symptom severity by Spearman test. Results: Symptoms followed airborne pollen concentrations in subjects with SAR, with a time lag between 0 and 13 days depending on the pollen type. Of the 7 subjects with nonallergy, 4 also exhibited in-season symptoms whereas 3 did not. Cumulative symptoms in those without allergy were lower than in those with SAR but followed the pollen exposure with similar kinetics. Nasal eotaxin-2, CCL22/MDC, and monocyte chemoattactant protein-1 (MCP-1) levels were higher in subjects with SAR, whereas IL-8 levels were higher in subjects without allergy. Principal component analysis and Spearman correlations identified nasal levels of IL-8, IL-33, and Betula verrucosa 1-specific IgG 4 (sIgG 4) and Betula verrucosa 1-specific IgE (sIgE) antibodies as predictive for seasonal symptom severity. Conclusions: Nasal pollen-specific IgA and IgG isotypes are potentially protective within the humoral compartment. Nasal levels of IL-8, IL-33, sIgG 4 and sIgE could be predictive biomarkers for pollen-specific symptom expression, irrespective of atopy.
Grasses have varying inflorescence shapes; however, little is known about the genetic mechanisms specifying such shapes among tribes. Here, we identify the grass-specific TCP transcription factor COMPOSITUM 1 (COM1) expressing in inflorescence meristematic boundaries of different grasses. COM1 specifies branch-inhibition in barley (Triticeae) versus branch-formation in non-Triticeae grasses. Analyses of cell size, cell walls and transcripts reveal barley COM1 regulates cell growth, thereby affecting cell wall properties and signaling specifically in meristematic boundaries to establish identity of adjacent meristems. COM1 acts upstream of the boundary gene Liguleless1 and confers meristem identity partially independent of the COM2 pathway. Furthermore, COM1 is subject to purifying natural selection, thereby contributing to specification of the spike inflorescence shape. This meristem identity pathway has conceptual implications for both inflorescence evolution and molecular breeding in Triticeae. Grasses have diverse inflorescence morphologies, but the underlying genetic mechanisms are unclear. Here, the authors report a TCP transcription factor COM1 affects cell growth through regulation of cell wall properties and promotes branch formation in non-Triticeae grasses but branch inhibition in barley (Triticeae).
BACKGROUND:Atopic eczema (atopic dermatitis, AD) is characterized by disrupted skin barrier associated with elevated skin pH and skin microbiome dysbiosis, due to high Staphylococcus aureus loads, especially during flares. Since S aureus shows optimal growth at neutral pH, we investigated the longitudinal interplay between these factors and AD severity in a pilot study.METHOD:Emollient (with either basic pH 8.5 or pH 5.5) was applied double-blinded twice daily to 6 AD patients and 6 healthy (HE) controls for 8 weeks. Weekly, skin swabs for microbiome analysis (deep sequencing) were taken, AD severity was assessed, and skin physiology (pH, hydration, transepidermal water loss) was measured.RESULTS:Physiological, microbiome, and clinical results were not robustly related to the pH of applied emollient. In contrast to longitudinally stable microbiome in HE, S aureus frequency significantly increased in AD over 8 weeks. High S aureus abundance was associated with skin pH 5.7-6.2. High baseline S aureus frequency predicted both increase in S aureus and in AD severity (EASI and local SCORAD) after 8 weeks.CONCLUSION:Skin pH is tightly regulated by intrinsic factors and limits the abundance of S aureus. High baseline S aureus abundance in turn predicts an increase in AD severity over the study period. This underlines the importance and potential of sustained intervention regarding the skin pH and urges for larger studies linking skin pH and skin S aureus abundance to understand driving factors of disease progression.
In plants, aerial organs originate continuously from stem cells in the center of the shoot apical meristem. Descendants of stem cells in the subepidermal layer are progenitors of germ cells, giving rise to male and female gametes. In these cells, mutations, including insertions of transposable elements or viruses, must be avoided to preserve genome integrity across generations. To investigate the molecular characteristics of stem cells in Arabidopsis, we isolated their nuclei and analyzed stage-specific gene expression and DNA methylation in plants of different ages. Stem cell expression signatures are largely defined by developmental stage but include a core set of stem cell-specific genes, among which are genes implicated in epigenetic silencing. Transiently increased expression of transposable elements in meristems prior to flower induction correlates with increasing CHG methylation during development and decreased CHH methylation, before stem cells enter the reproductive lineage. These results suggest that epigenetic reprogramming may occur at an early stage in this lineage and could contribute to genome protection in stem cells during germline development.
Today, high-throughput transcriptome sequencing and microarray experiments can generate enormous amounts of data to profile gene expression. This expression data often associates with phenotypic information and thus sheds light on the roles of particular traits on genes. Here, we present the user-interface TraitCorr to establish genotype-phenotype relationships and to determine genes, which correlate significantly with a selected trait, using simple correlation and regression analysis. TraitCorr also includes user-friendly visualization and export features. Furthermore, TraitCorr allows identifying genes that are shared among different traits and thus provides new aspects of gene functions which could be useful to better understand the relationship between gene products and phenotype.
Intrinsic biological fluctuation and/or measurement error can obscure the association of gene expression patterns between RNA and protein levels. Appropriate normalization of reverse-transcription quantitative PCR (RT-qPCR) data can reduce technical noise in transcript measurement, thus uncovering such relationships. The accuracy of gene expression measurement is often challenged in the context of cancer due to the genetic instability and “splicing weakness” involved. Here, we sequenced the poly(A) cancer transcriptome of canine osteosarcoma using mRNA-Seq. Expressed sequences were resolved at the level of two consecutive exons to enable the design of exon-border spanning RT-qPCR assays and ranked for stability based on the coefficient of variation ( CV ). Using the same template type for RT-qPCR validation, i.e. poly(A) RNA, avoided skewing of stability assessment by circular RNAs (circRNAs) and/or rRNA deregulation. The strength of the relationship between mRNA expression of the tumour marker S100A4 and its proportion score of quantitative immunohistochemistry (qIHC) was introduced as an experimental readout to fine-tune the normalization choice. Together with the essential logit transformation of qIHC scores, this approach reduced the noise of measurement as demonstrated by uncovering a highly significant, strong association between mRNA and protein expressions of S100A4 (Spearman’s coefficient ρ = 0.72 ( p = 0.006)). Key messages • RNA-seq identifies stable pairs of consecutive exons in a heterogeneous tumour. • Poly(A) RNA templates for RT-qPCR avoid bias from circRNA and rRNA deregulation. • HNRNPL is stably expressed across various cancer tissues and osteosarcoma. • Logit transformed qIHC score better associates with mRNA amount. • Quantification of minor S100A4 mRNA species requires poly(A) RNA templates and dPCR.
Genome-wide DNA methylation studies have quickly expanded due to advances in next-generation sequencing techniques along with a wealth of computational tools to analyze the data. Most of our knowledge about DNA methylation profiles, epigenetic heritability and the function of DNA methylation in plants derives from the model species Arabidopsis thaliana. There are increasingly many studies on DNA methylation in plants—uncovering methylation profiles and explaining variations in different plant tissues. Additionally, DNA methylation comparisons of different plant tissue types and dynamics during development processes are only slowly emerging but are crucial for understanding developmental and regulatory decisions. Translating this knowledge from plant model species to commercial crops could allow the establishment of new varieties with increased stress resilience and improved yield. In this review, we provide an overview of the most commonly applied bioinformatics tools for the analysis of DNA methylation data (particularly bisulfite sequencing data). The performances of a selection of the tools are analyzed for computational time and agreement in predicted methylated sites for A. thaliana, which has a smaller genome compared to the hexaploid bread wheat. The performance of the tools was benchmarked on five plant genomes. We give examples of applications of DNA methylation data analysis in crops (with a focus on cereals) and an outlook for future developments for DNA methylation status manipulations and data integration.
Orb-weaving spiders use a highly strong, sticky and elastic web to catch their prey. These web properties alone would be enough for the entrapment of prey; however, these spiders may be hiding venomous secrets in the web, which current research is revealing. Here, we provide strong proteotranscriptomic evidence for the presence of toxin/neurotoxin-like proteins, defensins, and proteolytic enzymes on the web silk from Nephila clavipes spider. The results from quantitative-based transcriptomic and proteomic approaches showed that silk-producing glands produce an extensive repertoire of toxin/neurotoxin-like proteins, similar to those already reported in spider venoms. Meanwhile, the insect toxicity results demonstrated that these toxic components can be lethal and/or paralytic chemical weapons used for prey capture on the web, and the presence of fatty acids in the web may be a responsible mechanism opening the way to the web toxins for accessing the interior of prey's body, as shown here. Comparative phylogenomic-level evolutionary analyses revealed orthologous genes among two spider groups, Araneomorphae and Mygalomorphae, and the findings showed protein sequences similar to toxins found in the taxa Scorpiones and Hymenoptera in addition to Araneae. Overall, these data represent a valuable resource to further investigate other spider web toxin systems and also suggest that N. clavipes web is not a passive mechanical trap for prey capture, but it exerts an active role in prey paralysis/killing using a series of neurotoxins.
The density of genomic elements such as genes or transposable elements along its consecutive sequence can provide an overview of a genomic sequence while in the detailed analysis of candidate genes it may depict enriched chromosomal hotspots harbouring genes that explain a certain trait. The herein presented python-based graphical user interface Gexplora allows to obtain more information about a genome by considering sequence-intrinsic information from external databases such as Ensembl, OMA and STRING database using REST API calls to retrieve sequence-intrinsic information, protein-protein datasets and orthologous groups. Gexplora is available under https://github.com/nthomasCUBE/Gexplora .
ABSTRACT The experimental and computational analysis of microbes is in the focus of basic and applied research for several decades already and has massively extended our knowledge of role and impact of microbes during that time. Amplicon sequencing is still widely adapted and standard procedure to determine key species in a biological sample and despite a high number of different datasets being steadily generated and fed into current data archives, there is a shortcoming in data interpretation. Furthermore, more computational tools are still required to guide the researcher towards an in-depth understanding of an experiment in a more systematic way. Towards that aim, MicrobComp was created as a graphical user interface which allows to compare two separate species frequency tables that contain the relative or absolute abundance of microbial species in an experiment. Therefore, this tool can be of great help when studying different amplicon experiments where studies were originally performed in different experimental labs.
Intra-species protein-protein interactions (PPI) provide valuable information about the systemic response of a model species when facing either abiotic and biotic stress conditions. Inter-species PPI can otherwise offer insights into how microbes interact with its host and can provide clues how early infection mechanism takes place. To understand these processes in a more comprehensive way and to compare it with experimental outcomes from omics studies, we require additional methods to analyse and visualize PPI data. We demonstrate the user-interface host_microbe_PPI that is implemented in R Shiny. It allows for interactively analysing inter-species and intra-species datasets from various published Arabidopsis thaliana datasets. It enables among other features comparisons of the centrality measurements (degree, betweenness and closeness) and analysis the existence of orthologous proteins in closely related genomes, e.g. when gene loss in host and non-host plants is compared. Arabidopsis was used even so the tool can be also applied in other host-microbe systems.
Cellulose is an important cell wall component and plays a key role in the environmental stress response. Cellulose synthase (CESA) genes along with cellulose synthase-like (CSL) genes are involved in cellulose biosynthesis and various studies have been carried out on phylogeny and clustering of the related cellulose synthase genes but there is no comparative investigation of factors that regulate the post-translation modification. Here, we have selected candidate genes in Arabidopsis and rice and investigated them by various bioinformatics tools to allow a precise dissection of the regulatory modification elements. According to multiple sequence alignments and the distribution of conserved motifs genes were classified into three major groups and LOC_Os11g13650, LOC_Os02g43440 and LOC_Os10g25740 were identified as members of Os-CSLC subfamily. The prominent differences between CESA and CSL genes were observed in terms of the post-translational modification and physicochemical properties of proteins in Arabidopsis (as dicotyledonous) and rice (as monocotyledonous). Moreover, the binding sites of common transcription factors such as MYB, WRKY and abscisic acid response element (ABRE) family were found in the promoter regions of CESA and CSL genes family, and expression patterns sharply showed differentially expression in Arabidopsis and rice tissues as response to biotic and abiotic stresses. The considerable outcome of the post-translation modification of studied genes demonstrated that CESA proteins were more phosphorylated than CSL proteins, and CESA proteins were more hydrophilic. It can be concluded that obtained results furnish further information about the genes which are involving in cellulose biosynthesis that will be useful for more dissection and functional analysis.
Today, transcriptomes and microarrays can be generated and analysed in high quantity. In addition, experiments often include descriptive information about each sample which needs to be compared to the gene expression profiles. The understanding of the relationships between gene expression and phenotype is introduced as new challenge in system biology. Combining expression (RNA-seq and microarray) and phenotype data could reveal the role or effects of each gene on traits. To address all these needs, the user-interface TraitCorr was developed which allows to determine genes that are significantly correlating with a selected trait. Furthermore, it allows to determine significantly correlated genes among different traits and provides visualisation and analysis possibilities.