BACKGROUND:The time between the appearance of successive leaves, or phyllochron, characterizes the vegetative development of annual plants. Hypothesis testing models, which allow the comparison of phyllochrons between genetic groups and/or environmental conditions, are usually based on regression of thermal time on the number of leaves; most of the time a constant leaf appearance rate is assumed. However regression models ignore auto-correlation of the leaf number process and may lead to biased testing procedures. Moreover, the hypothesis of constant leaf appearance rate may be too restrictive. METHODS:We propose a stochastic process model in which emergence of new leaves is considered to result from successive time-to-events. This model provides a flexible and more accurate modeling as well as unbiased testing procedures. It was applied to an original maize dataset collected in the field over three years on plants originating from two divergent selection experiments for flowering time in two maize inbred lines. RESULTS AND CONCLUSION:We showed that the main differences in phyllochron were not observed between selection populations but rather between ancestral lines, years of experimentation and leaf ranks. Our results highlight a strong departure from the assumption of a constant leaf appearance rate over a season which could be related to climate variations, even if the impact of individual climate variables could not be clearly determined.
Background With the emergence of metagenomic data, multiple links between the gut microbiome and the host health have been shown. Deciphering these complex interactions require evolved analysis methods focusing on the microbial ecosystem functions. Despite the fact that host or diet-derived fibres are the most abundant nutrients available in the gut, the presence of distinct functional traits regarding fibre and mucin hydrolysis, fermentation and hydrogenotrophic processes has never been investigated. Results After manually selecting 91 KEGG orthologies and 33 glycoside hydrolases further aggregated in 101 functional descriptors representative of fibre and mucin degradation pathways in the gut microbiome, we used nonnegative matrix factorization to mine metagenomic datasets. Four distinct metabolic profiles were further identified on a training set of 1153 samples, thoroughly validated on a large database of 2571 unseen samples from 5 external metagenomic cohorts and confirmed with metatranscriptomic data. Profiles 1 and 2 are the main contributors to the fibre-degradation-related metagenome: they present contrasted involvement in fibre degradation and sugar metabolism and are differentially linked to dysbiosis, metabolic disease and inflammation. Profile 1 takes over Profile 2 in healthy samples, and unbalance of these profiles characterize dysbiotic samples. Furthermore, high fibre diet favours a healthy balance between profiles 1 and profile 2. Profile 3 takes over profile 2 during Crohn’s disease, inducing functional reorientations towards unusual metabolism such as fucose and H2S degradation or propionate, acetone and butanediol production. Profile 4 gathers under-represented functions, like methanogenesis. Two taxonomic makes up of the profiles were investigated, using either the covariation of 203 prevalent genomes or metagenomic species, both providing consistent results in line with their functional characteristics. This taxonomic characterization showed that profiles 1 and 2 were respectively mainly composed of bacteria from the phyla Bacteroidetes and Firmicutes while profile 3 is representative of Proteobacteria and profile 4 of methanogens. Conclusions Integrating anaerobic microbiology knowledge with statistical learning can narrow down the metagenomic analysis to investigate functional profiles. Applying this approach to fibre degradation in the gut ended with 4 distinct functional profiles that can be easily monitored as markers of diet, dysbiosis, inflammation and disease.
One of the difficulties encountered in the statistical analysis of metaproteomics data is the high proportion of missing values, which are usually treated by imputation. Nevertheless, imputation methods are based on restrictive assumptions regarding missingness mechanisms, namely “at random” or “not at random”. To circumvent these limitations in the context of feature selection in a multi-class comparison, we propose a univariate selection method that combines a test of association between missingness and classes, and a test for difference of observed intensities between classes. This approach implicitly handles both missingness mechanisms. We performed a quantitative and qualitative comparison of our procedure with imputation-based feature selection methods on two experimental data sets, as well as simulated data with various scenarios regarding the missingness mechanisms and the nature of the difference of expression (differential intensity or differential presence). Whereas we observed similar performances in terms of prediction on the experimental data set, the feature ranking and selection from various imputation-based methods were strongly divergent. We showed that the combined test reaches a compromise by correlating reasonably with other methods, and remains efficient in all simulated scenarios unlike imputation-based feature selection methods.
ObjectivesBacillus cereus is responsible for food poisoning and rare but severe clinical infections. The pathogenicity of strains varies from harmless to lethal strains. However, there are currently no markers, either alone or in combination, to differentiate pathogenic from non-pathogenic strains. The objective of the study was to identify new genetic biomarkers to differentiate non-pathogenic from clinically relevant B. cereus strains.MethodsA first set of 15 B. cereus strains were compared by RNAseq. A logistic regression model with lasso penalty was applied to define combination of genes whose expression was associated with strain pathogenicity. The identified markers were checked for their presence/absence in a collection of 95 B. cereus strains with varying pathogenic potential (food-borne outbreaks, clinical and non-pathogenic). Receiver operating characteristic area under the curve (AUC) analysis was used to determine the combination of biomarkers, which best differentiate between the “disease” versus “non-disease” groups.ResultsSeven genes were identified during the RNAseq analysis with a prediction to differentiate between pathogenic and non-pathogenic strains. The validation of the presence/absence of these genes in a larger collection of strains coupled with AUC prediction showed that a combination of four biomarkers was sufficient to accurately discern clinical strains from harmless strains, with an AUC of 0.955, sensitivity of 0.9 and specificity of 0.86.ConclusionsThese new findings help in the understanding of B. cereus pathogenic potential and complexity and may provide tools for a better assessment of the risks associated with B. cereus contamination to improve patient health and food safety.
The gut microbiota are increasingly considered as a main partner of human health. Metaproteomics enables us to move from the functional potential revealed by metagenomics to the functions actually operating in the microbiome. However, metaproteome deciphering remains challenging. In particular, confident interpretation of a myriad of MS/MS spectra can only be pursued with smart database searches. Here, we compare the interpretation of MS/MS data sets from 48 individual human gut microbiomes using three interrogation strategies of the dedicated Integrated nonredundant Gene Catalog (IGC 9.9 million genes from 1267 individual fecal samples) together with the Homo sapiens database: the classical single-step interrogation strategy and two iterative strategies (in either two or three steps) aimed at preselecting a reduced-sized, more targeted search space for the final peptide spectrum matching. Both iterative searches outperformed the single-step classical search in terms of the number of peptides and protein clusters identified and the depth of taxonomic and functional knowledge, and this was the most convincing with the three-step approach. However, iterative searches do not help in reducing variability of repeated analyses, which is inherent to the traditional data-dependent acquisition mode, but this variability did not affect the hierarchical relationship between replicates and all other samples.
We propose a flexible statistical model for phyllochron that enables to seasonal variations analysis and hypothesis testing, and demonstrate its efficiency on a data set from a divergent selection experiment on maize. The time between appearance of successive leaves or phyllochron enables to characterize the vegetative development of maize plants which determines their flowering time. Phyllochron is usually considered as constant over the development of a given plant, even though studies have demonstrated response of growth parameters to environmental variables. In this paper, we proposed a novel statistical approach for phyllochron analysis based on a stochastic process, which combines flexibility and a more accurate modelling than existing regression models. The model enables accurate estimation of the phyllochron associated with each leaf rank and enables hypothesis testing. We applied the model on an original maize dataset collected in fields from plants belonging to closely related genotypes originated from divergent selection experiments for flowering time conducted on two maize inbred lines. We showed that the main differences in phyllochron were not observed between selection populations (Early or Late), but rather ancestral lines, years of experimentation, and leaf ranks. Finally, we showed that phyllochron variations through seasons could be related to climate variations, even if the impact of each climatic variables individually was not clearly elucidated. All script and data can be found at https://doi.org/10.15454/CUEHO6
An integrated analysis of gut microbiota, blood biochemical and metabolome in 52 endurance horses was performed. Clustering by gut microbiota revealed the existence of two communities mainly driven by diet as host properties showed little effect. Community 1 presented lower richness and diversity, but higher dominance and rarity of species, including some pathobionts. Moreover, its microbiota composition was tightly linked to host blood metabolites related to lipid metabolism and glycolysis at basal time. Despite the lower fiber intake, community type 1 appeared more specialized to produce acetate as a mean of maintaining the energy supply as glucose concentrations fell during the race. On the other hand, community type 2 showed an enrichment of fibrolytic and cellulolytic bacteria as well as anaerobic fungi, coupled to a higher production of propionate and butyrate. The higher butyrate proportion in community 2 was not associated with protective effects on telomere lengths but could have ameliorated mucosal inflammation and oxidative status. The gut microbiota was neither associated with the blood biochemical markers nor metabolome during the endurance race, and did not provide a biomarker for race ranking or risk of failure to finish the race.
Enterococci, in particular vancomycin-resistant enterococci (VRE), are a leading cause of hospital-acquired infections. Promoting intestinal resistance against enterococci could reduce the risk of VRE infections. We investigated the effects of two Lactobacillus strains to prevent intestinal VRE. We used an intestinal colonisation mouse model based on an antibiotic-induced microbiota dysbiosis to mimic enterococci overgrowth and VRE persistence. Each Lactobacillus spp. was administered daily to mice starting one week before antibiotic treatment until two weeks after antibiotic and VRE inoculation. Of the two strains, Lactobacillus paracasei CNCM I-3689 decreased significantly VRE numbers in the feces demonstrating an improvement of the reduction of VRE. Longitudinal microbiota analysis showed that supplementation with L. paracasei CNCM I-3689 was associated with a better recovery of members of the phylum Bacteroidetes. Bile salt analysis and expression analysis of selected host genes revealed increased level of lithocholate and of ileal expression of camp (human LL-37) upon L. paracasei CNCM I-3689 supplementation. Although a direct effect of L. paracasei CNCM I-3689 on the VRE reduction was not ruled out, our data provide clues to possible anti-VRE mechanisms supporting an indirect anti-VRE effect through the gut microbiota. This work sustains non-antibiotic strategies against opportunistic enterococci after antibiotic-induced dysbiosis.
Whole Genome Shotgun (WGS) metagenomics is increasingly used to study the structure and functions of complex microbial ecosystems, both from the taxonomic and functional point of view. Gene inventories of otherwise uncultured microbial communities make the direct functional profiling of microbial communities possible. The concept of community aggregated trait has been adapted from environmental and plant functional ecology to the framework of microbial ecology. Community aggregated traits are quantified from WGS data by computing the abundance of relevant marker genes. They can be used to study key processes at the ecosystem level and correlate environmental factors and ecosystem functions. In this paper we propose a novel model based approach to infer combinations of aggregated traits characterizing specific ecosystemic metabolic processes. We formulate a model of these Combined Aggregated Functional Traits (CAFTs) accounting for a hierarchical structure of genes, which are associated on microbial genomes, further linked at the ecosystem level by complex co-occurrences or interactions. The model is completed with constraints specifically designed to exploit available genomic information, in order to favor biologically relevant CAFTs. The CAFTs structure, as well as their intensity in the ecosystem, is obtained by solving a constrained Non-negative Matrix Factorization (NMF) problem. We developed a multicriteria selection procedure for the number of CAFTs. We illustrated our method on the modelling of ecosystemic functional traits of fiber degradation by the human gut microbiota. We used 1408 samples of gene abundances from several high-throughput sequencing projects and found that four CAFTs only were needed to represent the fiber degradation potential. This data reduction highlighted biologically consistent functional patterns while providing a high quality preservation of the original data. Our method is generic and can be applied to other metabolic processes in the gut or in other ecosystems.
In this paper, we propose a new method for inferring the metabolic potential of microbial ecosystems based on gene frequencies generated from shotgun metagenomic data. Our approach is based on Non-Negative Matrix Factorization with constraints accounting for prior biological knowledge of bacterial metabolism. The problem is solved using efficient accelerated projected gradient methods. The approach is illustrated on a toy model and on real data on fiber metabolism by the gut microbiota in humans. We show how this approach leads to the inference of biologically relevant gene clusters.
Background: The understanding of changes in temporal processes related to human carcinogenesis is limited. One approach for prospective functional genomic studies is to compile trajectories of differential expression of genes, based on measurements from many case-control pairs. We propose a new statistical method that does not assume any parametric shape for the gene trajectories.Methods: The trajectory of a gene is defined as the curve representing the changes in gene expression levels in the blood as a function of time to cancer diagnosis. In a nested case-control design it consists of differences in gene expression levels between cases and controls. Genes can be grouped into curve groups, each curve group corresponding to genes with a similar development over time. The proposed new statistical approach is based on a set of hypothesis testing that can determine whether or not there is development in gene expression levels over time, and whether this development varies among different strata. Curve group analysis may reveal significant differences in gene expression levels over time among the different strata considered. This new method was applied as a "proof of concept" to breast cancer in the Norwegian Women and Cancer (NOWAC) postgenome cohort, using blood samples collected prospectively that were specifically preserved for transcriptomic analyses (PAX tube). Cohort members diagnosed with invasive breast cancer through 2009 were identified through linkage to the Cancer Registry of Norway, and for each case a random control from the postgenome cohort was also selected, matched by birth year and time of blood sampling, to create a case-control pair. After exclusions, 441 case-control pairs were available for analyses, in which we considered strata of lymph node status at time of diagnosis and time of diagnosis with respect to breast cancer screening visits.Results: The development of gene expression levels in the NOWAC postgenome cohort varied in the last years before breast cancer diagnosis, and this development differed by lymph node status and participation in the Norwegian Breast Cancer Screening Program. The differences among the investigated strata appeared larger in the year before breast cancer diagnosis compared to earlier years.Conclusions: This approach shows good properties in term of statistical power and type 1 error under minimal assumptions. When applied to a real data set it was able to discriminate between groups of genes with non-linear similar patterns before diagnosis.
The adaptive response to extreme endurance exercise might involve transcriptional and translational regulation by microRNAs (miRNAs). Therefore, the objective of the present study was to perform an integrated analysis of the blood transcriptome and miRNome (using microarrays) in the horse before and after a 160 km endurance competition. A total of 2,453 differentially expressed genes and 167 differentially expressed microRNAs were identified when comparing pre- and post-ride samples. We used a hypergeometric test and its generalization to gain a better understanding of the biological functions regulated by the differentially expressed microRNA. In particular, 44 differentially expressed microRNAs putatively regulated a total of 351 depleted differentially expressed genes involved variously in glucose metabolism, fatty acid oxidation, mitochondrion biogenesis, and immune response pathways. In an independent validation set of animals, graphical Gaussian models confirmed that miR-21-5p, miR-181b-5p and miR-505-5p are candidate regulatory molecules for the adaptation to endurance exercise in the horse. To the best of our knowledge, the present study is the first to provide a comprehensive, integrated overview of the microRNA-mRNA co-regulation networks that may have a key role in controlling post-transcriptomic regulation during endurance exercise.
Motivation: Motility is a fundamental cellular attribute, which plays a major part in processes ranging from embryonic development to metastasis. Traditionally, single cell motility is often studied by live cell imaging. Yet, such studies were so far limited to low throughput. To systematically study cell motility at a large scale, we need robust methods to quantify cell trajectories in live cell imaging data. Results: The primary contribution of this article is to present Motility study Integrated Workflow (MotIW), a generic workflow for the study of single cell motility in high-throughput time-lapse screening data. It is composed of cell tracking, cell trajectory mapping to an original feature space and hit detection according to a new statistical procedure. We show that this workflow is scalable and demonstrates its power by application to simulated data, as well as large-scale live cell imaging data. This application enables the identification of an ontology of cell motility patterns in a fully unsupervised manner. Availability and implementation: Python code and examples are available online (http://cbio.ensmp.fr/∼aschoenauer/motiw.html) Contact: thomas.walter@mines-paristech.fr Supplementary information: Supplementary data are available at Bioinformatics online.
Traditionally, the prospective design has been chosen for risk factor analyses of lifestyle and cancer using mainly estimation by survival analysis methods. With new technologies, epidemiologists can expand their prospective studies to include functional genomics given either as transcriptomics, mRNA and microRNA, or epigenetics in blood or other biological materials. The novel functional analyses should not be assessed using classical survival analyses since the main goal is not risk estimation, but the analysis of functional genomics as part of the dynamic carcinogenic process over time, i.e., a "processual" approach. In the risk factor model, time to event is analysed as a function of exposure variables known at start of follow-up (fixed covariates) or changing over the follow-up period (time-dependent covariates). In the processual model, transcriptomics or epigenetics is considered as functions of time and exposures. The success of this novel approach depends on the development of new statistical methods with the capacity of describing and analysing the time-dependent curves or trajectories for tens of thousands of genes simultaneously. This approach also focuses on multilevel or integrative analyses introducing novel statistical methods in epidemiology. The processual approach as part of systems epidemiology might represent in a near future an alternative to human in vitro studies using human biological material for understanding the mechanisms and pathways involved in carcinogenesis.
Cellular motility is a fundamental biological process. Progress in the fields of gene silencing and high-throughput (HT) microscopy provide us with the tools to study its molecular basis and potential perturbators. The primary contribution of this paper is to present MotIW, a generic workflow for single cell motility study in HT time-lapse screening data. We successfully apply it to a simulated screen, as well as a genome-wide screen. Furthermore, MotIW enables the identification of eigth motility patterns into which all trajectories from this dataset divide up into, without any prior model of cell motion.
Characterization of blood biomarkers in endurance horses by integrating miRNA and mRNA expression profiling. 11. Dorothy Russel Havemeyer Foundation International Equine Genome Mapping Workshop
Background. With the increasing interest in post-GWAS research which represents a transition from genome-wide association discovery to analysis of functional mechanisms, attention has been lately focused on the potential of including various biological material in epidemiological studies. In particular, exploration of the carcinogenic process through transcriptional analysis at the epidemiological level opens up new horizons in functional analysis and causal inference, and requires a new design together with adequate analysis procedures. Results. In this article, we present the post-genome design implemented in the NOWAC cohort as an example of a prospective nested case-control study built for transcriptomics use, and discuss analytical strategies to explore the changes occurring in transcriptomics during the carcinogenic process in association with questionnaire information. We emphasize the inadequacy of survival analysis models usually considered in GWAS for post-genome design, and propose instead to parameterize the gene trajectories during the carcinogenic process. Conclusions. This novel approach, in which transcriptomics are considered as potential intermediate biomarkers of cancer and exposures, offers a flexible framework which can include various biological assumptions.