Summary Mass-spectrometry (MS) datasets present a unique set of challenges that make in-depth bioinformatics analysis non-trivial, with analysis requiring both expertise and time. Often these datasets have unique structures that need to be dealt with on an individual basis. Currently, tools providing a fast, interactive and guided way of exploring and analysing these data sets are not readily available. To this end, we have developed ThunderBolt: a highly interactive, point-and-click web-based application providing both bioinformaticians and biologists with a platform for i) searching and comparing multiple omics datasets, ii) fast data exploration and quality control, iii) interactive visualization, iv) pre-processing, v) statistical analysis and vi) functional and network enrichment analysis of large proteomics datasets using the Shiny framework. Availability ThunderBolt is a shiny-application accessible at Contact james.burchfield{at}sydney.edu.au ### Competing Interest Statement The authors have declared no competing interest.
Mass spectrometry (MS)-based phosphoproteomics enables the quantification of proteome-wide phosphorylation in cells and tissues. A major challenge in MS-based phosphoproteomics lies in identifying the substrates of kinases, as currently only a small fraction of substrates identified can be confidently linked with a known kinase. By leveraging large-scale phosphoproteomics data, machine learning has become an increasingly popular approach for computationally predicting substrates of kinases. However, the small number of high-quality experimentally validated kinase substrates (true positive) and the high data noise in many phosphoproteomics datasets together impact the performance of existing approaches. Here, we aim to develop advanced kinase-substrate prediction methods to address these challenges. Using a collection of seven large phosphoproteomics datasets, including six published datasets and a new muscle differentiation dataset, and both traditional and deep learning models, we first demonstrate that a ‘pseudo-positive’ learning strategy for alleviating small sample size is effective at improving model predictive performance. We next show that a data re-sampling based ensemble learning strategy is useful for improving model stability while further enhancing prediction. Lastly, we introduce an ensemble deep learning model (‘SnapKin’) incorporating the above two learning strategies into a ‘snapshot’ ensemble learning algorithm. We demonstrate that the SnapKin model achieves overall the best performance in kinase-substrate prediction. Together, we propose SnapKin as a promising approach for predicting substrates of kinases from large-scale phosphoproteomics data. SnapKin is freely available at https://github.com/PYangLab/SnapKin.
Abstract Genetic and environmental factors play a major role in metabolic health. However, they do not act in isolation, as a change in an environmental factor such as diet may exert different effects based on an individual’s genotype. Here, we sought to understand how such gene–diet interactions influenced nutrient storage and utilization, a major determinant of metabolic disease. We subjected 178 inbred strains from the Drosophila genetic reference panel (DGRP) to diets varying in sugar, fat, and protein. We assessed starvation resistance, a holistic phenotype of nutrient storage and utilization that can be robustly measured. Diet influenced the starvation resistance of most strains, but the effect varied markedly between strains such that some displayed better survival on a high carbohydrate diet (HCD) compared to a high-fat diet while others had opposing responses, illustrating a considerable gene × diet interaction. This demonstrates that genetics plays a major role in diet responses. Furthermore, heritability analysis revealed that the greatest genetic variability arose from diets either high in sugar or high in protein. To uncover the genetic variants that contribute to the heterogeneity in starvation resistance, we mapped 566 diet-responsive SNPs in 293 genes, 174 of which have human orthologs. Using whole-body knockdown, we identified two genes that were required for glucose tolerance, storage, and utilization. Strikingly, flies in which the expression of one of these genes, CG4607 a putative homolog of a mammalian glucose transporter, was reduced at the whole-body level, displayed lethality on a HCD. This study provides evidence that there is a strong interplay between diet and genetics in governing survival in response to starvation, a surrogate measure of nutrient storage efficiency and obesity. It is likely that a similar principle applies to higher organisms thus supporting the case for nutrigenomics as an important health strategy.
Insulin's activation of PI3K/Akt signaling, stimulates glucose uptake by enhancing delivery of GLUT4 to the cell surface. Here we examined the origins of intercellular heterogeneity in insulin signaling. Akt activation alone accounted for ~25% of the variance in GLUT4, indicating that additional sources of variance exist. The Akt and GLUT4 responses were highly reproducible within the same cell, suggesting the variance is between cells (extrinsic) and not within cells (intrinsic). Generalized mechanistic models (supported by experimental observations) demonstrated that the correlation between the steady-state levels of two measured signaling processes decreases with increasing distance from each other and that intercellular variation in protein expression (as an example of extrinsic variance) is sufficient to account for the variance in and between Akt and GLUT4. Thus, the response of a population to insulin signaling is underpinned by considerable single-cell heterogeneity that is largely driven by variance in gene/protein expression between cells.
The phosphoinositide 3-kinase (PI3K)-Akt network is tightly controlled by feedback mechanisms that regulate signal flow and ensure signal fidelity. A rapid overshoot in insulin-stimulated recruitment of Akt to the plasma membrane has previously been reported, which is indicative of negative feedback operating on acute timescales. Here, we show that Akt itself engages this negative feedback by phosphorylating insulin receptor substrate (IRS) 1 and 2 on a number of residues. Phosphorylation results in the depletion of plasma membrane-localised IRS1/2, reducing the pool available for interaction with the insulin receptor. Together these events limit plasma membrane-associated PI3K and phosphatidylinositol (3,4,5)-trisphosphate (PIP3) synthesis. We identified two Akt-dependent phosphorylation sites in IRS2 at S306 (S303 in mouse) and S577 (S573 in mouse) that are key drivers of this negative feedback. These findings establish a novel mechanism by which the kinase Akt acutely controls PIP3 abundance, through post-translational modification of the IRS scaffold.
The remarkable flexibility and adaptability of ensemble methods and deep learning models have led to the proliferation of their application in bioinformatics research. Traditionally, these two machine learning techniques have largely been treated as independent methodologies in bioinformatics applications. However, the recent emergence of ensemble deep learning—wherein the two machine learning techniques are combined to achieve synergistic improvements in model accuracy, stability and reproducibility—has prompted a new wave of research and application. Here, we share recent key developments in ensemble deep learning and look at how their contribution has benefited a wide range of bioinformatics research from basic sequence analysis to systems biology. While the application of ensemble deep learning in bioinformatics is diverse and multifaceted, we identify and discuss the common challenges and opportunities in the context of bioinformatics research. We hope this Review Article will bring together the broader community of machine learning researchers, bioinformaticians and biologists to foster future research and development in ensemble deep learning, and inspire novel bioinformatics applications that are unattainable by traditional methods.
Multi-modal profiling of single cells represents one of the latest technological advancements in molecular biology. Among various single-cell multi-modal strategies, cellular indexing of transcriptomes and epitopes by sequencing (CITE-seq) allows simultaneous quantification of two distinct species: RNA and surface marker proteins (ADT). Here, we introduce CiteFuse, a streamlined package consisting of a suite of tools for pre-processing, modality integration, clustering, differential RNA and ADT expression analysis, ADT evaluation, ligand-receptor interaction analysis, and interactive web-based visualization of CITE-seq data. We show the capacity of CiteFuse to integrate the two data modalities and its relative advantage against data generated from single modality profiling. Furthermore, we illustrate the pre-processing steps in CiteFuse and in particular a novel doublet detection method based on a combined index of cell hashing and transcriptome data. Collectively, we demonstrate the utility and effectiveness of CiteFuse for the integrative analysis of transcriptome and epitope profiles from CITE-seq data.
BACKGROUND:Single-cell RNA-sequencing (scRNA-seq) is a transformative technology, allowing global transcriptomes of individual cells to be profiled with high accuracy. An essential task in scRNA-seq data analysis is the identification of cell types from complex samples or tissues profiled in an experiment. To this end, clustering has become a key computational technique for grouping cells based on their transcriptome profiles, enabling subsequent cell type identification from each cluster of cells. Due to the high feature-dimensionality of the transcriptome (i.e. the large number of measured genes in each cell) and because only a small fraction of genes are cell type-specific and therefore informative for generating cell type-specific clusters, clustering directly on the original feature/gene dimension may lead to uninformative clusters and hinder correct cell type identification.RESULTS:Here, we propose an autoencoder-based cluster ensemble framework in which we first take random subspace projections from the data, then compress each random projection to a low-dimensional space using an autoencoder artificial neural network, and finally apply ensemble clustering across all encoded datasets to generate clusters of cells. We employ four evaluation metrics to benchmark clustering performance and our experiments demonstrate that the proposed autoencoder-based cluster ensemble can lead to substantially improved cell type-specific clusters when applied with both the standard k-means clustering algorithm and a state-of-the-art kernel-based clustering algorithm (SIMLR) designed specifically for scRNA-seq data. Compared to directly using these clustering algorithms on the original datasets, the performance improvement in some cases is up to 100%, depending on the evaluation metric used.CONCLUSIONS:Our results suggest that the proposed framework can facilitate more accurate cell type identification as well as other downstream analyses. The code for creating the proposed autoencoder-based cluster ensemble framework is freely available from https://github.com/gedcom/scCCESS.
Background Single-cell RNA-sequencing (scRNA-seq) is a fast emerging technology allowing global transcriptome profiling on the single cell level. Cell type identification from scRNA-seq data is a critical task in a variety of research such as developmental biology, cell reprogramming, and cancers. Typically, cell type identification relies on human inspection using a combination of prior biological knowledge (e.g. marker genes and morphology) and computational techniques (e.g. PCA and clustering). Due to the incompleteness of our current knowledge and the subjectivity involved in this process, a small amount of cells may be subject to mislabelling. Results Here, we propose a semi-supervised learning framework, named scReClassify, for ‘post hoc’ cell type identification from scRNA-seq datasets. Starting from an initial cell type annotation with potentially mislabelled cells, scReClassify first performs dimension reduction using PCA and next applies a semi-supervised learning method to learn and subsequently reclassify cells that are likely mislabelled initially to the most probable cell types. By using both simulated and real-world experimental datasets that profiled various tissues and biological systems, we demonstrate that scReClassify is able to accurately identify and reclassify misclassified cells to their correct cell types. Conclusions scReClassify can be used for scRNA-seq data as a post hoc cell type classification tool to fine-tune cell type annotations generated by any cell type classification procedure. It is implemented as an R package and is freely available from https://github.com/SydneyBioX/scReClassify
Exercise is extremely beneficial to whole body health reducing the risk of a number of chronic human diseases. Some of these physiological benefits appear to be mediated via the secretion of peptide/protein hormones into the blood stream. The plasma peptidome contains the entire complement of low molecular weight endogenous peptides derived from secretion, protease activity and PTMs, and is a rich source of hormones. In the current study we have quantified the effects of intense exercise on the plasma peptidome to identify novel exercise regulated secretory factors in humans. We developed an optimized 2D-LC-MS/MS method and used multiple fragmentation methods including HCD and EThcD to analyze endogenous peptides. This resulted in quantification of 5,548 unique peptides during a time course of exercise and recovery. The plasma peptidome underwent dynamic and large changes during exercise on a time-scale of minutes with many rapidly reversible following exercise cessation. Among acutely regulated peptides, many were known hormones including insulin, glucagon, ghrelin, bradykinin, cholecystokinin and secretogranins validating the method. Prediction of bioactive peptides regulated with exercise identified C-terminal peptides from Transgelins, which were increased in plasma during exercise. In vitro experiments using synthetic peptides identified a role for transgelin peptides on the regulation of cell-cycle, extracellular matrix remodeling and cell migration. We investigated the effects of exercise on the regulation of PTMs and proteolytic processing by building a site-specific network of protease/substrate activity. Collectively, our deep peptidomic analysis of plasma revealed that exercise rapidly modulates the circulation of hundreds of bioactive peptides through a network of proteases and PTMs. These findings illustrate that peptidomics is an ideal method for quantifying changes in circulating factors on a global scale in response to physiological perturbations such as exercise. This will likely be a key method for pinpointing exercise regulated factors that generate health benefits.