Abstract Motivation Splice variant neoantigens are a potential source of tumor-specific antigen (TSA) that are shared between patients in a variety of cancers, including acute myeloid leukemia. Current tools for genomic prediction of splice variant neoantigens demonstrate promise. However, many tools have not been well validated with simulated and/or wet lab approaches, with no studies published that have presented a targeted immunopeptidome mass spectrometry approach designed specifically for identification of predicted splice variant neoantigens. Results In this study, we describe NeoSplice, a novel computational method for splice variant neoantigen prediction based on (i) prediction of tumor-specific k-mers from RNA-seq data, (ii) alignment of differentially expressed k-mers to the splice graph and (iii) inference of the variant transcript with MHC binding prediction. NeoSplice demonstrates high sensitivity and precision (>80% on average across all splice variant classes) through in silico simulated RNA-seq data. Through mass spectrometry analysis of the immunopeptidome of the K562.A2 cell line compared against a synthetic peptide reference of predicted splice variant neoantigens, we validated 4 of 37 predicted antigens corresponding to 3 of 17 unique splice junctions. Lastly, we provide a comparison of NeoSplice against other splice variant prediction tools described in the literature. NeoSplice provides a well-validated platform for prediction of TSA vaccine targets for future cancer antigen vaccine studies to evaluate the clinical efficacy of splice variant neoantigens. Availability and implementation https://github.com/Benjamin-Vincent-Lab/NeoSplice Supplementary information Supplementary data are available at Bioinformatics Advances online.
Imagine a world in which robots are a part of everyday life, performing elegant and safe motions to accomplish complex tasks. To achieve this vision, robots will need access to extensive computational resources. Cloud-based computers have the potential to provide the needed computing power, while lowering robot cost, space, and energy requirements. Academia and industry are already exploring the cloud as a purveyor of data in a wide variety of applications, and have shown the benefit of the cloud for accelerating offline- and pre-computations.But what about interactive/online computation, as is often required by robot motion planning? This paper presents an economics-based argument that it is possible to extend a robot’s useful service life and battery-based operation time, improve its efficiency and profitability, and reduce its initial costs, by using the cloud in complex online and interactive computations. Gaining these benefits presents new, open research challenges, including: how to cost-effectively allocate cloud-based parallel computation, how to handle the unavoidable network-related bottlenecks, and how to design algorithms that distribute computation between the cloud and the robot.
Direct cellular reprogramming provides a powerful platform to study cell plasticity and dissect mechanisms underlying cell fate determination. Here, we report a single-cell transcriptomic study of human cardiac (hiCM) reprogramming that utilizes an analysis pipeline incorporating current data normalization methods, multiple trajectory prediction algorithms, and a cell fate index calculation we developed to measure reprogramming progression. These analyses revealed hiCM reprogramming-specific features and a decision point at which cells either embark on reprogramming or regress toward their original fibroblast state. In combination with functional screening, we found that immune-response-associated DNA methylation is required for hiCM induction and validated several downstream targets of reprogramming factors as necessary for productive hiCM reprograming. Collectively, this single-cell transcriptomics study provides detailed datasets that reveal molecular features underlying hiCM determination and rigorous analytical pipelines for predicting cell fate conversion.
Power is increasingly the limiting factor in High Performance Computing (HPC) at Exascale and will continue to influence future advancements in supercomputing. Recent processors equipped with on-board hardware counters allow real time monitoring of operating conditions such as energy and temperature, in addition to performance measures such as instructions retired and memory accesses. An experimental memory study presented on modern CPU architectures, Intel Sandybridge and Haswell, identifies a metric, TORo_core, that detects bandwidth saturation and increased latency. TORo-Core is used to construct a dynamic policy applied at coarse and fine-grained levels to modulate per-core power controls on Haswell machines. The coarse and fine-grained application of dynamic policy shows best energy savings of 32.1% and 19.5% with a 2% slowdown in both cases. On average for six MPI applications, the fine-grained dynamic policy speeds execution by 1% while the coarse-grained application results in a 3% slowdown. Energy savings through frequency reduction not only provide cost advantages, they also reduce resource contention and create additional thermal headroom for non-throttled cores improving performance.
Single cell genomic techniques promise to yield key insights into the dynamic interplay between gene expression and epigenetic modification. However, the experimental difficulty of performing multiple measurements on the same cell currently limits efforts to combine multiple genomic data sets into a united picture of single cell variation [1, 2]. The current understanding of epigenetic regulation suggests that any large changes in gene expression, such as those that occur during differentiation, are accompanied by epigenetic changes. This means that if cells undergoing a common process are sequenced using multiple genomic techniques, examining any of the genomic quantities should reveal the same underlying biological process. For example, the main difference among cells undergoing differentiation will be the extent of their differentiation progress, whether you look at the gene expression profiles or the chromatin accessibility profiles of the cells. We reasoned that this property of single cell data could be used to infer correspondence between different types of genomic data. To infer single cell correspondences, we use a technique called manifold alignment. Intuitively, manifold alignment constructs a low-dimensional representation (manifold) for each of the observed data types, then projects these representations into a common space (alignment) in which measurements of different types are directly comparable [3, 4]. To the best of our knowledge, manifold alignment has never been used in genomics. However, other application areas recognize the technique as a powerful tool for multimodal data fusion, such as retrieving images based on a text description, and multilingual search without direct translation [4]. We show for the first time that it is possible to construct cell trajectories, reflecting the changes that occur in a sequential biological process, from single cell epigenetic data. In addition, we present an approach called MATCHER that computationally circumvents the experimental difficulties of performing multiple genomic measurements on a single cell by inferring correspondence between single cell transcriptomic and epigenetic measurements performed on different cells of the same type. MATCHER works by first learning a separate manifold for the trajectory of each kind of genomic data, then aligning the manifolds to infer a shared trajectory in which cells measured using different techniques are directly comparable. Because there is, in general, no actual cell-to-cell correspondence
Single cell experimental techniques reveal transcriptomic and epigenetic heterogeneity among cells, but how these are related is unclear. We present MATCHER, an approach for integrating multiple types of single cell measurements. MATCHER uses manifold alignment to infer single cell multi-omic profiles from transcriptomic and epigenetic measurements performed on different cells of the same type. Using scM&T-seq and sc-GEM data, we confirm that MATCHER accurately predicts true single cell correlations between DNA methylation and gene expression without using known cell correspondences. MATCHER also reveals new insights into the dynamic interplay between the transcriptome and epigenome in single embryonic stem cells and induced pluripotent stem cells.
Single-cell transcriptomics analyses of cell intermediates during the reprogramming from fibroblast to cardiomyocyte were used to reconstruct the reprogramming trajectory and to uncover intermediate cell populations, gene pathways and regulators involved in this process. To elucidate the mechanistic underpinnings of fibroblasts reprogramming to cardiomyocytes, Li Qian and colleagues have used a single-cell RNA sequencing approach. They find that the initial steps that drive the global expression changes that are critical for reprogramming encompass the downregulation of factors involved in mRNA processing and splicing, and in particular the splicing factor Ptbp1. Downregulation of Ptbp1 is essential for cells to adopt a cardiac-specific splicing pattern. The approach also led to the identification of surface markers that allow enrichment of induced cardiomyocytes during reprogramming. Direct lineage conversion offers a new strategy for tissue regeneration and disease modelling. Despite recent success in directly reprogramming fibroblasts into various cell types, the precise changes that occur as fibroblasts progressively convert to the target cell fates remain unclear. The inherent heterogeneity and asynchronous nature of the reprogramming process renders it difficult to study this process using bulk genomic techniques. Here we used single-cell RNA sequencing to overcome this limitation and analysed global transcriptome changes at early stages during the reprogramming of mouse fibroblasts into induced cardiomyocytes (iCMs)1,2,3,4. Using unsupervised dimensionality reduction and clustering algorithms, we identified molecularly distinct subpopulations of cells during reprogramming. We also constructed routes of iCM formation, and delineated the relationship between cell proliferation and iCM induction. Further analysis of global gene expression changes during reprogramming revealed unexpected downregulation of factors involved in mRNA processing and splicing. Detailed functional analysis of the top candidate splicing factor, Ptbp1, revealed that it is a critical barrier for the acquisition of cardiomyocyte-specific splicing patterns in fibroblasts. Concomitantly, Ptbp1 depletion promoted cardiac transcriptome acquisition and increased iCM reprogramming efficiency. Additional quantitative analysis of our dataset revealed a strong correlation between the expression of each reprogramming factor and the progress of individual cells through the reprogramming process, and led to the discovery of new surface markers for the enrichment of iCMs. In summary, our single-cell transcriptomics approaches enabled us to reconstruct the reprogramming trajectory and to uncover intermediate cell populations, gene pathways and regulators involved in iCM induction.
Recent advance in technology enables researchers to gather and store enormous data sets with ultra high dimensionality. In bioinformatics, microarray and next generation sequencing technologies can produce data with tens of thousands of predictors of biomarkers. On the other hand, the corresponding sample sizes are often limited. For classification problems, to predict new observations with high accuracy, and to better understand the effect of predictors on classification, it is desirable, and often necessary, to train the classifier with variable selection. In the literature, sparse regularized classification techniques have been popular due to the ability of simultaneous classification and variable selection. Despite its success, such a sparse penalized method may have low computational speed, when the dimension of the problem is ultra high. To overcome this challenge, we propose a new sparse REgression based multicategory Classifier (REC). Our method uses a simplex to represent different categories of the classification problem. A major advantage of REC is that the optimization can be decoupled into smaller independent sparse penalized regression problems, and hence solved by using parallel computing. Consequently, REC enjoys an extraordinarily fast computational speed. Moreover, REC is able to provide class conditional probability estimation. Simulated examples and applications on microarray and next generation sequencing data suggest that REC is very competitive when compared to several existing methods.
Energy efficiency in high performance computing (HPC) will be critical to limit operating costs and carbon footprints in future supercomputing centers. Energy efficiency of a computation can be improved by reducing time to completion without a substantial increase in power drawn or by reducing power with a little increase in time to completion. We present an Adaptive Core-specific Runtime (ACR) that dynamically adapts core frequencies to workload characteristics, and show examples of both reductions in power and improvement in the average performance. This improvement in energy efficiency is obtained without changes to the application.The adaptation policy embedded in the runtime uses existing core-specific power controls like software-controlled clock modulation and per-core Dynamic Voltage Frequency Scaling (DVFS) introduced in Intel Haswell. Experiments on six standard MPI benchmarks and a real world application show an overall 20% improvement in energy efficiency with less than 1% increase in execution time on 32 nodes (1024 cores) using per-core DVFS.An improvement in energy efficiency of up to 42% is obtained with the real world application ParaDis through a combination of speedup and power reduction. For one configuration, ParaDis achieves an average speedup of 11%, while the power is lowered by about 31%. The average improvement in the performance seen is a direct result of the reduction in run-to-run variation and running at turbo frequencies.
We provide efficient implementations of common Fast Multipole Method (FMM) tasks for modern multi-core (Intel Xeon Haswell), many-core (Intel Xeon Phi Knights Landing) and Nvidia Pascal GPUs, offering optimization guidelines for each kernel and architecture, and exposing task granularity issues with evaluations on performance and scalability. These results motivate the use of hybrid execution models for FMM in heterogeneous architectures, in which per-kernel execution configurations are set by the kernel adaptability to the processor.
Genomic methods are used increasingly to interrogate the individual cells that compose specific tissues. However, current methods for single cell isolation struggle to phenotypically differentiate specific cells in a heterogeneous population and rely primarily on the use of fluorescent markers. Many cellular phenotypes of interest are too complex to be measured by this approach, making it difficult to connect genotype and phenotype at the level of individual cells. Here we demonstrate that microraft arrays, which are arrays containing thousands of individual cell culture sites, can be used to select single cells based on a variety of phenotypes, such as cell surface markers, cell proliferation and drug response. We then show that a common genomic procedure, RNA-seq, can be readily adapted to the single cells isolated from these rafts. We show that data generated using microrafts and our modified RNA-seq protocol compared favorably with the Fluidigm C1. We then used microraft arrays to select pancreatic cancer cells that proliferate in spite of cytotoxic drug treatment. Our single cell RNA-seq data identified several expected and novel gene expression changes associated with early drug resistance.
The simulation of multiscale physics is an important challenge for scientific computing. For this class of problem, large three-dimensional simulations are performed to advance scientific inquiry. On massively parallel computing systems, the volume of data generated by such approaches can become a productivity bottleneck if the raw data generated from the simulation is analyzed in a post-processing step. To address this, we present a physics-based framework for in situ data reduction that is theoretically grounded in multiscale averaging theory. We show how task parallelism can be exploited to concurrently perform a variety of analysis tasks with data-dependent costs, including the generation of iso-surfaces, morphological analyses, and connected components analysis. All analyses are performed in parallel using distributed memory and use the same domain decomposition as the simulation. A task management framework is constructed to leverage available parallelism within a node for analysis. The capabilities of the framework are to launch asynchronous analysis threads, manage dependencies between different tasks, promote data locality and minimize the impact of data transfers. The framework is applied to analyze GPU-based simulations of two-fluid-phase flow in porous media, generating a set of averaged measures that represents the overall system behavior. We demonstrate how the approach can be applied to perform physically-consistent analysis over fluid sub-regions determined from connected components analysis. Simulations performed on Oak Ridge National Lab's Titan supercomputer are profiled to demonstrate the performance of the associated multi-threaded in situ analysis approach for typical production simulation of two-fluid-phase flow.
Single cell RNA-seq experiments provide valuable insight into cellular heterogeneity but suffer from low coverage, 3' bias and technical noise. These unique properties of single cell RNA-seq data make study of alternative splicing difficult, and thus most single cell studies have restricted analysis of transcriptome variation to the gene level. To address these limitations, we developed SingleSplice, which uses a statistical model to detect genes whose isoform usage shows biological variation significantly exceeding technical noise in a population of single cells. Importantly, SingleSplice is tailored to the unique demands of single cell analysis, detecting isoform usage differences without attempting to infer expression levels for full-length transcripts. Using data from spike-in transcripts, we found that our approach detects variation in isoform usage among single cells with high sensitivity and specificity. We also applied SingleSplice to data from mouse embryonic stem cells and discovered a set of genes that show significant biological variation in isoform usage across the set of cells. A subset of these isoform differences are linked to cell cycle stage, suggesting a novel connection between alternative splicing and the cell cycle.
Triple-negative breast cancer (TNBC) accounts for 15-20% of the breast cancer cases. These tumors are heterogeneous and aggressive, fail to respond to targeted therapy, have poor prognosis, and a high risk of relapse. The molecular basis of this breast cancer subtype is currently unknown. Viruses, such as human cytomegalovirus (HCMV), Epstein-Barr virus (EBV), and human papillomavirus (HPV), are major suspects in the etiology of breast cancer. Since miRNAs have been implicated in the pathogenesis of breast cancer, including triple-negative breast cancer, we hypothesized that viral miRNAs may play a role in this aggressive disease. The objective of this project, therefore, was to determine the identity, prevalence, sequence variation, and differential expression of viral miRNAs in TNBC tumors as compared to control samples. We conducted a comprehensive profiling of viral miRNAs in 48 TNBC tumors as compared to 15 control normal breast tissues, utilizing deep sequencing analysis software and publically available deep sequencing data. Five novel putative HCMV miRNAs were found to be differentially expressed in TNBC tumors as compared to controls. Two of the putative miRNAs were differentially expressed in 60-79% of the TNBC tumors as compared to 0% of normal controls. Two additional putative miRNAs were expressed in 67-88% of TNBC tumors as compared to 25-31% of normal controls. One putative miRNA was expressed in 67% of TNBC tumors as compared to 88% of normal samples, while 1 putative miRNA in addition to HCMV-US25-1-3p displayed no significant differences. These putative HCMV miRNAs localize within genomic regions of HCMV that encode UL56 DNA packaging terminase subunit 2, single-stranded DNA-binding protein, noncoding RNA4.9, and enveloped glycoprotein M. Ongoing work has so far computationally validated one of these miRNAs as a novel HCMV miRNA. This is the first report on the differential expression of HCMV miRNAs in TNBC tumors. Since TNBC tumors are extremely heterogeneous, it is intriguing that 60-80% of TNBC tumors specifically express the same HCMV miRNA. Our findings suggest that these differentially expressed HCMV miRNAs may potentially play a role in the pathogenesis of TNBC. Citation Format: Tia Hudson, Scott Harrison, Dukka K C, Jan Prins, Perpetua M. Muganda. Differential expression of human cytomegalovirus microRNA in triple-negative breast cancer tumors. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 1120.
We introduce a method for splitting the computation of a robot’s motion plan between the robot’s low-power embedded computer, and a high-performance cloud-based compute service. To meet the requirements of an interactive and dynamic scenario, robot motion planning may need more computing power than is available on robots designed for reduced weight and power consumption (e.g., battery powered mobile robots). In our method, the robot communicates its configuration, its goals, and the obstacles to the cloud-based service. The cloud-based service takes into account the latency and bandwidth of the connection between it and the robot and computes and returns a motion plan within the time frame necessary for the robot to meet requirements of a dynamic and interactive scenario. The cloud-based service parallelizes construction of a roadmap, and returns a sparse subset of the roadmap giving the robot the ability to adapt to changes between updates from the server. In our results, we show that with typical latency and bandwidth limitations, our method gains significant improvement in the responsiveness and quality of motion plans in interactive scenarios.
Single cell experiments provide an unprecedented opportunity to reconstruct a sequence of changes in a biological process from individual "snapshots" of cells. However, nonlinear gene expression changes, genes unrelated to the process, and the possibility of branching trajectories make this a challenging problem. We develop SLICER (Selective Locally Linear Inference of Cellular Expression Relationships) to address these challenges. SLICER can infer highly nonlinear trajectories, select genes without prior knowledge of the process, and automatically determine the location and number of branches and loops. SLICER recovers the ordering of points along simulated trajectories more accurately than existing methods. We demonstrate the effectiveness of SLICER on previously published data from mouse lung cells and neural stem cells.
BACKGROUND:Recent studies have shown that some pseudogenes are transcribed and contribute to cancer when dysregulated. In particular, pseudogene transcripts can function as competing endogenous RNAs (ceRNAs). The high similarity of gene and pseudogene nucleotide sequence has hindered experimental investigation of these mechanisms using RNA-seq. Furthermore, previous studies of pseudogenes in breast cancer have not integrated miRNA expression data in order to perform large-scale analysis of ceRNA potential. Thus, knowledge of both pseudogene ceRNA function and the role of pseudogene expression in cancer are restricted to isolated examples.RESULTS:To investigate whether transcribed pseudogenes play a pervasive regulatory role in cancer, we developed a novel bioinformatic method for measuring pseudogene transcription from RNA-seq data. We applied this method to 819 breast cancer samples from The Cancer Genome Atlas (TCGA) project. We then clustered the samples using pseudogene expression levels and integrated sample-paired pseudogene, gene and miRNA expression data with miRNA target prediction to determine whether more pseudogenes have ceRNA potential than expected by chance.CONCLUSIONS:Our analysis identifies with high confidence a set of 440 pseudogenes that are transcribed in breast cancer tissue. Of this set, 309 pseudogenes exhibit significant differential expression among breast cancer subtypes. Hierarchical clustering using only pseudogene expression levels accurately separates tumor samples from normal samples and discriminates the Basal subtype from the Luminal and Her2 subtypes. Correlation analysis shows more positively correlated pseudogene-parent gene pairs and negatively correlated pseudogene-miRNA pairs than expected by chance. Furthermore, 177 transcribed pseudogenes possess binding sites for co-expressed miRNAs that are also predicted to target their parent genes. Taken together, these results increase the catalog of putative pseudogene ceRNAs and suggest that pseudogene transcription in breast cancer may play a larger role than previously appreciated.
Wei Wang (王薇)合作论文数Department of Computer Science, University of California at Los Angeles;Department of Computational Medicine, University of California at Los Angeles;Scalable Analytics Institute, University of California at Los Angeles19