Defining cell classes is central to the analysis of growing single-cell RNA sequencing (scRNA-seq) atlases. Marker genes are most often identified by differential expression (DE) methods that assess genes one at a time, ignoring the redundancy and complementarity revealed when genes are considered jointly. Working with binarized expression data, we instead seek discriminating panels of genes that together are specific to a cell type, framing marker-panel selection as a variant of the minimal set-covering problem in combinatorial optimization. This formulation efficiently searches the vast space of candidate panels, exploits the large cell numbers typical of scRNA-seq, and is robust to zero-inflation. Using blood and brain data, we show that our method, CellCover, reduces gene redundancy and captures cell-class-specific signals distinct from those found by DE. Transfer-learning experiments across mouse, primate, and human data demonstrate that CellCover identifies conserved cell classes in neocortical neurogenesis and tracks developmental progression in progenitors and neurons. Examining outer radial glia markers across mammals, we find that transcriptomic elements of this key cell type likely arose in rodent gliogenic precursors before the full program emerged in the primate lineage.
Definition of cell classes across the tissues of living organisms is central in the analysis of growing atlases of single-cell RNA sequencing (scRNA-seq) data across biomedicine. Marker genes for cell classes are most often defined by differential expression (DE) methods that serially assess individual genes across landscapes of diverse cells. This serial approach has been extremely useful, but is limited because it ignores possible redundancy or complementarity across genes that can only be captured by analyzing multiple genes simultaneously. We aim to identify discriminating panels of genes. To efficiently explore the vast space of possible marker panels, leverage the large number of cells often sequenced, and overcome zero-inflation in scRNA-seq data, we propose viewing gene panel selection as a variation of the "minimal set-covering problem" in combinatorial optimization. We show that this new method, CellCover, captures cell-class-specific signals in the developing mouse neocortex that are distinct from those defined by DE methods. Transfer learning experiments across mouse, primate, and human data demonstrate that CellCover identifies markers of conserved cell classes in neurogenesis, as well as temporal progression in both progenitors and neurons. Exploring markers of human outer radial glia (oRG, or basal RG) across mammals, we show that transcriptomic elements of this key cell type in the expansion of the human cortex appeared in gliogenic precursors of the rodent before the full program emerged in the primate lineage. We have assembled the public datasets we use in this report at NeMO analytics where the expression of individual genes {NeMO Individual Genes} and marker gene panels can be freely explored {NeMO: Telley 3 Sets Covering Panels}, {NeMO: Telley 12 Sets Covering Panels}, and {NeMO: Sorted Brain Cell Covering Panels}. CellCover is available in {CellCover R} and {CellCover Python}.
The selection of marker gene panels is critical for capturing the cellular and spatial heterogeneity in the expanding atlases of single-cell RNA sequencing and spatial transcriptomics data. We introduce geneCover, a label-free combinatorial method that selects an optimal panel of minimally redundant marker genes based on gene-gene correlations. Our method demonstrates excellent scalability to large datasets and identifies marker gene panels that capture distinct correlation structures across the transcriptome. This allows geneCover to distinguish cell states in various tissues of organisms effectively, including those associated with rare or difficult-to-identify cell types. We evaluate the performance of geneCover across various scRNA-seq and spatial transcriptomics datasets, comparing it to other label-free algorithms to highlight its utility in diverse biological contexts.
Spatial transcriptomics (ST) technologies measure gene expression at thousands of locations within a two-dimensional tissue slice, enabling the study of spatial gene expression patterns. Spatial variation in gene expression is characterized by spatial gradients , or the collection of vector fields describing the direction and magnitude in which the expression of each gene increases. However, the few existing methods that learn spatial gradients from ST data either make restrictive and unrealistic assumptions on the structure of the spatial gradients or do not accurately model discrete transcript locations/counts. We introduce SLOPER (for Score-based Learning Of Poisson-modeled Expression Rates), a generative model for learning spatial gradients (vector fields) from ST data. SLOPER models the spatial distribution of mRNA transcripts with an inhomogeneous Poisson point process (IPPP) and uses score matching to learn spatial gradients for each gene. SLOPER utilizes the learned spatial gradients in a novel diffusion-based sampling approach to enhance the spatial coherence and specificity of the observed gene expression measurements. We demonstrate that the spatial gradients and enhanced gene expression representations learned by SLOPER leads to more accurate identification of tissue organization, spatially variable gene modules, and continuous axes of spatial variation (isodepth) compared to existing methods. Software availability:SLOPER is available at https://github.com/chitra-lab/SLOPER .
The selection of marker gene panels is critical for capturing the cellular and spatial heterogeneity in the expanding atlases of single-cell RNA sequencing (scRNA-seq) and spatial transcriptomics data. Most current approaches to marker gene selection operate in a label-based framework, which is inherently limited by its dependency on predefined cell type labels or clustering results. In contrast, existing label-free methods often struggle to identify genes that characterize rare cell types or subtle spatial patterns, and they frequently fail to scale efficiently with large data sets. Here, we introduce geneCover, a label-free combinatorial method that selects an optimal panel of minimally redundant marker genes based on gene-gene correlations. Our method demonstrates excellent scalability to large data sets and identifies marker gene panels that capture distinct correlation structures across the transcriptome. This allows geneCover to distinguish cell states in various tissues of living organisms effectively, including those associated with rare or otherwise difficult-to-identify cell types. We evaluate the performance of geneCover across various scRNA-seq and spatial transcriptomics data sets, comparing it to other label-free algorithms to highlight its utility and potential in diverse biological contexts.
In many applications like experimental design, group testing, medical diagnosis, and active testing, the state of a random variable $Y$ is revealed by successively observing the outcomes of binary tests about $Y$, where new tests are selected adaptively based on the history of outcomes observed so far. If the number of states of $Y$ is finite, the process ends when $Y$ can be predicted with a desired level of confidence or all available tests have been used. Finding the strategy that minimizes the expected number of tests needed to predict $Y$ is virtually impossible in most real applications due to high dimensions. Therefore, the commonly used strategy is the greedy heuristic of information maximization that selects tests sequentially in order of information gain. However, this can be far from optimal for certain families of tests. In this paper, we argue that in most practical settings, for a given set of tests, there exists a $0 \ll \delta \ll \frac{1}{2}$, such that in every iteration of the greedy strategy, the selected binary test will have conditional probability of being `true', given the history, within $\delta$ units of one-half. Under this assumption, we first study the performance of the greedy strategy for the simpler case of oracle tests, that is, when all tests are functions of $Y$, and obtain tighter bounds than previously reported in literature. Subsequently, under the same assumption, we extend our analysis to incorporate noise in the test outcomes. In particular, we assume the outcomes are corrupted through a binary symmetric channel and obtain bounds on the expected number of tests needed to make accurate predictions.
Seven Tesla magnetic resonance spectroscopy (7T MRS) offers a precise measurement of metabolic levels in the human brain via a non-invasive approach. Studying longitudinal changes in brain metabolites could help evaluate the characteristics of disease over time. This approach may also shed light on how the age of study participants and duration of illness may influence these metabolites. This study used 7T MRS to investigate longitudinal patterns of brain metabolites in young adulthood in both healthy controls and patients. A four-year longitudinal cohort with 38 patients with first episode psychosis (onset within 2 years) and 48 healthy controls was used to examine 10 brain metabolites in 5 brain regions associated with the pathophysiology of psychosis in a comprehensive manner. Both patients and controls were found to have significant longitudinal reductions in glutamate in the anterior cingulate cortex (ACC). Only patients were found to have a significant decrease over time in γ-aminobutyric acid, N-acetyl aspartate, myo-inositol, total choline, and total creatine in the ACC. Together we highlight the ACC with dynamic changes in several metabolites in early-stage psychosis, in contrast to the other 4 brain regions that also are known to play roles in psychosis. Meanwhile, glutathione was uniquely found to have a near zero annual percentage change in both patients and controls in all 5 brain regions during a four-year follow-up in young adulthood. Given that a reduction of the glutathione in the ACC has been reported as a feature of treatment-refractory psychosis, this observation further supports the potential of glutathione as a biomarker for this subset of patients with psychosis.
There is a growing concern about typically opaque decision-making with high-performance machine learning algorithms. Providing an explanation of the reasoning process in domain-specific terms can be crucial for adoption in risk-sensitive domains such as healthcare. We argue that machine learning algorithms should be interpretable by design and that the language in which these interpretations are expressed should be domain- and task-dependent. Consequently, we base our model's prediction on a family of user-defined and task-specific binary functions of the data, each having a clear interpretation to the end-user. We then minimize the expected number of queries needed for accurate prediction on any given input. As the solution is generally intractable, following prior work, we choose the queries sequentially based on information gain. However, in contrast to previous work, we need not assume the queries are conditionally independent. Instead, we leverage a stochastic generative model (VAE) and an MCMC algorithm (Unadjusted Langevin) to select the most informative query about the input based on previous query-answers. This enables the online determination of a query chain of whatever depth is required to resolve prediction ambiguities. Finally, experiments on vision and NLP tasks demonstrate the efficacy of our approach and its superiority over post-hoc explanations.
There is a growing interest in the machine learning community in developing predictive algorithms that are "interpretable by design". Towards this end, recent work proposes to make interpretable decisions by sequentially asking interpretable queries about data until a prediction can be made with high confidence based on the answers obtained (the history). To promote short query-answer chains, a greedy procedure called Information Pursuit (IP) is used, which adaptively chooses queries in order of information gain. Generative models are employed to learn the distribution of query-answers and labels, which is in turn used to estimate the most informative query. However, learning and inference with a full generative model of the data is often intractable for complex tasks. In this work, we propose Variational Information Pursuit (V-IP), a variational characterization of IP which bypasses the need for learning generative models. V-IP is based on finding a query selection strategy and a classifier that minimizes the expected cross-entropy between true and predicted labels. We then demonstrate that the IP strategy is the optimal solution to this problem. Therefore, instead of learning generative models, we can use our optimal strategy to directly pick the most informative query given any history. We then develop a practical algorithm by defining a finite-dimensional parameterization of our strategy and classifier using deep networks and train them end-to-end using our objective. Empirically, V-IP is 10-100x faster than IP on different Vision and NLP tasks with competitive performance. Moreover, V-IP finds much shorter query chains when compared to reinforcement learning which is typically used in sequential-decision-making problems. Finally, we demonstrate the utility of V-IP on challenging tasks like medical diagnosis where the performance is far superior to the generative modelling approach.
Many gene signatures have been developed by applying machine learning (ML) on omics profiles, however, their clinical utility is often hindered by limited interpretability and unstable performance. Here, we show the importance of embedding prior biological knowledge in the decision rules yielded by ML approaches to build robust classifiers. We tested this by applying different ML algorithms on gene expression data to predict three difficult cancer phenotypes: bladder cancer progression to muscle-invasive disease, response to neoadjuvant chemotherapy in triple-negative breast cancer, and prostate cancer metastatic progression. We developed two sets of classifiers: mechanistic, by restricting the training to features capturing specific biological mechanisms; and agnostic, in which the training did not use any a priori biological information. Mechanistic models had a similar or better testing performance than their agnostic counterparts, with enhanced interpretability. Our findings support the use of biological constraints to develop robust gene signatures with high translational potential.
ABSTRACT Biopsy is crucial in clinical medicine to obtain tissues and cells that directly reflect the pathological changes of each disease. However, the brain is an exception due to ethical and practical challenges. Nasal biopsy, which captures the olfactory neuronal epithelium, has been considered as an alternative method of obtaining neuronal cells from living patients. Multiple groups have enriched olfactory neuronal cells (ONCs) from biopsied nasal tissue. ONCs can be obtained from repeated biopsies in a longitudinal study, providing mechanistic insight associated with dynamic changes along the disease trajectory and treatment response. Nevertheless, molecular characterization of biopsied nasal cells/tissue has been insufficient. Taking advantage of recent advances in next-generation sequencing technologies at the single-cell resolution and related rich public databases, we aimed to define the neuronal characteristics, homogeneity, and utility of ONCs. We applied single-cell and bulk RNA sequencing for ONCs, analyzing and comparing the data with multiple public datasets. We observed that the molecular signatures of ONCs are similar to those of neurons, distinct from major glial cells. The signatures of ONCs resemble those of developing neurons and share features of excitatory neurons in the prefrontal and cingulate cortex. The high homogeneity of ONCs is advantageous in pharmacological, functional, and protein studies. Accordingly, we provide two proof-of-concept examples for functional and protein studies, solidifying the utility of ONCs in studying objective biomarkers and molecular mechanisms for brain disorders. The ONCs may also be useful in the studies for the olfactory epithelium impairment and the resultant mental dysfunction elicited by SARS-CoV-2. SIGNIFICANCE STATEMENT To study dynamic changes and underlying mechanisms along disease trajectory and treatment response in neuropsychiatric disorders, olfactory neuronal cells (ONCs) enriched from biopsied nasal tissue may provide a crucial tool. Because ONCs can be obtained from repeated biopsies in a longitudinal study, this tool has been believed to be useful and complementary to postmortem brains and induced pluripotent stem cell-derived neurons. Nevertheless, molecular characterization of biopsied nasal cells/tissue has been insufficient, which hampers a broader use of this resource. Taking advantage of recent advances in next-generation sequencing technologies at the single-cell resolution and related rich public databases, the present study defines ONCs’ neuronal characteristics, homogeneity, and unique utility for the first time.
Cancer cells are adept at reprogramming energy metabolism, and the precise manifestation of this metabolic reprogramming exhibits heterogeneity across individuals (and from cell to cell). In this study, we analyzed the metabolic differences between interpersonal heterogeneous cancer phenotypes. We used divergence analysis on gene expression data of 1156 breast normal and tumor samples from The Cancer Genome Atlas (TCGA) and integrated this information with a genome-scale reconstruction of human metabolism to generate personalized, context-specific metabolic networks. Using this approach, we classified the samples into four distinct groups based on their metabolic profiles. Enrichment analysis of the subsystems indicated that amino acid metabolism, fatty acid oxidation, citric acid cycle, androgen and estrogen metabolism, and reactive oxygen species (ROS) detoxification distinguished these four groups. Additionally, we developed a workflow to identify potential drugs that can selectively target genes associated with the reactions of interest. MG-132 (a proteasome inhibitor) and OSU-03012 (a celecoxib derivative) were the top-ranking drugs identified from our analysis and known to have anti-tumor activity. Our approach has the potential to provide mechanistic insights into cancer-specific metabolic dependencies, ultimately enabling the identification of potential drug targets for each patient independently, contributing to a rational personalized medicine approach.
Interactive face retrieval aims at finding target subjects in face databases through human and machine interaction, which involves user feedback based on human perception and machine similarity measure in feature spaces. In this article, we propose an attribute prototype learning method to tackle the semantic gap between human and machine in face perception for fast interactive face retrieval. We reformulate the theoretical explanation of the interactive retrieval model and develop the algorithm of the heuristic solution of the model. Each module of the prototype model is learned with a set of identity-related facial attributes. The outputs of the prototype modules form the semantic representation. To adapt the prototype models across different databases, we propose a transfer selection algorithm based on the coherence measurements in interactive face retrieval. Coherence analysis proves that the proposed attribute prototype representation can effectively narrow down the semantic gap even in the case of cross-database transfer learning. The prototype representation can effectively reduce the feature dimension in the retrieval process. Real user retrieval with the Bayesian relevance feedback model shows that attribute prototype space is superior to low-level feature space and proves that interactive retrieval with attribute prototype representation can converge fast in large face databases.
Cancer cells display massive dysregulation of key regulatory pathways due to now well-catalogued mutations and other DNA-related aberrations. Moreover, enormous heterogeneity has been commonly observed in the identity, frequency and location of these aberrations across individuals with the same cancer type or subtype, and this variation naturally propagates to the transcriptome, resulting in myriad types of dysregulated gene expression programs. Many have argued that a more integrative and quantitative analysis of heterogeneity of DNA and RNA molecular profiles may be necessary for designing more systematic explorations of alternative therapies and improving predictive accuracy. We introduce a representation of multi- omics profiles which is sufficiently rich to account for observed heterogeneity and support the construction of quantitative, integrated, metrics of variation. Starting from the network of interactions existing in Reactome, we build a library of “paired DNA-RNA aberrations” that represent prototypical and recurrent patterns of dysregulation in cancer; each two-gene “Source-Target Pair” (STP) consists of a “source” regulatory gene and a “target” gene whose expression is plausibly “controlled” by the source gene. The STP is then “aberrant” in a joint DNA-RNA profile if the source gene is DNA-aberrant ( e.g ., mutated, deleted, or duplicated), and the downstream target gene is “RNA-aberrant”, meaning its expression level is outside the normal, baseline range. With M STPs, each sample profile has exactly one of the 2 M possible configurations. We concentrate on subsets of STPs, and the corresponding reduced configurations, by selecting tissue-dependent minimal coverings, defined as the smallest family of STPs with the property that every sample in the considered population displays at least one aberrant STP within that family. These minimal coverings can be computed with integer programming. Given such a covering, a natural measure of cross-sample diversity is the extent to which the particular aberrant STPs composing a covering vary from sample to sample; this variability is captured by the entropy of the distribution over configurations. We apply this program to data from TCGA for six distinct tumor types (breast, prostate, lung, colon, liver, and kidney cancer). This enables an efficient simplification of the complex landscape observed in cancer populations, resulting in the identification of novel signatures of molecular alterations which are not detected with frequency-based criteria. Estimates of cancer heterogeneity across tumor phenotypes reveals a stable pattern: entropy increases with disease severity. This framework is then well-suited to accommodate the expanding complexity of cancer genomes and epigenomes emerging from large consortia projects.
How can we measure the “complexity” of a learning task so that we can compare one task to another? From classical information theory, we know that entropy is a useful measure of the complexity of a random variable and provides a lower bound on the minimum expected number of bits needed for transmitting its state. In this paper, we propose to measure the complexity of a learning task by the minimum expected number of questions that need to be answered to solve the task. For example, the minimum expected number of patches that need to be observed to classify FashionMNIST images. We prove several properties of the proposed complexity measure, including connections with classical entropy and sub-additivity for multiple tasks. As the computation of the minimum expected number of questions is generally intractable, we propose a greedy procedure called “information pursuit” (IP), which selects one question at a time depending on previous questions and their answers. This requires learning a probabilistic generative model relating data and questions to the task, for which we employ variational autoencoders and normalizing flows. We illustrate the usefulness of the proposed measure on various binary image classification tasks using image patches as the query set. Our results indicate that the complexity of a classification task increases as signal-to-noise ratio decreases, and that classification of the KMNIST dataset is more complex than classification of the FashionMNIST dataset. As a byproduct of choosing patches as queries, our approach also provides a principled way of determining which pixels in an image are most informative for a task.
We consider the problem of uncovering an unknown attributed graph, where both its edges and vertices are hidden from view, through a sequence of binary questions about it. In order to select questions efficiently, we define a probability distribution over graphs, with randomness not just over edges, but over vertices as well. We then sequentially select questions so as to: (1) minimize the expected entropy of the random graph, given the answers to the previous questions in the sequence; and (2) instantiate the vertices that compose the graph. We propose some basic question spaces, from which to select questions, that vary in their capacity. We apply this framework to the problem of test generation in Visual Question Answering (VQA), where semantic questions are used to evaluate vision systems over rich image representations. To do this, we use a restricted question vocabulary, resulting in image representations that take the form of scene graphs; by defining a distribution over them, a consistent set of probabilities is associated with the questions, and used in their selection.
Given the ever-increasing amount of high-dimensional and complex omics data becoming available, it is increasingly important to discover simple but effective methods of analysis. Divergence analysis transforms each entry of a high-dimensional omics profile into a digitized (binary or ternary) code based on the deviation of the entry from a given baseline population. This is a novel framework that is significantly different from existing omics data analysis methods: it allows digitization of continuous omics data at the univariate or multivariate level, facilitates sample level analysis, and is applicable on many different omics platforms. The divergence package, available on the R platform through the Bioconductor repository collection, provides easy-to-use functions for carrying out this transformation. Here we demonstrate how to use the package with data from the Cancer Genome Atlas.
Abstract In this work we develop a framework which allows for a systematic analysis of joint DNA and putative downstream RNA effects in cancer data cohorts. Using the Reactome database, we extract gene pairs that are linked by known mechanistic connections. Such pairs, which we refer to as 'Source Target Pairs' or STPs, consist of a source gene for which we examine aberrant activity in the DNA profile, and a target gene that is affected by said source gene, for which we examine aberrant activity in the RNA profile. Using TCGA data for six different cancer types (breast, colon, kidney, liver, lung and prostate), we use mutation and copy number variation information to compile DNA aberrant activity data. For the same cancer cohorts, we use RNASeq gene expression data to quantify RNA aberrant activity via the previous 'divergence' method we have developed. In the divergence framework, normal samples from the same cancer are used to estimate a normal range of expression for target genes of interest and deviation from the normal range is assumed to indicate aberrant activity which may result from upstream DNA aberrations. Then for a given sample, an STP can be represented as a binary variable, indicating presence or absence of joint DNA-RNA aberrant activity. We utilize integer programming to discover a small set of such STPs for each cancer type such that every sample displays aberrant activity in at least one STP. We refer to these reduced STP configurations as 'minimal coverings' of that cancer. These configurations then allow for the quantification of heterogeneity for that cancer type, as well as for phenotypical groups of interest. This is made possible due to the fact that sample to sample variability can be compared via the entropy of the distribution of the minimal covering, where the small number of STPs in such a configuration makes the computation more tractable. Our results reveal many known putative drivers of cancer, as well as identify some novel genes of interest for further consideration. Comparison of heterogeneity across phenotypes of interest show higher entropy in more pathological phenotypes, indicating increasing heterogeneity with severity of disease. Citation Format: Qian Ke, Wikum Dinalankara, Laurent Younes, Donald Geman, Luigi Marchionni. Efficient representations of tumor diversity with paired DNA-RNA aberrations [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2021; 2021 Apr 10-15 and May 17-21. Philadelphia (PA): AACR; Cancer Res 2021;81(13_Suppl):Abstract nr 173.
We analyze control of the familywise error rate (FWER) in a multiple testing scenario with a great many null hypotheses about the distribution of a high-dimensional random variable among which only a very small fraction are false, or active. In order to improve power relative to conservative Bonferroni bounds, we explore a coarse-to-fine procedure adapted to a situation in which tests are partitioned into subsets, or cells, and active hypotheses tend to cluster within cells. We develop procedures for a standard linear model with Gaussian data and a non-parametric case based on generalized permutation testing, and demonstrate considerably higher power than Bonferroni estimates at the same FWER when the active hypotheses do cluster. The main technical difficulty arises from the correlation between the test statistics at the individual and cell levels, which increases the likelihood of a hypothesis being falsely discovered when the cell that contains it is falsely discovered (survivorship bias). This requires sharp estimates of certain quadrant probabilities when a cell is inactive.
Yali Amit合作论文数Department of Statistics, Physical Sciences Division, The University of Chicago;Department of Computer Science, Physical Sciences Division, The University of Chicago9