While numerous methods have been developed for analyzing scRNA-seq data, benchmarking various methods remains challenging. There is a lack of ground truth datasets for evaluating novel gene selection and/or clustering methods. We propose the use of crafted experiments, a new approach based upon perturbing signals in a real dataset for comparing analysis methods. We demonstrate the effectiveness of crafted experiments for evaluating new univariate distribution-oriented suite of feature selection methods, called GOF. We show GOF selects features that robustly identify crafted features and perform well on real non-crafted data sets. Using varying ways of crafting, we also show the context in which each GOF method performs the best. GOF is implemented as an open-source R package and freely available under GPL-2 license at https://github.com/siyao-liu/GOF. Source code, including all functions for constructing crafted experiments and benchmarking feature selection methods, are publicly available at https://github.com/siyao-liu/CraftedExperiment.
Clustering methods are popular for revealing structure in data, particularly in the high-dimensional setting common to contemporary data science. A central statistical question is, "are the clusters really there?" One pioneering method in statistical cluster validation is SigClust, but it is severely underpowered in the important setting where the candidate clusters have unbalanced sizes, such as in rare subtypes of disease. We show why this is the case, and propose a remedy that is powerful in both the unbalanced and balanced settings, using a novel generalization of k-means clustering. We illustrate the value of our method using a high-dimensional dataset of gene expression in kidney cancer patients. A Python implementation is available at https://github.com/thomaskeefe/sigclust.
Abstract Single-cell RNA sequencing (scRNA-seq) technologies have revolutionized our understanding of cellular compositions and tumor behavior at the single-cell level, and have been widely used in many biomedical applications, such as identifying and characterizing novel cell types/states. While numerous computational methods have been developed for analyzing scRNA-seq data, benchmarking various analytical methods remains a key challenge for two reasons. First, there is a lack of ground truth datasets where unambiguous cell type identities are known using external information. Second, simulation of scRNA-seq data tends to be overly artificial and simplistic, which may not well mimic the underlying biological complexity as well as high level of technical variability. Here, we propose the use of crafted experiments, a new approach based upon perturbing signals in a real dataset for comparing different scRNA-seq analytical methodologies. We demonstrate the effectiveness of crafted experiments in the context of a novel univariate distribution-oriented suite of feature selection methods, called GOF (Goodness of Fit). We show that GOF more frequently selects features that robustly identify crafted artificial clusters in crafted experiments and achieves similar performance compared to the field standard method using real datasets. Importantly, we show the use of crafted experiments for identifying the contexts in which each method performs the best. Crafted experiments offer valuable comparisons of scRNA-seq analysis methods, and the crafted datasets generated from this study provide a useful resource for the single-cell community. Citation Format: Siyao Liu, David Corcoran, Susana Garcia-Recio, Charles Perou, J.S. Marron. Crafted experiments to evaluate feature selection methods for single cell RNA-seq data [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 2 (Late-Breaking, Clinical Trial, and Invited Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(7_Suppl):Abstract nr LB253.
This paper is motivated by the joint analysis of genetic, imaging, and clinical (GIC) data collected in the Alzheimer's Disease Neuroimaging Initiative (ADNI) study. We propose a regression framework based on partially functional linear regression models to map high-dimensional GIC-related pathways for Alzheimer's Disease (AD). We develop a joint model selection and estimation procedure by embedding imaging data in the reproducing kernel Hilbert space and imposing the L0 penalty for the coefficients of genetic variables. We apply the proposed method to the ADNI dataset to identify important features from tens of thousands of genetic polymorphisms (reduced from millions using a preprocessing step) and study the effects of a certain set of informative genetic variants and the baseline hippocampus surface on thirteen future cognitive scores measuring different aspects of cognitive function. We explore the shared and different heritability patterns of these cognitive scores. Analysis results suggest that both the hippocampal and genetic data have heterogeneous effects on different scores, with the trend that the value of both hippocampi is negatively associated with the severity of cognition deficits. Polygenic effects are observed for all thirteen cognitive scores. The well-known APOE4 genotype only explains a small part of cognitive function. Shared genetic etiology exists, however, greater genetic heterogeneity exists within disease classifications after accounting for the baseline diagnosis status. These analyses are useful in further investigation of functional mechanisms for AD evolution.
High intratumoral heterogeneity is thought to be a poor prognostic indicator. However, the source of heterogeneity may also be important, as genomic heterogeneity is not always reflected in histologic or ‘visual’ heterogeneity. We aimed to develop a predictor of histologic heterogeneity and evaluate its association with outcomes and molecular heterogeneity. We used VGG16 to train an image classifier to identify unique, patient-specific visual features in 1655 breast tumors (5907 core images) from the Carolina Breast Cancer Study (CBCS). Extracted features for images, as well as the epithelial and stromal image components, were hierarchically clustered, and visual heterogeneity was defined as a greater distance between images from the same patient. We assessed the association between visual heterogeneity, clinical features, and DNA-based molecular heterogeneity using generalized linear models, and we used Cox models to estimate the association between visual heterogeneity and tumor recurrence. Basal-like and ER-negative tumors were more likely to have low visual heterogeneity, as were the tumors from younger and Black women. Less heterogeneous tumors had a higher risk of recurrence (hazard ratio = 1.62, 95% confidence interval = 1.22–2.16), and were more likely to come from patients whose tumors were comprised of only one subclone or had a TP53 mutation. Associations were similar regardless of whether the image was based on stroma, epithelium, or both. Histologic heterogeneity adds complementary information to commonly used molecular indicators, with low heterogeneity predicting worse outcomes. Future work integrating multiple sources of heterogeneity may provide a more comprehensive understanding of tumor progression.
BACKGROUND:Breast cancer subtypes Luminal A and Luminal B are classified by the expression of PAM50 genes and may benefit from different treatment strategies. Machine learning models based on H&E images may contain features associated with subtype, allowing early identification of tumors with higher risk of recurrence. METHODS:H&E images (n = 630 ER+/HER2-breast cancers) were pixel-level segmented into epithelium and stroma. Convolutional neural network and multiple instance learning were used to extract image features from original and segmented images. Patient-level classification models were trained to discriminate Luminal A versus B image features in tenfold cross-validation, with or without grade adjustment. The best-performing visual classifier was incorporated into envisioned diagnostic protocols as an alternative to genomic testing (PAM50). The protocols were then compared in time-to-recurrence models. RESULTS:Among ER+/HER2-tumors, the image-based protocol differentiated recurrence times with a hazard ratio (HR) of 2.81 (95% CI: 1.73-4.56), which was similar to the HR for PAM50 (2.66, 95% CI: 1.65-4.28). Grade adjustment did not improve subtype prediction accuracy, but did help balance sensitivity and specificity. Among high grade participants, sensitivity and specificity (0.734 and 0.474, respectively) became more similar (0.732 and 0.624, respectively) in grade-adjusted models. The original and epithelium-specific images had similar performance and highest accuracy, followed by stroma or binarized images showing only the epithelial-stromal interface. CONCLUSIONS:Given low rates of genomic testing uptake nationally, image-based methods may help identify ER+/HER2-patients who could benefit from testing.
Objective:To investigate the relationship between measures of radiographic joint space width (JSW) loss and magnetic resonance imaging (MRI)-based cartilage thickness loss in the medial weight-bearing region of the tibiofemoral joint over 12-24 months. To stratify this relationship by clinically meaningful subgroups (sex and pain status). Design:We analyzed a subset of knees (n = 256) from the Osteoarthritis Initiative (OAI) likely in early stage OA based on joint space narrowing (JSN) measurements. Natural logarithm transformation was used to approximate near normal distributions for JSW loss. Pearson Correlation coefficients described the relationship between ln-transformed JSW loss and several versions of deep learning-derived MRI-based cartilage thickness loss parameters (minimum, maximum, and mean) in subregions of the femoral condyle, tibial plateau, and combined femoral and tibial regions. Linear mixed-effects models evaluated the associations between the ln-transformed radiographic and MRI-derived measures including potential confounders. Results:We found weak correlations between ln-transformed JSW loss and MRI-based cartilage thickness ranging from R = -0.13 (p = 0.20) to R = 0.26 (p < 0.01). Correlations were higher (still poor) among females compared to males and painful compared to non-painful knees. Model results showed weak associations for nearly all MRI-based measures, ranging from no association to β (95% CI) = 0.25 (0.11, 0.39). Associations were higher among females compared to males and minimal differences between painful and non-painful knees. Conclusions:Despite its recommended use in disease-modifying OA drug clinical trials, results suggest that JSW loss is an ineffective proxy measure of cartilage thickness loss over 12-24 months and within a localized region of the tibiofemoral joint.
Automated region of interest detection in histopathological image analysis is a challenging and important topic with tremendous potential impact on clinical practice. The deep learning methods used in computational pathology may help us to reduce costs and increase the speed and accuracy of cancer diagnosis. We started with the UNC Melanocytic Tumor Dataset cohort which contains 160 hematoxylin and eosin whole slide images of primary melanoma (86) and nevi (74). We randomly assigned 80% (134) as a training set and built an in-house deep learning method to allow for classification, at the slide level, of nevi and melanoma. The proposed method performed well on the other 20% (26) test dataset; the accuracy of the slide classification task was 92.3% and our model also performed well in terms of predicting the region of interest annotated by the pathologists, showing excellent performance of our model on melanocytic skin tumors. Even though we tested the experiments on a skin tumor dataset, our work could also be extended to other medical image detection problems to benefit the clinical evaluation and diagnosis of different tumors.
BackgroundModeling of single cell RNA-sequencing (scRNA-seq) data remains challenging due to a high percentage of zeros and data heterogeneity, so improved modeling has strong potential to benefit many downstream data analyses. The existing zero-inflated or over-dispersed models are based on aggregations at either the gene or the cell level. However, they typically lose accuracy due to a too crude aggregation at those two levels.ResultsWe avoid the crude approximations entailed by such aggregation through proposing an independent Poisson distribution (IPD) particularly at each individual entry in the scRNA-seq data matrix. This approach naturally and intuitively models the large number of zeros as matrix entries with a very small Poisson parameter. The critical challenge of cell clustering is approached via a novel data representation as Departures from a simple homogeneous IPD (DIPD) to capture the per-gene-per-cell intrinsic heterogeneity generated by cell clusters. Our experiments using real data and crafted experiments show that using DIPD as a data representation for scRNA-seq data can uncover novel cell subtypes that are missed or can only be found by careful parameter tuning using conventional methods.ConclusionsThis new method has multiple advantages, including (1) no need for prior feature selection or manual optimization of hyperparameters; (2) flexibility to combine with and improve upon other methods, such as Seurat. Another novel contribution is the use of crafted experiments as part of the validation of our newly developed DIPD-based clustering pipeline. This new clustering pipeline is implemented in the R (CRAN) package scpoisson.
In the age of big data, data integration is a critical step especially in the understanding of how diverse data types work together and work separately. Among data integration methods, the Angle-Based Joint and Individual Variation Explained (AJIVE) approach is particularly attractive because it not only studies joint behavior but also individual behavior. Typically AJIVE scores indicate important relationships between data objects, such as clusters. An important challenge is understanding which features, i.e. variables, are associated with those relationships. This challenge is addressed by the proposal of a hypothesis test for assessing statistical significance of features. The new test is inspired by the related jackstraw method developed for Principal Component Analysis. We use a high-dimensional multi-genomic cancer data set as our strong motivation and deep illustration of the methodology.
Advanced genomic and molecular profiling technologies accelerated the enlightenment of the regulatory mechanisms behind cancer development and progression, and the targeted therapies in patients. Along this line, intense studies with immense amounts of biological information have boosted the discovery of molecular biomarkers. Cancer is one of the leading causes of death around the world in recent years. Elucidation of genomic and epigenetic factors in Breast Cancer (BRCA) can provide a roadmap to uncover the disease mechanisms. Accordingly, unraveling the possible systematic connections between-omics data types and their contribution to BRCA tumor progression is crucial. In this study, we have developed a novel machine learning (ML) based integrative approach for multi-omics data analysis. This integrative approach combines information from gene expression (mRNA), microRNA (miRNA) and methylation data. Due to the complexity of cancer, this integrated data is expected to improve the prediction, diagnosis and treatment of disease through patterns only available from the 3-way interactions between these 3-omics datasets. In addition, the proposed method bridges the interpretation gap between the disease mechanisms that drive onset and progression. Our fundamental contribution is the 3 Multi-omics integrative tool (3Mint). This tool aims to perform grouping and scoring of groups using biological knowledge. Another major goal is improved gene selection via detection of novel groups of cross-omics biomarkers. Performance of 3Mint is assessed using different metrics. Our computational performance evaluations showed that the 3Mint classifies the BRCA molecular subtypes with lower number of genes when compared to the miRcorrNet tool which uses miRNA and mRNA gene expression profiles in terms of similar performance metrics (95% Accuracy). The incorporation of methylation data in 3Mint yields a much more focused analysis. The 3Mint tool and all other supplementary files are available at https://github.com/malikyousef/3Mint/.
Data matrix centering is an ever-present yet under-examined aspect of data analysis. Functional data analysis (FDA) often operates with a default of centering such that the vectors in one dimension have mean zero. We find that centering along the other dimension identifies a novel useful mode of variation beyond those familiar in FDA. We explore ambiguities in both matrix orientation and nomenclature. Differences between centerings and their potential interaction can be easily misunderstood. We propose a unified framework and new terminology for centering operations. We clearly demonstrate the intuition behind and consequences of each centering choice with informative graphics. We also propose a new direction energy hypothesis test as part of a series of diagnostics for determining which choice of centering is best for a data set. We explore the application of these diagnostics in several FDA settings.
Objective: To employ novel methodologies to identify phenotypes in knee OA based on variation among three baseline data blocks: 1) femoral cartilage thickness, 2) tibial cartilage thickness, and 3) participant characteristics and clinical features. Methods: Baseline data were from 3321 Osteoarthritis Initiative (OAI) participants with available cartilage thickness maps (6265 knees) and 77 clinical features. Cartilage maps were obtained from 3D DESS MR images using a deep-learning based segmentation approach and an atlas-based analysis developed by our group. Anglebased Joint and Individual Variation Explained (AJIVE) was used to capture and quantify variation, both shared among multiple data blocks and individual to each block, and to determine statistical significance. Results: Three major modes of variation were shared across the three data blocks. Mode 1 reflected overall thicker cartilage among men, those with higher education, and greater knee forces; Mode 2 showed associations between worsening Kellgren-Lawrence Grade, medial cartilage thinning, and worsening symptoms; and Mode 3 contrasted lateral and medial-predominant cartilage loss associated with BMI and malalignment. Each data block also demonstrated individual, independent modes of variation consistent with the known discordance between symptoms and structure in knee OA and reflecting the importance of features such as physical function, symptoms, and comorbid conditions independent of structural damage. Conclusions: This exploratory analysis, combining the rich OAI dataset with novel methods for determining and visualizing cartilage thickness, reinforces known associations in knee OA while providing insights into the potential for data integration in knee OA phenotyping.
Supplementary Figure S1 from A 10-Gene Classifier for Distinguishing Head and Neck Squamous Cell Carcinoma and Lung Squamous Cell Carcinoma
For measuring the strength of visually-observed subpopulation differences, the Population Difference Criterion is proposed to assess the statistical significance of visually observed subpopulation differences. It addresses the following challenges: in high-dimensional contexts, distributional models can be dubious; in high-signal contexts, conventional permutation tests give poor pairwise comparisons. We also make two other contributions: Based on a careful analysis we find that a balanced permutation approach is more powerful in high-signal contexts than conventional permutations. Another contribution is the quantification of uncertainty due to permutation variation via a bootstrap confidence interval. The practical usefulness of these ideas is illustrated in the comparison of subpopulations of modern cancer data.
Approaches for rapidly identifying patients at high risk of early breast cancer recurrence are needed. Image-based methods for prescreening hematoxylin and eosin (H&E) stained tumor slides could offer temporal and financial efficiency. We evaluated a data set of 704 1-mm tumor core H&E images (2–4 cores per case), corresponding to 202 participants (101 who recurred; 101 non-recurrent matched on age and follow-up time) from breast cancers diagnosed between 2008–2012 in the Carolina Breast Cancer Study. We leveraged deep learning to extract image information and trained a model to identify recurrence. Cross-validation accuracy for predicting recurrence was 62.4% [95% CI: 55.7, 69.1], similar to grade (65.8% [95% CI: 59.3, 72.3]) and ER status (66.3% [95% CI: 59.8, 72.8]). Interestingly, 70% (19/27) of early-recurrent low-intermediate grade tumors were identified by our image model. Relative to existing markers, image-based analyses provide complementary information for predicting early recurrence.
Binary expansion testing (BET) provides powerful detection of interesting nonlinear dependence among pairs of variables in the exploratory data analysis of large-scale data sets. However, the Bonferroni adjusted p-values can be overly conservative when used to determine the significant testing pairs. A novel contribution of this paper is the extreme value theory analysis of BET. This results in a potentially powerful new significance threshold for the maximal BET z-statistics.
Correlated shape features involving nearby objects often contain important anatomic information. However, it is difficult to capture shape information within and between objects for a joint analysis of multi-object complexes. This paper proposes (1) capturing between-object shape based on an explicit mathematical model called a linking structure, (2) capturing shape features that are invariant to rigid transformation using local affine frames and (3) capturing Correlation of Within- and Between-Object (CoWBO) shape features using a statistical method called NEUJIVE. The resulting correlated shape features give comprehensive understanding of multi-object complexes from various perspectives. First, these features explicitly account for the positional and geometric relations between objects that can be anatomically important. Second, the local affine frames give rise to rich interior geometric features that are invariant to global alignment. Third, the joint analysis of within- and between-object shape yields robust and useful features. To demonstrate the proposed methods, we classify individuals with autism and controls using the extracted shape features of two functionally related brain structures, the hippocampus and the caudate. We found that the CoWBO features give the best classification performance among various choices of shape features. Moreover, the group difference is statistically significant in the feature space formed by the proposed methods.
High throughput -omics technologies facilitate the investigation of regulatory mechanisms of complex diseases. Along this line, scientists develop promising tools and methods to extend our understanding at the molecular and functional levels. To this end, miRcorrNet tool performs integrative analysis of microRNA (miRNA) and gene expression profiles via machine learning (ML) approach to identify significant miRNA groups and their associated target genes. In this study, we propose miRcorrNetPro tool, which extends miRcorrNet by tracking group scoring, ranking and other information through the cross-validation iterations. Heatmap visualizations enable deep novel insights into the collective behavior of clusters of groups in cellular signaling and hence facilitate detection of potential biomarkers for the disease under investigation. Although miRcorrNetPro is designed as a generic tool, here we present our findings and potential miRNA biomarkers for Breast Cancer (BRCA). The miRcorrNetPro tool and all other supplementary files are available at https://github.com/Miray-Unlu/miRcorrNetPro.
In The Cancer Genome Atlas (TCGA) data set, there are many interesting nonlinear dependencies between pairs of genes that reveal important relationships and subtypes of cancer. Such genomic data analysis requires a rapid, powerful and interpretable detection process, especially in a high-dimensional environment. We study the nonlinear patterns among the expression of pairs of genes from TCGA using a powerful tool called Binary Expansion Testing. We find many nonlinear patterns, some of which are driven by known cancer subtypes, some of which are novel.
Fred Godtliebsen合作论文数Department of Mathematics, University of Tromsø9