Modern biological research is increasingly data-intensive, leading to a growing demand for effective training in biological data science. In this article, we provide an overview of key resources and best practices available within the Bioconductor project - an open-source software community focused on omics data analysis. This guide serves as a valuable reference for both learners and educators in the field.
BACKGROUND:While coronavirus disease 2019 (COVID-19) is primarily a respiratory infection, few studies have characterised the immune response to COVID-19 in lung tissue. We sought to understand the pathogenic role of microenvironmental interactions and the extracellular matrix in post-mortem COVID-19 lung using an integrative multi-omic approach. METHODS:Post-mortem formalin-fixed paraffin-embedded lung tissue from fatal COVID-19 and nonrespiratory death control lung underwent multi-omic evaluation by Quantseq Bulk RNA sequencing, Nanostring GeoMx spatial transcriptomics, RNAscope, multiplex immunofluorescence and immunohistochemistry, to evaluate virus distribution, immune composition and the extracellular matrix. Markers of extracellular synthesis and breakdown were measured in the serum of 215 patients with COVID-19 and 54 healthy volunteer controls using ELISA. RESULTS:We found that severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) infection was restricted to the pneumocytes and macrophages of early-stage disease. Spatial analyses revealed an immunosuppressive virus microenvironment, enriched for PDL1+IDO1+ macrophages and depleted of T-cells. Oligoclonal T-cells in COVID-19 lung showed no enrichment of SARS-CoV-2 specific T-cell receptors. Collagen VI was upregulated and contributed to alveolar wall thickening and impaired gas exchange in COVID-19 lung. Serum from COVID-19 patients showed increased levels of PRO-C6, a marker of collagen VI synthesis, predicted mortality in hospitalised patients. CONCLUSIONS:Our data refine the current model of respiratory COVID-19 with regard to virus distribution, immune niches and the role of the noncellular microenvironment in pathogenesis and risk stratification in COVID-19. We show that collagen deposition is an early event in the course of the disease.
Defining minimal standards for data collection is key to creating interoperative, searchable genomic and clinical databases. We highlight here the 1+Million Genomes Minimal Dataset for Cancer, encompassing 140 items in 8 domains to foster the collection of cancer data, inform transnational cooperation and advance precision cancer medicine.
Background Multiplexing single-cell RNA sequencing experiments reduces sequencing cost and facilitates larger scale studies. However, factors such as cell hashing quality and class size imbalance impact demultiplexing algorithm performance, reducing cost effectivenessFindings We propose a supervised algorithm, demuxSNP, leveraging both cell hashing and genetic variation between individuals (SNPs). The supervised algorithm addresses fundamental limitations in demultiplexing with only one data modality. The genetic variants (SNPs) of the subset of cells assigned with high confidence using a probabilistic hashing algorithm are used to train a KNN classifier that predicts the demultiplexing classes of unassigned or uncertain cells. We benchmark demuxSNP against hashing (HTODemux, cellhashR, GMM-demux, demuxmix) and genotype-free SNP (souporcell) methods on simulated and real data from renal cell cancer. Our results demonstrate that demuxSNP outperformed standalone hashing methods on low quality hashing data, improving overall classification accuracy and allowing more high RNA quality cells to be recovered. Through varying simulated doublet rates, we show genotype-free SNP methods are unable to identify biological samples with low cell counts at high doublet rates. When compared to unsupervised SNP demultiplexing methods, demuxSNP’s supervised approach was more robust to doublet rate in experiments with class size imbalance.Conclusions demuxSNP is a performant demultiplexing approach that uses hashing and SNP data to demultiplex datasets with low hashing quality where biological samples are genetically distinct. Unassigned cells (negatives) with high RNA quality can be recovered, making more cells available for analysis, especially when applied to data with low hashing quality or suspected misassigned cells. Pipelines for simulated data and processed benchmarking data for 5-50% doublets are publicly available. demuxSNP is available as an R/Bioconductor package ([https://doi.org/doi:10.18129/B9.bioc.demuxSNP][1]).### Competing Interest StatementThe authors have declared no competing interest.* scRNASeq : single-cell RNA sequencing SNP : single nucleotide polymorphism KNN : k-nearest neighbours [1]: https://doi.org/10.18129/B9.bioc.demuxSNP
From selfish silo to collaborative culture – embracing data-enabled cancer research Aedin Culhane and Mark Lawler, Co-Leads of the eHealth Hub for Cancer, reflect on their data-enabled cancer research journeys, how their collaborative team science approach has reaped significant dividends in cancer research and policy and how the hub is inducing a paradigm shift in how health data are deployed on the island of Ireland. An article in The Wall Street Journal (not the normal reading material for scientists) in 2011 highlighted a new approach to performing scientific research that was gaining significant credence at the time. Entitled ‘The New Einsteins Will Be Scientists Who Share’ (with the strapline ‘From cancer to cosmology, researchers could race ahead by working together – online and in the open’), the article presaged an unprecedented change in how scientists interact with each other, ushering in a culture of collaboration and cross-disciplinary research. Nowhere was this change more obvious than in the genomics and data science community, where a bottom-up movement led to the creation of the Global Alliance for Genomics and Health (GA4GH), bringing together researchers from around the world to work together to address some of human health’s greatest challenges through the deployment of data and data tools.
Background: Multiplexing single-cell RNA sequencing experiments reduces sequencing cost and facilitates larger-scale studies. However, factors such as cell hashing quality and class size imbalance impact demultiplexing algorithm performance, reducing cost-effectiveness. Findings: We propose a supervised algorithm, demuxSNP, which leverages both cell hashing and genetic variation between individuals (single-nucletotide polymorphisms [SNPs]). demuxSNP addresses fundamental limitations in demultiplexing methods that use only one data modality. Some cells may be confidently demultiplexed using probabilistic hashing methods. demuxSNP uses these data to infer the genotype of singlet and doublet clusters and predict on cells assigned as negative, uncertain, or doublet using a nearest-neighbor approach adapted for missing data. We benchmarked demuxSNP against hashing, genotype-free SNP and hybrid methods on simulated and real data from renal cell cancer. demuxSNP outperformed standalone hashing methods on low-quality hashing data benchmark, improved overall classification accuracy, and allowed more high RNA quality cells to be recovered. Through varying simulated doublet rates, we showed that genotype-free SNP and hybrid methods that leverage them were impacted by class size imbalance and doublet rate. demuxSNP's supervised approach was more robust to doublet rate in experiments with class size imbalance. Conclusions: demuxSNP uses hashing and SNP data to demultiplex datasets with low hashing quality where biological samples are genetically distinct. Unassigned or negative cells with high RNA quality are recovered, making more cells available for analysis. Data simulation and benchmarking pipelines as well as processed benchmarking data for 5-50% doublets are publicly available. demuxSNP is available as an R/Bioconductor package (https://doi.org/doi:10.18129/B9.bioc.demuxSNP).
One of the major barriers that have restricted successful use of chimeric antigen receptor (CAR) T cells the treatment of solid tumors is an unfavorable tumor microenvironment (TME). We engineered CAR cells targeting carbonic anhydrase IX (CAIX) to secrete anti -PD -L1 monoclonal antibody (mAb), termed mune -restoring (IR) CAR G36-PDL1. We tested CAR -T cells in a humanized clear cell renal cell carcinoma (ccRCC) orthotopic mouse model with reconstituted human leukocyte antigen (HLA) partially matched man leukocytes derived from fetal CD34* hematopoietic stem cells (HSCs) and bearing human ccRCC skrc59 cells under the kidney capsule. G36-PDL1 CAR -T cells, haploidentical to the tumor cells, had a potent antitumor effect compared to those without immune -restoring effect. Analysis of the TME revealed that G36-PDL1 CAR -T cells restored active antitumor immunity by promoting tumor -killing cytotoxicity, reducing immunosuppressive cell components such as M2 macrophages and exhausted CD8* T cells, and enhancing T follicular helper (Tfh)-B cell crosstalk.
2023 marks the 25th anniversary of the Good Friday Agreement, which led peace in Northern Ireland. As well as its impact on peace and reconciliation, the Good Friday Agreement has also had a lasting positive impact on cancer research and cancer care across the island of Ireland. Pursuant to the Good Friday Agreement, a Memorandum of Understanding (MOU) was signed between the respective Departments of Health in Ireland, Northern Ireland and the US National Cancer Institute (NCI), giving rise to the Ireland - Northern Ireland - National Cancer Institute Cancer Consortium, an unparalleled tripartite agreement designed to nurture and develop linkages between cancer researchers, physicians and allied healthcare professionals across Ireland, Northern Ireland and the US, delivering world class research and better care for cancer patients on the island of Ireland and driving research and innovation in the US.
Supplemental Materials & Methods and Figures 1-6. Sup. Figure 1. The effect of antiestrogens against ER+ BrCa spheroid cultures is attenuated by BMSCs. Sup. Figure 2. BMSCs attenuate the antiestrogen response of BrCa xenografts in vivo. Sup. Figure 3. BMSCs attenuate hormone-induced gene expression in cancer spheroids with varying effect on HR expression. Sup. Figure 4. BMSCs paracrine effect induces HT resistance in PrCa/BrCa 3D spheroids via IL-6 secretion or in IL-6-independent manner. Sup. Figure 5. Transcriptional signatures of genes downregulated (log2FC< -1) or upregulated (log2FC > 1) in MCF7 spheroids cocultured (8 days) with HS5 (vs. monoculture) do not correlate with relapse-free survival of 342 ER-negative BrCa patients (stratified using the upper tertile of the respective signature as cutoff). Sup. Figure 6. BrCa spheroids in coculture with BMSC acquire increased dependence on growth factor receptor and downstream signaling.
Supplementary Figures 1-3 from Epithelial Progeny of Estrogen-Exposed Breast Progenitor Cells Display a Cancer-like Methylome
Supplementary Data from Identification of Novel Kinase Targets for the Treatment of Estrogen Receptor–Negative Breast Cancer
Supplementary Data from Altered Cytoplasmic-to-Nuclear Ratio of Survivin Is a Prognostic Indicator in Breast Cancer
The therapeutic targeting of tumor cells with loss of function (LOF) for tumor suppressor genes (TSGs) is challenging across cancers, including hematologic neoplasias, because pharmacological mechanisms to restore the function(s) of such genes are not readily feasible, in contrast to e.g., inhibition of oncogenic drivers. We reasoned, however, that, although LOF for TSGs leads to de-repressed growth of neoplastic hematopoietic cells, it may not necessarily protect them from immune attack. We thus explored the hypothesis that CRISPR-based loss of function of TSGs in cells from multiple myeloma (MM), leukemias or lymphoma may still be associated with substantial response to immune effector cells such as NK cells, which have the advantage to kill tumor cells across HLA barriers. We have conducted CRISPR-based studies in 7 cell lines that represent different hematologic malignancies and different levels of sensitivity to natural killer (NK) cells, namely the B-cell lymphoma (SUDHL4), precursor B cell acute lymphoblastic leukemia (NALM6), multiple myeloma (MM1.S, LP1, KMS11), chronic myeloid leukemia (K562), and acute myeloid leukemia (MOLM14). In these genome-scale or focused CRISPR screens for LOF (CRISPR-based gene editing) or gain of function (GOF, CRISPR activation), the blood cancer lines were exposed to allogeneic donor-derived NK cells (vs. control cultures without NK cells). We evaluated the performance of genes known to represent recurrent TSGs based on genomic data of patient samples or cell lines; as well as candidate TSGs, based on results from genome-scale CRISPR gene editing screens (e.g., CERES or CHRONOS scores >0.4 in multiple DepMap releases and TPM>1 [RNA-seq]) in the same cell lines as the NK cell resistance screens. These analyses sought to identify any TSGs whose LOF may potentially alter the response of blood cancer cells to NK cells. We also evaluated genes identified as top recurrent TSGs in patient samples from MM and other hematologic neoplasias (e.g., based on prior genomic studies). On aggregate, we evaluated a collection of known and recurrent TSGs (e.g., PTEN, TP53, RB1, CDKN2C, CDKN1B, TENT5C/FAM46C) as well as other, previously underappreciated candidate genes with TSG properties in hematologic neoplasias (e.g., HIF1A, DEPDC5). Perturbation of none of these genes was identified to meet criteria for association with significant resistance to allogeneic donor-derived NK cells (e.g., log2FC>1.0, at least 3-4 sgRNAs with enrichment upon CRISPR KO or depletion with CRISPR activation, p-value <0.05, enrichment rank <100, based on rank aggregation algorithm) in any of the MM, leukemia or lymphoma cell lines examined in LOF or GOF CRISPR screens for NK cell resistance. In fact, for a limited set of cases (e.g., PTEN in KMS11 cells), KO of a TSG was associated with sensitization to NK cell treatment. To probe the mechanistic basis for this lack of effect of TSG loss, we examined the molecular sequelae of CRISPR-based KO for one of these recurrent TSGs, PTEN, by performing scRNAseq using the CROP-seq platform. Pools of MM1.S and LP1 cells expressing sgRNAs targeting select hits from our CRISPR screens and also PTEN were co-cultured with NK cells for 24 h or left untreated, followed by scRNA-seq and sgRNA detection, differential gene-expression analysis. We observed in these studies that the transcriptional changes induced in MM cells by KO of PTEN included limited, if any, changes in the expression of key genes involved in regulation of NK cell responses of tumor cells, e.g., ligands for activating or inhibitor receptors in NK cells, death receptors (e.g., TRAIL, Fas) or their downstream effectors/regulators or other molecules identified from our aforementioned genome-scale and focused CRISPR LOF and GOF studies. Notably, our in-house results from studies of NK exposure of cell lines from hematologic neoplasias are concordant with results for these TSGs in CRISPR screens of non-hematological cancer cells treated with cytotoxic T-cells (e.g., 4T1 or RENCA cells; GSE149933). Overall, our observations indicate that MM, leukemia or lymphoma cells that become deficient for diverse TSGs based on CRISPR-based gene editing are equally responsive to NK cells as their TSG-proficient counterparts. NK cell-based therapies may thus be a promising approach to target TSG-deficient hematologic malignancies for which specific pharmacological therapies are not currently available.
XLS file, 509K, supplementary Tables 1-3: 1. Information from the Boston cohort, 2. Information from AOCS cohort, 3. Information from TCGA.
PDF file, 80K, Supplementary Tables 4-8: 4. LOH patterns in three cohorts, 5A&B. FLOH and chemotherapy resistance in three cohorts, 6. BRCA mutations in the AOCS, 7. Univariate and multivariate analysis of platinum resistantce, 8. PFS in LOH subclusters and patients with BRCA mutations.
Supplementary Tables, Figures and Methods from Confounding Effects in “A Six-Gene Signature Predicting Breast Cancer Lung Metastasis”
PDF file, 366K, Supplementary Figures 1-9: 1. LOH patterns on different array platforms, 2. Allelic imbalance in LOH clusters, 3. LOH clustering breast cancer, 4. LOH and copy number prevalence in breast cancer, 5. Chemotherapy resistance and FLOH in three cohorts, 6. Chemotherapy resistance and FLOH with or without BRCA mutations, 7. PFS regardless of stage and residual disease, 8. PFS with or without BRCA mutations, 9. BRCA1 expression and FLOH.
Effective dimension reduction is essential for single cell RNA-seq (scRNAseq) analysis. Principal component analysis (PCA) is widely used, but requires continuous, normally-distributed data; therefore, it is often coupled with log-transformation in scRNAseq applications, which can distort the data and obscure meaningful variation. We describe correspondence analysis (CA), a count-based alternative to PCA. CA is based on decomposition of a chi-squared residual matrix, avoiding distortive log-transformation. To address overdispersion and high sparsity in scRNAseq data, we propose five adaptations of CA, which are fast, scalable, and outperform standard CA and glmPCA, to compute cell embeddings with more performant or comparable clustering accuracy in 8 out of 9 datasets. In particular, we find that CA with Freeman–Tukey residuals performs especially well across diverse datasets. Other advantages of the CA framework include visualization of associations between genes and cell populations in a “CA biplot,” and extension to multi-table analysis; we introduce corralm for integrative multi-table dimension reduction of scRNAseq data. We implement CA for scRNAseq data in corral , an R/Bioconductor package which interfaces directly with single cell classes in Bioconductor. Switching from PCA to CA is achieved through a simple pipeline substitution and improves dimension reduction of scRNAseq datasets.