Tumor microenvironments (TMEs) are compositionally and functionally heterogeneous, making it challenging to discover organizing structural principles. Through a study of 262 solid tumors profiled by spatial transcriptomics, we identify a conserved architecture where TMEs are partitioned into discrete, hierarchically organized multicellular sub-regions, which we term “spatial groups” (SGs). As indicated by orthogonal spatial measurements and expert pathologist review, SGs associate with recognizable biological domains spanning global tissue context to local cellular neighborhoods. Comparing tumors through SGs reveals a pan-tumor classification where the dominant axis of variation is spatial heterogeneity of immune biology. In an independent, retrospective cohort of non-small cell lung cancer patients treated with immune checkpoint blockade (ICB; n = 16), pan-tumor spatial biology classification distinguishes clinical response and captures structural and biological hallmarks associated with ICB sensitivity. Together, these findings suggest that SGs may be important organizing domains of the TME that relate spatial structure, biological function, and response to therapy.
The human gut microbiome contains many bacterial strains of the same species (“strain-level variants”) that shape microbiome function. The tremendous scale and molecular resolution at which microbial communities are being interrogated motivates addressing how to describe strain-level variants. We introduce the “Spectral Tree”—an inferred tree of relatedness built from patterns of co-evolutionary constraint between greater than 7,000 diverse bacteria. Using the Spectral Tree to describe over 600 diverse gut commensal strains that we isolated, whole-genome sequenced, and metabolically profiled revealed (1) widespread phylogenetic structure among strain-level variants, (2) the origins of subspecies phylogeny as a shared history of phage infections across humans, and (3) the key role of inter-human strain variation in predicting strain-level metabolic qualities. Overall, our work demonstrates the existence and metabolic importance of structured phylogeny below the level of species for commensal gut bacteria, motivating a redefinition of individual strains according to their evolutionary context. A record of this paper’s transparent peer review process is included in the supplemental information.
Hematoxylin and eosin (H&E) is a common and inexpensive histopathology assay. Though widely used and information-rich, it cannot directly inform about specific molecular markers, which require additional experiments to assess. To address this gap, we present ROSIE, a deep-learning framework that computationally imputes the expression and localization of dozens of proteins from H&E images. Our model is trained on a dataset of over 1300 paired and aligned H&E and multiplex immunofluorescence (mIF) samples from over a dozen tissues and disease conditions, spanning over 16 million cells. Validation of our in silico mIF staining method on held-out H&E samples demonstrates that the predicted biomarkers are effective in identifying cell phenotypes, particularly distinguishing lymphocytes such as B cells and T cells, which are not readily discernible with H&E staining alone. Additionally, ROSIE facilitates the robust identification of stromal and epithelial microenvironments and immune cell subtypes like tumor-infiltrating lymphocytes (TILs), which are important for understanding tumor-immune interactions and can help inform treatment strategies in cancer research.
Immune checkpoint blockade (ICB) has transformed cancer therapy and is now approved in every type of cancer, yet it remains difficult to identify the ∼ 20% of patients that will benefit from such therapy. Explainable models that predict ICB response can potentially spare ‘responders’ from toxic chemotherapies while identifying alternative options for ‘non-responders’. The only widely approved ICB biomarker of response, PD-L1, treats TME cells as individual components and is only weakly predictive of outcomes. Recent studies have suggested that biomarkers built on cell-to-cell spatial proximity, such as tumor cell to CD8+ T cell distance, can improve upon PD-L1 in predicting ICB efficacy. Yet, there is a lack of generalized approaches for identifying tumor spatial biology relationships relevant to ICB response, especially considering the wealth of information generated by high-resolution spatial biology techniques. Our approach was (1) to integrate spatial transcriptomics datasets from 96 tumors across twelve tumor types and (2) to apply statistical covariation approaches to identify organizing principles of spatial biology that would facilitate the comparison of these datasets. The basis of this approach is our identification of a conserved unit of pan-tumor spatial organization that we termed ‘Spatial Groups’. Spatial Groups (SGs) are multi-cellular units demonstrating coordinated transcriptional programs contained within spatial domains ranging in size from tens of microns (dozens of cells) to mm-scale (large regions of a tumor). SGs are hierarchically organized such that large-scale SGs exert contextual field effects on their nested small-scale SGs. For example, we found that T cell presence in large-scale SGs impacted metabolic pathway shifts occurring within nested medium-scale SGs, which in turn impacted cell adhesion programs within nested small-scale SGs. We found that SGs revealed immune-related spatial biology patterns that we hypothesized would serve as a genome-wide spatial biomarker of ICB response. To test this idea, we performed a retrospective study of 16 patients with metastatic non-small cell lung cancer treated with ICB or chemotherapy-ICB in the frontline setting. Unlike PD-L1 status, our SG-derived biomarker was highly significant in predicting progression-free survival after therapy initiation. Additionally, analysis of SGs in non-responder patients provided insights into patient-specific spatial biology architectures limiting ICB efficacy. In summary, our discovery of Spatial Groups provides a general framework for using spatial biology to identify pan-tumor ICB response. Vivek Behera, Hannah Giba, Ue-Yu Pen, Anna Di Lello, Benjamin A. Doran, Alessandra Esposito, Apameh Pezeshk, Christine Bestvina, Justin P. Kline, Marina C. Garassino, Arjun S. Raman. Spatial Groups provide a pan-tumor spatial transcriptomic biomarker of immunotherapy response [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 7480.
Heterogeneity of immune checkpoint blockade (ICB) response amongst cancer patients has a profound impact on cancer-related mortality. It is now well understood that immunotherapy efficacy and resistance relate to a wide spectrum of cellular and non-cellular components in the tumor microenvironment (TME). Although the only widely used ICB biomarker, PD-L1, treats TME cells as individual components, recent work on possible biomarkers predictive of ICB response have had success by considering interactions within the TME, such as tumor cell to CD8+ T cell distances. We hypothesized that application of generalized statistical interaction approaches to tumor spatial transcriptomic datasets, which contain genome-wide information at nearly single-cell resolution, might yield novel predictive biomarkers of ICB response through leveraging TME spatial biology. Our approach was (1) to integrate spatial transcriptomics datasets from 96 tumors across twelve tumor types and (2) to apply statistical covariation approaches to identify organizing principles of spatial biology that would facilitate comparison of these datasets. Our work has identified ‘Spatial Groups’ as a conserved unit of pan-tumor spatial organization. Spatial Groups are multi-cellular units demonstrating coordinated transcriptional programs contained within spatial domains ranging in size from tens of microns (dozens of cells) to mm-scale (large regions of a tumor). These domains are hierarchically organized such that large-scale domains exert contextual field effects on nested smaller-scale domains. We found that Spatial Groups revealed immune-related spatial biology that we hypothesized might serve as a genome-wide spatial biomarker of ICB response. To test this idea, we performed a retrospective study of 16 patients with metastatic non-small cell lung cancer treated with ICB or chemotherapy-ICB in the frontline setting. We found that our Spatial Group-derived biomarker, unlike PD-L1 status, was highly significant in predicting progression-free survival after therapy initiation. Additionally, analysis of Spatial Groups in non-responder patients provided insights into patient-specific spatial biology architectures limiting ICB efficacy. In summary, our discovery of Spatial Groups provides a general framework for using spatial biology to identify pan-tumor ICB response. Citation Format: Vivek Behera, Hannah Giba, Ue-Yu Pen, Anna Di Lello, Benjamin A. Doran, Alessandra Esposito, Apameh Pezeshk, Christine M. Bestvina, Justin Kline, Marina C. Garassino, Arjun S. Raman. Spatial groups provide a pan-tumor spatial transcriptomic biomarker of immunotherapy response [abstract]. In: Proceedings of the AACR Special Conference in Cancer Research: Functional and Genomic Precision Medicine in Cancer: Different Perspectives, Common Goals; 2025 Mar 11-13; Boston, MA. Philadelphia (PA): AACR; Cancer Res 2025;85(5 Suppl):Abstract nr A038.
Abstract Heterogeneity of therapeutic response amongst cancer patients has a profound impact on cancer-related mortality. It has been widely hypothesized that diversity in clinical outcomes is driven by variability in tumor biology at scales ranging from inter-patient to inter-cellular. Indeed, the advent of technologies describing the tumor microenvironment (TME) at cellular resolution has demonstrated that TME organization is heterogeneous with respect to intra-tumor spatial location, body tumor location (primary vs metastatic), patient, and tumor type. However, the organizing principles that govern TME heterogeneity remain unclear. Without such organizing principles, it remains an arduous task to link descriptions of tumor biology heterogeneity to clinical outcomes in a mechanistically explainable manner. Through the application of network statistical inference models to spatial transcriptomic data, we have identified that nested spatial units (NSUs) are a core organizing principle of the TME. We find that NSUs are widely conserved across 96 patient biopsies of 14 distinct tumor types and 3 body locations (primary, abdominal metastases, brain metastases). Moreover, NSUs represent a transferrable organizing blueprint that allows for the spatial location of individual spots (5-50 cells) in a tumor biopsy to be predicted by the NSUs of a secondary, unrelated biopsy (Pearson correlation of spot-spot distance: 0.04 - 0.71, median 0.25). By using NSUs as the organizing unit of comparison, we identify variation in gene expression, biological pathways, and multicellular structures occurring across tumors in a distance-dependent manner. In summary, NSUs provide a foundation for the comparative analysis of tumor spatial biology and a path by which spatial biology can be mechanistically linked to patient outcomes. Citation Format: Vivek Behera, Hannah Giba, Benjamin A. Doran, Justin P. Kline, Arjun S. Raman. Identifying nested spatial units as a conserved organizing principle of the spatial tumor microenvironment [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2024; Part 1 (Regular Abstracts); 2024 Apr 5-10; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2024;84(6_Suppl):Abstract nr 6226.
The tumor microenvironment (TME) is an immensely complex ecosystem1,2. This complexity underlies difficulties in elucidating principles of spatial organization and using molecular profiling of the TME for clinical use3. Through statistical analysis of 96 spatial transcriptomic (ST–seq) datasets spanning twelve diverse tumor types, we found a conserved distribution of multicellular, transcriptionally covarying units termed ′Spatial Groups′ (SGs). SGs were either dependent on a hierarchical local spatial context – enriched for cell–extrinsic processes such as immune regulation and signal transduction – or independent from local spatial context – enriched for cell–intrinsic processes such as protein and RNA metabolism, DNA repair, and cell cycle regulation. We used SGs to define a measure of gene spatial heterogeneity – ′spatial lability′ – and categorized all 96 tumors by their TME spatial lability profiles. The resulting classification captured spatial variation in cell–extrinsic versus cell–intrinsic biology and motivated class–specific strategies for therapeutic intervention. Using this classification to characterize pre–treatment biopsy samples of 16 non–small cell lung cancer (NSCLC) patients outside our database distinguished responders and non–responders to immune checkpoint blockade while programmed death–ligand 1 (PD–L1) status and spatially unaware bulk transcriptional markers did not. Our findings show conserved principles of TME spatial biology that are both biologically and clinically significant. ### Competing Interest Statement The authors have declared no competing interest.
The human gut microbiome contains many bacterial strains of the same species ('strain-level variants'). Describing strains in a biologically meaningful way rather than purely taxonomically is an important goal but challenging due to the genetic complexity of strain-level variation. Here, we measured patterns of co-evolution across >7,000 strains spanning the bacterial tree-of-life. Using these patterns as a prior for studying hundreds of gut commensal strains that we isolated, sequenced, and metabolically profiled revealed widespread structure beneath the phylogenetic level of species. Defining strains by their co-evolutionary signatures enabled predicting their metabolic phenotypes and engineering consortia from strain genome content alone. Our findings demonstrate a biologically relevant organization to strain-level variation and motivate a new schema for describing bacterial strains based on their evolutionary history.
The widely expressed bromodomain and extraterminal motif (BET) proteins bromodomain-containing protein 2 (BRD2), BRD3, and BRD4 are multifunctional transcriptional regulators that bind acetylated chromatin via their conserved tandem bromodomains. Small molecules that target BET bromodomains are being tested for various diseases but typically do not discern between BET family members. Genomic distributions and protein partners of BET proteins have been described, but the basis for differences in BET protein function within a given lineage remains unclear. By establishing a gene knockout-rescue system in a Brd2-null erythroblast cell line, here we compared a series of mutant and chimeric BET proteins for their ability to modulate cell growth, differentiation, and gene expression. We found that the BET N-terminal halves bearing the bromodomains convey marked differences in protein stability but do not account for specificity in BET protein function. Instead, when BET proteins were expressed at comparable levels, their specificity was largely determined by the C-terminal half. Remarkably, a chimeric BET protein comprising the N-terminal half of the structurally similar short BRD4 isoform (BRD4S) and the C-terminal half of BRD2 functioned similarly to intact BRD2. We traced part of the BRD2-specific activity to a previously uncharacterized short segment predicted to harbor a coiled-coil (CC) domain. Deleting the CC segment impaired BRD2's ability to restore growth and differentiation, and the CC region functioned in conjunction with the adjacent ET domain to impart BRD2-like activity onto BRD4S. In summary, our results identify distinct BET protein domains that regulate protein turnover and biological activities.
Global changes in chromatin organization and the cessation of transcription during mitosis are thought to challenge the resumption of appropriate transcription patterns after mitosis. The acetyl-lysine binding protein BRD4 has been previously suggested to function as a transcriptional "bookmark" on mitotic chromatin. Here, genome-wide location analysis of BRD4 in erythroid cells, combined with data normalization and peak characterization approaches, reveals that BRD4 widely occupies mitotic chromatin. However, removal of BRD4 from mitotic chromatin does not impair post-mitotic activation of transcription. Additionally, histone mass spectrometry reveals global preservation of most posttranslational modifications (PTMs) during mitosis. In particular, H3K14ac, H3K27ac, H3K122ac, and H4K16ac widely mark mitotic chromatin, especially at lineage-specific genes, and predict BRD4 mitotic binding genome wide. Therefore, BRD4 is likely not a mitotic bookmark but only a "passenger." Instead, mitotic histone acetylation patterns may constitute the actual bookmarks that restore lineage-specific transcription patterns after mitosis.
Single-nucleotide variants that underlie phenotypic variation can affect chromatin occupancy of transcription factors (TFs). To delineate determinants of in vivo TF binding and chromatin accessibility, we introduce an approach that compares ChIP-seq and DNase-seq data sets from genetically divergent murine erythroid cell lines. The impact of discriminatory single-nucleotide variants on TF ChIP signal enables definition at single base resolution of in vivo binding characteristics of nuclear factors GATA1, TAL1, and CTCF. We further develop a facile complementary approach to more deeply test the requirements of critical nucleotide positions for TF binding by combining CRISPR-Cas9-mediated mutagenesis with ChIP and targeted deep sequencing. Finally, we extend our analytical pipeline to identify nearby contextual DNA elements that modulate chromatin binding by these three TFs, and to define sequences that impact kb-scale chromatin accessibility. Combined, our approaches reveal insights into the genetic basis of TF occupancy and their interplay with chromatin features.
Erythroid transcription factors (TFs) control gene expression programs, lineage decisions, and disease outcomes. How transcription factors contact DNA has been studied extensively in vitro, but in vivo binding characteristics are less well understood as they are influenced in a reciprocal manner by chromatin accessibility and neighboring transcription factors. Here, we present a comparative analysis approach that takes advantage of non-coding sequence variation between functionally equivalent erythroid cell lines to conduct an in-depth analysis of erythroid TF binding profiles and chromatin features.
The unlimited growth that occurs in tumors requires telomere maintenance. Yet, a portion of human tumors lack telomerase, and maintain telomeres using recombination-based mechanisms. Studies in other model organisms indicate that two different pathways of recombination-based mechanisms impact telomere maintenance and rely on the DNA repair proteins Rad50 and Rad51. In the Rad50-dependent pathway telomere recombination occurs within the telomere repeats. In contrast, recombination using the Rad51-dependent pathway occurs within repetitive sequences in the subtelomeres. Using a mouse B-cell lymphoma model lacking telomerase, Eμmyc+mTR-/-, and immortalized fibroblast cells lacking the RNA component of telomerase (mTR-/-) we have examined the impact of inhibiting Rad50 and Rad51a on telomere recombination. We find inhibiting Rad50 or Rad51a in Eμmyc+mTR-/- B-cell lymphomas, and in mTR-/- immortalized fibroblasts, has a synergistic effect on DNA damage sensitivity to mitomycin but not camptothecin. Inhibiting Rad50 in telomerase deficient cells also results in telomere shortening and in some tumors, reduced growth. In contrast, when Rad50 or Rad51a is inhibited in cells with telomerase, DNA damage sensitivity from mitomycin is not observed when compared to cells expressing a control shRNA. In addition inhibiting Rad50 in cells with telomerase does not significantly impact telomere length or recombination. Next we developed a comparative genomic hybridization (aCGH) approach that detects recombination events in the subtelomeres. Using these subtelomere arrays we find B-cell lymphomas lacking telomerase exhibit a significant increase in subtelomere recombination compared to primary cells. We also examined the impact of inhibiting Rad50 on subtelomere recombination events. Our findings using aCGH suggest that inhibiting Rad50 does not impact subtelomere recombination in Eμmyc+mTR-/- B-cell lymphomas. Overall, our findings suggest that inhibiting either Rad50 or Rad51a in mTR-/- cells has a synergistic impact on the sensitivity to DNA damaging agents in contrast to cells with mTR+/+. Currently we are testing the impact of inhibiting Rad51a on subtelomere recombination. In addition these results further support that Rad50 contributes to telomere recombination mechanisms in tumors lacking telomerase and will provide insight into the mechanism of subtelomere recombination in mammalian cells.
The unlimited growth of tumors requires telomere maintenance. Telomerase typically maintains telomeres. Yet some tumors lack telomerase and instead use recombination‐based mechanisms for telomere maintenance. Yeast deleted for telomerase use multiple mechanisms of break‐induced replication (BIR) that can occur within the telomeres or subtelomeres. To study BIR in mammalian cells, we are using a mouse model, myc+mTR−/− that generates B‐cell lymphomas lacking telomerase. Using this model we detected subtelomere recombination in tumors lacking telomerase. To examine and quantify subtelomere recombination, we are using comparative genomic hybridization (aCGH) to quantify changes in subtelomere copy numbers. Using these arrays we find frequent copy number changes within the subtelomeres of some tumors lacking telomerase, consistent with a mechanism of BIR. To examine the mechanism of subtelomere recombination we are testing whether Rad50 and Rad51a contribute to subtelomere recombination in tumors lacking telomerase. These studies address the role of different recombination genes during telomere maintenance and subtelomere recombination in tumors lacking telomerase. NIH/NCI K99/R00
Cells with unlimited growth potential such as tumors and stem cells require telomere maintenance mechanisms. Some human tumors lack telomerase and instead likely utilize recombination‐based mechanisms for telomere maintenance. Telomere maintenance in telomerase deficient yeast occurs by several distinct recombination mechanisms that allow break‐induced replication (BIR). In previous studies using (E(mu)myc+mTR−/−) mice with short telomeres, we observed frequent telomere recombination in primary bone marrow cells and tumors. These recombination events sometimes included subtelomeric regions as previously seen in yeast. To determine the role of recombination in telomere maintenance, we are measuring the frequency of subtelomere recombination in mTR+/+ and mTR−/− primary fibroblasts and B‐cell tumors lacking telomerase. We are examining genes that contribute to tumor growth, telomere length, and telomere recombination by using shRNAs targeting specific recombination genes. We found that shRNAs directed against Rad50 reduced tumor growth and the frequency of telomere recombination, implicating a role for Rad50 in tumor growth in the absence of telomerase.
Due to the energetic frustration of RNA folding, tertiary structured RNA is typically characterized by a rugged folding free energy landscape where deep kinetic barriers separate numerous misfolded states from one or more native states. While most in vitro studies of RNA rely on (re)folding chemically and/or enzymatically synthesized RNA in its entirety, which frequently leads into kinetic traps, nature reduces the complexity of the RNA folding problem by segmental, co-transcriptional folding starting from the 5′ end. We here have developed a simplified, general, nondenaturing purification protocol for RNA to ask whether avoiding denaturation of a co-transcriptionally folded RNA can reduce commonly observed in vitro folding heterogeneity. Our protocol bypasses the need for large-scale auxiliary protein purification and expensive chromatographic equipment and involves rapid affinity capture with magnetic beads and removal of chemical heterogeneity by cleavage of the target RNA from the beads using the ligand-induced glmS ribozyme. For two disparate model systems, the Varkud satellite (VS) and hepatitis delta virus (HDV) ribozymes, we achieve >95% conformational purity within one hour of enzymatic transcription, without the need for any folding chaperones. We further demonstrate that in vitro refolding introduces severe conformational heterogeneity into the natively-purified VS ribozyme but not into the compact, double-nested pseudoknot fold of the HDV ribozyme. We conclude that conformational heterogeneity in complex RNAs can be avoided by co-transcriptional folding followed by nondenaturing purification, providing rapid access to chemically and conformationally pure RNA for biologically relevant biochemical and biophysical studies.