Cas9 is a programmable nuclease that has furnished transformative technologies, including base editors and transcription modulators (e.g., CRISPRi/a), but several applications of these technologies, including therapeutics, mandatorily require precision control of their half-life. For example, such control can help avert any potential immunological and adverse events in clinical trials. Current genome editing technologies to control the half-life of Cas9 are slow, have lower activity, involve fusion of large response elements (> 230 amino acids), utilize expensive controllers with poor pharmacological attributes, and cannot be implemented in vivo on several CRISPR-based technologies. We report a general platform for half-life control using the molecular glue, pomalidomide, that binds to a ubiquitin ligase complex and a response-element bearing CRISPR-based technology, thereby causing the latter's rapid ubiquitination and degradation. Using pomalidomide, we were able to control the half-life of large CRISPR-based technologies (e.g., base editors, CRISPRi) and small anti-CRISPRs that inhibit such technologies, allowing us to build the first examples of on-switch for base editors. The ability to switch on, fine-tune and switch-off CRISPR-based technologies with pomalidomide allowed complete control over their activity, specificity, and genome editing outcome. Importantly, the miniature size of the response element and favorable pharmacological attributes of the drug pomalidomide allowed control of activity of base editor in vivo using AAV as the delivery vehicle. These studies provide methods and reagents to precisely control the dosage and half-life of CRISPR-based technologies, propelling their therapeutic development.
With the growing number of single-cell analysis tools, benchmarks are increasingly important to guide analysis and method development. However, a lack of standardisation and extensibility in current benchmarks limits their usability, longevity, and relevance to the community. We present Open Problems, a living, extensible, community-guided benchmarking platform including 10 current single-cell tasks that we envision will raise standards for the selection, evaluation, and development of methods in single-cell analysis.
Transcriptional coregulators and transcription factors (TFs) contain intrinsically disordered regions (IDRs) that are critical for their association and function in gene regulation. More recently, IDRs have been shown to promote multivalent protein-protein interactions between coregulators and TFs to drive their association into condensates. By contrast, here we demonstrate how the IDR of the corepressor LSD1 excludes TF association, acting as a dynamic conformational switch that tunes repression of active cis-regulatory elements. Hydrogen-deuterium exchange shows that the LSD1 IDR interconverts between transient open and closed conformational states, the latter of which inhibits partitioning of the protein’s structured domains with TF condensates. This autoinhibitory switch controls leukemic differentiation by modulating repression of active cis-regulatory elements bound by LSD1 and master hematopoietic TFs. Together, these studies unveil alternative mechanisms by which disordered regions and their dynamic crosstalk with structured regions can shape coregulator-TF interactions to control cis-regulatory landscapes and cell fate.
Hematopoietic malignancies arise from driver alterations acquired during specific, often intermediate and transient, differentiation states. When adequately sampled, “snapshot” single-cell measurements can capture dynamic biological processes, including transient states. While there are effective computational modeling solutions to order and capture the relationships among these cell states (often projecting onto a lower dimensional manifold with directionality), methods to infer the causal gene regulatory mechanisms driving cell state transitions remain limited. Recently, a mathematical framework for modeling cell dynamics from snapshot data using a physics-based drift-diffusion equation was proposed. This drift-diffusion equation framework describes the dynamics of cell distributions with respect to time, wherein stereotypical driving forces (e.g., hematopoietic development) are captured by the drift term, and biological stochasticity is attributed to the diffusion term. The existing modeling solutions within this framework are forced to assume a fixed diffusion parameter across cell states, ascribing all observed dynamics to drift alone. Here, we expand on current drift-diffusion models to learn both the drift and diffusion terms that describe every observed cell state. These models are constructed from stochastic neural differential equations, continuous deep neural networks that approximate the theoretically complex differential equations underlying gene regulatory networks in hematopoiesis. Lineage-barcoded, multi-time point single-cell RNA sequencing experiments can approximate ground truth cellular dynamics. Using such data, we demonstrate the predictive accuracy of our proposed approach through benchmarking against the state-of-the-art methods. Through our drift-diffusion model, we can now distinguish, for the first time, between the deterministic and stochastic contributions to cell fate decision making. In doing so, we identify genes whose expression is associated with cellular drift and indeed show that they correspond to well-known markers of specific cell types; moreover, we identify genes that are associated with cell states that exhibit notable diffusion. Finally, we demonstrate the potential of such models to capture the essence of the dynamic processes through prediction on out-of-distribution data (referred to as “transfer learning” in the machine learning field). For example, when the model is trained on in vitro hematopoiesis data, our model is able to accurately predict cell trajectories in the corresponding in vivo mouse model. We anticipate that such insights will be useful in identifying key dynamic and time-dependent regulatory processes in normal development and cancer. Citation Format: Michael E. Vinyard, Anders W. Rasmussen, Ruitong Li, Luca Pinello, Gad Getz. Modeling single-cell dynamics using stochastic generative models based on neural differential equations. [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2023; Part 1 (Regular and Invited Abstracts); 2023 Apr 14-19; Orlando, FL. Philadelphia (PA): AACR; Cancer Res 2023;83(7_Suppl):Abstract nr 5371.
Single-cell sequencing measurements facilitate the reconstruction of dynamic biology by capturing snapshots of the molecular profiles of individual cells. Cell fate decisions in development and disease are orchestrated through an intricate balance of deterministic and stochastic regulatory events. Drift-diffusion equations are effective in modeling single-cell dynamics from high-dimensional single-cell measurements. While existing solutions describe the deterministic dynamics associated with the drift term of these equations at the level of cell state, the diffusion is modeled as a constant across cell states. To fully understand the dynamic regulatory logic in development and disease, models explicitly attuned to the balance between deterministic and stochastic biology are required. Addressing these limitations, we introduce scDiffEq, a generative framework for learning neural stochastic differential equations that approximate the deterministic and stochastic dynamics in biology. Using lineage-traced single-cell data, we demonstrate that scDiffEq offers improved reconstruction of held-out cell states and prediction of cell fate from multipotent progenitors during hematopoiesis. By imparting in silico perturbations to multipotent progenitor cells, we find that scDiffEq accurately recapitulates the dynamics of CRISPR-perturbed hematopoiesis. Using scDiffEq, we simulate high-resolution developmental cell trajectories, modeling their drift and diffusion, enabling us to study their time-dependent gene-level dynamics.### Competing Interest StatementThe authors have declared no competing interest.
Lysine-specific histone demethylase 1 (LSD1/KDM1A) governs hematopoiesis by regulating hematopoietic stem cell renewal and proper differentiation. Additionally, LSD1 is required for the maintenance of acute myeloid leukemia (AML) cells and has emerged as a key therapeutic target. Prior work from our group and others have revealed that the LSD1 demethylase activity is not required for AML proliferation. Instead, the interaction between LSD1 and GFI1/GFI1B is necessary to sustain AML cell survival. Thus, beyond its canonical demethylase activity, LSD1 has nonenzymatic, transcription factor (TF)-specific scaffolding functions critical for its function. Most LSD1 inhibitors function by blocking protein-protein interactions between LSD1 and TFs, even though they were originally designed as enzyme inhibitors. While studies indicate that LSD1 has important scaffolding functions, how LSD1 can coordinate with TFs to control gene expression programs is not well understood. Here, we used CRISPR-suppressor scanning to identify regions of LSD1 outside the active site that when mutated can lead to resistance to the covalent LSD1 inhibitor, GSK-LSD1. The understudied N-terminal intrinsically disordered region (IDR) of LSD1 was among the most enriched regions. To explore this further, we generated clonal AML cell lines harboring in-frame deletions within the IDR. These IDR-mutant cell lines were drug resistant and failed to differentiate upon GSK-LSD1 treatment. Moreover, knockdown of endogenous LSD1 followed by overexpression of LSD1 IDR-mutants rescued growth in the presence of GSK-LSD1, confirming that a fully functional IDR is necessary for differentiation. We performed LSD1 co-immunoprecipitation followed by mass spectrometry in wild-type and IDR-mutant cells to measure changes in LSD1 interactors upon GSK-LSD1 treatment. GSK-LSD1 treatment disrupted the LSD1-GFI1B interaction in both wild-type and mutant cells at comparable levels. Since GSK-LSD1 still induces dissociation of the mutant complex but has minimal effect on proliferation, our results suggest that the IDR-mutant cells no longer require GFI1B. In support, knockout of GFI1B inhibited growth of wild-type but not IDR-mutant cells. Strikingly, GSK-LSD1 treatment promoted LSD1's interaction with several master regulator TFs in hematopoiesis, including C/EBPα, PU.1, and RUNX1. Notably, C/EBPα was the most up-regulated interactor in both cell lines, with modest enrichment in IDR-mutant over wild-type. Using live-cell imaging assays, we found that deletions within the LSD1 IDR enhance association with these TFs. These results suggest that GSK-LSD1 reprograms LSD1-TF interactions, disrupting LSD1-GFI1B to induce LSD1-C/EBPα association, and that the IDR mutations alters this LSD1 redistribution to other TF partners. LSD1 is known to silence gene enhancers during differentiation in a process termed “enhancer decommissioning”. The mechanisms by which LSD1 occupies and controls enhancers remain elusive. To test if the LSD1 IDR regulates enhancer elements, we profiled LSD1 and H3K27ac using ChIP-seq in wild-type and IDR-mutant cells upon GSK-LSD1 treatment. Upon GSK-LSD1 treatment, disruption of LSD1-GFI1B binding was accompanied by increased LSD1 levels at other TF and H3K27ac sites, supporting the notion that loss of LSD1-GFI1B interaction induces LSD1 redistribution to other TF partner sites. A concomitant increase in H3K27ac levels was observed across LSD1 sites co-bound by C/EBPα, PU.1, or RUNX1 in wild-type cells. By contrast, in IDR-mutant cells, increases in H3K27ac levels were blocked at these sites upon GSK-LSD1 treatment, suggesting that full enhancer activation is impaired by mutations in the IDR, potentially due to altered association with master regulator TFs. This points to a function of LSD1 IDR in controlling enhancer activation for differentiation. Collectively, our study reveals a role for IDR in regulating LSD1-TF interactions, controlling enhancer activation that is necessary for AML cell differentiation. These data provide deep functional and mechanistic insights into the role of LSD1 IDR and offer important considerations for pharmacological targeting of protein-protein interactions.
fname = "adata.Weinreb2020.in_vitro.task_01.timepoint_recovery.h5ad"
Most current single-cell analysis pipelines are limited to cell embeddings and rely heavily on clustering, while lacking the ability to explicitly model interactions between different feature types. Furthermore, these methods are tailored to specific tasks, as distinct single-cell problems are formulated differently. To address these shortcomings, here we present SIMBA, a graph embedding method that jointly embeds single cells and their defining features, such as genes, chromatin-accessible regions and DNA sequences, into a common latent space. By leveraging the co-embedding of cells and features, SIMBA allows for the study of cellular heterogeneity, clustering-free marker discovery, gene regulation inference, batch effect removal and omics data integration. We show that SIMBA provides a single framework that allows diverse single-cell problems to be formulated in a unified way and thus simplifies the development of new analyses and extension to new single-cell modalities. SIMBA is implemented as a comprehensive Python library ( https://simba-bio.readthedocs.io ).
Although vast numbers of putative gene regulatory elements have been cataloged, the sequence motifs and individual bases that underlie their functions remain largely unknown. Here, we combine epigenetic perturbations, base editing, and deep learning to dissect regulatory sequences within the exemplar immune locus encoding CD69. We converge on a ∼170 base interval within a differentially accessible and acetylated enhancer critical for CD69 induction in stimulated Jurkat T cells. Individual C-to-T base edits within the interval markedly reduce element accessibility and acetylation, with corresponding reduction of CD69 expression. The most potent base edits may be explained by their effect on regulatory interactions between the transcriptional activators GATA3 and TAL1 and the repressor BHLHE40. Systematic analysis suggests that the interplay between GATA3 and BHLHE40 plays a general role in rapid T cell transcriptional responses. Our study provides a framework for parsing regulatory elements in their endogenous chromatin contexts and identifying operative artificial variants.
Background: Multiple Myeloma (MM) is a hematological malignancy characterized by abnormal proliferation of terminally differentiated plasma cells (PCs) in the bone marrow (BM). MM is almost always preceded by the precursor stage smoldering multiple myeloma (SMM). BM biopsies are useful to monitor disease progression, but they are invasive and not routinely collected from patients for disease monitoring during precursor stages. Profiling circulating tumor cells (CTCs) from peripheral blood (PB) could aid early detection, disease monitoring, and biomarker identification to predict patients at high risk of progression that may benefit from early therapeutic intervention. Methods: Paired PB and BM aspirates were collected from 40 SMM patients enrolled in the PCROWD study (IRB #14-174) at Dana-Farber Cancer Institute. Malignant PCs were enriched by magnetic bead-based methods and underwent 5’ single-cell RNA sequencing (scRNA-seq) and single-cell B-cell receptor sequencing (scBCR-seq) (10x Genomics). Results: We analyzed 105,246 BM PCs and 33,234 PB PCs from 15 patients. To differentiate malignant from normal PCs, we used clonal V(D)J rearrangements, assessed by concurrent scBCR-seq. A total of 86,986 BM tumor cells and 8,718 CTCs were captured. A median of 5, 26, and 47 CTCs were present per mL of blood from low, intermediate, and high-risk SMM patients as defined by the International Myeloma Working Group (IMWG) “20/2/20” criteria, suggesting sequencing-based CTC enumeration corresponds to prognosis. High levels of driver genes commonly upregulated in patients with specific translocations, including CCND1 and MAF, were detected in both BM tumor and CTC clusters in 3 patients with t(11;14) and t(14;16) confirmed by fluorescence in situ hybridization (FISH) clinical testing, and 2 additional patients with inconclusive FISH results (Wilcoxon, q <10-3), supporting the idea of CTC-based prognostication. Differential expression (DE) analysis revealed 8 genes that were significantly upregulated and 3 genes that were significantly downregulated in CTCs compared to BM tumor cells robustly across 15 paired samples. Gene set enrichment analysis (GSEA) revealed genes DE in CTCs are associated with TNF-α and NF-κB signaling, which are commonly induced by extrinsic factors in the bone marrow milieu, providing insight into the biology of tumor cell circulation. Conclusions: This study highlights the utility of scRNA-seq for molecular profiling of CTCs, even in asymptomatic low tumor burden disease. Additional analyses are ongoing in the expanded cohort of 40 patients with paired samples to help gain further insight into CTC heterogeneity. Overall, this study will help enable the design of new molecular liquid biopsy-based approaches to diagnosis, disease monitoring, and biological insights to improve treatment strategies for precursor myeloma patients. Citation Format: Elizabeth D. Lightbody, Danielle T. Firer, Romanos Sklavenitis-Pistofidis, Michael Agius, Ankit K. Dutta, Michelle Aranha, Jean-Baptiste Alberge, Laura Hevenor, Nang Kham Su, Cody Boehner, Erica Horowitz, Jacqueline Perry, Anna Cowan, Hadley Barr, Anna Justis, Daniel Auclair, Catherine R. Marinac, Gad Getz, Irene Ghobrial. Single-cell RNA sequencing of rare circulating tumor cells in precursor myeloma patients reveals molecular underpinnings of tumor cell circulation [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2022; 2022 Apr 8-13. Philadelphia (PA): AACR; Cancer Res 2022;82(12_Suppl):Abstract nr 641.
Single-cell assays have transformed our ability to model heterogeneity within cell populations. As these assays have advanced in their ability to measure various aspects of molecular processes in cells, computational methods to analyze and meaningfully visualize such data have required matched innovation. Independently, Virtual Reality (VR) has recently emerged as a powerful technology to dynamically explore complex data and shows promise for adaptation to challenges in single-cell data visualization. However, adopting VR for single-cell data visualization has thus far been hindered by expensive prerequisite hardware or advanced data preprocessing skills. To address current shortcomings, we present singlecellVR, a user-friendly web application for visualizing single-cell data, designed for cheap and easily available virtual reality hardware (e.g., Google Cardboard, ∼$8). singlecellVR can visualize data from a variety of sequencing-based technologies including transcriptomic, epigenomic, and proteomic data as well as combinations thereof. Analysis modalities supported include approaches to clustering as well as trajectory inference and visualization of dynamical changes discovered through modelling RNA velocity. We provide a companion software package, scvr to streamline data conversion from the most widely-adopted single-cell analysis tools as well as a growing database of pre-analyzed datasets to which users can contribute.
Rapid technological advances in transcriptomics and lineage tracing technologies provide new opportunities to understand organismal development at the single-cell level. Building on these advances, various computational methods have been proposed to infer developmental trajectories and to predict cell fate. These methods have unveiled previously uncharacterized transitional cell types and differentiation processes. Importantly, the ability to recover cell states and trajectories has been evolving hand-in-hand with new technologies and diverse experimental designs; more recent methods can capture complex trajectory topologies and infer short- and long-term cell fate dynamics. Here, we summarize and categorize the most recent and popular computational approaches for trajectory inference based on the information they leverage and describe future challenges and opportunities for the development of new methods for reconstructing differentiation trajectories and inferring cell fates.
Recent advances in single-cell omics technologies enable the individual and joint profiling of cellular measurements. Currently, most single-cell analysis pipelines are cluster-centric and cannot explicitly model the interactions between different feature types. In addition, single-cell methods are generally designed for a particular task as distinct single-cell problems are formulated differently. To address these current shortcomings, we present SIMBA , a graph embedding method that jointly embeds single cells and their defining features, such as genes, chromatin accessible regions, and transcription factor binding sequences into a common latent space. By leveraging the co-embedding of cells and features, SIMBA allows for the study of cellular heterogeneity, clustering-free marker discovery, gene regulation inference, batch effect removal, and omics data integration. SIMBA has been extensively applied to scRNA-seq, scATAC-seq, and dual-omics data. We show that SIMBA provides a single framework that allows diverse single-cell analysis problems to be formulated in a unified way and thus simplifies the development of new analyses and integration of other single-cell modalities. SIMBA is implemented as an efficient, comprehensive, and extensible Python library ( https://simba-bio.readthedocs.io ) for the analysis of single-cell omics data using graph embedding.
Background Recent innovations in single-cell Assay for Transposase Accessible Chromatin using sequencing (scATAC-seq) enable profiling of the epigenetic landscape of thousands of individual cells. scATAC-seq data analysis presents unique methodological challenges. scATAC-seq experiments sample DNA, which, due to low copy numbers (diploid in humans) lead to inherent data sparsity (1-10% of peaks detected per cell) compared to transcriptomic (scRNA-seq) data (20-50% of expressed genes detected per cell). Such challenges in data generation emphasize the need for informative features to assess cell heterogeneity at the chromatin level. Results We present a benchmarking framework that was applied to 10 computational methods for scATAC-seq on 13 synthetic and real datasets from different assays, profiling cell types from diverse tissues and organisms. Methods for processing and featurizing scATAC-seq data were evaluated by their ability to discriminate cell types when combined with common unsupervised clustering approaches. We rank evaluated methods and discuss computational challenges associated with scATAC-seq analysis including inherently sparse data, determination of features, peak calling, the effects of sequencing coverage and noise, and clustering performance. Running times and memory requirements are also discussed. Conclusions This reference summary of scATAC-seq methods offers recommendations for best practices with consideration for both the non-expert user and the methods developer. Despite variation across methods and datasets, SnapATAC, Cusanovich2018 , and cisTopic outperform other methods in separating cell populations of different coverages and noise levels in both synthetic and real datasets. Notably, SnapATAC was the only method able to analyze a large dataset (> 80,000 cells).
Understanding the mechanism of small molecules is a critical challenge in chemical biology and drug discovery. Medicinal chemistry is essential for elucidating drug mechanism, enabling variation of small molecule structure to gain structure-activity relationships (SARs). However, the development of complementary approaches that systematically vary target protein structure could provide equally informative SARs for investigating drug mechanism and protein function. Here we explore the ability of CRISPR-Cas9 mutagenesis to profile the interactions between lysine-specific histone demethylase 1 (LSD1) and chemical inhibitors in the context of acute myeloid leukemia (AML). Through this approach, termed CRISPR-suppressor scanning, we elucidate drug mechanism of action by showing that LSD1 enzyme activity is not required for AML survival and that LSD1 inhibitors instead function by disrupting interactions between LSD1 and the transcription factor GFI1B on chromatin. Our studies clarify how LSD1 inhibitors mechanistically operate in AML and demonstrate how CRISPR-suppressor scanning can uncover novel aspects of target biology.