Hematopoiesis is a highly dynamic and hierarchical process that gives rise to all blood cell types from multipotent hematopoietic stem cells. This system must simultaneously maintain self-renewal, allow for multilineage output, and respond rapidly to environmental changes. Over the last decade, single-cell transcriptomic studies have provided a high-resolution view of differentiation trajectories, showing that lineage commitment occurs gradually and not as discrete steps. However, by focusing solely on RNA expression, these studies have blurred the molecular definition of classical cell populations, traditionally identified by functional assays and surface markers. Moreover, scRNA-seq cannot capture epigenetic priming events that often precede transcriptional change. Lineage tracing studies have shown that cell fate decisions are made earlier than transcriptomic data alone can predict, suggesting that integrating chromatin accessibility is essential for a comprehensive understanding of differentiation. To overcome these limitations, we applied a trimodal single-cell assay, TEA-seq, which simultaneously profiles gene expression, chromatin accessibility, and surface protein levels in the same individual cells. We generated a high-resolution map of human hematopoiesis using TEA-seq on hematopoietic stem and progenitor cells isolated from umbilical cord blood. Cells were collected at four time points across a fourteen-day erythroid differentiation culture system. Each sample was stained with 158 barcoded antibodies targeting surface proteins, followed by joint measurement of the transcriptome and epigenome. Data from all three modalities were integrated using MultiVI, a probabilistic model that learns a shared latent space across cells, accounting for batch effects. The resulting map reconstructs the erythroid trajectory and links progressive fate acquisition with specific chromatin and protein signatures. Early in differentiation, transcriptional changes are primarily driven by fluctuations in transcription factor abundance, with minimal changes in chromatin accessibility. In intermediate progenitors, chromatin remodeling becomes the dominant regulatory mechanism. At later stages, transcriptional dynamics once again take precedence, acting on a largely pre-established chromatin landscape. These observations suggest a model in which chromatin accessibility serves as a scaffold for future transcriptional activation, decoupling stemness loss from immediate lineage commitment. To further investigate how these molecular changes are orchestrated, we reconstructed enhancer-based gene regulatory networks using SCENIC+, which integrates TF expression, chromatin accessibility, motif enrichment, and enhancer-gene associations. We identified dynamic rewiring of gene regulatory networks (GRNs) along the differentiation trajectory. In early hematopoietic stem and progenitor cells, GRN transitions were driven by transcriptional changes. In intermediate populations, chromatin remodeling was the major driver. In late-stage progenitors, transcription factor abundance once again played a critical role. These distinct regulatory phases indicate that cell identity is shaped by the dynamic rewiring of the GRN, involving a complex interplay between transcription factor abundance and chromatin accessibility. Notably, stemness-related GRNs remain active well into the erythroid trajectory, consistent with retained plasticity in late erythroid progenitors. These findings demonstrate that loss of stemness and acquisition of lineage identity are not synchronous events but are governed by temporally and mechanistically distinct programs. Chromatin remodeling at cell-specific enhancers occurs gradually and establishes a permissive state, while changes in transcription factor availability trigger abrupt transcriptional shifts. This sequential architecture likely serves to buffer against stochastic activation and ensures robust lineage commitment. Together, our work establishes the first dynamic, enhancer-based gene regulatory networks built from simultaneous single-cell measurements of chromatin accessibility, gene expression, and protein abundance in human hematopoiesis. This trimodal approach unifies classical immunophenotyping with transcriptomic and epigenetic profiles. Our interactive web portal enables further exploration of the GRN and provides a resource for understanding cell fate decisions in human hematopoiesis.
Hematopoiesis requires the coordinated loss of stemness and acquisition of lineage identity, yet the regulatory logic and principles underlying these transitions has remained elusive. Single-cell studies suggest that hematopoietic stem and progenitor cells progress along continuous trajectories, but this view conflicts with the existence of discrete, functionally validated populations. Here, we establish the first dynamic, enhancer-based gene regulatory networks (eGRNs) that resolve the molecular programs underlying early human hematopoietic fate decisions. Built on a high-resolution trimodal framework generated with TEA-seq, these networks integrate simultaneously measured transcription factor abundance, enhancer and promoter accessibility, and gene expression within single-cell trajectories anchored to immunophenotypically defined populations. Our framework reveals that stemness loss and lineage acquisition are temporally and mechanistically uncoupled: stemness programs decline gradually through reduced TF abundance long before chromatin closure, whereas lineage identity emerges stepwise through enhancer reconfiguration and activation of lineage-defining eGRNs. This process generates discrete regulatory states that align with immunophenotypically defined populations. Together, these findings reconcile continuous and discrete models of hematopoiesis and establish eGRNs as a powerful framework for defining cell types by their regulatory logic. In addition, we provide an interactive web-based resource to facilitate further investigation of eGRNs and trajectories during early human hematopoiesis.
Dynamic conformational and structural changes in proteins and protein complexes play a central and ubiquitous role in the regulation of protein function, yet it is very challenging to study these changes, especially for large protein complexes, under physiological conditions. Here, we introduce a novel isobaric crosslinker, Qlinker, for studying conformational and structural changes in proteins and protein complexes using quantitative crosslinking mass spectrometry. Qlinkers are small and simple, amine-reactive molecules with an optimal extended distance of ~10 Å, which use MS2 reporter ions for relative quantification of Qlinker-modified peptides derived from different samples. We synthesized the 2-plex Q2linker and showed that the Q2linker can provide quantitative crosslinking data that pinpoints key conformational and structural changes in biosensors, binary and ternary complexes composed of the general transcription factors TBP, TFIIA, and TFIIB, and RNA polymerase II complexes.
The human nucleosome acetyltransferase of histone H4 (NuA4)/Tat-interactive protein, 60 kilodalton (TIP60) coactivator complex, a fusion of the yeast switch/sucrose nonfermentable related 1 (SWR1) and NuA4 complexes, both incorporates the histone variant H2A.Z into nucleosomes and acetylates histones H4, H2A, and H2A.Z to regulate gene expression and maintain genome stability. Our cryo–electron microscopy studies show that, within the NuA4/TIP60 complex, the E1A binding protein P400 (EP400) subunit serves as a scaffold holding the different functional modules in specific positions, creating a distinct arrangement of the actin-related protein (ARP) module. EP400 interacts with the transformation/transcription domain-associated protein (TRRAP) subunit by using a footprint that overlaps with that of the Spt-Ada-Gcn5 acetyltransferase (SAGA) complex, preventing the formation of a hybrid complex. Loss of the TRRAP subunit leads to mislocalization of NuA4/TIP60, resulting in the redistribution of H2A.Z and its acetylation across the genome, emphasizing the dual functionality of NuA4/TIP60 as a single macromolecular assembly.
The Mixed Lineage Leukemia 1 (MLL) gene, a histone-3-lysine-4-methyltransferase, is a critical transcription factor in cell fate determination during hematopoiesis and is frequently mutated in patients with myelodysplastic syndrome (MDS) and acute myeloid leukemia (AML). While the mechanism of MLL chimeric fusion proteins has been extensively studied, the mechanism through which another oncogenic mutant, the Partial Tandem Duplication (PTD) of the MLL N-terminal DNA binding domain promotes leukemogenesis, remains unknown. Most commonly, the PTD is formed through an in-frame duplication of exons 2 through 6, although other rearrangements have also been identified. However, MLLPTD was recently recognized by the Molecular International Prognostic Scoring System Model as one of the strongest predictors of adverse outcomes in MDS, pressing the need to unravel its role in the leukemogenic process.Our approach to understanding how MLLPTD drives leukemogenesis involved different techniques. We did targeted QconCAT mass spectrometry, which utilizes a known amount of peptides incorporated in a hybrid protein named QconCAT. Once it is included in the nuclear extract, we detected copy numbers of the endogenous protein. At the molecular level, we demonstrated that MLL-PTD cannot interact with the WRAD complex, a complex of proteins required for MLL enzymatic activity, suggesting a defect of MLL function. Instead, MLL-PTD associates with distinct chromatin-modifying enzymes. In addition, to further study MLLPTD, we developed and validated an antibody that recognizes the exon 6-2 peptide. Using this antibody, we performed CUT&Tag revealing for the first time the genome-wide binding of MLLPTD.Taken together our results provide critical insights into the mechanism through which MLL-PTD leads to the deregulation of gene expression contributing to the progression of MDS to AML.
The SWI/SNF ATP-dependent chromatin remodeler is a master regulator of the epigenome, controlling pluripotency and differentiation. Towards the C-terminus of the catalytic subunit of SWI/SNF is a motif called the AT-hook that is evolutionary conserved. The AT-hook is present in many chromatin modifiers and generally thought to help anchor them to DNA. We observe however that the AT-hook regulates the intrinsic DNA-stimulated ATPase activity aside from promoting SWI/SNF recruitment to DNA or nucleosomes by increasing the reaction velocity a factor of 13 with no accompanying change in substrate affinity (K M ). The changes in ATP hydrolysis causes an equivalent change in nucleosome movement, confirming they are tightly coupled. The catalytic subunit’s AT-hook is required in vivo for SWI/SNF remodeling activity in yeast and mouse embryonic stem cells. The AT-hook in SWI/SNF is required for transcription regulation and activation of stage-specific enhancers critical in cell lineage priming. Similarly, growth assays suggest the AT-hook is required in yeast SWI/SNF for activation of genes involved in amino acid biosynthesis and metabolizing ethanol. Our findings highlight the importance of studying SWI/SNF attenuation versus eliminating the catalytic subunit or completely shutting down its enzymatic activity.
TFIIH is an evolutionarily conserved complex that plays central roles in both RNA polymerase II (pol II) transcription and DNA repair. As an integral component of the pol II preinitiation complex, TFIIH regulates pol II enzyme activity in numerous ways. The TFIIH subunit XPB/Ssl2 is an ATP-dependent DNA translocase that stimulates promoter opening prior to tran-scription initiation. Crosslinking-mass spectrometry and cryo-EM results have shown a conserved interaction network involving XPB/Ssl2 and the C-terminal Hub region of the TFIIH p52/Tfb2 subunit, but the functional significance of specific residues is unclear. Here, we systematically mutagenized the HubA region of Tfb2 and screened for growth phenotypes in a TFB6 deletion background in Saccharomyces cerevisiae. We identified six lethal and 12 conditional mutants. Slow growth phenotypes of all but three conditional mutants were relieved in the presence of TFB6, thus identifying a functional interaction between Tfb2 HubA mutants and Tfb6, a protein that dissociates Ssl2 from TFIIH. Our biochemical analysis of Tfb2 mutants with severe growth phenotypes revealed defects in Ssl2 association, with similar results in human cells. Further characterization of these tfb2 mutant cells revealed defects in GAL gene induction, and reduced occupancy of TFIIH and pol II at GAL gene pro-moters, suggesting that functionally competent TFIIH is required for proper pol II recruitment to preinitiation complexes in vivo. Consistent with recent structural models of TFIIH, our results identify key residues in the p52/Tfb2 HubA domain that are required for stable incorporation of XPB/Ssl2 into TFIIH and for pol II transcription.
Pyruvate dehydrogenase (PDH) is the gatekeeper enzyme of the tricarboxylic acid (TCA) cycle. Here we show that the deglycase DJ-1 (encoded by PARK7, a key familial Parkinson's disease gene) is a pacemaker regulating PDH activity in CD4+ regulatory T cells (Treg cells). DJ-1 binds to PDHE1-β (PDHB), inhibiting phosphorylation of PDHE1-α (PDHA), thus promoting PDH activity and oxidative phosphorylation (OXPHOS). Park7 (Dj-1) deletion impairs Treg survival starting in young mice and reduces Treg homeostatic proliferation and cellularity only in aged mice. This leads to increased severity in aged mice during the remission of experimental autoimmune encephalomyelitis (EAE). Dj-1 deletion also compromises differentiation of inducible Treg cells especially in aged mice, and the impairment occurs via regulation of PDHB. These findings provide unforeseen insight into the complicated regulatory machinery of the PDH complex. As Treg homeostasis is dysregulated in many complex diseases, the DJ-1-PDHB axis represents a potential target to maintain or re-establish Treg homeostasis.
The human general transcription factor TFIID is composed of the TATA-binding protein (TBP) and 13 TBP-associated factors (TAFs). In eukaryotic cells, TFIID is thought to nucleate RNA polymerase II (Pol II) preinitiation complex formation on all protein coding gene promoters and thus, be crucial for Pol II transcription. TFIID is composed of three lobes, named A, B and C. Structural studies showed that TAF8 forms a histone fold pair with TAF10 in lobe B and participates in connecting lobe B to lobe C. In the present study, we have investigated the requirement of the different regions of TAF8 for in vitro TFIID assembly, and the importance of certain TAF8 regions for mouse embryonic stem cell (ESC) viability. We have identified a TAF8 region, different from the histone fold domain of TAF8, important for assembling with the 5TAF core complex in lobe B, and four regions of TAF8 each individually required for interacting with TAF2 in lobe C. Moreover, we show that the 5TAF coreinteracting TAF8 domain, and the proline rich domain of TAF8 that interacts with TAF2, are both required for mouse embryonic stem cell survival. Thus, our study demonstrates that distinct TAF8 regions involved in connecting lobe B to lobe C are crucial for TFIID function and consequent ESC survival.
Leishmania parasites cause a variety of serious human diseases, with no effective vaccine and emerging resistance to current drug therapy. We have previously shown that a novel DNA base called J is critical for transcription termination at the ends of the polycistronic gene clusters that are a hallmark of Leishmania and related trypanosomatids.
PURPOSE OF REVIEW:Erythropoiesis is a hierarchical process by which hematopoietic stem cells give rise to red blood cells through gradual cell fate restriction and maturation. Deciphering this process requires the establishment of dynamic gene regulatory networks (GRNs) that predict the response of hematopoietic cells to signals from the environment. Although GRNs have historically been derived from transcriptomic data, recent proteomic studies have revealed a major role for posttranscriptional mechanisms in regulating gene expression during erythropoiesis. These new findings highlight the need to integrate proteomic data into GRNs for a refined understanding of erythropoiesis.RECENT FINDINGS:Here, we review recent proteomic studies that have furthered our understanding of erythropoiesis with a focus on quantitative mass spectrometry approaches to measure the abundance of transcription factors and cofactors during differentiation. Furthermore, we highlight challenges that remain in integrating transcriptomic, proteomic, and other omics data into a predictive model of erythropoiesis, and discuss the future prospect of single-cell proteomics.SUMMARY:Recent proteomic studies have considerably expanded our knowledge of erythropoiesis beyond the traditional transcriptomic-centric perspective. These findings have both opened up new avenues of research to increase our understanding of erythroid differentiation, while at the same time presenting new challenges in integrating multiple layers of information into a comprehensive gene regulatory model.
Quantitative changes in transcription factor (TF) abundance regulate dynamic cellular processes, including cell fate decisions. Protein copy number provides information about the relative stoichiometry of TFs that can be used to determine how quantitative changes in TF abundance influence gene regulatory networks. In this protocol, we describe a targeted selected reaction monitoring (SRM)-based mass-spectrometry method to systematically measure the absolute protein concentration of nuclear TFs as human hematopoietic stem and progenitor cells differentiate along the erythropoietic lineage. For complete details on the use and execution of this protocol, please refer to Gillespie et al. (2020).
Mammalian SWI/SNF complexes are ATP-dependent chromatin remodeling complexes that regulate genomic architecture. Here, we present a structural model of the endogenously purified human canonical BAF complex bound to the nucleosome, generated using cryoelectron microscopy (cryo-EM), cross-linking mass spectrometry, and homology modeling. BAF complexes bilaterally engage the nucleosome H2A/H2B acidic patch regions through the SMARCB1 C-terminal α-helix and the SMARCA4/2 C-terminal SnAc/post-SnAc regions, with disease-associated mutations in either causing attenuated chromatin remodeling activities. Further, we define changes in BAF complex architecture upon nucleosome engagement and compare the structural model of endogenous BAF to those of related SWI/SNF-family complexes. Finally, we assign and experimentally interrogate cancer-associated hot-spot mutations localizing within the endogenous human BAF complex, identifying those that disrupt BAF subunit-subunit and subunit-nucleosome interfaces in the nucleosome-bound conformation. Taken together, this integrative structural approach provides important biophysical foundations for understanding the mechanisms of BAF complex function in normal and disease states.
Dynamic cellular processes such as differentiation are driven by changes in the abundances of transcription factors (TFs). However, despite years of studies, our knowledge about the protein copy number of TFs in the nucleus is limited. Here, by determining the absolute abundances of 103 TFs and co-factors during the course of human erythropoiesis, we provide a dynamic and quantitative scale for TFs in the nucleus. Furthermore, we establish the first gene regulatory network of cell fate commitment that integrates temporal protein stoichiometry data with mRNA measurements. The model revealed quantitative imbalances in TFs' cross-antagonistic relationships that underlie lineage determination. Finally, we made the surprising discovery that, in the nucleus, co-repressors are dramatically more abundant than co-activators at the protein level, but not at the RNA level, with profound implications for understanding transcriptional regulation. These analyses provide a unique quantitative framework to understand transcriptional regulation of cell differentiation in a dynamic context.
U6 snRNA is transcribed by RNA polymerase III (Pol III) and has an external upstream promoter that consists of a TATA sequence recognized by the TBP subunit of the Pol III basal transcription factor IIIB and a proximal sequence element (PSE) recognized by the small nuclear RNA activating protein complex (SNAPc). Previously, we found that Drosophila melanogaster SNAPc (DmSNAPc) bound to the U6 PSE can recruit the Pol III general transcription factor Bdp1 to form a stable complex with the DNA. Here, we show that DmSNAPc-Bdp1 can recruit TBP to the U6 promoter, and we identify a region of Bdp1 that is sufficient for TBP recruitment. Moreover, we find that this same region of Bdp1 cross-links to nucleotides within the U6 PSE at positions that also cross-link to DmSNAPc. Finally, cross-linking mass spectrometry reveals likely interactions of specific DmSNAPc subunits with Bdp1 and TBP. These data, together with previous findings, have allowed us to build a more comprehensive model of the DmSNAPc-Bdp1-TBP complex on the U6 promoter that includes nearly all of DmSNAPc, a portion of Bdp1, and the conserved region of TBP.
Synaptic activity in neurons leads to the rapid activation of genes involved in mammalian behavior. ATP-dependent chromatin remodelers such as the BAF complex contribute to these responses and are generally thought to activate transcription. However, the mechanisms keeping such "early activation" genes silent have been a mystery. In the course of investigating Mendelian recessive autism, we identified six families with segregating loss-of-function mutations in the neuronal BAF (nBAF) subunit ACTL6B (originally named BAF53b). Accordingly, ACTL6B was the most significantly mutated gene in the Simons Recessive Autism Cohort. At least 14 subunits of the nBAF complex are mutated in autism, collectively making it a major contributor to autism spectrum disorder (ASD). Patient mutations destabilized ACTL6B protein in neurons and rerouted dendrites to the wrong glomerulus in the fly olfactory system. Humans and mice lacking ACTL6B showed corpus callosum hypoplasia, indicating a conserved role for ACTL6B in facilitating neural connectivity. Actl6b knockout mice on two genetic backgrounds exhibited ASD-related behaviors, including social and memory impairments, repetitive behaviors, and hyperactivity. Surprisingly, mutation of Actl6b relieved repression of early response genes including AP1 transcription factors (Fos, Fosl2, Fosb, and Junb), increased chromatin accessibility at AP1 binding sites, and transcriptional changes in late response genes associated with early response transcription factor activity. ACTL6B loss is thus an important cause of recessive ASD, with impaired neuron-specific chromatin repression indicated as a potential mechanism.
Eukaryotic DNA is packaged into nucleosome arrays, which are repositioned by chromatin remodeling complexes to control DNA accessibility. The Saccharomyces cerevisiae RSC (Remodeling the Structure of Chromatin) complex, a member of the SWI/SNF chromatin remodeler family, plays critical roles in genome maintenance, transcription, and DNA repair. Here, we report cryo-electron microscopy (cryo-EM) and crosslinking mass spectrometry (CLMS) studies of yeast RSC complex and show that RSC is composed of a rigid tripartite core and two flexible lobes. The core structure is scaffolded by an asymmetric Rsc8 dimer and built with the evolutionarily conserved subunits Sfh1, Rsc6, Rsc9 and Sth1. The flexible ATPase lobe, composed of helicase subunit Sth1, Arp7, Arp9 and Rtt102, is anchored to this core by the N-terminus of Sth1. Our cryo-EM analysis of RSC bound to a nucleosome core particle shows that in addition to the expected nucleosome-Sth1 interactions, RSC engages histones and nucleosomal DNA through one arm of the core structure, composed of the Rsc8 SWIRM domains, Sfh1 and Npl6. Our findings provide structural insights into the conserved assembly process for all members of the SWI/SNF family of remodelers, and illustrate how RSC selects, engages, and remodels nucleosomes.
Chemical cross-linking combined with mass spectrometry (CL-MS) is a powerful method for characterizing the architecture of protein assemblies and for mapping protein–protein interactions. Despite its proven utility, confident identification of cross-linked peptides remains a formidable challenge, especially when the peptides are derived from complex mixtures. MS cleavable cross-linkers are gaining importance for CL-MS as they permit reliable identification of cross-linked peptides by whole proteome database searching using MS/MS information. Here we introduce a novel class of MS cleavable cross-linkers called isotopomeric cross-linkers (ICLs), which allow for confident and efficient identification of cross-linked peptides by whole proteome database searching. ICLs are simple, symmetrical molecules that asymmetrically incorporate heavy and light stable isotopes into the two arms of the cross-linker. As a result of this property, ICLs automatically generate pairs of isotopomeric cross-linked peptides, which differ only by the positions of the heavy and light isotopes. Upon fragmentation during MS analysis, these isotopomeric cross-linked peptides generate unique isotopic doublet ions that correspond to the individual peptides in the cross-link. The doublet ion information is used to determine the masses of the two cross-linked peptides from the same MS2 spectrum that is also used for peptide spectrum matching (PSM) by sequence database searching. Here we present the rationale for and mechanism of cross-linked peptide identification by ICL-MS. We describe the synthesis of the ICL-1 reagent, the ICL-MS workflow, and the performance characteristics of ICL-MS for identifying cross-linked peptides derived from increasingly complex mixtures by whole proteome database searching.