Pooled CRISPR screening has emerged as a powerful method of mapping gene functions thanks to its scalability, affordability, and robustness against well or plate-specific confounders present in array-based screening 1–6 . Most pooled CRISPR screens assay for low dimensional phenotypes (e.g. fitness, fluorescent markers). Higher-dimensional assays such as perturb-seq are available but costly and only applicable to transcriptomics readouts 7–11 . Recently, pooled optical screening, which combines pooled CRISPR screening and microscopy-based assays, has been demonstrated in the studies of the NFkB pathway, essential human genes, cytoskeletal organization and antiviral response 12–15 . While the pooled optical screening methodology is scalable and information-rich, the applications thus far employ hypothesis-specific assays. Here, we enable hypothesis-free reverse genetic screening for generic morphological phenotypes by re-engineering the Cell Painting 16 technique to provide compatibility with pooled optical screening. We validated this technique using well-defined morphological genesets (124 genes), compared classical image analysis and self-supervised learning methods using a mechanism-of-action (MoA) library (300 genes), and performed discovery screening with a druggable genome library (1640 genes) 17 . Across these three experiments we show that the combination of rich morphological data and deep learning allows gene networks to emerge without the need for target-specific biomarkers, leading to better discovery of gene functions.
Narcolepsy type 1 (NT1) is caused by a loss of hypocretin/orexin transmission. Risk factors include pandemic 2009 H1N1 influenza A infection and immunization with Pandemrix®. Here, we dissect disease mechanisms and interactions with environmental triggers in a multi-ethnic sample of 6,073 cases and 84,856 controls. We fine-mapped GWAS signals within HLA (DQ0602, DQB1*03:01 and DPB1*04:02) and discovered seven novel associations (CD207, NAB1, IKZF4-ERBB3, CTSC, DENND1B, SIRPG, PRF1). Significant signals at TRA and DQB1*06:02 loci were found in 245 vaccination-related cases, who also shared polygenic risk. T cell receptor associations in NT1 modulated TRAJ*24, TRAJ*28 and TRBV*4-2 chain-usage. Partitioned heritability and immune cell enrichment analyses found genetic signals to be driven by dendritic and helper T cells. Lastly comorbidity analysis using data from FinnGen, suggests shared effects between NT1 and other autoimmune diseases. NT1 genetic variants shape autoimmunity and response to environmental triggers, including influenza A infection and immunization with Pandemrix®.
Despite much research, our understanding of the architecture and cis-regulatory elements of human promoters is still lacking. Here, we devised a high-throughput assay to quantify the activity of approximately 15,000 fully designed sequences that we integrated and expressed from a fixed location within the human genome. We used this method to investigate thousands of native promoters and preinitiation complex (PIC) binding regions followed by in-depth characterization of the sequence motifs underlying promoter activity, including core promoter elements and TF binding sites. We find that core promoters drive transcription mostly unidirectionally and that sequences originating from promoters exhibit stronger activity than those originating from enhancers. By testing multiple synthetic configurations of core promoter elements, we dissect the motifs that positively and negatively regulate transcription as well as the effect of their combinations and distances, including a 10-bp periodicity in the optimal distance between the TATA and the initiator. By comprehensively screening 133 TF binding sites, we find that in contrast to core promoters, TF binding sites maintain similar activity levels in both orientations, supporting a model by which divergent transcription is driven by two distinct unidirectional core promoters sharing bidirectional TF binding sites. Finally, we find a striking agreement between the effect of binding site multiplicity of individual TFs in our assay and their tendency to appear in homotypic clusters throughout the genome. Overall, our study systematically assays the elements that drive expression in core and proximal promoter regions and sheds light on organization principles of regulatory regions in the human genome.
BIOPHYSICS AND COMPUTATIONAL BIOLOGY Correction for “Deciphering the rules by which 5′-UTR sequences affect protein expression in yeast,” by Shlomi Dvir, Lars Velten, Eilon Sharon, Danny Zeevi, Lucas B. Carey, Adina Weinberger, and Eran Segal, which was first published July 5, 2013; 10.1073/pnas.1222534110 (Proc. Natl. Acad. Sci. U.S.A. 110, E2792–E2801). The authors note that the following statement should be added to the Acknowledgments: “This work was supported by a grant from the European Research Council (ERC) to E.S.”
A major challenge in genetics is to identify genetic variants driving natural phenotypic variation. However, current methods of genetic mapping have limited resolution. To address this challenge, we developed a CRISPR-Cas9-based high-throughput genome editing approach that can introduce thousands of specific genetic variants in a single experiment. This enabled us to study the fitness consequences of 16,006 natural genetic variants in yeast. We identified 572 variants with significant fitness differences in glucose media; these are highly enriched in promoters, particularly in transcription factor binding sites, while only 19.2% affect amino acid sequences. Strikingly, nearby variants nearly always favor the same parent's alleles, suggesting that lineage-specific selection is often driven by multiple clustered variants. In sum, our genome editing approach reveals the genetic architecture of fitness variation at single-base resolution and could be adapted to measure the effects of genome-wide genetic variation in any screen for cell survival or cell-sortable markers.
Type 1 narcolepsy (T1N) is a neurological condition, in which the death of hypocretin-producing neurons in the lateral hypothalamus leads to excessive daytime sleepiness and symptoms of abnormal Rapid Eye Movement (REM) sleep. Known triggers for narcolepsy are influenza-A infection and associated immunization during the 2009 H1N1 influenza pandemic. Here, we genotyped all remaining consented narcolepsy cases worldwide and assembled this with the existing genotyped individuals. We used this multi-ethnic sample in genome wide association study (GWAS) to dissect disease mechanisms and interactions with environmental triggers (5,339 cases and 20,518 controls). Overall, we found significant associations with HLA (2 GWA significant subloci) and 11 other loci. Six of these other loci have been previously reported ( TRA , TRB , CTSH , IFNAR1 , ZNF365 and P2RY11 ) and five are new ( PRF1 , CD207 , SIRPG , IL27 and ZFAND2A) . Strikingly, in vaccination-related cases GWA significant effects were found in HLA , TRA, and in a novel variant near SIRPB1 . Furthermore, IFNAR1 associated polymorphisms regulated dendritic cell response to influenza-A infection in vitro (p-value =1.92*10 −25 ). A partitioned heritability analysis indicated specific enrichment of functional elements active in cytotoxic and helper T cells. Furthermore, functional analysis showed the genetic variants in TRA and TRB loci act as remarkable strong chain usage QTLs for TRAJ*24 ( p-value = 0.0017 ) , TRAJ*28 (p-value = 1.36*10 −10 ) and TRBV*4-2 (p-value = 3.71*10 −117 ). This was further validated in TCR sequencing of 60 narcolepsy cases and 60 DQB1*06:02 positive controls, where chain usage effects were further accentuated. Together these findings show that the autoimmune component in narcolepsy is defined by antigen presentation, mediated through specific T cell receptor chains, and modulated by influenza-A as a critical trigger.
Despite its pivotal role in regulating transcription, our understanding of core promoter function, architecture, and cis-regulatory elements is lacking. Here, we devised a highthroughput assay to quantify the activity of ∼15,000 fully designed core promoters that we integrated and expressed from a fixed location within the human genome. We find that core promoters drive transcription unidirectionally, and that sequences originating from promoters exhibit stronger activity than sequences originating from enhancers. Testing multiple combinations and distances of core promoter elements, we observe a positive effect of TATA and Initiator, a negative effect of BREu and BREd, and a 10bp periodicity in the optimal distance between the TATA and the Initiator. By comprehensively screening TF binding-sites, we show that site orientation has little effect, that the effect of binding site number on expression is factor-specific, and that there is a striking agreement between the effect of binding site multiplicity in our assay and the tendency of the TF to appear in homotypic clusters throughout the genome. Overall, our results systematically assay the elements that drive expression in core- and proximal-promoter regions and shed light on organization principles of regulatory regions in the human genome.
Quantification of cell-free DNA (cfDNA) in circulating blood derived from a transplanted organ is a powerful approach to monitoring post-transplant injury. Genome transplant dynamics (GTD) quantifies donor-derived cfDNA (dd-cfDNA) by taking advantage of single-nucleotide polymorphisms (SNPs) distributed across the genome to discriminate donor and recipient DNA molecules. In its current implementation, GTD requires genotyping of both the transplant recipient and donor. However, in practice, donor genotype information is often unavailable. Here, we address this issue by developing an algorithm that estimates dd-cfDNA levels in the absence of a donor genotype. Our algorithm predicts heart and lung allograft rejection with an accuracy that is similar to conventional GTD. We furthermore refined the algorithm to handle closely related recipients and donors, a scenario that is common in bone marrow and kidney transplantation. We show that it is possible to estimate dd-cfDNA in bone marrow transplant patients that are unrelated or that are siblings of the donors, using a hidden Markov model (HMM) of identity-by-descent (IBD) states along the genome. Last, we demonstrate that comparing dd-cfDNA to the proportion of donor DNA in white blood cells can differentiate between relapse and the onset of graft-versus-host disease (GVHD). These methods alleviate some of the barriers to the implementation of GTD, which will further widen its clinical application.
Transcription factors (TFs) are key mediators that propagate extracellular and intracellular signals through to changes in gene expression profiles. However, the rules by which promoters decode the amount of active TF into target gene expression are not well understood. To determine the mapping between promoter DNA sequence, TF concentration, and gene expression output, we have conducted in budding yeast a large-scale measurement of the activity of thousands of designed promoters at six different levels of TF. We observe that maximum promoter activity is determined by TF concentration and not by the number of binding sites. Surprisingly, the addition of an activator site often reduces expression. A thermodynamic model that incorporates competition between neighboring binding sites for a local pool of TF molecules explains this behavior and accurately predicts both absolute expression and the amount by which addition of a site increases or reduces expression. Taken together, our findings support a model in which neighboring binding sites interact competitively when TF is limiting but otherwise act additively.
In each individual, a highly diverse T cell receptor (TCR) repertoire interacts with peptides presented by major histocompatibility complex (MHC) molecules. Despite extensive research, it remains controversial whether germline-encoded TCR-MHC contacts promote TCR-MHC specificity and, if so, whether differences exist in TCR V gene compatibilities with different MHC alleles. We applied expression quantitative trait locus (eQTL) mapping to test for associations between genetic variation and TCR V gene usage in a large human cohort. We report strong trans associations between variation in the MHC locus and TCR V gene usage. Fine-mapping of the association signals identifies specific amino acids from MHC genes that bias V gene usage, many of which contact or are spatially proximal to the TCR or peptide in the TCR-peptide-MHC complex. Hence, these MHC variants, several of which are linked to autoimmune diseases, can directly affect TCR-MHC interaction. These results provide the first examples of trans-QTL effects mediated by protein-protein interactions and are consistent with intrinsic TCR-MHC specificity.
Genome Research 25: 1018–1029 (2015) The name of the sixth author was originally misspelled in the author line of this article. Please note the correct spelling as Maya Lotan-Pompan. The file has already been corrected online. doi: 10.1101/gr.196725.115
The 3'end genomic region encodes a wide range of regulatory process including mRNA stability, 3' end processing and translation. Here, we systematically investigate the sequence determinants of 3' end mediated expression control by measuring the effect of 13,000 designed 3' end sequence variants on constitutive expression levels in yeast. By including a high resolution scanning mutagenesis of more than 200 native 3' end sequences in this designed set, we found that most mutations had only a mild effect on expression, and that the vast majority (~90%) of strongly effecting mutations localized to a single positive TA-rich element, similar to a previously described 3' end processing efficiency element, and resulted in up to ten-fold decrease in expression. Measurements of 3' UTR lengths revealed that these mutations result in mRNAs with aberrantly long 3'UTRs, confirming the role for this element in 3' end processing. Interestingly, we found that other sequence elements that were previously described in the literature to be part of the polyadenylation signal had a minor effect on expression. We further characterize the sequence specificities of the TA-rich element using additional synthetic 3' end sequences and show that its activity is sensitive to single base pair mutations and strongly depends on the A/T content of the surrounding sequences. Finally, using a computational model, we show that the strength of this element in native 3' end sequences can explain some of their measured expression variability (R = 0.41). Together, our results emphasize the importance of efficient 3' end processing for endogenous protein levels and contribute to an improved understanding of the sequence elements involved in this process.
The response of gene expression to intra-and extra-cellular cues is largely mediated through changes in the activity of transcription factors (TFs), whose sequence specificities are largely known. However, the rules by which promoters decode the amount of active TF into gene expression are not well understood. Here, we measure the activity of 6500 designed promoters at six different levels of TF activity in budding yeast. We observe that maximum promoter activity is determined by TF activity and not by the number of sites. Surprisingly, the addition of an activator-binding site often reduces expression. A thermodynamic model that incorporates competition between neighboring binding sites for a local pool of TF molecules explains this behavior and accurately predicts both absolute expression and the amount by which addition of a site increases or reduces expression. Taken together, our findings support a model in which neighboring binding sites interact competitively when TF is limiting but otherwise act additively Significance Statement In response to intracellular and extracellular signals organisms alter the concentration and activity of transcription factors (TFs), proteins that regulate gene expression. However, the molecular mechanisms that determine the response of a target promoter to changes in the number of active TF molecules are not well understood. By combining mathematical modeling with measurements of TF dose-response curves for thousands of designed promoters, we show that competition for active TF molecules is a major factor in determining gene expression. At low TF concentrations additional activator-binding sites within a promoter can actually reduce expression. Thermodynamic modeling suggests that steric hindrance between neighboring binding sites cannot explain this behavior, but that competition for limiting TF molecules can.
Binding of transcription factors (TFs) to regulatory sequences is a pivotal step in the control of gene expression. Despite many advances in the characterization of sequence motifs recognized by TFs, our ability to quantitatively predict TF binding to different regulatory sequences is still limited. Here, we present a novel experimental assay termed BunDLE-seq that provides quantitative measurements of TF binding to thousands of fully designed sequences of 200 bp in length within a single experiment. Applying this binding assay to two yeast TFs, we demonstrate that sequences outside the core TF binding site profoundly affect TF binding. We show that TF-specific models based on the sequence or DNA shape of the regions flanking the core binding site are highly predictive of the measured differential TF binding. We further characterize the dependence of TF binding, accounting for measurements of single and co-occurring binding events, on the number and location of binding sites and on the TF concentration. Finally, by coupling our in vitro TF binding measurements, and another application of our method probing nucleosome formation, to in vivo expression measurements carried out with the same template sequences serving as promoters, we offer insights into mechanisms that may determine the different expression outcomes observed. Our assay thus paves the way to a more comprehensive understanding of TF binding to regulatory sequences and allows the characterization of TF binding determinants within and outside of core binding sites.
Daphne Koller合作论文数Computer Science Department, Stanford University;Insitro;Engageli5