On-policy distillation (OPD) is a widely used technique to transfer capabilities from capable teacher language models to the base student models, and can be formulated in a reinforcement learning style objective using student generated rollouts. Yet, despite the divergence reward being dependent on student model likelihood, existing works usually adopt a stop gradient design primarily for stability, which makes the resulting advantage estimation questionable. In this work, we provide a generic optimization framework based on f-divergence between the student and teacher, and mathematically revisit whether such design space is valid. We prove that general stop-gradient operation would lead to biased estimates of the reward objective and corresponding gradient for general divergence functions. We propose OPD+, the corrected version of OPD that demonstrates improved performance over the baseline KL approach and also supports the choice of various f-divergence. We validate our findings on mathematical reasoning and tool-use benchmarks.
Heterochromatin Protein 1α (HP1α) is a fundamental component of constitutive heterochromatin, forming subnuclear condensates whose regulation and function remain poorly understood. Here, we present an image-based CRISPR screen targeting nuclear factors that identifies splicing as a pivotal pathway regulating HP1α condensates. We discovered that unspliced intronic RNA modulates HP1α condensates by interacting co-transcriptionally with HP1α. By modulating the intron content, RNA processing restricts HP1α-RNA interactions at chromatin, thus enabling heterochromatin organization. Disruption of HP1α condensates due to enhanced interactions with unspliced RNA leads to loss of heterochromatin and the activation of stress response protective genes. We propose that RNA is a central component of heterochromatin that modulates HP1α condensates, and that RNA processing enzymes act as a surveillance mechanism for condensates by dynamically regulating the network of multi-valent interactions between RNA and chromatin factors. This model underscores the crosstalk between chromatin organization, transcription, and RNA processing, potentially governing broader nuclear functions.
The assembly of cortical circuits involves the generation and migration of cortical interneurons from the ventral to the dorsal forebrain, which has been challenging to study in humans as these processes take place at inaccessible stages of late gestation and early postnatal development. Autism spectrum disorder (ASD) and other neurodevelopmental disorders (NDDs) have been associated with abnormal cortical interneuron development, but which of the hundreds of NDD genes impact interneuron generation and migration into circuits and how they mediate these effects remain unknown. We previously developed a stem cell-based platform to study human cortical interneurons in self-organizing organoids resembling the ventral forebrain and their migration using forebrain assembloids. Here, we integrate assembloid technology with CRISPR screening to systematically investigate the involvement of 425 NDD genes in human interneuron development. The first screen aimed at interneuron generation revealed 13 candidate genes, including the RNA-binding protein CSDE1 and the canonical TGFβ signaling activator SMAD4. Then, we ran an interneuron migration screen in ~1,000 forebrain assembloids that identified 33 candidate genes, including cytoskeleton-related genes and, notably, the endoplasmic reticulum (ER)-related gene LNPK. Interestingly, we discovered that, during interneuron migration, the ER is displaced along the leading neuronal branch prior to nuclear translocation. Deletion of LNPK interfered with this ER displacement and resulted in reduced interneuron saltation length, indicating a critical role for the ER in this migratory process. Taken together, these results highlight how this versatile CRISPR-assembloid platform can be used to systematically map disease genes onto early stages of human neural development and to reveal novel mechanisms regulating interneuron development.
Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the effect of human genetic variation on disease1-6. Here we create and evaluate a resource of more than 92 million enhancer-gene regulatory interactions across 1,458 biosamples covering 369 cell types and tissues, by integrating predictive models, chromatin states, three-dimensional contacts and large-scale genetic perturbations generated by the ENCODE Consortium7. We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,356 element-gene pairs measured in CRISPR perturbation experiments, more than 30,000 fine-mapped expression quantitative trait loci and 569 fine-mapped genome-wide association study (GWAS) variants linked to a probable causal gene. Using this framework, we develop ENCODE-rE2G, a predictive model achieving state-of-the-art performance across several prediction tasks, demonstrating that iterative perturbations and supervised machine learning can build increasingly accurate predictive models of enhancer regulation. Using ENCODE-rE2G, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, revealing global properties of enhancer networks, identifying differences in regulatory complexity across genes and improving analyses linking noncoding variants to target genes and cell types for common complex diseases. By interpreting the model, we find that beyond enhancer activity and three-dimensional enhancer-promoter contacts, additional features that guide enhancer-promoter communication include promoter class and enhancer-enhancer synergy. These genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.
The field of simulation optimization (SO) encompasses various methods developed to optimize complex, expensive-to-sample stochastic systems. Established methods include, but are not limited to, ranking-and-selection for finite alternatives and surrogate-based methods for continuous domains, with broad applications in engineering and operations management. The recent advent of large language models (LLMs) offers a new paradigm for exploiting system structure and automating the strategic selection and composition of these established SO methods into a tailored optimization procedure. This work introduces SOCRATES (Simulation Optimization with Correlated Replicas and Adaptive Trajectory Evaluations), a novel two-stage procedure that leverages LLMs to automate the design of tailored SO algorithms. The first stage constructs an ensemble of digital replicas of the real system. An LLM is employed to implement causal discovery from a textual description of the system, generating a structural `skeleton' that guides the sample-efficient learning of the replicas. In the second stage, this replica ensemble is used as an inexpensive testbed to evaluate a set of baseline SO algorithms. An LLM then acts as a meta-optimizer, analyzing the performance trajectories of these algorithms to iteratively revise and compose a final, hybrid optimization schedule. This schedule is designed to be adaptive, with the ability to be updated during the final execution on the real system when the optimization performance deviates from expectations. By integrating LLM-driven reasoning with LLM-assisted trajectory-aware meta-optimization, SOCRATES creates an effective and sample-efficient solution for complex SO optimization problems.
We propose DiFFPO, Diffusion Fast and Furious Policy Optimization, a unified framework for training masked diffusion large language models (dLLMs) to reason not only better (furious), but also faster via reinforcement learning (RL). We first unify the existing baseline approach such as d1 by proposing to train surrogate policies via off-policy RL, whose likelihood is much more tractable as an approximation to the true dLLM policy. This naturally motivates a more accurate and informative two-stage likelihood approximation combined with importance sampling correction, which leads to generalized RL algorithms with better sample efficiency and superior task performance. Second, we propose a new direction of joint training efficient samplers/controllers of dLLMs policy. Via RL, we incentivize dLLMs' natural multi-token prediction capabilities by letting the model learn to adaptively allocate an inference threshold for each prompt. By jointly training the sampler, we yield better accuracies with lower number of function evaluations (NFEs) compared to training the model only, obtaining the best performance in improving the Pareto frontier of the inference-time compute of dLLMs. We showcase the effectiveness of our pipeline by training open source large diffusion language models over benchmark math and planning tasks.
TP53 , the most frequently mutated gene in human cancer, encodes a transcriptional activator that induces myriad downstream target genes. Despite the importance of p53 in tumor suppression, the specific p53 target genes important for tumor suppression remain unclear. Recent studies have identified the p53-inducible gene Zmat3 as a critical effector of tumor suppression, but many questions remain regarding its p53-dependence, activity across contexts, and mechanism of tumor suppression alone and in cooperation with other p53-inducible genes. To address these questions, we used Tuba-seq Ultra somatic genome editing and tumor barcoding in a mouse lung adenocarcinoma model, combinatorial in vivo CRISPR/Cas9 screens, meta-analyses of gene expression and Cancer Dependency Map data, and integrative RNA-sequencing and shotgun proteomic analyses. We established Zmat3 as a core component of p53-mediated tumor suppression and identified Cdkn1a as the most potent cooperating p53-induced gene in tumor suppression. We discovered that ZMAT3/CDKN1A serve as near-universal effectors of p53-mediated tumor suppression that regulate cell division, migration, and extracellular matrix organization. Accordingly, combined Zmat3 - Cdkn1a inactivation dramatically enhanced cell proliferation and migration compared to controls, akin to p53 inactivation. Together, our findings place ZMAT3 and CDKN1A as hubs of a p53-induced gene program that opposes tumorigenesis across various cellular and genetic contexts.
Although critical for tuning the timing and level of transcription, enhancer communication with distal promoters is not well understood. Here, we bypass the need for sequence-specific transcription factors (TFs) and recruit activators directly using a chimeric array of gRNA oligos to target dCas9 fused to the activator VP64-p65-Rta (CARGO-VPR). We show that this approach achieves effective activator recruitment to arbitrary genomic sites, even those inaccessible when targeted with a single guide. We utilize CARGO-VPR across the Prdm8-Fgf5 locus in mouse embryonic stem cells (mESCs), where neither gene is expressed. Although activator recruitment to any tested region results in the transcriptional induction of at least one gene, the expression level strongly depends on the genomic distance between the promoter and activator recruitment site. However, the expression-distance relationship for each gene scales distinctly in a manner not attributable to differences in 3D contact frequency, promoter DNA sequence, or the presence of repressive chromatin marks at the locus.
The ENCODE Consortium’s efforts to annotate noncoding cis -regulatory elements (CREs) have advanced our understanding of gene regulatory landscapes. Pooled, noncoding CRISPR screens offer a systematic approach to investigate cis -regulatory mechanisms. The ENCODE4 Functional Characterization Centers conducted 108 screens in human cell lines, comprising >540,000 perturbations across 24.85 megabases of the genome. Using 332 functionally confirmed CRE–gene links in K562 cells, we established guidelines for screening endogenous noncoding elements with CRISPR interference (CRISPRi), including accurate detection of CREs that exhibit variable, often low, transcriptional effects. Benchmarking five screen analysis tools, we find that CASA produces the most conservative CRE calls and is robust to artifacts of low-specificity single guide RNAs. We uncover a subtle DNA strand bias for CRISPRi in transcribed regions with implications for screen design and analysis. Together, we provide an accessible data resource, predesigned single guide RNAs for targeting 3,275,697 ENCODE SCREEN candidate CREs with CRISPRi and screening guidelines to accelerate functional characterization of the noncoding genome.
Transcriptional effectors are protein domains known to activate or repress gene expression; however, a systematic understanding of which effector domains regulate transcription across genomic, cell type and DNA-binding domain (DBD) contexts is lacking. Here we develop dCas9-mediated high-throughput recruitment (HT-recruit), a pooled screening method for quantifying effector function at endogenous target genes and test effector function for a library containing 5,092 nuclear protein Pfam domains across varied contexts. We also map context dependencies of effectors drawn from unannotated protein regions using a larger library tiling chromatin regulators and transcription factors. We find that many effectors depend on target and DBD contexts, such as HLH domains that can act as either activators or repressors. To enable efficient perturbations, we select context-robust domains, including ZNF705 KRAB, that improve CRISPRi tools to silence promoters and enhancers. We engineer a compact human activator called NFZ, by combining NCOA3, FOXO3 and ZNF473 domains, which enables efficient CRISPRa with better viral delivery and inducible control of chimeric antigen receptor T cells. Improved effectors for CRISPRi/CRISPRa are developed following high-throughput screening of transcriptional domains.
Identifying transcriptional enhancers and their target genes is essential for understanding gene regulation and the impact of human genetic variation on disease 1–6 . Here we create and evaluate a resource of >13 million enhancer-gene regulatory interactions across 352 cell types and tissues, by integrating predictive models, measurements of chromatin state and 3D contacts, and large-scale genetic perturbations generated by the ENCODE Consortium 7 . We first create a systematic benchmarking pipeline to compare predictive models, assembling a dataset of 10,411 element-gene pairs measured in CRISPR perturbation experiments, >30,000 fine-mapped eQTLs, and 569 fine-mapped GWAS variants linked to a likely causal gene. Using this framework, we develop a new predictive model, ENCODE-rE2G, that achieves state-of-the-art performance across multiple prediction tasks, demonstrating a strategy involving iterative perturbations and supervised machine learning to build increasingly accurate predictive models of enhancer regulation. Using the ENCODE-rE2G model, we build an encyclopedia of enhancer-gene regulatory interactions in the human genome, which reveals global properties of enhancer networks, identifies differences in the functions of genes that have more or less complex regulatory landscapes, and improves analyses to link noncoding variants to target genes and cell types for common, complex diseases. By interpreting the model, we find evidence that, beyond enhancer activity and 3D enhancer-promoter contacts, additional features guide enhancer-promoter communication including promoter class and enhancer-enhancer synergy. Altogether, these genome-wide maps of enhancer-gene regulatory interactions, benchmarking software, predictive models, and insights about enhancer function provide a valuable resource for future studies of gene regulation and human genetics.
The assembly of cortical circuits involves the generation and migration of interneurons from the ventral to the dorsal forebrain 1 – 3 , which has been challenging to study at inaccessible stages of late gestation and early postnatal human development 4 . Autism spectrum disorder and other neurodevelopmental disorders (NDDs) have been associated with abnormal cortical interneuron development 5 , but which of these NDD genes affect interneuron generation and migration, and how they mediate these effects remains unknown. We previously developed a platform to study interneuron development and migration in subpallial organoids and forebrain assembloids 6 . Here we integrate assembloids with CRISPR screening to investigate the involvement of 425 NDD genes in human interneuron development. The first screen aimed at interneuron generation revealed 13 candidate genes, including CSDE1 and SMAD4 . We subsequently conducted an interneuron migration screen in more than 1,000 forebrain assembloids that identified 33 candidate genes, including cytoskeleton-related genes and the endoplasmic reticulum-related gene LNPK . We discovered that, during interneuron migration, the endoplasmic reticulum is displaced along the leading neuronal branch before nuclear translocation. LNPK deletion interfered with this endoplasmic reticulum displacement and resulted in abnormal migration. These results highlight the power of this CRISPR-assembloid platform to systematically map NDD genes onto human development and reveal disease mechanisms.
Selectively ablating damaged cells is an evolving therapeutic approach for age-related disease. Current methods for genome-wide screens to identify genes whose deletion might promote the death of damaged or senescent cells are generally underpowered because of the short timescales of cell death as well as the difficulty of scaling non-dividing cells. Here, we establish "Death-seq,"a positive-selection CRISPR screen optimized to identify enhancers and mechanisms of cell death. Our screens identified synergistic enhancers of cell death induced by the known senolytic ABT-263. The screen also identified inducers of cell death and senescent cell clearance in models of age-related diseases by a related compound, ABT-199, which alone is not senolytic but exhibits less toxicity than ABT-263. Death-seq enables the systematic screening of cell death pathways to uncover molecular mechanisms of regulated cell death subroutines and identifies drug targets for the treatment of diverse pathological states such as senescence, cancer, and fibrosis.
Human nuclear proteins contain >1000 transcriptional effector domains that can activate or repress transcription of target genes. We lack a systematic understanding of which effector domains regulate transcription robustly across genomic, cell-type, and DNA-binding domain (DBD) contexts. Here, we developed dCas9-mediated high-throughput recruitment (HT-recruit), a pooled screening method for quantifying effector function at endogenous targets, and tested effector function for a library containing 5092 nuclear protein Pfam domains across varied contexts. We find many effectors depend on target and DBD contexts, such as HLH domains that can act as either activators or repressors. We then confirm these findings and further map context dependencies of effectors drawn from unannotated protein regions using a larger library containing 114,288 sequences tiling chromatin regulators and transcription factors. To enable efficient perturbations, we select effectors that are potent in diverse contexts, and engineer (1) improved ZNF705 KRAB CRISPRi tools to silence promoters and enhancers, and (2) a compact human activator combination NFZ for better CRISPRa and inducible circuit delivery. Together, this effector-by-context functional map reveals context-dependence across human effectors and guides effector selection for robustly manipulating transcription.
We study reinforcement learning (RL) in the setting of continuous time and space, for an infinite horizon with a discounted objective and the underlying dynamics driven by a stochastic differential equation. Built upon recent advances in the continuous approach to RL, we develop a notion of occupation time (specifically for a discounted objective), and show how it can be effectively used to derive performance-difference and local-approximation formulas. We further extend these results to illustrate their applications in the PG (policy gradient) and TRPO/PPO (trust region policy optimization/ proximal policy optimization) methods, which have been familiar and powerful tools in the discrete RL setting but under-developed in continuous RL. Through numerical experiments, we demonstrate the effectiveness and advantages of our approach.
Selectively ablating senescent cells (“senolysis”) is an evolving therapeutic approach for age-related diseases. Current senolytics are limited to local administration by potency and side effects. While genetic screens could identify senolytics, current screens are underpowered for identifying genes that regulate cell death due to limitations in screen methodology. Here, we establish Death-seq, a positive selection CRISPR screen optimized to identify enhancers and mechanisms of cell death. Our screens identified synergistic enhancers of cell death induced by the known senolytic ABT-263, a BH3 mimetic. SMAC mimetics, enhancers in our screens, synergize with ABT-199, another BH3 mimetic that is not senolytic alone, clearing senescent cells in models of age-related disease while sparing human platelets, avoiding the thrombocytopenia associated with ABT-263. Death-seq enables the systematic screening of cell death pathways to uncover molecular mechanisms of regulated cell death subroutines and identify drug targets for diverse pathological states such as senescence, cancer, and neurodegeneration.
The ENCODE Consortium’s efforts to annotate non-coding, cis -regulatory elements (CREs) have advanced our understanding of gene regulatory landscapes which play a major role in health and disease. Pooled, non-coding CRISPR screens are a promising approach for systematically investigating gene regulatory mechanisms. Here, the ENCODE Functional Characterization Centers report 109 screens comprising 346,970 individual perturbations across 13.3Mb of the genome, using a variety of methods, readouts, and statistical analyses. Across 332 functionally confirmed CRE-gene links, we identify principles for screening endogenous, non-coding elements for causal regulatory mechanisms. Nearly all CREs show strong evidence of open chromatin, and targeting accessibility peak summits is a critical component of our proposed sgRNA design rules. We provide experimental guidelines to accurately detect CREs with variable, often low, transcriptional effects. We discover a previously undescribed DNA strand-bias for CRISPRi in transcribed regions with implications for screen design and analysis. Benchmarking five screen analysis tools, we find CASA produces the most conservative CRE calls and is robust to artifacts of low-specificity sgRNAs. Together, we provide an accessible data resource, predesigned sgRNAs targeting 3,275,697 ENCODE SCREEN candidate CREs, and screening guidelines to accelerate functional characterization of the non-coding genome.
Decision trees are robust modeling tools in machine learning with human-interpretable representations. The curse of dimensionality of Markov Decision Process (MDP) makes exact solution methods computationally intractable in practice for large state-action spaces. In this paper, we show that even for problems with large state space, when the solution policy of the MDP can be represented by a tree-like structure, our proposed algorithm retrieves a tree of the solution policy of the MDP in computationally tractable time. Our algorithm uses a tree growing strategy to incrementally disaggregate the state space solving smaller MDP instances with Linear Programming. These ideas can be extended to experience based RL problems as an alternative to black-box based policies.
The trafficking of specific protein cohorts to correct subcellular locations at correct times is essential for every signaling and regulatory process in biology. Gene perturbation screens could provide a powerful approach to probe the molecular mechanisms of protein trafficking, but only if protein localization or mislocalization can be tied to a simple and robust phenotype for cell selection, such as cell proliferation or fluorescence-activated cell sorting (FACS). To empower the study of protein trafficking processes with gene perturbation, we developed a genetically encoded molecular tool named HiLITR (High-throughput Localization Indicator with Transcriptional Readout). HiLITR converts protein colocalization into proteolytic release of a membrane-anchored transcription factor, which drives the expression of a chosen reporter gene. Using HiLITR in combination with FACS-based CRISPRi screening in human cell lines, we identified genes that influence the trafficking of mitochondrial and ER tail-anchored proteins. We show that loss of the SUMO E1 component SAE1 results in mislocalization and destabilization of many mitochondrial tail-anchored proteins. We also demonstrate a distinct regulatory role for EMC10 in the ER membrane complex, opposing the transmembrane-domain insertion activity of the complex. Through transcriptional integration of complex cellular functions, HiLITR expands the scope of biological processes that can be studied by genetic perturbation screening technologies.