Scientific progress depends not only on finding solutions, but on learning the rules that explain why they work and using that understanding to design better experiments. We introduce science sandboxes, a framework for studying this capability in AI agents through repeated cycles of experimentation, feedback, and hypothesis revision. Science sandboxes invite an agent to query the natural world in different ways, ranging from "wet" physical experiments, to "damp" predictive models trained on empirical data, to "dry" invented rules. By establishing a common experimental loop and a protocol for evaluating agents within it, science sandboxes allow assessment of both quantitative performance on specific metrics and qualitative scientific reasoning, across a spectrum of empirical verifiability. Here, we instantiate this framework in two biological settings, models of regulatory genomics and protein fitness prediction, and examine the capabilities of frontier agents. Across these settings, we could see when agents successfully optimized a quantitative metric without understanding the rules underlying the system. In particular, their scientific reasoning deteriorated when they encountered systems whose rules fell outside familiar biological priors. By highlighting such failure modes, science sandboxes make the frontier of scientific capability measurable and provide a controlled setting in which to study and ultimately expand it.
Identifying the causal variants and mechanisms that drive complex traits and diseases remains a core problem in human genetics1-5. Most of these variants individually have weak effects6 and lie in non-coding gene-regulatory elements7-10, for which we lack a complete understanding of how single-nucleotide alterations modulate transcriptional processes to affect human phenotypes5,11-15. To address this problem, we measured the activity of 221,412 fine-mapped trait-associated variants using a massively parallel reporter assay16-20 in 5 diverse cell types. We show that this assay effectively discriminates between likely causal variants and controls, and identified 13,121 regulatory variants with high precision. Although the effects of these variants largely agree with orthogonal measures of function, only 69% of them can plausibly be explained by the disruption of a known transcription factor binding motif. We investigated the mechanisms of 136 variants using saturation mutagenesis and assigned affected transcription factors for 91% of variants without a clear canonical mechanism. Finally, we detected regulatory epistasis at 11% of tested regulatory variants in close proximity and identified multiple functional variants on the same haplotype at a small, but important, subset of trait-associated loci. Overall, our study provides a systematic functional characterization of likely causal common variants that underlie complex and molecular human traits, enabling new insights into the regulatory grammar underlying disease risk.
The inability to interpret the functional impact of non-coding variants has been a major impediment in the promise of precision medicine. While high-throughput experimental approaches such as Massively Parallel Reporter Assays (MPRAs) have made major progress in identifying causal variants and their underlying molecular mechanisms, these tools cannot exhaustively measure variant effects genome-wide. Here we present MPAC, an ensemble of machine-learning models trained on MPRA data that provides accurate and scalable prediction of the cis-regulatory impact of non-coding variants. Using MPAC we predict allelic effects for 575M single nucleotide variants (SNVs) across diverse applications, including complex trait genetics, clinical and tumor sequencing, evolutionary analyses, and saturation mutagenesis. We find MPAC predictions match the performance of empirical MPRAs in identifying causal complex trait-associated alleles. We demonstrate the utility of MPAC by applying it to ClinVar, identifying non-coding pathogenic variation with higher accuracy than other sequence-to-function models. We also nominate 1,892 candidate non-coding cancer drivers by predicting the functional effects of somatic SNVs in the COSMIC database. Next, we evaluate population-level genetic variation by predicting effects for all 514M non-coding SNVs in gnomAD, quantifying the relationship between regulatory function and evolutionary constraint. Finally, we generate prospective functional maps using in-silico saturation mutagenesis across 18,658 human promoters, observing widespread selection against variants predicted to disrupt promoter activity. Collectively, this study establishes the value of non-coding functional predictions and provides a comprehensive, publicly available resource for variant interpretation. ### Competing Interest Statement P.C.S. and R.T. have filed intellectual property related to MPRA. S.J.G., P.C.S., R.I.C., R.T., and S.K.R., have filed intellectual property related to MPRA models. P.C.S. is a co-founder and shareholder of Delve Bio, and was formerly a co-founder and shareholder of Sherlock Biosciences and Board Member and shareholder of Danaher Corporation. The other authors declare no competing interests.
Identifying the causal variants and mechanisms that drive complex traits and diseases remains a core problem in human genetics. The majority of these variants have individually weak effects and lie in non-coding gene-regulatory elements where we lack a complete understanding of how single nucleotide alterations modulate transcriptional processes to affect human phenotypes. To address this, we measured the activity of 221,412 trait-associated variants that had been statistically fine-mapped using a Massively Parallel Reporter Assay (MPRA) in 5 diverse cell-types. We show that MPRA is able to discriminate between likely causal variants and controls, identifying 12,025 regulatory variants with high precision. Although the effects of these variants largely agree with orthogonal measures of function, only 69% can plausibly be explained by the disruption of a known transcription factor (TF) binding motif. We dissect the mechanisms of 136 variants using saturation mutagenesis and assign impacted TFs for 91% of variants without a clear canonical mechanism. Finally, we provide evidence that epistasis is prevalent for variants in close proximity and identify multiple functional variants on the same haplotype at a small, but important, subset of trait-associated loci. Overall, our study provides a systematic functional characterization of likely causal common variants underlying complex and molecular human traits, enabling new insights into the regulatory grammar underlying disease risk.
The ENCODE Consortium’s efforts to annotate noncoding cis -regulatory elements (CREs) have advanced our understanding of gene regulatory landscapes. Pooled, noncoding CRISPR screens offer a systematic approach to investigate cis -regulatory mechanisms. The ENCODE4 Functional Characterization Centers conducted 108 screens in human cell lines, comprising >540,000 perturbations across 24.85 megabases of the genome. Using 332 functionally confirmed CRE–gene links in K562 cells, we established guidelines for screening endogenous noncoding elements with CRISPR interference (CRISPRi), including accurate detection of CREs that exhibit variable, often low, transcriptional effects. Benchmarking five screen analysis tools, we find that CASA produces the most conservative CRE calls and is robust to artifacts of low-specificity single guide RNAs. We uncover a subtle DNA strand bias for CRISPRi in transcribed regions with implications for screen design and analysis. Together, we provide an accessible data resource, predesigned single guide RNAs for targeting 3,275,697 ENCODE SCREEN candidate CREs with CRISPRi and screening guidelines to accelerate functional characterization of the noncoding genome.
Cis-regulatory elements (CREs) control gene expression, orchestrating tissue identity, developmental timing, and stimulus responses, which collectively define the thousands of unique cell types in the body. While there is great potential for strategically incorporating CREs in therapeutic or biotechnology applications that require tissue specificity, there is no guarantee that an optimal CRE for an intended purpose has arisen naturally through evolution. Here, we present a platform to engineer and validate synthetic CREs capable of driving gene expression with programmed cell type specificity. We leverage innovations in deep neural network modeling of CRE activity across three cell types, efficient in silico optimization, and massively parallel reporter assays (MPRAs) to design and empirically test thousands of CREs. Through in vitro and in vivo validation, we show that synthetic sequences outperform natural sequences from the human genome in driving cell type-specific expression. Synthetic sequences leverage unique sequence syntax to promote activity in the on-target cell type and simultaneously reduce activity in off-target cells. Together, we provide a generalizable framework to prospectively engineer CREs and demonstrate the required literacy to write regulatory code that is fit-for-purpose in vivo across vertebrates.
The ENCODE Consortium’s efforts to annotate non-coding, cis -regulatory elements (CREs) have advanced our understanding of gene regulatory landscapes which play a major role in health and disease. Pooled, non-coding CRISPR screens are a promising approach for systematically investigating gene regulatory mechanisms. Here, the ENCODE Functional Characterization Centers report 109 screens comprising 346,970 individual perturbations across 13.3Mb of the genome, using a variety of methods, readouts, and statistical analyses. Across 332 functionally confirmed CRE-gene links, we identify principles for screening endogenous, non-coding elements for causal regulatory mechanisms. Nearly all CREs show strong evidence of open chromatin, and targeting accessibility peak summits is a critical component of our proposed sgRNA design rules. We provide experimental guidelines to accurately detect CREs with variable, often low, transcriptional effects. We discover a previously undescribed DNA strand-bias for CRISPRi in transcribed regions with implications for screen design and analysis. Benchmarking five screen analysis tools, we find CASA produces the most conservative CRE calls and is robust to artifacts of low-specificity sgRNAs. Together, we provide an accessible data resource, predesigned sgRNAs targeting 3,275,697 ENCODE SCREEN candidate CREs, and screening guidelines to accelerate functional characterization of the non-coding genome.
Effective interpretation of genome function and genetic variation requires a shift from epigenetic mapping of cis-regulatory elements (CREs) to characterization of endogenous function. We developed hybridization chain reaction fluorescence in situ hybridization coupled with flow cytometry (HCR–FlowFISH), a broadly applicable approach to characterize CRISPR-perturbed CREs via accurate quantification of native transcripts, alongside CRISPR activity screen analysis (CASA), a hierarchical Bayesian model to quantify CRE activity. Across >325,000 perturbations, we provide evidence that CREs can regulate multiple genes, skip over the nearest gene and display activating and/or silencing effects. At the cholesterol-level-associated FADS locus, we combine endogenous screens with reporter assays to exhaustively characterize multiple genome-wide association signals, functionally nominate causal variants and, importantly, identify their target genes. HCR–FlowFISH is a new approach to characterize CRISPR-perturbed cis-regulatory elements (CREs) via accurate quantification of native transcripts, alongside CRISPR activity screen analysis (CASA), a hierarchical Bayesian model to quantify CRE activity.
CRISPR screens for cis-regulatory elements (CREs) have shown unprecedented power to endogenously characterize the non-coding genome. To characterize CREs we developed HCR-FlowFISH (Hybridization Chain Reaction Fluorescent In-Situ Hybridization coupled with Flow Cytometry), which directly quantifies native transcripts within their endogenous loci following CRISPR perturbations of regulatory elements, eliminating the need for restrictive phenotypic assays such as growth or transcript-tagging. HCR-FlowFISH accurately quantifies gene expression across a wide range of transcript levels and cell types. We also developed CASA (CRISPR Activity Screen Analysis), a hierarchical Bayesian model to identify and quantify CRE activity. Using >270,000 perturbations, we identified CREs for GATA1, HDAC6, ERP29, LMO2, MEF2C, CD164, NMU, FEN1 and the FADS gene cluster. Our methods detect subtle gene expression changes and identify CREs regulating multiple genes, sometimes at different magnitudes and directions. We demonstrate the power of HCR-FlowFISH to parse genome-wide association signals by nominating causal variants and target genes.
Current genomics methods are designed to handle tens to thousands of samples but will need to scale to millions to match the pace of data and hypothesis generation in biomedical science. Here, we show that high efficiency at low cost can be achieved by leveraging general-purpose libraries for computing using graphics processing units (GPUs), such as PyTorch and TensorFlow. We demonstrate > 200-fold decreases in runtime and ~ 5–10-fold reductions in cost relative to CPUs. We anticipate that the accessibility of these libraries will lead to a widespread adoption of GPUs in computational genomics.
N-6-methyladenosine (m(6)A) is a dynamic, reversible, covalently modified ribonucleotide that occurs predominantly toward 3' ends of eukaryotic mRNAs and is essential for their proper function and regulation. In Arabidopsis thaliana, many RNAs contain at least one m(6)A site, yet the transcriptome-wide function of m(6)A remains mostly unknown. Here, we show that many m(6)A-modified mRNAs in Arabidopsis have reduced abundance in the absence of this mark. The decrease in abundance is due to transcript destabilization caused by cleavage occurring 4 or 5 nt directly upstream of unmodified m(6)A sites. Importantly, we also find that, upon agriculturally relevant salt treatment, m(6)A is dynamically deposited on and stabilizes transcripts encoding proteins required for salt and osmotic stress response. Overall, our findings reveal that m(6)A generally acts as a stabilizing mark through inhibition of site-specific cleavage in plant transcriptomes, and this mechanism is required for proper regulation of the salt-stress-responsive transcriptome.
In the version of this article initially published online, the fifth author's name was given as Alexander Amlie-Wolf. The correct name is Alexandre Amlie-Wolf. The error has been corrected in the print, PDF and HTML versions of this article.
Nonconserved linc-ADAL interacts with distinct cytoplasmic and nuclear factors to regulate human adipocyte differentiation and metabolism.
MicroRNA precursors (pre-miRNAs) are short hairpin RNAs that are rapidly processed into mature microRNAs (miRNAs) in the cytoplasm. Due to their low abundance in cells, sequencing-based studies of pre-miRNAs have been limited. We successfully enriched for and deep sequenced pre-miRNAs in human cells by capturing these RNAs during their interaction with Argonaute (AGO) proteins. Using this approach, we detected > 350 pre-miRNAs in human cells and > 250 pre-miRNAs in a reanalysis of a similar study in mouse cells. We uncovered widespread trimming and non-templated additions to the 3’ ends of pre- and mature miRNAs. Additionally, we created an index for microRNA precursor processing efficiency. This analysis revealed a subset of pre-miRNAs that produce low levels of mature miRNAs despite abundant precursors, including an annotated miRNA in the 5’ UTR of the DiGeorge syndrome critical region 8 mRNA transcript. This led us to search for AGO-associated stem-loops originating from other mRNA species, which identified hundreds of putative pre-miRNAs derived from human and mouse mRNAs. In summary, we provide a wealth of information on mammalian pre-miRNAs, and identify novel microRNA and microRNA-like elements localized in mRNAs.
Type II topoisomerases orchestrate proper DNA topology, and they are the targets of anti-cancer drugs that cause treatment-related leukemias with balanced translocations. Here, we develop a high-throughput sequencing technology to define TOP2 cleavage sites at single-base precision, and use the technology to characterize TOP2A cleavage genome-wide in the human K562 leukemia cell line. We find that TOP2A cleavage has functionally conserved local sequence preferences, occurs in cleavage cluster regions (CCRs), and is enriched in introns and lincRNA loci. TOP2A CCRs are biased toward the distal regions of gene bodies, and TOP2 poisons cause a proximal shift in their distribution. We find high TOP2A cleavage levels in genes involved in translocations in TOP2 poison–related leukemia. In addition, we find that a large proportion of genes involved in oncogenic translocations overall contain TOP2A CCRs. The TOP2A cleavage of coding and lincRNA genes is independently associated with both length and transcript abundance. Comparisons to ENCODE data reveal distinct TOP2A CCR clusters that overlap with marks of transcription, open chromatin, and enhancers. Our findings implicate TOP2A cleavage as a broad DNA damage mechanism in oncogenic translocations as well as a functional role of TOP2A cleavage in regulating transcription elongation and gene activation.
The giant, single-celled organism Stentor coeruleus has a long history as a model system for studying pattern formation and regeneration in single cells. Stentor [1, 2] is a heterotrichous ciliate distantly related to familiar ciliate models, such as Tetrahymena or Paramecium. The primary distinguishing feature of Stentor is its incredible size: a single cell is 1 mm long. Early developmental biologists, including T.H. Morgan [3], were attracted to the system because of its regenerative abilities if large portions of a cell are surgically removed, the remnant reorganizes into a normal-looking but smaller cell with correct proportionality [2, 3]. These biologists were also drawn to Stentor because it exhibits a rich repertoire of behaviors, including light avoidance, mechanosensitive contraction, food selection, and even the ability to habituate to touch, a simple form of learning usually seen in higher organisms [4]. While early microsurgical approaches demonstrated a startling array of regenerative and morphogenetic processes in this single-celled organism, Stentor was never developed as a molecular model system. We report the sequencing of the Stentor coeruleus macronuclear genome and reveal key features of the genome. First, we find that Stentor uses the standard genetic code, suggesting that ciliate-specific genetic codes arose after Stentor branched from other ciliates. We also discover that ploidy correlates with Stentor's cell size. Finally, in the Stentor genome, we discover the smallest spliceosomal introns reported for any species. The sequenced genome opens the door to molecular analysis of single-cell regeneration in Stentor.
Eukaryotic transcriptomes contain a major non-protein-coding component that includes precursors of small RNAs as well as long noncoding RNA (lncRNAs). Here, we utilized the mapping of ribosome footprints on RNAs to explore translational regulation of coding and noncoding RNAs in roots of Arabidopsis thaliana shifted from replete to deficient phosphorous (Pi) nutrition. Homodirectional changes in steady-state mRNA abundance and translation were observed for all but 265 annotated protein-coding genes. Of the translationally regulated mRNAs, 30% had one or more upstream ORF (uORF) that influenced the number of ribosomes on the principal protein-coding region. Nearly one-half of the 2,382 lncRNAs detected had ribosome footprints, including 56 with significantly altered translation under Pi-limited nutrition. The prediction of translated small ORFs (sORFs) by quantitation of translation termination and peptidic analysis identified lncRNAs that produce peptides, including several deeply evolutionarily conserved and significantly Pi-regulated lncRNAs. Furthermore, we discovered that natural antisense transcripts (NATs) frequently have actively translated sORFs, including five with low-Pi up-regulation that correlated with enhanced translation of the sense protein-coding mRNA. The data also confirmed translation of miRNA target mimics and lncRNAs that produce trans-acting or phased small-interfering RNA (tasiRNA/phasiRNAs). Mutational analyses of the positionally conserved sORF of TAS3a linked its translation with tasiRNA biogenesis. Altogether, this systematic analysis of ribosome-associated mRNAs and lncRNAs demonstrates that nutrient availability and translational regulation controls protein and small peptide-encoding mRNAs as well as a diverse cadre of regulatory RNAs.
The Arabidopsis thaliana root epidermis is comprised of two cell types, hair and nonhair cells, which differentiate from the same precursor. Although the transcriptional programs regulating these events are well studied, post-transcriptional factors functioning in this cell fate decision are mostly unknown. Here, we globally identify RNA-protein interactions and RNA secondary structure in hair and nonhair cell nuclei. This analysis reveals distinct structural and protein binding patterns across both transcriptomes, allowing identification of differential RNA binding protein (RBP) recognition sites. Using these sequences, we identify two RBPs that regulate hair cell development. Specifically, we find that SERRATE functions in a microRNA-dependent manner to inhibit hair cell fate, while also terminating growth of root hairs mostly independent of microRNA biogenesis. In addition, we show that GLYCINE-RICH PROTEIN 8 promotes hair cell fate while alleviating phosphate starvation stress. In total, this global analysis reveals post-transcriptional regulators of plant root epidermal cell fate.
SummaryThe impact of metabolic engineering on nontarget pathways and outcomes of metabolic engineering from different genomes are poorly understood questions. Therefore, squalene biosynthesis genes FARNESYL DIPHOSPHATE SYNTHASE (FPS) and SQUALENE SYNTHASE (SQS) were engineered via the Nicotiana tabacum chloroplast (C), nuclear (N) or both (CN) genomes to promote squalene biosynthesis. SQS levels were ~4300‐fold higher in C and CN lines than in N, but all accumulated ~150‐fold higher squalene due to substrate or storage limitations. Abnormal leaf and flower phenotypes, including lower pollen production and reduced fertility, were observed regardless of the compartment or level of transgene expression. Substantial changes in metabolomes of all lines were observed: levels of 65–120 unrelated metabolites, including the toxic alkaloid nicotine, changed by as much as 32‐fold. Profound effects of transgenesis on nontarget gene expression included changes in the abundance of 19 076 transcripts by up to 2000‐fold in CN; 7784 transcripts by up to 1400‐fold in N; and 5224 transcripts by as much as 2200‐fold in C. Transporter‐related transcripts were induced, and cell cycle‐associated transcripts were disproportionally repressed in all three lines. Transcriptome changes were validated by qRT‐PCR. The mechanism underlying these large changes likely involves metabolite‐mediated anterograde and/or retrograde signalling irrespective of the level of transgene expression or end product, due to imbalance of metabolic pools, offering new insight into both anticipated and unanticipated consequences of metabolic engineering.
Aging is a major risk factor for many neurodegenerative disorders. A key feature of aging biology that may underlie these diseases is cellular senescence. Senescent cells accumulate in tissues with age, undergo widespread changes in gene expression, and typically demonstrate altered, pro-inflammatory profiles. Astrocyte senescence has been implicated in neurodegenerative disease, and to better understand senescence-associated changes in astrocytes, we investigated changes in their transcriptome using RNA sequencing. Senescence was induced in human fetal astrocytes by transient oxidative stress. Brain-expressed genes, including those involved in neuronal development and differentiation, were downregulated in senescent astrocytes. Remarkably, several genes indicative of astrocytic responses to injury were also downregulated, including glial fibrillary acidic protein and genes involved in the processing and presentation of antigens by major histocompatibility complex class II proteins, while pro-inflammatory genes were upregulated. Overall, our findings suggest that senescence-related changes in the function of astrocytes may impact the pathogenesis of age-related brain disorders.