Gene editing has the potential to solve fundamental challenges in agriculture, biotechnology and human health. CRISPR-based gene editors derived from microorganisms, although powerful, often show notable functional tradeoffs when ported into non-native environments, such as human cells1. Artificial-intelligence-enabled design provides a powerful alternative with the potential to bypass evolutionary constraints and generate editors with optimal properties. Here, using large language models2 trained on biological diversity at scale, we demonstrate successful precision editing of the human genome with a programmable gene editor designed with artificial intelligence. To achieve this goal, we curated a dataset of more than 1 million CRISPR operons through systematic mining of 26 terabases of assembled genomes and metagenomes. We demonstrate the capacity of our models by generating 4.8× the number of protein clusters across CRISPR-Cas families found in nature and tailoring single-guide RNA sequences for Cas9-like effector proteins. Several of the generated gene editors show comparable or improved activity and specificity relative to SpCas9, the prototypical gene editing effector, while being 400 mutations away in sequence. Finally, we demonstrate that an artificial-intelligence-generated gene editor, denoted as OpenCRISPR-1, exhibits compatibility with base editing. We release OpenCRISPR-1 to facilitate broad, ethical use across research and commercial applications.
Disease resistance genes in livestock provide health benefits to animals and opportunities for farmers to meet the growing demand for affordable, high-quality protein. Previously, researchers used gene editing to modify the porcine CD163 gene and demonstrated resistance to a harmful virus that causes porcine reproductive and respiratory syndrome (PRRS). To maximize potential benefits, this disease resistance trait needs to be present in commercially relevant breeding populations for multiplication and distribution of pigs. Toward this goal, a first-of-its-kind, scaled gene editing program was established to introduce a single modified CD163 allele into four genetically diverse, elite porcine lines. This effort produced healthy pigs that resisted PRRS virus infection as determined by macrophage and animal challenges. This founder population will be used for additional disease and trait testing, multiplication, and commercial distribution upon regulatory approval. Applying CRISPR-Cas to eliminate a viral disease represents a major step toward improving animal health.
Gene editing has the potential to solve fundamental challenges in agriculture, biotechnology, and human health. CRISPR-based gene editors derived from microbes, while powerful, often show significant functional tradeoffs when ported into non-native environments, such as human cells. Artificial intelligence (AI) enabled design provides a powerful alternative with potential to bypass evolutionary constraints and generate editors with optimal properties. Here, using large language models (LLMs) trained on biological diversity at scale, we demonstrate the first successful precision editing of the human genome with a programmable gene editor designed with AI. To achieve this goal, we curated a dataset of over one million CRISPR operons through systematic mining of 26 terabases of assembled genomes and meta-genomes. We demonstrate the capacity of our models by generating 4.8x the number of protein clusters across CRISPR-Cas families found in nature and tailoring single-guide RNA sequences for Cas9-like effector proteins. Several of the generated gene editors show comparable or improved activity and specificity relative to SpCas9, the prototypical gene editing effector, while being 400 mutations away in sequence. Finally, we demonstrate an AI-generated gene editor, denoted as OpenCRISPR-1, exhibits compatibility with base editing. We release OpenCRISPR-1 publicly to facilitate broad, ethical usage across research and commercial applications.
Abstract (Note: a correction was made to the Index Reverse primer 20 Sep2023) CRISPR-Cas9 RNA-guided endonucleases are widely used in genome engineering, yet information on biochemical and cellular off-target cleavage activity is lacking. Here, we present a biochemical method, based on the selective enrichment and identification of adapter-tagged DNA ends by sequencing \(SITE-Seq). SITE-Seq can be used to identify off-target cleavage sites within a genomic DNA sample. This protocol details the preparation of SITE-Seq libraries for high throughput Next Generation Sequencing on the Illumina platform.
RNA-guided CRISPR-Cas9 endonucleases are widely used for genome engineering, but our understanding of Cas9 specificity remains incomplete. Here, we developed a biochemical method (SITE-Seq), using Cas9 programmed with single-guide RNAs (sgRNAs), to identify the sequence of cut sites within genomic DNA. Cells edited with the same Cas9-sgRNA complexes are then assayed for mutations at each cut site using amplicon sequencing. We used SITE-Seq to examine Cas9 specificity with sgRNAs targeting the human genome. The number of sites identified depended on sgRNA sequence and nuclease concentration. Sites identified at lower concentrations showed a higher propensity for off-target mutations in cells. The list of off-target sites showing activity in cells was influenced by sgRNP delivery, cell type and duration of exposure to the nuclease. Collectively, our results underscore the utility of combining comprehensive biochemical identification of off-target sites with independent cell-based measurements of activity at those sites when assessing nuclease activity and specificity.
The target DNA specificity of the CRISPR-associated genome editor nuclease Cas9 is determined by complementarity to a 20-nucleotide segment in its guide RNA. However, Cas9 can bind and cleave partially complementary off-target sequences, which raises safety concerns for its use in clinical applications. Here, we report crystallographic structures of Cas9 bound to bona fide off-target substrates, revealing that off-target binding is enabled by a range of noncanonical base-pairing interactions within the guide:off-target heteroduplex. Off-target substrates containing single-nucleotide deletions relative to the guide RNA are accommodated by base skipping or multiple noncanonical base pairs rather than RNA bulge formation. Finally, PAM-distal mismatches result in duplex unpairing and induce a conformational change in the Cas9 REC lobe that perturbs its conformational activation. Together, these insights provide a structural rationale for the off-target activity of Cas9 and contribute to the improved rational design of guide RNAs and off-target prediction algorithms.
The off-target activity of the CRISPR-associated nuclease Cas9 is a potential concern for therapeutic genome editing applications. Although high-fidelity Cas9 variants have been engineered, they exhibit varying efficiencies and have residual off-target effects, limiting their applicability. Here, we show that CRISPR hybrid RNA-DNA (chRDNA) guides provide an effective approach to increase Cas9 specificity while preserving on-target editing activity. Across multiple genomic targets in primary human T cells, we show that 2'-deoxynucleotide (dnt) positioning affects guide activity and specificity in a target-dependent manner and that this can be used to engineer chRDNA guides with substantially reduced off-target effects. Crystal structures of DNA-bound Cas9-chRDNA complexes reveal distorted guide-target duplex geometry and allosteric modulation of Cas9 conformation. These structural effects increase specificity by perturbing DNA hybridization and modulating Cas9 activation kinetics to disfavor binding and cleavage of off-target substrates. Overall, these results pave the way for utilizing customized chRDNAs in clinical applications.
Type I CRISPR–Cas systems are the most abundant adaptive immune systems in bacteria and archaea 1 , 2 . Target interference relies on a multi-subunit, RNA-guided complex called Cascade 3 , 4 , which recruits a trans-acting helicase-nuclease, Cas3, for target degradation 5 – 7 . Type I systems have rarely been used for eukaryotic genome engineering applications owing to the relative difficulty of heterologous expression of the multicomponent Cascade complex. Here, we fuse Cascade to the dimerization-dependent, non-specific FokI nuclease domain 8 – 11 and achieve RNA-guided gene editing in multiple human cell lines with high specificity and efficiencies of up to ~50%. FokI–Cascade can be reconstituted via an optimized two-component expression system encoding the CRISPR-associated (Cas) proteins on a single polycistronic vector and the guide RNA (gRNA) on a separate plasmid. Expression of the full Cascade–Cas3 complex in human cells resulted in targeted deletions of up to ~200 kb in length. Our work demonstrates that highly abundant, previously untapped type I CRISPR–Cas systems can be harnessed for genome engineering applications in eukaryotic cells.
Table S8. Primer sequences used for on- and off-target deep sequencing analysis for all genomic editing sites analyzed by SITE-Seq. (XLSX 48 kb)
Background: The development of CRISPR genome editing has transformed biomedical research. Most applications reported thus far rely upon the Cas9 protein from Streptococcus pyogenes SF370 (SpyCas9). With many RNA guides, wildtype SpyCas9 can induce significant levels of unintended mutations at near-cognate sites, necessitating substantial efforts toward the development of strategies to minimize off-target activity. Although the genome-editing potential of thousands of other Cas9 orthologs remains largely untapped, it is not known how many will require similarly extensive engineering to achieve single-site accuracy within large genomes. In addition to its off-targeting propensity, SpyCas9 is encoded by a relatively large open reading frame, limiting its utility in applications that require size-restricted delivery strategies such as adeno-associated virus vectors. In contrast, some genome-editing-validated Cas9 orthologs are considerably smaller and therefore better suited for viral delivery. Results: Here we show that wildtype NmeCas9, when programmed with guide sequences of the natural length of 24 nucleotides, exhibits a nearly complete absence of unintended editing in human cells, even when targeting sites that are prone to off-target activity with wildtype SpyCas9. We also validate at least six variant protospacer adjacent motifs (PAMs), in addition to the preferred consensus PAM (5'-N(4)GATT-3'), for NmeCas9 genome editing in human cells. Conclusions: Our results show that NmeCas9 is a naturally high-fidelity genome-editing enzyme and suggest that additional Cas9 orthologs may prove to exhibit similarly high accuracy, even without extensive engineering.