Allele conversion describes a process where a heterozygous variant is made homozygous. Recently, it has been shown that allele conversion can be triggered by DNA damage at the heterozygous site. This process has the potential to repair pathogenic heterozygous mutations; however, the efficiency is low. Here, we endeavoured to understand the mechanism underlying allele conversion, ultimately to raise allele conversion efficiency to functionally relevant levels. To test this, we developed a Compound Heterozygous Allele Conversion Reporter (CHACR) cell line. This line comprises knocked-in fluorescent protein encoding genes, with heterozygous inactivating mutations resulting in different fluorescence profiles from each allele. These mutations create protospacer adjacent motifs (PAM) for Cas9 recognition, where allele-specific gRNAs (AS-gRNAs) target the heterozygous mutations. We showed that applying these AS-gRNAs with either Cas9 nuclease or Cas9(D10A) nickase can recover mCherry fluorescence. Sorting and sequencing these fluorescent cells revealed wild-type sequences, suggesting allele conversion repaired the mutation using the homologous allele as a template. Allele conversion can also be triggered using an adenine base editor with an AS-gRNA, and this allele conversion mechanism can be manipulated by inhibiting DNA-PKcs or overexpressing RAD51. This work introduces a model for measuring allele conversion, and modifiers of this mechanism.
CRISPR/Cas genome editing technologies enable effective and controlled genetic modifications; however, off-target effects remain a significant concern, particularly in clinical applications. Experimental and in silico methods are developed to predict potential off-target sites (OTS), including deep learning based methods, which can automatically and comprehensively learn sequence features, offer a promising tool for OTS prediction. Here, this work reviews the existing OTS prediction tools with an emphasis on deep learning methods, characterizes datasets used for deep learning training and testing, and evaluates six deep learning models -CRISPR-Net, CRISPR-IP, R-CRISPR, CRISPR-M, CrisprDNT, and Crispr-SGRU -using six public datasets and validates OTS data from the CRISPRoffT database. Performance of these models is assessed using standardized metrics, such as Precision, Recall, F1 score, MCC, AUROC and PRAUC. This work finds that incorporating validated OTS datasets into model training enhanced overall model performance, and improved robustness of prediction, particularly with highly imbalanced datasets. While no model consistently outperforms other models across all scenarios, CRISPR-Net, R-CRISPR, and Crispr-SGRU show strong overall performance. This analysis demonstrates the importance of integrating high-quality validated OTS data with advanced deep learning architectures to improve CRISPR/Cas off-target site predictions, ensuring safer genome editing applications.
Design of guide RNA (gRNA) with high efficiency and specificity is vital for successful application of the CRISPR gene editing technology. Although many machine learning (ML) and deep learning (DL)-based tools have been developed to predict gRNA activities, a systematic and unbiased evaluation of their predictive performance is still needed. Here, we provide a brief overview of in silico tools for CRISPR design and assess the CRISPR datasets and statistical metrics used for evaluating model performance. We benchmark seven ML and DL-based CRISPR-Cas9 editing efficiency prediction tools across nine CRISPR datasets covering six cell types and three species. The DL models CRISPRon and DeepHF outperform the other models exhibiting greater accuracy and higher Spearman correlation coefficient across multiple datasets. We compile all CRISPR datasets and in silico prediction tools into a GuideNet resource web portal, aiming to facilitate and streamline the sharing of CRISPR datasets. Furthermore, we summarize features affecting CRISPR gene editing activity, providing important insights into model performance and the further development of more accurate CRISPR prediction models.
Genome-wide association studies typically identify hundreds to thousands of loci, many of which harbor multiple independent peaks, each parsimoniously assumed to be due to the activity of a single causal variant. Fine-mapping of such variants has become a priority and since most associations are located within regulatory regions, it is also assumed that they colocalize with regulatory variants that influence the expression of nearby genes. Here we examine these assumptions by using a moderate throughput expression CROPseq protocol in which Cas9 nuclease is used to induce small insertions and deletions across the credible set of SNPs that may account for expression quantitative trait loci (eQTL) for genes associated with inflammatory bowel disease (IBD). Of the 4,384 SNPs targeted in 88 loci (an average of 50 per locus), 439 were significant and further examined for validation. From these, 98 significantly altered target gene expression in HL-60 myeloid cell line, 74 in induced macrophages from these HL-60 cells, and 78 in induced neutrophils for a total of 201 validated effects (46%), 43 of which were observed in at least two of the cell types. Considering the observed sensitivity and specificity of the controls, we estimate that there are at least 150 true positives per cell type, an average of almost 2.4 for each of the 64 eQTL for which putative causal variants have been fine-mapped. This implies that haplotype effects are likely to explain many of the associations. We also demonstrate that the same approach can be used to investigate the activity of very rare variants in regulatory regions for 89 genes, providing a rapid strategy for establishing clinical relevance of non-coding mutations.
Recent studies have mapped substantial regional differences between human proximal and distal airway cell types, including basal cells (BCs) - the primary stem cell population in human adult airways. Regionally distinct airway basal stem cells derived from human induced pluripotent stem cells (hiPSCs) have broad applications in regenerative medicine and airway disease studies. Here, we report that the NOGGIN-BMP signaling axis is critical for proximal-distal regional patterning of hiPSC-derived lung progenitors. Continuous BMP inhibition (through NOGGIN) and tapering WNT activation efficiently generate BCs that resemble those in human proximal airways at both molecular and functional levels. These hiPSC-derived proximal basal cells (defined as proximal iBCs) are capable of self-renewal and competent differentiation, both in vitro and in vivo , into the full repertoire of specialized cell types found in normal human proximal airways, including ionocytes and pulmonary neuroendocrine cells. Conversely, BMP activation and tapering WNT generate airway and BCs resembling those in human distal airways, with limited ionocyte differentiation potential. The progeny of proximal iBCs derived from hiPSCs with G551D mutation in the CFTR gene recapitulate the ionocyte phenotype characteristic of cystic fibrosis.
Bacteriocins are broad or narrow-spectrum antimicrobial compounds that have received significant scientific attention due to their potential to treat infections caused by antibiotic-resistant pathogenic bacteria. The genome of Bifidobacterium pseudocatenulatum MM0196, an antimicrobial-producing, fecal isolate from a healthy pregnant woman, was shown to contain a gene cluster predicted to encode Pseudocin 196, a novel lantibiotic, in addition to proteins involved in its processing, transport and immunity. Following antimicrobial assessment against various indicator strains, protease-sensitive Pseudocin 196 was purified to homogeneity from cell-free supernatant. MALDI TOF mass spectrometry confirmed that the purified antimicrobial compound corresponds to a molecular mass of 2679 Da, which is consistent with that deduced from its genetic origin. Pseudocin 196 is classified as a lantibiotic based on its similarity to lacticin 481, a lanthionine ring-containing lantibiotic produced by Lactococcus lactis. Pseudocin 196, the first reported bacteriocin produced by a B. pseudocatenulatum species of human origin, was shown to inhibit clinically relevant pathogens, such as Clostridium spp. and Streptococcus spp. thereby highlighting the potential application of this strain as a probiotic to treat and prevent bacterial infections.
Bifidobacteria are commensal microorganisms that typically inhabit the mammalian gut, including that of humans. As they may be vertically transmitted, they commonly colonize the human intestine from the very first day following birth and may persist until adulthood and old age, although generally at a reduced relative abundance and prevalence compared to infancy. The ability of bifidobacteria to persist in the human intestinal environment has been attributed to genes involved in adhesion to epithelial cells and the encoding of complex carbohydrate-degrading enzymes. Recently, a putative mucin-degrading glycosyl hydrolase belonging to the GH136 family and encoded by the perB gene has been implicated in gut persistence of certain bifidobacterial strains. In the current study, to better characterize the function of this gene, a comparative genomic analysis was performed, revealing the presence of perB homologues in just eight bifidobacterial species known to colonize the human gut, including Bifidobacterium bifidum and Bifidobacterium longum subsp. longum strains, or in non-human primates. Mucin-mediated growth and adhesion to human intestinal cells, in addition to a rodent model colonization assay, were performed using B. bifidum PRL2010 as a perB prototype and its isogenic perB-insertion mutant. These results demonstrate that perB inactivation reduces the ability of B. bifidum PRL2010 to grow on and adhere to mucin, as well as to persist in the rodent gut niche. These results corroborate the notion that the perB gene is one of the genetic determinants involved in the persistence of B. bifidum PRL2010 in the human gut.
Members of the genus Bifidobacterium are commonly found in the human gut and are known to utilize complex carbohydrates that are indigestible by the human host. Members of the Bifidobacterium longum subsp. longum taxon can metabolize various plant-derived carbohydrates common to the human diet. To metabolize such polysaccharides, which include arabinoxylan, bifidobacteria need to encode appropriate carbohydrate-active enzymes in their genome. In the current study, we describe two GH43 family enzymes, denoted here as AxuA and AxuB, which are encoded by B. longum subsp. longum NCIMB 8809 and are shown to be required for cereal-derived arabinoxylan metabolism by this strain. Based on the observed hydrolytic activity of AxuA and AxuB, assessed by employing various synthetic and natural substrates, and based on in silico analyses, it is proposed that both AxuA and AxuB represent extracellular α-L-arabinofuranosidases with distinct substrate preferences. The variable presence of the axuA and axuB genes and other genes previously described to be involved in the metabolism of arabinose-containing glycans can in the majority cases explain the (in)ability of individual B. longum subsp. longum strains to grow on cereal-derived arabinoxylans and arabinan.
Wiskott-Aldrich syndrome (WAS) is a severe X-linked primary immunodeficiency resulting from a diversity of mutations distributed across all 12 exons of the WAS gene. WAS encodes a hematopoietic-specific and developmentally regulated cytoplasmic protein (WASp). The objective of this study was to develop a gene correction strategy potentially applicable to most WAS patients by employing nuclease-mediated, site-specific integration of a corrective WAS gene sequence into the endogenous WAS chromosomal locus. In this study, we demonstrate the ability to target the integration of WAS2-12-containing constructs into intron 1 of the endogenous WAS gene of primary CD34+ hematopoietic stem and progenitor cells (HSPCs), as well as WASp-deficient B cell lines and WASp-deficient primary T cells. This intron 1 targeted integration (TI) approach proved to be quite efficient and restored WASp expression in treated cells. Furthermore, TI restored WASp-dependent function to WAS patient T cells. Edited CD34+ HSPCs exhibited the capacity for multipotent differentiation to various hematopoietic lineages in vitro and in transplanted immunodeficient mice. This methodology offers a potential editing approach for treatment of WAS using patient’s CD34+ cells.
The ability to answer causal questions is crucial in many domains, as causal inference allows one to understand the impact of interventions. In many applications, only a single intervention is possible at a given time. However, in some important areas, multiple interventions are concurrently applied. Disentangling the effects of single interventions from jointly applied interventions is a challenging task -- especially as simultaneously applied interventions can interact. This problem is made harder still by unobserved confounders, which influence both treatments and outcome. We address this challenge by aiming to learn the effect of a single-intervention from both observational data and sets of interventions. We prove that this is not generally possible, but provide identification proofs demonstrating that it can be achieved under non-linear continuous structural causal models with additive, multivariate Gaussian noise -- even when unobserved confounders are present. Importantly, we show how to incorporate observed covariates and learn heterogeneous treatment effects. Based on the identifiability proofs, we provide an algorithm that learns the causal model parameters by pooling data from different regimes and jointly maximizing the combined likelihood. The effectiveness of our method is empirically demonstrated on both synthetic and real-world data.
Cystic fibrosis (CF) is a monogenic disease caused by impaired production and/or function of the cystic fibrosis transmembrane conductance regulator (CFTR) protein. Although we have previously shown correction of the most common pathogenic mutation, there are many other pathogenic mutations throughout the CF gene. An autologous airway stem cell therapy in which the CFTR cDNA is precisely inserted into the CFTR locus may enable the development of a durable cure for almost all CF patients, irrespective of the causal mutation. Here, we use CRISPR/Cas9 and two adeno-associated viruses (AAV) carrying the two halves of the CFTR cDNA to sequentially insert the full CFTR cDNA along with a truncated CD19 (tCD19) enrichment tag in upper airway basal stem cells (UABCs) and human bronchial basal stem cells (HBECs). The modified cells were enriched to obtain 60-80% tCD19 + UABCs and HBECs from 11 different CF donors with a variety of mutations. Differentiated epithelial monolayers cultured at air-liquid interface showed restored CFTR function that was >70% of the CFTR function in non-CF controls. Thus, our study enables the development of a therapy for almost all CF patients, including patients who cannot be treated using recently approved modulator therapies.
CRISPR-Cas technology has revolutionized gene editing, but concerns remain due to its propensity for off-target interactions. This, combined with genotoxicity related to both CRISPR-Cas9-induced double-strand breaks and transgene delivery, poses a significant liability for clinical genome-editing applications. Current best practice is to optimize genome-editing parameters in preclinical studies. However, quantitative tools that measure off-target interactions and genotoxicity are costly and time-consuming, limiting the practicality of screening large numbers of potential genome-editing reagents and conditions. Here, we show that flow-based imaging facilitates DNA damage characterization of hundreds of human hematopoietic stem and progenitor cells per minute after treatment with CRISPR-Cas9 and recombinant adeno-associated virus serotype 6. With our web-based platform that leverages deep learning for image analysis, we find that greater DNA damage response is observed for guide RNAs with higher genome-editing activity, differentiating even single on-target guide RNAs with different levels of off-target interactions. This work simplifies the characterization and screening process of genome-editing parameters toward enabling safer and more effective gene-therapy applications.
Adeno-associated virus vectors are the most used delivery method for liver-directed gene editing. Still, they are associated with significant disadvantages that can compromise the safety and efficacy of therapies. Here, we investigate the effects of electroporating CRISPR-Cas9 as mRNA and ribonucleoproteins (RNPs) into primary hepatocytes regarding on-target activity, specificity, and cell viability. We observed a transfection efficiency of >60% and on-target insertions/deletions (indels) of up to 95% in primary mouse hepatocytes electroporated with Cas9 RNPs targeting Hpd, the gene encoding hydroxyphenylpyruvate dioxygenase. In primary human hepatocytes, we observed on-target indels of 52.4% with Cas9 RNPs and >65% viability after electroporation. These results establish the impact of using electroporation to deliver Cas9 RNPs into primary hepatocytes as a highly efficient and potentially safe approach for therapeutic liver-directed gene editing and the production of liver disease models.
Targeted genome editing in hematopoietic stem and progenitor cells (HSPCs) using CRISPR/Cas9 can potentially provide a permanent cure for hematologic diseases. However, the utility of CRISPR/Cas9 systems for therapeutic genome editing can be compromised by their off-target effects. In this chapter, we outline the procedures for CRISPR/Cas9 off-target identification and validation in HSPCs. This method is broadly applicable to diverse CRISPR/Cas9 systems and cell types. Using this protocol, researchers can perform computational prediction and experimental identification of potential off-target sites followed by off-target activity quantification by next-generation sequencing.
An amendment to this paper has been published and can be accessed via a link at the top of the paper.
Preclinical efficacy and safety data in mice provide support for ex vivo β-globin gene correction to treat patients with sickle cell disease.
Targeted DNA correction of disease-causing mutations in hematopoietic stem and progenitor cells (HSPCs) may enable the treatment of genetic diseases of the blood and immune system. It is now possible to correct mutations at high frequencies in HSPCs by combining CRISPR/Cas9 with homologous DNA donors. Because of the precision of gene correction, these approaches preclude clonal tracking of gene-targeted HSPCs. Here, we describe Tracking Recombination Alleles in Clonal Engraftment using sequencing (TRACE-Seq), a methodology that utilizes barcoded AAV6 donor template libraries, carrying in-frame silent mutations or semi-randomized nucleotides outside the coding region, to track the in vivo lineage contribution of gene-targeted HSPC clones. By targeting the HBB gene with an AAV6 donor template library consisting of ~20,000 possible unique exon 1 in-frame silent mutations, we track the hematopoietic reconstitution of HBB targeted myeloid-skewed, lymphoid-skewed, and balanced multi-lineage repopulating human HSPC clones in mice. We anticipate this methodology could potentially be used for HSPC clonal tracking of Cas9 RNP and AAV6-mediated gene targeting outcomes in translational and basic research settings.
Rewiring of host cytokine networks is a key feature of inflammatory bowel diseases (IBD) such as Crohn's disease (CD). Th1-type cytokines-IFN-γ and TNF-α-occupy critical nodes within these networks and both are associated with disruption of gut epithelial barrier function. This may be due to their ability to synergistically trigger the death of intestinal epithelial cells (IECs) via largely unknown mechanisms. In this study, through unbiased kinome RNAi and drug repurposing screens we identified JAK1/2 kinases as the principal and nonredundant drivers of the synergistic killing of human IECs by IFN-γ/TNF-α. Sensitivity to IFN-γ/TNF-α-mediated synergistic IEC death was retained in primary patient-derived intestinal organoids. Dependence on JAK1/2 was confirmed using genetic loss-of-function studies and JAK inhibitors (JAKinibs). Despite the presence of biochemical features consistent with canonical TNFR1-mediated apoptosis and necroptosis, IFN-γ/TNF-α-induced IEC death was independent of RIPK1/3, ZBP1, MLKL or caspase activity. Instead, it involved sustained activation of JAK1/2-STAT1 signalling, which required a nonenzymatic scaffold function of caspase-8 (CASP8). Further modelling in gut mucosal biopsies revealed an intercorrelated induction of the lethal CASP8-JAK1/2-STAT1 module during ex vivo stimulation of T cells. Functional studies in CD-derived organoids using inhibitors of apoptosis, necroptosis and JAKinibs confirmed the causative role of JAK1/2-STAT1 in cytokine-induced death of primary IECs. Collectively, we demonstrate that TNF-α synergises with IFN-γ to kill IECs via the CASP8-JAK1/2-STAT1 module independently of canonical TNFR1 and cell death signalling. This non-canonical cell death pathway may underpin immunopathology driven by IFN-γ/TNF-α in diverse autoinflammatory diseases such as IBD, and its inhibition may contribute to the therapeutic efficacy of anti-TNFs and JAKinibs.