Abstract RNA modifications are important for RNA structure, stability, and ribosome function, but their identification and localisation remains challenging. Oxford Nanopore direct RNA sequencing (DRS) enables modification-agnostic detection in native RNA, but existing tool benchmarks have focused almost exclusively on m6A in eukaryotic mRNA, leaving multi-modification tool performance in bacterial systems largely untested. Here, we benchmark ten RNA modification detection tools spanning signal-comparison, error-rate, and hybrid approaches on Escherichia coli K-12 MG1655 16S and 23S rRNA, which harbour 11 and 25 known modified sites, respectively, across 17 modification types. Using native RNA and in vitro transcribed (IVT) unmodified RNA, we evaluate performance across 25 coverage levels (5× to 1000×). DiffErr and JACUSA2 showed the strongest discrimination performance (AUROC >0.9 on both 16S and 23S rRNA), with DiffErr achieving the highest F1 score on 16S and JACUSA2 showing the most consistent precision-recall balance across both rRNAs. Both tools achieved full transcript-wide scoring and, along with DRUMMER, exact positional localisation. Several other tools produced no output at many rRNA positions, and restricting evaluation to reported positions inflated apparent performance. Signal-based tools showed a systematic 1-4 nucleotide 5ʹ offset from known modified positions, consistent with the ∼5-mer nucleotide stretch present in the read head of the nanopore; applying tool-specific offset corrections substantially improved per-site recovery and reduced false positives, substantially improving the performance of tools such as EpiNano and nanoDoc. At single-site resolution, no known modified site was recovered by all tools, and several m5C, m5U, and m6A sites were missed by the majority of tools. Tool combination analysis showed that pairing error-rate-based tools with offset-corrected signal-based tools improved site recovery beyond any individual tool, with the best three-tool combination recovering 30 of the 36 known sites while maintaining low false positive rates. These results establish that discrimination metrics (e.g. AUROC) alone are insufficient to evaluate modification detection tools: output completeness, positional precision, and per-modification-type sensitivity should be reported alongside standard benchmarking metrics.
Random Forest models are widely used in genomic data analysis and can offer insights into complex biological mechanisms, particularly when features influence the target in interactive, nonlinear, or nonadditive ways. Currently, some of the most efficient Random Forest methods in terms of computational speed are implemented in Python. However, many biologists use R for genomic data analysis, as R offers a unified platform for performing additional statistical analysis and visualization. Here, we present an R package, pyRforest, which integrates Python scikit-learn "RandomForestClassifier" algorithms into the R environment. pyRforest inherits the efficient memory management and parallelization of Python, and is optimized for classification tasks on large genomic datasets, such as those from RNA-seq. pyRforest offers several additional capabilities, including a novel rank-based permutation method for biomarker identification. This method can be used to estimate and visualize P-values for individual features, allowing the researcher to identify a subset of features for which there is robust statistical evidence of an effect. In addition, pyRforest includes methods for the calculation and visualization of SHapley Additive exPlanations values. Finally, pyRforest includes support for comprehensive downstream analysis for gene ontology and pathway enrichment. pyRforest thus improves the implementation and interpretability of Random Forest models for genomic data analysis by merging the strengths of Python with R. pyRforest can be downloaded at: https://www.github.com/tkolisnik/pyRforest with an associated vignette at https://github.com/tkolisnik/pyRforest/blob/main/vignettes/pyRforest-vignette.pdf.
The hourglass dolphin (Lagenorhynchus cruciger) is a small cetacean species of the Southern Ocean, with significance to iwi Māori (Māori tribes) of Aotearoa New Zealand as taonga (treasured/valued). Due to the remoteness and difficulty of surveying Antarctic waters, it remains one of the least-studied dolphin species. A recent stranding of an hourglass dolphin represented a rare opportunity to generate a genome assembly as a resource for future study into the conservation and evolutionary biology of this species. In this study, we present a high-quality genome assembly of an hourglass dolphin individual using a single sequencing platform, Oxford Nanopore Technologies, coupled with computationally efficient assembly methods. Our assembly strategy yielded a genome of high contiguity (N50 of 8.07 Mbp) and quality (98.3% BUSCO completeness). Compared to other Delphinoidea reference genomes, this assembly has fewer missing BUSCOs than any except Orcinus orca, more single-copy complete BUSCOs than any except Phocoena sinus, and 20% fewer duplicated BUSCOs than the average Delphinoidea reference genome. This suggests that it is one of the most complete and accurate marine mammal genomes to date. This study showcases the feasibility of a cost-effective mammalian genome assembly method, allowing for genomic data generation outside the traditional confines of academia and/or resource-rich genome assembly hubs, and facilitating the ability to uphold Indigenous data sovereignty. In the future, the genome assembly presented here will allow valuable insights into the past population size changes, adaptation, vulnerability to future climate change of the hourglass dolphin and related species.
Immune responses can have opposing effects in colorectal cancer (CRC), the balance of which may determine whether a cancer regresses, progresses, or potentially metastasizes. These effects are evident in CRC consensus molecular subtypes (CMS) where both CMS1 and CMS4 contain immune infiltrates yet have opposing prognoses. The microbiome has previously been associated with CRC and immune response in CRC but has largely been ignored in the CRC subtype discussion. We used CMS subtyping on surgical resections from patients and aimed to determine the contributions of the microbiome to the pleiotropic effects evident in immune-infiltrated subtypes. We integrated host gene-expression and meta-transcriptomic data to determine the link between immune characteristics and microbiome contributions in these subtypes and identified lipopolysaccharide (LPS) binding as a potential functional mechanism. We identified candidate bacteria with LPS properties that could affect immune response, and tested the effects of their LPS on cytokine production of peripheral blood mononuclear cells (PBMCs). We focused on Fusobacterium periodonticum and Bacteroides fragilis in CMS1, and Porphyromonas asaccharolytica in CMS4. Treatment of PBMCs with LPS isolated from these bacteria showed that F. periodonticum stimulates cytokine production in PBMCs while both B. fragilis and P. asaccharolytica had an inhibitory effect. Furthermore, LPS from the latter two species can inhibit the immunogenic properties of F. periodonticum LPS when co-incubated with PBMCs. We propose that different microbes in the CRC tumor microenvironment can alter the local immune activity, with important implications for prognosis and treatment response.
Parāoa Rēwena is a traditional sourdough bread that has been produced and consumed by the indigenous Māori of Aotearoa New Zealand for more than a century. We previously optimised the fermentation conditions (bulk fermentation temperature, bulk fermentation time, starter culture concentration, retarding time) of Rēwena sourdough bread using response surface methodology (RSM) to produce sourdough bread with high acidity, high specific volume, and low hardness. The present study aimed to characterise the sourdough during fermentation and the resultant Rēwena sourdough bread products by analysing the levels of organic acids (lactic acid, acetic acid), sugars (maltose, glucose, fructose), and the FODMAPs (total fructan content). The texture profile, specific volume, colour, water activity, and moisture content were also determined. The overall acceptability of the Rēwena sourdough bread was evaluated by consumer sensory panellists. As expected, organic acids increased during fermentation, likely due to the metabolism of sugars by the starter microorganisms and the long fermentation time (retarding) increased the levels of acids. Fructose and glucose levels decreased during sourdough fermentation while no changes were observed for maltose. The total fructan content in the sourdough bread decreased by approximately 74% after fermentation to 0.216 ± 0.016 g/100 g, resulting in the Rēwena sourdough bread being sufficiently low in total fructan content to be categorised as a low FODMAPs product. Texture analysis of the sourdough bread revealed that the acidification during fermentation resulted in hard sourdough bread. During storage, the acidity (pH, total titratable acidity), moisture content, and water activity of all tested formulations decreased. There were no marked changes in specific volume of the tested formulations of sourdough bread after storage for three days. Sensory evaluation indicated that panellists preferred Rēwena sourdough bread fermented at 30°C with 20% potato starter culture.
Accurate determination of animal diets is difficult.Methods such as molecular barcoding or metagenomics offer a promising approach allowing quantitative and sensitive detection of different taxa.Here we show that rapid and inexpensive quantification of animal, plant, and fungal content from stomach contents is possible through metagenomic sequencing with the portable Oxford Nanopore Technologies (ONT) MinION.Using an amplification-free approach, we profile the stomach contents from 24 wild-caught rats.We conservatively identify stomach contents from over 50 taxonomic orders, ranging across nine phyla, including plants, vertebrates, invertebrates, and fungi.This highlights the wide range of taxa that can be identified using this approach.We calibrate the accuracy of this method by comparing the characteristics of reads matching the ground-truth host genome (rat) to those matching non-rat non-microbial taxa (i.e.stomach content) and show that at the family-level taxon assignments are approximately 97.5% accurate.Some inaccuracies may arise from biases in sequence databases, for example due to overrepresentation of DNA sequences from commonly studied species.We suggest a means to decrease the effects of database biases on inferring taxon membership when using metagenomic approaches.Finally, we implement a constrained ordination analysis and show that it is possible to identify the sampling location of an individual rat within tens of kilometres based on stomach contents alone.This work establishes proof-of-principle for long-read metagenomic methods in quantitative analysis of the stomach contents of a terrestrial mammal.We show that stomach content can be quantified even with limited expertise using a simple, amplification free workflow and a relatively inexpensive and accessible next generation sequencing method.Continued increases in the accuracy and throughput of ONT sequencing, along with improved genomic databases, suggests that a metagenomic approach for quantification of stomach contents, and by proxy animal diets, will become an important method in the future.
In New Zealand, traditional Rēwena sourdough bread has been consumed among Māori for >100 years. Rēwena bread uses back-slopping fermentation with an indigenous potato starter culture (PSC). The PSC comprises the mother sourdough (RT1) which is refreshed through back-slopping by mixing wheat flour, mashed boiled potatoes, and water recovered from boiled potatoes. The mixture used to refresh the culture (RT1) is called RH1 by the indigenous Rēwena producers. The optimal growth conditions of RT1 were optimised for Rēwena sourdough bread using response surface methodology (RSM). The ideal activation conditions of RT1 comprised dough yield (DY) 200% with 10% (w/w) sugar fermented at 25°C/24 h. Production of Rēwena bread was optimised by evaluating the fermentation conditions for low pH, high total titratable acidity (TTA), low hardness, and high specific volume of the bread using RSM. The (fermentation) conditions evaluated were bulk fermentation temperature, fermentation time, starter concentration, and retarding time. Bulk fermentation temperature, fermentation time, starter culture, and retarding time impacted on bread acidity, hardness, and specific volume (p<0.05). The aw of bread decreased with increased fermentation temperature and retarding time. The moisture content of the bread increased with increased starter culture. Long retarding resulted in prolonged shelf-life of bread.
BACKGROUND:The gold standard treatment for locally advanced rectal cancer is total mesorectal excision after preoperative chemoradiotherapy. Response to chemoradiotherapy varies, with some patients completely responding to the treatment and some failing to respond at all. Identifying biomarkers of response to chemoradiotherapy could allow patients to avoid unnecessary treatment-associated morbidity rate. While previous studies have attempted to identify such biomarkers, none have reached clinical utility, which may be due to heterogeneity of the cancer. In this study, potential human gene and microbial biomarkers were explored in a cohort of rectal cancer patients who underwent chemoradiotherapy.METHODS:RNA sequencing was carried out on matched tumour and adjacent normal rectum biopsies from patients with rectal cancer with varying chemoradiotherapy responses treated between 2016 and 2019 at two institutions. Enriched genes and microbes from tumours of complete responders were compared with those from tumours of others with lesser response.RESULTS:In 39 patients analysed, enriched gene sets in complete responders indicate involvement of immune responses, including immunoglobulin production, B cell activation and response to bacteria (adjusted P values <0.050). Bacteria such as Ruminococcaceae bacterium and Bacteroides thetaiotaomicron were documented to be abundant in tumours of complete responders compared with all other patients (adjusted P value <0.100).CONCLUSION:These results identify potential genetic and microbial biomarkers of response to chemoradiotherapy in rectal cancer, as well as suggesting a potential mechanism of complete response to chemoradiotherapy that may benefit further testing in the laboratory.
BACKGROUND:Colorectal cancer (CRC) is a heterogeneous disease, with subtypes that have different clinical behaviours and subsequent prognoses. There is a growing body of evidence suggesting that right-sided colorectal cancer (RCC) and left-sided colorectal cancer (LCC) also differ in treatment success and patient outcomes. Biomarkers that differentiate between RCC and LCC are not well-established. Here, we apply random forest (RF) machine learning methods to identify genomic or microbial biomarkers that differentiate RCC and LCC.METHODS:RNA-seq expression data for 58,677 coding and non-coding human genes and count data for 28,557 human unmapped reads were obtained from 308 patient CRC tumour samples. We created three RF models for datasets of human genes-only, microbes-only, and genes-and-microbes combined. We used a permutation test to identify features of significant importance. Finally, we used differential expression (DE) and paired Wilcoxon-rank sum tests to associate features with a particular side.RESULTS:RF model accuracy scores were 90%, 70%, and 87% with area under curve (AUC) of 0.9, 0.76, and 0.89 for the human genomic, microbial, and combined feature sets, respectively. 15 features were identified as significant in the model of genes-only, 54 microbes in the model of microbes-only, and 28 genes and 18 microbes in the model with genes-and-microbes combined. PRAC1 expression was the most important feature for differentiating RCC and LCC in the genes-only model, with HOXB13, SPAG16, HOXC4, and RNLS also playing a role. Ruminococcus gnavus and Clostridium acetireducens were the most important in the microbial-only model. MYOM3, HOXC4, Coprococcus eutactus, PRAC1, lncRNA AC012531.25, Ruminococcus gnavus, RNLS, HOXC6, SPAG16 and Fusobacterium nucleatum were most important in the combined model.CONCLUSIONS:Many of the identified genes and microbes among all models have previously established associations with CRC. However, the ability of RF models to account for inter-feature relationships within the underlying decision trees may yield a more sensitive and biologically interconnected set of genomic and microbial biomarkers.
Quantifying SARS-like coronavirus (SL-CoV) evolution is critical to understanding the origins of SARS-CoV-2 and the molecular processes that could underlie future epidemic viruses. While genomic analyses suggest recombination was a factor in the emergence of SARS-CoV-2, few studies have quantified recombination rates among SL-CoVs. Here, we infer recombination rates of SL-CoVs from correlated substitutions in sequencing data using a coalescent model with recombination. Our computationally-efficient, non-phylogenetic method infers recombination parameters of both sampled sequences and the unsampled gene pools with which they recombine. We apply this approach to infer recombination parameters for a range of positive-sense RNA viruses. We then analyze a set of 191 SL-CoV sequences (including SARS-CoV-2) and find that ORF1ab and S genes frequently undergo recombination. We identify which SL-CoV sequence clusters have recombined with shared gene pools, and show that these pools have distinct structures and high recombination rates, with multiple recombination events occurring per synonymous substitution. We find that individual genes have recombined with different viral reservoirs. By decoupling contributions from mutation and recombination, we recover the phylogeny of non-recombined portions for many of these SL-CoVs, including the position of SARS-CoV-2 in this clonal phylogeny. Lastly, by analyzing >400,000 SARS-CoV-2 whole genome sequences, we show current diversity levels are insufficient to infer the within-population recombination rate of the virus since the pandemic began. Our work offers new methods for inferring recombination rates in RNA viruses with implications for understanding recombination in SARS-CoV-2 evolution and the structure of clonal relationships and gene pools shaping its origins.
Most animal mitochondrial genomes are small, circular and structurally conserved. However, recent work indicates that diverse taxa possess unusual mitochondrial genomes. In Isopoda, species in multiple lineages have atypical and rearranged mitochondrial genomes. However, more species of this speciose taxon need to be evaluated to understand the evolutionary origins of atypical mitochondrial genomes in this group. In this study, we report the presence of an atypical mitochondrial structure in the New Zealand endemic marine isopod, Isocladus armatus. Data from long- and short-read DNA sequencing suggest that I. armatus has two mitochondrial chromosomes. The first chromosome consists of two mitochondrial genomes that have been inverted and fused together in a circular form, and the second chromosome consists of a single mitochondrial genome in a linearized form. This atypical mitochondrial structure has been detected in other isopod lineages, and our data from an additional divergent isopod lineage (Sphaeromatidae) lends support to the hypothesis that atypical structure evolved early in the evolution of Isopoda. Additionally, we find that an asymmetrical site previously observed across many species within Isopoda is absent in I. armatus, but confirm the presence of two asymmetrical sites recently reported in two other isopod species.
Bacterial cells often respond to changes in the environment by modifying protein expression. This can be achieved through changes in transcriptional or translational activity, or both. Recent research has shed some light on how natural selection shapes overall protein expression. Still, little is known about how selection acts on transcription or translation individually. To address part of this question, we implement an experimental system which allows us to measure how genetic changes affect transcription only, excluding the effects on translation. We use this system to quantify changes in three regulatory phenotypes of the lacZ promoter: transcriptional activity, plasticity, and cell-to-cell variability. We compare these phenotypes from segregating variants that have been subject to natural selection, and random variants that have never been subjected to natural selection. We show that natural selection filters out mutations causing large changes in transcriptional levels from the lacZ promoter. Further, we detect directional selection acting on transcriptional plasticity in combinations of glucose, galactose and lactose environments. Focusing on cell-to-cell variability in transcription, we describe both directional and diversifying selection acting on this phenotype depending on the environment used. We also observe a link between the phylogeny of the environmental E. coli strains and high and low transcriptional noise levels in glucose which are mediated by just one or two SNPs. Our results thus provide new insight into how one of the most well-characterized bacterial promoters is shaped in nature by selection.
Parāroa Rēwena is a traditional Māori sourdough produced by fermentation using a potato starter culture. The microbial composition of the starter culture is not well characterised, despite the long history of this product. The morphological, physiological, biochemical and genetic tests were conducted to characterise 26 lactic acid bacteria (LAB) and 15 yeast isolates from a Parāroa Rēwena potato starter culture. The results of sugar fermentation tests, API 50 CHL tests, and API ID 32 C tests suggest the presence of four different LAB phenotypes and five different yeast phenotypes. 16S rRNA and 26S rRNA sequencing identified the LAB as Lacticaseibacillus paracasei and the yeast isolates as Saccharomyces cerevisiae, respectively. Multilocus sequence typing (MLST) of the L. paracasei isolates indicated that they had identical genotypes at the MLST loci, to L. paracasei subsp. paracasei IBB 3423 or L. paracasei subsp. paracasei F19. This study provides new insights into the microbial composition of the traditional sourdough Parāroa Rēwena starter culture.
Bacteria often respond to dynamically changing environments by regulating gene expression. Despite this regulation being critically important for growth and survival, little is known about how selection shapes gene regulation in natural populations. To better understand the role natural selection plays in shaping bacterial gene regulation, here we compare differences in the regulatory behaviour of naturally segregating promoter variants from Escherichia coli (which have been subject to natural selection) to randomly mutated promoter variants (which have never been exposed to natural selection). We quantify gene expression phenotypes (expression level, plasticity and noise) for hundreds of promoter variants across multiple environments and show that segregating promoter variants are enriched for mutations with minimal effects on expression level. In many promoters, we infer that there is strong selection to maintain high levels of plasticity, and direct selection to decrease or increase cell-to-cell variability in expression. Taken together, these results expand our knowledge of how gene regulation is affected by natural selection and highlight the power of comparing naturally segregating polymorphisms to de novo random mutations to quantify the action of selection.
DNA methylation in bacteria frequently serves as a simple immune system, allowing recognition of DNA from foreign sources, such as phages or selfish genetic elements. However, DNA methylation also affects other cell phenotypes in a heritable manner (i.e. epigenetically). While there are several examples of methylation affecting transcription in an epigenetic manner in highly localized contexts, it is not well-established how frequently methylation serves a more general epigenetic function over larger genomic scales. To address this question, here we use Oxford Nanopore sequencing to profile DNA modification marks in three natural isolates of Escherichia coli. We first identify the DNA sequence motifs targeted by the methyltransferases in each strain. We then quantify the frequency of methylation at each of these motifs across the entire genome in different growth conditions. We find that motifs in specific regions of the genome consistently exhibit high or low levels of methylation. Furthermore, we show that there are replicable and consistent differences in methylated regions across different growth conditions. This suggests that during growth, E. coli transiently differentiate into distinct methylation states that depend on the growth state, raising the possibility that measuring DNA methylation alone can be used to infer bacterial growth states without additional information such as transcriptome or proteome data. These results show the utility of using Oxford Nanopore sequencing as an economic means to infer DNA methylation status. They also provide new insights into the dynamics of methylation during bacterial growth and provide evidence of differentiated cell states, a transient analog to what is observed in the differentiation of cell types in multicellular organisms.
Genomes of 135 environmental Escherichia coli strains.Strain isolation:https://doi.org/10.1128/AEM.72.1.612-621.2006Whole genome sequencing:https://doi.org/10.7554/eLife.65366https://doi.org/10.1128/MRA.00222-20
The expanding knowledge of the variety of synthetic genetic elements has enabled the construction of new and more efficient genetic circuits and yielded novel insights into molecular mechanisms. However, context dependence, in which interactions between proximal (cis) or distal (trans) elements affect the behaviour of these elements, can reduce their general applicability or predictability. Genetic insulators, which mitigate unintended context-dependent cis-interactions, have been used to address this issue. One of the most commonly used genetic insulators is a self-splicing ribozyme called RiboJ, which can be used to decouple upstream 5’ UTR in mRNA from downstream sequences (e.g., open reading frames). Despite its general use as an insulator, there has been no systematic study quantifying the efficiency of RiboJ splicing or whether this autocatalytic activity is robust to trans- and cis-genetic context. Here, we determine the robustness of RiboJ splicing in the genetic context of six widely divergent E. coli strains. We also check for possible cis-effects by assessing two SNP versions close to the catalytic site of RiboJ. We show that mRNA molecules containing RiboJ are rapidly spliced even during rapid exponential growth and high levels of gene expression, with a mean efficiency of 98%. We also show that neither the cis- nor trans-genetic context has a significant impact on RiboJ activity, suggesting this element is robust to both cis- and trans-genetic changes.
Late in 2020, two genetically-distinct clusters of severe acute respiratory syndrome coronavirus 2 (SARS-CoV-2) with mutations of biological concern were reported, one in the United Kingdom and one in South Africa. Using a combination of data from routine surveillance, genomic sequencing and international travel we track the international dispersal of lineages B.1.1.7 and B.1.351 (variant 501Y-V2). We account for potential biases in genomic surveillance efforts by including passenger volumes from location of where the lineage was first reported, London and South Africa respectively. Using the software tool grinch (global report investigating novel coronavirus haplotypes), we track the international spread of lineages of concern with automated daily reports, Further, we have built a custom tracking website (cov-lineages.org/global_report.html) which hosts this daily report and will continue to include novel SARS-CoV-2 lineages of concern as they are detected.
Long read sequencing technologies now allow routine highly contiguous assembly of bacterial genomes. However, because of the lower accuracy of some long read data, it is often combined with short read data (e.g. Illumina), to improve assembly quality. There are a number of methods available for producing such hybrid assemblies. Here we use Illumina and Oxford Nanopore (ONT) data from 49 natural isolates of Escherichia coli to characterise differences in assembly accuracy for five assembly methods (Canu, Unicycler, Raven, Flye, and Redbean). We evaluate assembly accuracy using five metrics designed to measure structural accuracy and sequence accuracy (indel and substitution frequency). We assess structural accuracy by quantifying (1) the contiguity of chromosomes and plasmids; (2) the fraction of concordantly mapped Illumina reads withheld from the assembly; and (3) whether rRNA operons are correctly oriented. We assess indel and substitution frequency by quantifying (1) the fraction of open reading frames that appear truncated and (2) the number of variants that are called using Illumina reads only. Applying these assembly metrics to a large number of E. coli strains, we find that different assembly methods offer different advantages. In particular, we find that Unicycler assemblies have the highest sequence accuracy in non-repetitive regions, while Flye and Raven tend to be the most structurally accurate. In addition, we find that there are unidentified strain-specific characteristics that affect ONT consensus accuracy, despite individual reads having similar levels of accuracy. The differences in consensus accuracy of the ONT reads can preclude accurate assembly regardless of assembly method. These results provide quantitative insight into the best approaches for hybrid assembly of bacterial genomes and the expected levels of structural and sequence accuracy. They also show that there are intrinsic idiosyncratic strain-level differences that inhibit accurate long read bacterial genome assembly. However, we also show it is possible to diagnose problematic assemblies, even in the absence of ground truth, by comparing long-read first and short-read first assemblies. Author Notes All supporting data, code and protocols have been provided within the article or through supplementary data files. The supporting code is available from the GitHub repository https://github.com/GeorgiaBreckell/assembly_pipeline . nine supplementary figures and three supplementary tables are available with the online version of this article. Data summary Sequence data and genome assemblies for the natural isolates are available at https://www.ebi.ac.uk/ena/browser/view/PRJEB36951 . Genome assemblies for additional E. coli strains used here are available from NCBI: ( MG1655 , SE11 , REL606 , CFT073 , W , IA136 , O157:H7-EDL933 )
In this protocol we describe the steps for the design and production of dual guide RNAs (dgRNAs) from DNA oligos for the enrichment of specific sequences from mixed samples. These enriched molecules can then be sequenced using long read Oxford Nanopore Sequencing. The method as outlined and applied here is termed Bac-PULCE (Bacterial strain and antimicrobial resistance Profiling Using Long reads via CRISPR Enrichment) but can be applied to enrich specific DNA molecules from any samples for downstream ONT sequencing. For CRISPR-cas9 protocols including Bac-PULCE, we use T7 polymerase to transcribe the crRNA and tracrRNA to make dgRNA for Cas9. The two components of the dual guides are the crRNA (containing your variable 20 bp target plus a 22 bp constant region) and the tracrRNA (a 72 bp constant region). The protocol outlined here is a combination of the two protocols below: https://www.protocols.io/view/in-vitro-transcription-for-dgrna-3bpgimn /dx.doi.org/10.17504/protocols.io.3bpgimn https://international.neb.com/protocols/2013/04/02/standard-rna-synthesis-e2050 And the end product of this protocol is used for Nanopore protocol Cas9-mediated PCR-free enrichment as outlined by Oxford Nanopore - please refer to this protocol in particular for further important detail around sequencing the DNA library. The aim of this protocol is to design and produce crRNAs and tracRNAs from DNA oligos, and combine these with Cas9 to cut specific DNA sequences. We will add a T7 RNA polymerase binding site at the 5’ end of the crRNA and tracRNA to allow in vitro transcription, as well as a 3’ tracrRNA binding site to the crRNA. To allow crRNA and tracRNA production from DNA oligos, we order our DNA oligos of the reverse complement (see below for more details).