Abstract Due to recent scientific advances, there is a growing demand for new indices with high fidelity to eliminate errors associated with certain index combinations and sequencing chemistries. These errors are due to insufficient edit distances (the minimum number of changes required to transform one index sequence to the other) between index sequences or index hopping. Recent publications have highlighted how sequencing reads are misassigned due to "index hopping" on Illumina patterned flow cells, such as the recently launched Illumina Novaseq which can generate billions of reads (>100 exomes) in a single run. This misassignment canlead to false positives in ultra-sensitive assays where low frequency variants or nucleic acid species are monitored. We have observed misassignment due to insufficient edit distance among Illumina TruSeq HT i7 indices at a frequency up to 1.5%. Therefore, we developed 968nt i7 indices that can be paired with the existing TruSeq HT i5 indices to achieve 768 high throughput dual combinations, which were validated using a novel method on both Illumina's 2- and 4-channel technologies as both single and dual indices. This method involved thepreparation of 96 libraries with unique, non-overlapping inserts to facilitate tracking of index misassignment. This allowed us to assess not only which index has misassigned library molecules, but also pinpoint the origin and rate of misassignment within a single run. With our 96 i7 indices, misassignment was observed at rates <0.1%. For even higher fidelity de-multiplexing, we have paired our 96 indices in a non-tandem manner for single use in both the i5 and i7 positions, known as Unique Dual Indices (UDIs). The use of UDIs further eliminates index read errors that misassign reads, enabling increased confidence in calling low frequency variants. UDIs can also eliminate PCR-induced chimerisms, which can significantly improve data from a variety of assays. We are validating 96 new indices as UDIs for avoidance of index hopping and for eliminating PCR-induced chimerism during multiplexed library amplification. Citation Format: Jonathan C. Irish, Jordan RoseFigura, Sukhinder K. Sandhu, Bita Carrion, Laurie Kurihara, Vladimir Makarov. Improved sample indexing for high fidelity demultiplexing to increase confidence in low frequency variant calling [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 1430.
Abstract The growing use of liquid biopsy for early detection and monitoring of disease necessitates accurate variant detection at <1% allele frequencies due to a low population of disease DNA within circulating, cell-free DNA (cfDNA). Reliable, low-frequency variant detection by next-generation sequencing (NGS) is challenging due to background noise from PCR and sequencing errors. We employed molecular identifiers (MIDs) to uniquely label individual DNA molecules prior to amplification, facilitating the distinction of true variants from PCR and sequencing errors. We incorporated MIDs in both our amplicon library prep that uses multiplex PCR for targeted NGS and our whole genome library prep followed by targeting with hybridization capture using an 800kb pan-cancer panel. We performed low frequency spike-in experiments at <1% allele frequencies. We prepared MID libraries with various amplicon panels including a 17 amplicon EGFR pathway panel and a 104 amplicon SNP panel. Deep sequencing to >30,000x was done to maximize MID family size (number of PCR duplicates) and optimize generation of a consensus sequence. This analysis identified all known variants present at 1%, 0.5%, and 0.25% allele frequencies. Next, the hybridization capture libraries were prepared with low-frequency spike-in samples, sequenced to >8000x, and all known variants at 1% and 0.5% allele frequencies were maintained in the consensus data. In both cases, the number of false positives was reduced, resulting in improved specificity. Further, EGFR amplicon libraries and hybridization capture libraries were prepared using cfDNA samples from lung, ovarian, liver, stomach, and colon cancers. Variant calling based on MID generated consensus sequences identified mutations in cfDNA samples as well as corresponding tumor and normal samples when available. This study highlights the ability of MID technology to enable low frequency variant detection, critical to track known variants and identify novel pathogenic mutations in cfDNA samples. Citation Format: Ashley Wood, Sukhinder Sandhu, Mida Pezeshkian, Vanessa Kelchner, Jordan RoseFigura, Justin Lenhart, Laurie Kurihara, Vladimir Makarov. Low frequency variant detection in cell free DNA by applying molecular identifiers to targeted NGS [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2018; 2018 Apr 14-18; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2018;78(13 Suppl):Abstract nr 2230.
The promise of liquid biopsy assays lies in the non-invasive monitoring of diseases, such as cancer, through circulating, cell-free DNA (cfDNA) or circulating tumor cell DNA. This may assist in advancing early-stage diagnosis while simultaneously monitoring treatment response over time. Since these materials are often limited, most liquid biopsy assays incorporate targeted sequencing to enable cost-effective deep coverage of target loci for detection of low frequency pathogenic variants. Yet a critical aspect in attaining the necessary sensitivity is an assay that produces uniform, comprehensive coverage from low DNA input quantities. We have developed a liquid biopsy workflow to enable low frequency variant detection from a 10 mL blood draw using the Promega Maxwell RSC combined with Swift Biosciences Accel-NGS 2S® library preparation methodologies. In addition, we explore whether methylation patterns of the extracted cfDNA possess information of the tissue-of-origin. Whole blood samples were collected in Streck cell-free DNA BCT vials from patients with late stage cancer and cfDNA was extracted with the Promega Maxwell RSC. This instrument yielded DNA outputs ranging from 8-32 ng, with a size profile defined by a predominant peak of ~170bp and a mean Alu repeat qPCR integrity score of 0.22, characteristic of high quality cfDNA lacking cellular DNA content. A total of 20 ng cfDNA was used to make an Accel-NGS 2S Hyb library followed by hybridization capture using Agilent SureSelect Human All Exon probes. The Accel-NGS 2S Hyb Kit exhibits a 90% library conversion rate with cfDNA and provides high complexity libraries with uniform target coverage. In addition, molecular barcodes were incorporated to label each library molecule uniquely prior to PCR amplification. These molecular barcodes were utilized for accurate removal of PCR duplicates while simultaneously preserving naturally occurring fragmentation and strand duplicates to maximize data recovery. Secondly, these barcoded molecules were grouped to generate consensus sequences after removal of false positives originating from PCR and sequencing errors. Variant calling was performed using Vardict and Lofreq enabling highly sensitive and precise detection of variants down to a 0.5% allele frequency. In parallel, we have developed a workflow to determine if the epigenetic status of cfDNA can identify tissue-of-origin. This workflow utilizes the Accel-NGS Methyl-Seq DNA Library Kit to enable unbiased characterization from low (5 ng) cfDNA inputs. Through whole genome bisulfite sequencing, using a priori knowledge of differentially methylated regions characteristic of different human tissues, we can identify the predominant tissue source of cfDNA in blood. Citation Format: Justin S. Lenhart, Ashley Wood, Sukhinder Sandhu, Cassie Schumacher, Laurie Kurihara, Vladimir Makarov, Tim Harkins. Low frequency variant detection and tissue-of-origin exploration using liquid biopsies [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 5391. doi:10.1158/1538-7445.AM2017-5391
Abstract Tumor heterogeneity creates significant challenges for the generation and monitoring of treatment strategies. Circulating, cell-free DNA (cfDNA) can facilitate these processes by providing a noninvasive method to detect and track mutations without the need for multiple biopsies. To study both tumor heterogeneity and its relationship to cfDNA, we used samples taken from multiple positions on two metastatic tumors from a stage 3B ovarian carcinosarcoma patient as well as a cfDNA sample that was collected at the time of surgery. To maximize low frequency variant detection from next-generation sequencing (NGS) data, molecular identifiers (MIDs) were employed, allowing accurate distinction of true variants from PCR-induced errors along with sequencing artifacts. The use of MIDs to uniquely label input DNA generates tagged library molecules for the detection and removal of PCR duplicates while simultaneously preserving fragmentation and strand duplicates. By over-sequencing the MID tagged library, sequencing reads from PCR duplicates can be grouped based on their shared MID tag, generating a consensus sequence to identify and subsequently remove errors. To validate this technology, we performed low frequency spike-in experiments to 0.5% and 1% by combining Coriell NA12878 and HG005 genomic DNA as well as different cfDNA samples. Libraries were sequenced to 8000x coverage and a consensus sequence was generated with BMFtools (ARUP labs). All known variants present at 1% and 0.5% allele frequencies were maintained in the resulting data set. Further, true variants were preserved while sequencing and PCR-induced errors were removed, demonstrating improved sensitivity but also improved specificity using MIDs. After validation of the technology, libraries with MIDs were prepared from the ovarian carcinosarcoma FFPE and cfDNA samples. Libraries were sequenced on an Illumina® HiSeq® to a minimum of 13,000x coverage, and we determined data retention after de-duplication with and without the use of MIDs. We observed an increase in data retention for both Covaris-sheared FFPE and cfDNA libraries that led to a 2 to 3-fold increase in coverage using MIDs. Variant calling depicted significant inter-tumor heterogeneity in these samples as well as low frequency variants that represent intra-tumor heterogeneity. Specifically, we identified a pathogenic TP53 mutation present in one tumor sample and the cfDNA but absent from the second tumor. Additionally, we observed a significant loss of heterozygosity for only one of the tumors. We also identified low frequency, pathogenic mutations (as low as 1%) that are either unique to one tumor or unique to one sample from a tumor. This study highlights the use of MID technology to enable low frequency variant calling during disease diagnosis and tracking to facilitate precise, personalized treatment of complex cancers. Citation Format: Ashley Wood, Sukhinder Sandhu, Matthew Dashkoff, Olga Camacho-Vanegas, Laurie Kurihara, Timothy Harkins, John Martignetti, Vladimir Makarov, Peter Dottino. Characterizing tumorigenesis of ovarian cancer across metastatic tumors and circulating, cell-free DNA [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr LB-276. doi:10.1158/1538-7445.AM2017-LB-276
Targeted sequencing of cell-free (cfDNA) and circulating tumor DNA (ctcDNA) from blood enables detection of cancer-related mutations using minimally-invasive sample collection methods, and may make early detection of cancer possible, as well as improve monitoring of disease burden in translational research studies. We have developed a series of targeted panels for detection of multiple cancer-related mutations. The panels are designed to efficiently amplify damaged or short fragments of DNA derived from FFPE and cfDNA/ctcDNA, where hundreds of primer pairs can be amplified in a single tube from overlapping targets using only 10 ng input material, making these panels ideal for limiting liquid biopsy samples. A panel which covers known “hotspot” mutations in 56 oncology-related genes has been used in a pilot research study to monitor gynecological cancer in 11 women in a longitudinal study, which found a correlation between the presence of cancer mutations and morbidity and mortality. In 2 of 11 women, the initial absence of mutations above 1% allele-frequency was followed by the appearance of mutations in 1-3 genes at allele frequencies of 5-78% in the later time point. These 2 patients experienced increased morbidity or mortality. In 9 of the 11 women, no mutations were observed, and 6 remain in remission, while 3 are living with cancer. In an effort to further improve both workflow and performance, we are developing two technologies to incorporate into the panel design for future studies. The first will normalize library yield during PCR amplification for simple library pooling, which eliminates the requirement for library quantification and minimizes the time from sample to sequence. The second technology is a molecular ID (MID) system to tag each amplicon uniquely to allow data tracking to individual DNA fragments from the sample, and to increase confidence in variant calling by filtering PCR and sequencing errors. By incorporating technologies that reduce steps in the workflow, the likelihood of error is minimized, and combined with methods that increase confidence in low frequency variant calling, an ideal workflow for liquid biopsy samples is created. Citation Format: Jonathan C. Irish, Cassie A. Schumacher, Navya Nair, Olga Camacho-Vanegas, Jordan Rose- Figura, Ashley Wood, Sukhinder Sandhu, Sushma Chaluvadi, Sergey Chupreta, Laurie Kurihara, Timothy Harkins, John A. Martignetti, Vladimir Makarov. Targeted next-generation sequencing of cell-free tumor DNA to longitudinally monitor cancer burden and progression [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2017; 2017 Apr 1-5; Washington, DC. Philadelphia (PA): AACR; Cancer Res 2017;77(13 Suppl):Abstract nr 5392. doi:10.1158/1538-7445.AM2017-5392
Pyrodictium delaneyi strain Hulk is a newly sequenced strain isolated from chimney samples collected from the Hulk sulfide mound on the main Endeavour Segment of the Juan de Fuca Ridge (47.9501 latitude, −129.0970 longitude, depth 2200 m) in the Northeast Pacific Ocean. The draft genome of strain Hulk shared 99.77% similarity with the complete genome of the type strain Su06 T , which shares with strain Hulk the ability to reduce iron and nitrate for respiration. The annotation of the genome of strain Hulk identified genes for the reduction of several sulfur-containing electron acceptors, an unsuspected respiratory capability in this species that was experimentally confirmed for strain Hulk. This makes P. delaneyi strain Hulk the first hyperthermophilic archaeon known to gain energy for growth by reduction of iron, nitrate, and sulfur-containing electron acceptors. Here we present the most notable features of the genome of P. delaneyi strain Hulk and identify genes encoding proteins critical to its respiratory versatility at high temperatures. The description presented here corresponds to a draft genome sequence containing 2,042,801 bp in 9 contigs, 2019 protein-coding genes, 53 RNA genes, and 1365 hypothetical genes.
Abstract Introduction Hybridization-based target enrichment techniques coupled with Next Generation Sequencing (NGS) provide a useful and cost-efficient means to study disease specific target regions including whole exomes and gene panels. More than 500 whole exome analyses are behind The Cancer Genome Atlas (TCGA), which is the largest knowledge base for cancer studies including research, prevention, treatment, and care. Clinical samples may be limited in input and of compromised quality due to formalin fixation, however most current NGS library preparation methods require 50-100 ng high quality DNA for such studies. Here we present an efficient method which enables high quality target enrichment and variant calling from inputs as low as 1-25 ng. Method NGS libraries for hybridization capture were made from 1 to 100 ng of Coriell Hapmap samples (NA12878, etc), clinical Formalin Fixed Paraffin Embedded (FFPE) and Horizon Discovery (HDx) reference DNA (HDx 701) using the Accel-NGS® 2S Hyb DNA Library Kit. Amplified libraries were then enriched using specific targeted panels (xGen® Pan-Cancer and xGen® AML) or SeqCap™ EZ MedExome using manufacturer's specifications (IDT™ and Roche NimbleGen™). Enriched libraries were captured using streptavidin beads. Captured libraries were then amplified according to the manufacturer's specifications. Targeted panels were sequenced on an Illumina® MiSeq® using V2 chemistry and MedExome on a HiSeq® using V4 chemistry. Sequence analysis was performed with custom pipelines using BWA for alignment and GATK, Samtools, and Freebayes for variant calling. Results Sequencing analyses yielded a minimum average coverage of 30x and more than thirty fold enrichment. Enriched libraries exhibited significant sequence complexity with minimal duplicates and without any base-composition bias. The percent on-target varied with type of input and ranged from 50-80%. Germline variant calling results had > 99% concordance with the NIST GIAB truth list with sensitivity and precision of > 98%, even from inputs as low as 1 ng of DNA. Somatic variant calling down to 1-5% allele frequency was also evaluated at various DNA input quantities and greater depth of sequencing; results for FFPE and DNA standards will be presented. Conclusions Accel-NGS 2S Hyb technology can be used for whole exome and targeted enrichment studies from low quantity and low quality FFPE clinical samples. High complexity libraries with minimal bias yield high quality sequence data which enables discovery and detection of somatic variants in tumor samples to identify molecular drivers associated with different cancer types. The technique facilitates better understanding of complex cancer genomes and guide precision medicine. Citation Format: Sukhinder K. Sandhu, Cassie Schumacher, Laurie Kurihara, Tim Harkins, Vladimir Makarov. Targeted exome and panel analysis from low input and FFPE DNA using hybridization capture for cancer genome studies. [abstract]. In: Proceedings of the 107th Annual Meeting of the American Association for Cancer Research; 2016 Apr 16-20; New Orleans, LA. Philadelphia (PA): AACR; Cancer Res 2016;76(14 Suppl):Abstract nr 3637.
O1 The metabolomics approach to autism: identification of biomarkers for early detection of autism spectrum disorder
Detection of low abundant alterations in oncogenes from circulating cell free DNA (cfDNA) samples or formaldehyde fixed-paraffin embedded samples (FFPE) has been a challenge for the community. Additionally, cfDNA and FFPE samples are typically limited in the amount of material that can be obtained, and often the FFPE samples are highly degraded, compounding the ability to sequence them and to perform accurate variant calling. To overcome these challenges, we have developed a single tube, multiplexed amplicon sequencing technology that employ hundreds of primer pairs for amplification of targeted loci, producing ready-to-run libraries for Illumina sequencing platforms. The resulting amplicons are less than 150 bp in length, enabling amplification and sensitive detection of mutations from cfDNA-sized DNA fragments or highly damaged DNA. A 200+ amplicon panel across 55 genes was developed to target known clinically relevant oncological mutations. The panel design encompasses single exons (e.g. BRAF) as well as comprehensive exon coverage of entire genes (e.g. TP53). DNA was extracted from FFPE, and cfDNA was extracted from fresh plasma. DNA was quantified by qPCR using 2 primer pairs targeting ALU repeat regions. 10ng of DNA input was used to perform the multiplexed PCR. Libraries were quantified by qPCR and sequenced with MiSEQ V2 reagents, paired end reads. Alignment was performed using Burrows-wheeler Aligner (BWA) and variant calling was performed with at least two validated publicly available tools and confirmed manually by IGV. A variety of tools were used to validate allele technology was used to validate the sequencing data. We interrogated 80 FFPE samples using our amplicon technology.The percent of on-target bases was over 95% and the coverage uniformity was over 98%, where uniformity is the percent based covered over 20% of the mean coverage. Out of the 40 colorectal cancer samples tested, 17 showed mutations in KRAS. Out of the 20 melanoma samples (Cure line) also analyzed, 9 were positive for BRAF V600E. The presence of those mutations was confirmed using myTprimerTM KRAS and BRAF qPCR assay. The limit of detection of the assay was set at 3% due to the inherent noise caused by the damage from the FFPE samples. For cell lines with known mutations, the sensitivity was 0.5%. No false positive or false negative were reported. 20 FFPE samples from different cancer types were also assessed. Mutations were identified in ERBB2 (cervical cancer), ALK (lung cancer) and TP53 (colon cancer) genes, as well as MET amplification (lung cancer). Cell free DNA provides a non-invasive tool to monitor cancer progression and treatment efficacy. Tumor DNA, matching blood DNA and cfDNA from 10 cancer patients will be subjected to our multiplex PCR oncology panel. Data of this study will be presented. Our results show that this unique multiplex PCR panel is an excellent tool to assess multiple oncogenes in limiting clinical samples, enabling high throughput, cost effective NGS analysis. Citation Format: Julie Laliberte, Catherine Couture, Vladimir Makarov, Cassie Schumacher, Sukhinder Sandhu, Laurie Kurihara, Timothy Harkins. Detecting low abundant mutations in circulating cell free DNA and FFPE samples. [abstract]. In: Proceedings of the 106th Annual Meeting of the American Association for Cancer Research; 2015 Apr 18-22; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2015;75(15 Suppl):Abstract nr 4902. doi:10.1158/1538-7445.AM2015-4902
Abstract Circulating, cell free DNA (cfDNA) is a non-invasive sample source that contains tumor associated DNA. Next generation sequencing (NGS) of cfDNA has shown that tumor specific mutations can be detected, providing an effective means to monitor disease, treatment efficacy, and characterize the genome wide methylation state of cfDNA. While hypermethylation is observed for specific gene promoters, on a genome wide perspective, hypomethylation of cfDNA is observed in cancer, making cfDNA an attractive target to assess cancer burden. The challenges in deep sequencing of cfDNA include: mandatory fast turnaround times due to sample degradation, limited sample material, short DNA fragments, and limit of detection issues caused by the presence of both normal and tumor DNA. This study describes a novel library preparation for the Illumina platforms that is cost-effective, sensitive, and specific to assess the methylation status of cfDNA. To characterize the methylation status of cfDNA, NGS libraries were generated utilizing a chemistry that sequentially ligates the adapters to each end of the DNA molecules. Since the library is generated after bisulfite treatment, a high recovery of DNA library molecules is observed, thus enabling high complexity library preparation from 5 ng of cfDNA. Upon establishment of this technique, we obtained cfDNA from healthy subjects as well as from subjects with a spectrum of cancers. Following bisulfite-conversion and library preparation of these samples, we sequenced each sample on the Illumina MiSeq to a depth of 10 million reads. A minimum threshold of 1.1% was used to determine significant hypomethylation. Upon analysis, preliminary analysis of the hypomethylation status of the cfDNA from 4 of the 5 cancer subjects ranged from 3% to 9% when compared to the healthy controls. The cfDNA sample which was negative for hypomethylation originated from the plasma of a subject with a high grade serous adenocarconima in the fallopian tube. The most hypomethylated cfDNA came from the plasma of subject with metastatic adenocarcinoma of the colon which had metastasized to the liver. Other tumor types which resulted in cfDNA hypomethylation between these two extremes included invasive breast carcinoma as well as pancreatic ductal carcinoma. This method provides the basis for a technique to reliably prepare and characterize cfDNA via NGS deep sequencing. We obtained results 5 days after resection and reliably made high complexity library using only 5 ng of cfDNA. We have demonstrated reproducibility and sensitivity for genome wide methylation status, while also observing that the utility of this assay may vary with cancer type. To further define biologically significant thresholds for methylation status, we will present a longitudinal study examining the cfDNA methylation status of individuals before, during, and after treatment, thereby generating a methylation gradient as a function of time and treatment. Citation Format: Cassie A. Schumacher, Sukhinder Sandhu, Vladimir Makarov. An assay to detect hypomethylation in circulating cell free DNA and monitor cancer burden. [abstract]. In: Proceedings of the 106th Annual Meeting of the American Association for Cancer Research; 2015 Apr 18-22; Philadelphia, PA. Philadelphia (PA): AACR; Cancer Res 2015;75(15 Suppl):Abstract nr 1051. doi:10.1158/1538-7445.AM2015-1051