Abstract Background: Microsatellite instability (MSI) refers to the accumulation of somatic mutations in simple repeat regions and is a hallmark of mismatch repair deficiency as well as a predictive biomarker for immunotherapy response. Most existing MSI callers are built for short-read sequencing and typically require paired tumor-normal data or high sequencing depth. These methods quantify repeat-length variability and classify genomes as MSI high when 10-30 percent of markers are unstable. Long-read sequencing (LRS) enables haplotype-resolved repeat profiling, allowing accurate MSI detection from tumor-only genomes at standard (∼30×) coverage. Methods: We developed Owl, a MSI detection software tool for PacBio HiFi data. Owl interrogates more than 160,000 simple repeats using a wrap-around dynamic programming algorithm and calculates the coefficient of variation (CV) in repeat length across phased haplotypes. Genome-wide MSI scores are then reported as the fraction of markers exceeding a parametrically derived CV threshold (CV > 5). Results: We first applied Owl to 131 healthy controls from the Human Pangenome Reference Consortium and observed that MSI-stable genomes typically fall between 2-6%. We then profiled 26 additional cancer genomes, all of which had Owl scores between 2-3%, consistent with microsatellite-stable profiles. Finally, we identified five MSI-high samples that exceeded our 10% threshold, with scores ranging from 15-18%. These included two gastric cancers, an astrocytoma sample, and two Ewing sarcoma cell lines. Only one sample had both HiFi and Illumina data available for comparison; in that astrocytoma sample, Owl (17.1%) and Illumina DRAGEN (20.0%) produced concordant MSI-high classifications.By measuring MSI at >160,000 repeats, we can also detect motif-specific instability in tumor samples. For example, the two Ewing sarcoma cell lines (TC32 and CHLA10) showed a two fold increase of GGAA motif instability (23-26% MSI) compared to other motifs, an interesting and potentially disease-relevant pattern. The EWS::FLI1 fusion is known to bind GGAA-rich regulatory elements, and the observed instability at these motifs suggests that repeat variation itself may play an important role in Ewing sarcoma biology. Conclusions: Owl enables robust, quantitative detection of MSI from long-read whole-genome data, requiring only tumor samples. Integrated into the PacBio HiFi Somatic workflow, Owl extends MSI profiling to long-read sequencing, revealing motif-specific instability patterns not captured by short-read approaches. Citation Format: Zev Kronenberg, Khi Pin Chua, Mark J. Chaisson, Byunggil Yoo, Lisa Lansdon, William J. Rowell, Egor Dolzhenko, Kie Kyon Huang, Patrick Tan, Shruti S. Bhise, Everett Fan, Mark Mendoza, Emily O'donnell, Tomi Pastinen, Elizabeth R. Lawlor, Scott N. Furlan, Midhat S. Farooqi, Michael A. Eberle. Hunting for microsatellite instability in long-read data with Owl [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5510.
Variant benchmarking is critical in assessing the accuracy of genomic secondary pipelines. However, traditional benchmarking tools that require exact genotype matches inject biases from variant representation and are ill-suited for tandem repeat or structural variation. We describe Aardvark, a variant benchmarking tool that introduces the basepair score to directly compare haplotype sequences, reducing representation biases while allowing for partial credit scoring. The tool also includes the traditional genotype score and supports separate or joint benchmarking of small variants, tandem repeats, and structural variants (<10 kb). Aardvark accepts standard inputs, runs ≈ 18x faster than hap.py, and is open source (https://github.com/PacificBiosciences/aardvark or https://doi.org/10.5281/zenodo.20271610).
Tandem repeats (TRs) exhibit high levels of somatic mosaicism, which is increasingly recognized as an important modifier of repeat expansion disorders. Long-read sequencing can capture full-length repeat alleles, yet robust frameworks for quantifying instability across TRs genome-wide are still needed. Here, we introduce a general-purpose model for quantifying TR instability in a given long-read sequencing dataset, without explicitly distinguishing biological mosaicism from technical noise, and which is broadly applicable to both simple and structurally complex loci. This model accurately characterizes allelic instability at each TR locus by representing the distribution of read-to-consensus deviations for each allele. Using HiFi sequencing data from 256 HPRC cell line samples, we fitted models for 617,007 TR loci, including known pathogenic repeats. We observe that instability levels are generally low, but vary substantially across individual TRs, and are driven more strongly by repeat composition than overall repeat length. Furthermore, we applied our method to targeted PureTarget long-read data from samples with known repeat expansions and identified significant mosaicism in the majority of expanded alleles. Our model offers a practical way to quantify instability of tandem repeats across the genome and to detect unusually unstable repeat alleles.
Short-read sequencing (SRS) methods have improved the detection of small genetic variants but remain limited in highly homologous genomic regions, such as segmental duplications with gene-pseudogene pairs. These paralogous regions often require complex, locus-specific assays for accurate analysis. Long-read genome sequencing (lrGS) technologies, such as PacBio HiFi sequencing, can span these regions but still face challenges in variant calling due to alignment ambiguities. Here, we evaluated PacBio HiFi lrGS combined with Paraphase, a dedicated haplotype-based variant caller, in 86 individuals with 125 known clinically relevant variants across 11 paralogous loci. Standard HiFi variant callers detected 95/125 variants, while the remaining 30 variants were only identified by Paraphase. Together, the standard variant callers and Paraphase detected all known variants, including single-nucleotide variants (SNVs), insertions or deletions (InDels), copy-number variants (CNVs), structural variants (SVs), and gene conversions. In addition, lrGS allowed for accurate phasing and gene-pseudogene copy-number detection. We demonstrate that PacBio HiFi lrGS, particularly when integrated with Paraphase, enables comprehensive variant detection in previously difficult-to-assess genomic regions. These results also suggest that lrGS is ready for a wider implementation, possibly as a first-tier diagnostic approach for individuals with suspected variants in these paralogous regions. Ideally, clinical adoption is guided by prospectively designed clinical utility studies, alongside evaluation of sensitivity, specificity, and cost-effectiveness.
Tandem repeat (TR) catalogs are important components of repeat genotyping studies because they define the genomic coordinates and expected motifs of all TR loci being analyzed. In recent years, genome-wide studies have used catalogs ranging in size from fewer than 200,000 to over 7 million loci. These catalogs differed not only in which loci were included but often also in their definitions of TR locus boundaries and motifs. Now, with multiple groups developing public databases of TR variation in large population cohorts, there is a risk that the use of divergent repeat catalogs will lead to confusion, fragmentation, and incompatibility across resources. Here, we compare existing TR catalogs and discuss desirable features of a comprehensive genome-wide catalog. We then present a new, richly annotated catalog designed for genome-wide analyses and population datasets based on short-read or long-read samples. Additionally, using an algorithm that leverages long-read HiFi sequencing data, our catalog stratifies TRs into (1) isolated repeats suitable for repeat copy-number analysis and (2) variation clusters where TRs are embedded within wider polymorphic regions best studied through sequence-level analysis. We share the TR catalog, variation clusters, and annotations through the TRExplorer portal in order to support both the initial selection of TR loci for inclusion in an analysis and the subsequent interpretation of results.
Leveraging new sequencing and omic technologies to enhance the detection of pathogenic variants in known disease genes is a key step toward increasing the likelihood of a precise genetic diagnosis for affected individuals. Short-read sequencing is widely used in clinical laboratories for multi-gene panels and exome and genome sequencing, but this technology has inherent limitations in detecting certain classes of genetic variation. As a result, diagnostic laboratories continue to offer complementary assays, often used sequentially, reducing efficiency and speed in providing a diagnosis. We applied PacBio long-read genome sequencing (HiFi) to samples from 191 probands previously tested with short-read sequencing alone and/or other diagnostic technologies and enriched for pathogenic variants difficult to detect (VDDs). HiFi’s pipeline automatically detected 479 of 481 (99.6%) disease-causing variants, many of which were called in samples not optimized for long-read genome sequencing (such as buccal samples or low-molecular-weight DNA). The two variants not automatically detected were a mosaic trisomy 18 (23% mosaicism) and a 5,594-bp mosaic deletion (13% mosaicism). However, other mosaic variants were detected, indicating that HiFi at ∼30× genome coverage is sensitive to the degree of mosaicism. Of 481 variants, 49 were suspected based on the clinical report but not confirmed molecularly prior to HiFi. Our findings demonstrate that HiFi sequencing detects a wide range of VDDs in real-world clinical laboratory samples, highlighting a key advantage of HiFi as a potential first-tier test over the myriad of complementary technologies currently used to detect VDDs.
Tandem repeats (TRs) are among the most mutable loci in the human genome, but the genomic determinants of TR mutagenesis remain mysterious. We used PacBio HiFi long-read sequencing to profile nearly eight million TR loci in 28 members of a large, four-generation CEPH/Utah family designated K1463. We identified 1,270 de novo TR expansions and contractions across 20 children in the pedigree. De novo mutations (DNMs) were more likely to occur at loci that were longer, composed of uninterrupted motif sequences, and heterozygous in the parental germline. Children born to older fathers also exhibited more de novo mutations at short tandem repeats (STRs). A total of 43 TR loci were hyper-mutable in K1463, expanding or contracting up to twelve times across the pedigree. Among hyper-mutable loci that comprised multiple motifs (i.e., "complex" loci), specific motifs expanded and contracted more often than others; for example, all ten DNMs at a complex, hyper-mutable locus near the non-coding RNA LINC03021 involved the same 19bp motif. The mutability of particular motifs may be attributable to allele length, as 95% of DNMs at complex loci were expansions and contractions of the most abundant motif on a parental haplotype. However, future work will be required to disentangle the effects of nucleotide content and allele length on motif-specific mutability, especially at hyper-mutable TRs. Overall, this study combines long-read sequencing technologies with new software tools to comprehensively investigate the factors that influence TR mutagenesis.
Abstract The D4Z4 macrosatellite repeat encompasses some of the most difficult-to-resolve disease-related variations in the human genome. D4Z4 has a repeat unit of 3.3 kb (encoding the DUX4 gene) that is present in up to 100 copies on two chromosomes (4 and 10), while DUX4 can only be expressed in somatic cells from the permissive A haplotype that usually occurs on chromosome 4. Facioscapulohumeral muscular dystrophy (FSHD) is caused by chromatin relaxation and ectopic expression of DUX4 in skeletal muscle, mediated by contraction of D4Z4 to 1-10 copies (FSHD1, 95% of FSHD cases) or mutations in chromatin factor genes such as SMCHD1 (FSHD2, 5% of FSHD cases). Due to its large size, disease specific haplotypes and sequence homology between chromosomes, D4Z4 is challenging to resolve by current sequencing technologies. We report a computational tool, Kivvi, to genotype D4Z4 using PacBio whole-genome long-read sequence data. Kivvi detects all D4Z4 alleles in a sample, reporting the repeat size, chromosome (4 vs. 10), distal haplotype (A vs. non-permissive haplotypes) and the methylation level of each allele. We validated Kivvi against gold standard assays for FSHD diagnostics, detecting 100% of contracted alleles and correctly classifying 90% of noncontracted alleles. We showed differential methylation signals between FSHD1 and candidate FSHD2 samples. We profiled D4Z4 across 601 individuals from five ancestral populations, revealing extensive genetic diversity. We identified common haplotypes of D4Z4 alleles and characterized hybrid repeat units, hybrid repeat arrays, and translocation alleles. Combined with HiFi long reads, Kivvi enables the consolidation of multiple FSHD assays into a single workflow and facilitates the discovery of novel genetic modifiers of FSHD through population-scale studies.
Tandem repeat (TR) catalogs are important components of repeat genotyping studies as they define the genomic coordinates and expected motifs of all TR loci being analyzed. In recent years, genome-wide studies have used catalogs ranging in size from fewer than 200,000 to over 7 million loci. Where these catalogs overlapped, they often disagreed on locus boundaries, hindering the comparison and reuse of results across studies. Now, with multiple groups developing public databases of TR variation in large population cohorts, there is a risk that, without sufficient consensus in the choice of locus definitions, the use of divergent repeat catalogs will lead to confusion, fragmentation, and incompatibility across resources. In this paper, we compare existing TR catalogs and discuss desirable features of a comprehensive genome-wide catalog. We then present a new, richly annotated catalog designed for large-scale analyses and population databases. This new catalog, which we call the TRExplorer catalog v1.0, contains 4.86 million TR loci and, unlike most catalogs, is designed to be useful for both short-read and long-read analyses. It consists of 4,803,366 STRs and 59,675 VNTRs, of which 780,607 STRs and 21,888 VNTRs are both polymorphic and entirely absent from widely-used catalogs previously developed for short-read analyses. Additionally, our catalog stratifies TRs into two groups: 1) isolated TRs suitable for repeat copy number analysis using short-read or long-read data and 2) so-called variation clusters that contain TRs within wider polymorphic regions that are best studied through sequence-level analysis. To define variation clusters, we present a novel algorithm that leverages long-read HiFi sequencing data to group repeats with surrounding polymorphisms. We show that the human genome contains at least 25,000 complex variation clusters, most of which span over 120 bp and contain five or more TRs. Resolving the sequence of entire variation clusters instead of individually genotyping constituent TRs leads to a more accurate analysis of these regions and enables us to profile variation that would have been missed otherwise. We also share the trexplorer.broadinstitute.org portal which allows anyone to search, visualize, and download the catalog along with variation clusters and annotations.
Microsatellite instability (MSI) is a key biomarker of mismatch repair deficiency and response to immunotherapy, yet most existing genomic detection methods are optimized for short-read sequencing and rely on a panel of homopolymer markers, limiting the ability to characterize genome-wide and motif-specific patterns of instability. Here we present Owl, a bioinformatic tool for quantifying MSI from long-read (PacBio) genomic data. Owl leverages a genome-wide marker set of more than 140,000 microsatellite repeats ranging from 1-6 bp in length to measure MSI across a phased genome. Using a wrap-around alignment algorithm, Owl constructs repeat-length distributions at each marker site and flags somatic instability using the coefficient of variation. We applied Owl to screen for markers with stable coverage, phasing, and baseline variation across 131 diverse genomes from the Human Pangenome Reference Consortium, where Owl scores ranged from 1.4% to 5.4% of markers exceeding the instability threshold. When applied to cancer cell lines and one diffuse astrocytoma tumor-normal pair, Owl identified six MSI genomes with 10-27% unstable markers and showed close concordance with an Illumina DRAGEN MSI assay for the astrocytoma sample. Motif-level analyses revealed shared enrichment of short homopolymer and dinucleotide (A- and AT-rich) repeats across MSI cancers. Owl is implemented in Rust and integrated into the PacBio HiFi Somatic workflow, providing a scalable framework for MSI analysis from long-read sequencing focused on repeat instability specifically in tumor samples.
Clinical short-read exome and genome sequencing approaches have positively impacted diagnostic testing for rare diseases. Yet, technical limitations associated with short reads challenge their use for the detection of disease-associated variation in complex regions of the genome. Long-read sequencing (LRS) technologies may overcome these challenges, potentially qualifying as a first-tier test for all rare diseases. To test this hypothesis, we performed LRS (30× high-fidelity [HiFi] genomes) for 100 samples with 145 known clinically relevant germline variants that are challenging to detect using short-read sequencing and necessitate a broad range of complementary test modalities in diagnostic laboratories. We show that relevant variant callers readily re-identified the majority of variants (120/145, 83%), including ∼90% of structural variants, SNVs/insertions or deletions (indels) in homologous sequences, and expansions of short tandem repeats. Another 10% (n = 14) was visually apparent in the data but not automatically detected. Our analyses also identified systematic challenges for the remaining 7% (n = 11) of variants, such as the detection of AG-rich repeat expansions. Titration analysis showed that 90% of all automatically called variants could also be identified using 15-fold coverage. Long-read genomes thus identified 93% of challenging pathogenic variants from our dataset. Even with reduced coverage, the vast majority of variants remained detectable, possibly enhancing cost-effective diagnostic implementation. Most importantly, we show the potential to use a single technology to accurately identify all types of clinically relevant variants.
Variant calling is hindered in segmental duplications by sequence homology. We developed Paraphase, a HiFi-based informatics method that resolves highly similar genes by phasing all haplotypes of paralogous genes together. We applied Paraphase to 160 long (>10 kb) segmental duplication regions across the human genome with high (>99%) sequence similarity, encoding 316 genes. Analysis across five ancestral populations revealed highly variable copy numbers of these regions. We identified 23 paralog groups with exceptionally low within-group diversity, where extensive gene conversion and unequal crossing over contribute to highly similar gene copies. Furthermore, our analysis of 36 trios identified 7 de novo SNVs and 4 de novo gene conversion events, 2 of which are non-allelic. Finally, we summarized extensive genetic diversity in 9 medically relevant genes previously considered challenging to genotype. Paraphase provides a framework for resolving gene paralogs, enabling accurate testing in medically relevant genes and population-wide studies of previously inaccessible genes. Here, the authors present Paraphase, a HiFi-based informatics method that resolves highly similar genes located in segmental duplications. They apply Paraphase to 316 paralogous genes and summarize extensive genetic diversity across populations.
Understanding the human de novo mutation (DNM) rate requires complete sequence information1. Here using five complementary short-read and long-read sequencing technologies, we phased and assembled more than 95% of each diploid human genome in a four-generation, twenty-eight-member family (CEPH 1463). We estimate 98-206 DNMs per transmission, including 74.5 de novo single-nucleotide variants, 7.4 non-tandem repeat indels, 65.3 de novo indels or structural variants originating from tandem repeats, and 4.4 centromeric DNMs. Among male individuals, we find 12.4 de novo Y chromosome events per generation. Short tandem repeats and variable-number tandem repeats are the most mutable, with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 16% of de novo single-nucleotide variants are postzygotic in origin with no paternal bias, including early germline mosaic mutations. We place all this variation in the context of a high-resolution recombination map (~3.4 kb breakpoint resolution) and find no correlation between meiotic crossover and de novo structural variants. These near-telomere-to-telomere familial genomes provide a truth set to understand the most fundamental processes underlying human genetic variation.
Structural variants are genomic variants that impact at least 50 nucleotides. Structural variants can play major roles in diversity and human health. Many structural variants are difficult to interpret and understand with existing visualization tools, especially when comprised of inverted sequences or multiple breakend pairs. We present SVTopo, a tool to visualize germline structural variants with supporting evidence from high-accuracy long reads in easily understood figures. We include examples of 101 visually complex structural variants from seven unrelated human genomes, manually assigned to ten categories. These demonstrate a broad spectrum of rearrangement and showcase the frequency of complex structural variants in human genomes. SVTopo shows breakpoint evidence in ways that aid reasoning about the impact of multi-breakpoint rearrangements. The images created aid human reasoning about the result of structural variation on gene and regulatory regions.
Tandem repeats are a highly polymorphic class of genomic variation that play causal roles in rare diseases but are notoriously difficult to sequence using short-read techniques1,2. Most previous studies profiling tandem repeats genome-wide have reduced the description of each locus to the singular value of the length of the entire repetitive locus3,4. Here we introduce a comprehensive database of 3.6 billion tandem repeat allele sequences from over one thousand individuals using HiFi long-read sequencing. We show that the previously identified pathogenic loci are among the most variable tandem repeat loci in the genome, when incorporating nucleotide resolution sequence content to measure the longest pure motif segment. More broadly, we introduce a novel measure, 'tandem repeat constraint', that assists in distinguishing potentially pathogenic from benign loci. Our approach of measuring variation as 'the length of the longest pure segment' successfully prioritizes pathogenic repeats within their previously published linkage regions. We also present evidence for two novel pathogenic repeat expansion candidates. In summary, this analysis significantly clarifies the potential for short tandem repeat pathogenicity at over 1.7 million tandem repeat loci and will aid the identification of disease-causing repeat expansions.
Recent advances in genome sequencing have improved variant calling in complex regions of the human genome. However, it is difficult to quantify variant calling performance because existing standards often focus on specificity, neglecting completeness in difficult-to-analyze regions. To create a more comprehensive truth set, we used Mendelian inheritance in a large pedigree (CEPH-1463) to filter variants across PacBio high-fidelity (HiFi), Illumina and Oxford Nanopore Technologies platforms. This generated a variant map with over 4.7 million single-nucleotide variants, 767,795 insertions and deletions (indels), 537,486 tandem repeats and 24,315 structural variants, covering 2.77 Gb of the GRCh38 genome. This work adds 200 Mb of high-confidence regions, including 8
The characterization of somatic variation, especially in complex genomic regions, is crucial for understanding the molecular drivers of cancer progression. Accurate PacBio long-read sequencing (HiFi) enables detection of all variant classes, from simple SNVs and INDELs up to complex structural variation, tandem repeats, and changes in epigenetic signatures. Complex and repetitive regions, while fully sequenced by HiFi reads, remain bioinformatically challenging, requiring tailored solutions. Here, we describe new tools to genotype understudied repetitive regions in cancer genomes, a task that has historically posed significant challenges to short-read sequencing.Our software tool, TRGT, resolves the sequence and methylation patterns of complex repetitive regions at a single-nucleotide resolution. A recent algorithmic improvement (vclust module) enables us to type expanded hyper-polymorphic regions, termed variation clusters. Variation clusters usually result in many incorrect variant calls in computational pipelines. By analyzing a variation cluster as a single region, rather than many small variants, we can accurately measure genetic variation. To demonstrate the importance of variation clusters in cancer, we comprehensively resolved repeat regions in five cell lines and 50 clinical tumor samples (with matched healthy tissue). We identified 5,500 new variation clusters across the tumors and an average of 46,220 (out of 4 million, 1%) loci that are contracted or expanded in a tumor compared to its matched normal tissue. We find 3% of these loci can be located within genes found in the Compendium of Cancer Genes, suggesting potentially oncogenic significance. In addition, using TRGT, we were able to detect microsatellite instability (MSI) in a tumor with a compound heterozygous coding mutation in MSH2. Lastly, we further analyzed haplotypes in selected variation clusters to detect and resolve sequences in somatic mosaicism. By incorporating this new class of variants—complex repeats—into our analysis, we significantly enhanced the capability of HiFi reads to detect a comprehensive spectrum of genomic variations. The inclusion of complex repeat genotyping with PacBio HiFi reads will enable a more detailed understanding of cancer genetics and open new avenues for research and clinical applications. Khi Pin Chua, Egor Dolzhenko, Zev N. Kronenberg, Seiya Imoto, Seiichi Mori, Michael A. Eberle. Precise characterization of complex repeats in cancer genomes [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2025; Part 1 (Regular Abstracts); 2025 Apr 25-30; Chicago, IL. Philadelphia (PA): AACR; Cancer Res 2025;85(8_Suppl_1):Abstract nr 6288.