Abstract Background: Determining the parent of origin (PofO) of a variant in hereditary cancer guides counseling, risk management, recurrence risk assessment, and variant classification. In conditions involving genes with PofO effects, such as SDHD, SDHAF2 and MAX, this information can determine whether disease will manifest. Current approaches rely on family-based testing, yet uptake among eligible first-degree relatives remains low, with fewer than 30% undergoing testing, creating a major barrier in hereditary cancer.To address this gap, we developed Parent-of-Origin-Aware Genomic Analysis (POAga), a method that integrates chromosome-scale haplotyping with DNA methylation at differentially imprinted regions to assign PofO without parental data. To validate POAga, we applied it to individuals with hereditary cancer and known segregation of their pathogenic variants, and compared the predicted PofO with the established segregation to assess concordance and limitations. Methods: Blood samples are being collected from individuals with pathogenic variants in hereditary cancer genes, representing broad ranges of ages, ancestries, and cancer histories. Parental segregation was previously known or established through confirmatory testing. Predicted PofO is compared with true segregation to assess concordance. All samples undergo Strand-seq and long-read sequencing under an REB-approved protocol. Results: To date, 285 samples with 290 pathogenic variants have been analyzed across the following genes: BRCA2 (n=46), BRCA1 (n=42), MSH2 (n=34), SDHD (n=29), MLH1 (n=27), MSH6 (n=26), PMS2 (n=20), PALB2 (n=15), TP53 (n=14), ATM (n=14), CDH1 (n=9), CHEK2 (n=3), EPCAM (n=2), SDHAF2 (n=2), MUTYH (n=2), CDKN2A (n=1), POT1 (n=1), RAD51D (n=1), and SDHC (n=1). PofO was assigned for 250 variants, with 98.4% concordance (246/250). PofO could not be determined for 40 variants (13.8%, 40/290), mainly due to insufficient allele-specific methylation at imprinted regions or extended homozygosity that impeded phasing. Misassignments were rare and mainly due to stochastic phasing errors, unresolved inversions, or random allelic methylation at imprinted regions. Conclusion: POAga achieves clinical-grade accuracy in assigning PofO from a single blood sample in hereditary cancer. This directly addresses a major barrier in clinical genetics, particularly when parental samples are unavailable. For genes with PofO effects, this information can determine whether disease will manifest. By enabling reliable segregation without parental testing, POAga helps direct clinical efforts toward those truly at risk and improves the clinical interpretation of variants. Ongoing analyses will refine its performance and support its adoption as a transformative tool in hereditary cancer genomics. Citation Format: Lilian Cordova, Vahid Akbari, Tiffany Leung, Kieran O’Neill, Katherine Dixon, Alexandra Roston, Eugene Cheung, Chuyi Zheng, Millicent Sharman, Alshanee Sharma, Steve Bilobram, Yaoqing Shen, Janine Senz, Yanni Wang, Daniel Chan, Alexandra Fok, Jennifer Nuk, Quang Hong, Robin Coope, Eric Chuah, Simon Chan, Hyun-Wu Lee, Yongjun Zhao, Miruna Bala, Karen Mungall, Andrew Mungall, Richard Moore, Nur Diana Binte Ishak, Siao Ting Chong, Ee Ling Chew, Ashley McDonald, Anna Martinez, Gregory Kelly, Rosella Delgado, Caitlin Orr, Joanne Yuen Yie Ngeow, Kara N. Maxwell, Stephen B. Gruber, Dean Regier, Alice Virani, Louis Lefebvre, Fabio Feldman, Marco Marra, Sophie Sun, Stephen Yip, Peter Lansdorp, Steven John Jones, Kasmintan Schrader. Parent-of-Origin-Aware genomic analysis in hereditary cancer: identifying the side of the family at risk using only the proband’s blood sample [abstract]. In: Proceedings of the American Association for Cancer Research Annual Meeting 2026; Part 1 (Regular Abstracts); 2026 Apr 17-22; San Diego, CA. Philadelphia (PA): AACR; Cancer Res 2026;86(7 Suppl):Abstract nr 5287.
10595 Background: Determining the parental origin of germline variants is a critical gap in clinical genetics, essential for risk management, variant classification, and cascade genetic testing. Traditional methods rely on testing family members, which can be time-consuming and impractical when relatives are unavailable, deceased, or unwilling to participate. Parent-of-Origin-Aware Genomic Analysis (POAga) offers a transformative solution by enabling accurate assignment of any autosomal variant to either parent with 99% accuracy using only a blood sample from the proband. This method integrates methylation and sequence data from Oxford Nanopore long-read sequencing with chromosome-length haplotypes generated from Strand-seq, leveraging the accurate phasing of imprinted differentially methylated regions (iDMRs) that occur on each autosome to infer the parent of origin (PofO) of variants across the genome. This study aims to validate POAga across multiple hereditary cancer syndromes, including high-penetrance conditions such as hereditary breast and ovarian cancer (HBOC) and Lynch syndrome, as well as rarer syndromes with PofO effects and other genes associated with breast and gastrointestinal malignancies. Methods: Blood samples from carriers of pathogenic variants in ATM , BRCA1 , BRCA2 , CDH1 , MLH1 , MSH2 , MSH6 , PMS2 , EPCAM , PALB2 , SDHD , SDHAF2 and TP53 with known parental segregation, are currently being ascertained and undergoing whole-genome analysis to determine the analytic validity of POAga. These samples span diverse demographics, including variations in age, sex, ethnicity, and cancer status. PofO predictions are made according to previously described methods (Akbari V, Hanlon VCT, et al . Cell Genom. 2022 Dec 21;3(1):100233) under an REB-approved protocol. Results: To date, 188 individuals carrying 189 pathogenic variants with known parental segregation have been analyzed. The distribution of variants includes BRCA2 (n=31), MLH1 (n=23), MSH2 (n=22), BRCA1 (n=22), SDHD (n=21), MSH6 (n=20), PALB2 (n=14), PMS2 (n=13), ATM (n=9), CDH1 (n=9), SDHAF2 (n=2), EPCAM (n=2) and TP53 (n=1). PofO assignment was successful for 172 of 189 (91%) variants. Only one sample with an MLH1 variant was misassigned, while all other cases demonstrated concordance between the predicted and known parental origin (188 of 189, 99.5% accuracy). Conclusions: These results support the ability of POAga to accurately infer the parental origin of pathogenic variants in diverse hereditary cancer syndromes using only blood sample from the proband. Ongoing validation will further assess its feasibility in real-world clinical settings and refine its clinical translation. POAga represents a powerful advancement in hereditary cancer genetics, with the potential transform how we conduct genetic cancer risk assessments for patients and families.
Understanding the human de novo mutation (DNM) rate requires complete sequence information1. Here using five complementary short-read and long-read sequencing technologies, we phased and assembled more than 95% of each diploid human genome in a four-generation, twenty-eight-member family (CEPH 1463). We estimate 98-206 DNMs per transmission, including 74.5 de novo single-nucleotide variants, 7.4 non-tandem repeat indels, 65.3 de novo indels or structural variants originating from tandem repeats, and 4.4 centromeric DNMs. Among male individuals, we find 12.4 de novo Y chromosome events per generation. Short tandem repeats and variable-number tandem repeats are the most mutable, with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 16% of de novo single-nucleotide variants are postzygotic in origin with no paternal bias, including early germline mosaic mutations. We place all this variation in the context of a high-resolution recombination map (~3.4 kb breakpoint resolution) and find no correlation between meiotic crossover and de novo structural variants. These near-telomere-to-telomere familial genomes provide a truth set to understand the most fundamental processes underlying human genetic variation.
The most common genomic disorder, chromosome 22q11.2 microdeletion syndrome (22q11.2DS), is mediated by highly identical and polymorphic segmental duplications (SDs) known as low copy repeats (LCRs; regions A-D) that have been challenging to sequence and characterize. Here, we report the sequence-resolved genomic architecture of 135 chromosome 22q11.2 haplotypes from diverse 1000 Genomes Project samples. We find that more than 90% of the copy number variation is polarized to the most proximal LCR region A (LCRA) where 50 distinct structural configurations are observed (~189 kbp to ~2.15 Mbp or 11-fold length variation). A higher-order SD cassette structure of 105 kbp in length, flanked by 25 kbp long inverted repeats, drives this variation and emerged in the human-chimpanzee ancestral lineage later expanding in humans ~1.0 [0.8-1.2] million years ago. African LCRA haplotypes are significantly longer (p=0.0047) when compared to non-Africans yet are predicted to be more protected against recurrent microdeletions (p=0.00053) due to a preponderance of flanking SDs in an inverted orientation. Conversely, we identified nine distinct inversion polymorphisms, including five recurrent ~2.28 Mbp inversions extending across the critical region (LCRA-D) and four smaller inversions (two LCRA-B, one LCRC-D, and one LCRB-D); 7/9 of these events were identified in haplotypes of African and admixed American ancestry. Finally, we sequence and assemble four families and show that LCRA-D deletion breakpoints map to the 105 kbp repeat unit while inversion breakpoints associate with the 25 kbp repeats adjacent to palindromic AT-rich regions. In one family, we observe evidence of more complex unequal crossover events associated with gene conversion and multiple breakpoints. Our findings suggest that specific haplotype configurations are protective and susceptible to chromosome 22q11.2DS while recurrent large-scale inversions help to explain why this syndrome is less prevalent among individuals of African descent.
Recent advances in genome sequencing have improved variant calling in complex regions of the human genome. However, it is difficult to quantify variant calling performance because existing standards often focus on specificity, neglecting completeness in difficult-to-analyze regions. To create a more comprehensive truth set, we used Mendelian inheritance in a large pedigree (CEPH-1463) to filter variants across PacBio high-fidelity (HiFi), Illumina and Oxford Nanopore Technologies platforms. This generated a variant map with over 4.7 million single-nucleotide variants, 767,795 insertions and deletions (indels), 537,486 tandem repeats and 24,315 structural variants, covering 2.77 Gb of the GRCh38 genome. This work adds 200 Mb of high-confidence regions, including 8
Using five complementary short- and long-read sequencing technologies, we phased and assembled >95% of each diploid human genome in a four-generation, 28-member family (CEPH 1463) allowing us to systematically assess de novo mutations (DNMs) and recombination. From this family, we estimate an average of 192 DNMs per generation, including 75.5 de novo single-nucleotide variants (SNVs), 7.4 non-tandem repeat indels, 79.6 de novo indels or structural variants (SVs) originating from tandem repeats, 7.7 centromeric de novo SVs and SNVs, and 12.4 de novo Y chromosome events per generation. STRs and VNTRs are the most mutable with 32 loci exhibiting recurrent mutation through the generations. We accurately assemble 288 centromeres and six Y chromosomes across the generations, documenting de novo SVs, and demonstrate that the DNM rate varies by an order of magnitude depending on repeat content, length, and sequence identity. We show a strong paternal bias (75-81%) for all forms of germline DNM, yet we estimate that 17% of de novo SNVs are postzygotic in origin with no paternal bias. We place all this variation in the context of a high-resolution recombination map (~3.5 kbp breakpoint resolution). We observe a strong maternal recombination bias (1.36 maternal:paternal ratio) with a consistent reduction in the number of crossovers with increasing paternal (r=0.85) and maternal (r=0.65) age. However, we observe no correlation between meiotic crossover locations and de novo SVs, arguing against non-allelic homologous recombination as a predominant mechanism. The use of multiple orthogonal technologies, near-telomere-to-telomere phased genome assemblies, and a multi-generation family to assess transmission has created the most comprehensive, publicly available "truth set" of all classes of genomic variants. The resource can be used to test and benchmark new algorithms and technologies to understand the most fundamental processes underlying human genetic variation.
10516 Background: Predicting which side of the family a germline variant comes from is a critical gap in current clinical practice, vital for risk management, variant curation, and cascade genetic testing. Assignment of any autosomal variant to either parent with 99% accuracy is now possible with only a blood sample from the proband. Parent-of-Origin-Aware genomic analysis (POAga) is achieved by combining methylation and sequence data from Oxford Nanopore long-read sequencing with chromosome-length haplotypes generated from Strand-seq to infer parent of origin of any variant along the length of a chromosome, due to accurate phasing of imprinted differentially methylated regions that occur on each autosome. We sought to validate POAga in common, high penetrant hereditary cancer conditions such as hereditary breast and ovarian cancer (HBOC) and Lynch syndrome, rarer syndromes with parent-of-origin-effects and other genes predisposing to breast cancer and gastrointestinal malignancies that are associated with genes across multiple chromosomes. Methods: Blood samples from carriers of pathogenic variants in ATM, BRCA1, BRCA2, CDH1, MLH1, MSH2, MSH6, PMS2, EPCAM, PALB2, SDHD and SDHAF2, with known parental segregation, are currently being ascertained to determine the analytic validity of POAga in real world samples of differing age, sex, ethnicity, and cancer status. Samples are undergoing whole genome analysis by long and short read sequencing. Parent-of-origin of the pathogenic or likely pathogenic variant or variant of uncertain significance is predicted according to previously described methods (Akbari V, Hanlon VCT, et al. Cell Genom. 2022 Dec 21;3(1):100233.) under an REB approved protocol. Results: To date, analysis is complete for 100 individuals with a total of 107 rare or pathogenic germline variants with known or presumed parental segregation. Germline variants are in SDHD (n = 18), BRCA2 (n = 17), BRCA1 (n = 14), MLH1 (n = 10), PALB2 (n = 9), PMS2 (n = 8), MSH2 (n = 9), MSH6 (n = 7), ATM (n = 5), CDH1 (n = 5), SDHAF2 (n = 1), DICER1 (n = 1), MUTYH (n = 1), RET (n = 1), and EPCAM (n = 1). Of variants able to be assigned a parent-of-origin (n = 104 of 107, 97%), there was complete concordance between the predicted parent-of-origin and known clinical segregation (n = 104, 100%). Conclusions: Results to date support the ability of POAga to accurately infer parent-of-origin of rare or pathogenic variants with known parental segregation using only a blood sample from carriers of diverse hereditary cancer syndromes. Ongoing validation of POAga will continue to test its feasibility in real-world samples and inform its path towards clinically translation. Parent-of-Origin-Aware genomic analysis is a powerful technology that could improve our understanding of hereditary cancer syndromes and transform our ability to conduct genetic cancer risk assessments for patients and families.
Encounters with pathogens and other molecules can imprint long-lasting effects on our immune system, influencing future physiological outcomes. Given the wide range of microbes to which humans are exposed, their collective impact on health is not fully understood. To explore relations between exposures and biological aging and inflammation, we profiled an antibody-binding repertoire against 2,815 microbial, viral, and environmental peptides in a population cohort of 1,443 participants. Utilizing antibody-binding as a proxy for past exposures, we investigated their impact on biological aging, cell composition and inflammation. Immune response against cytomegalovirus (CMV), rhinovirus and gut bacteria relates with telomere length. Single-cell expression measurements identified an effect of CMV infection on the transcriptional landscape of subpopulations of CD8 and CD4 T-cells. This examination of the relationship between microbial exposures and biological aging and inflammation highlights a role for chronic infections (CMV and Epstein-Barr Virus) and common pathogens (rhinoviruses and adenovirus C).
Alternative Lengthening of Telomeres (ALT) is an aberrant DNA recombination pathway which grants replicative immortality to approximately 10% of all cancers. Despite this high prevalence of ALT in cancer, the mechanism and genetics by which cells activate this pathway remain incompletely understood. A major challenge in dissecting the events that initiate ALT is the extremely low frequency of ALT induction in human cell systems. Guided by the genetic lesions that have been associated with ALT from cancer sequencing studies, we genetically engineered primary human pluripotent stem cells to deterministically induce ALT upon differentiation. Using this genetically defined system, we demonstrate that disruption of the p53 and Rb pathways in combination with ATRX loss-of-function is sufficient to induce all hallmarks of ALT and results in functional immortalization in a cell type-specific manner. We further demonstrate that ALT can be induced in the presence of telomerase, is neither dependent on telomere shortening nor crisis, but is rather driven by continuous telomere instability triggered by the induction of differentiation in ATRX-deficient stem cells.
Previous studies indicate that genomic loci harboring G-quadruplexes (G4s)—stacked structures that can form in single-stranded DNA—can be linked to epigenetic instability. However, the role of chromatin redistribution and the genome-wide nature of this process need further investigation. Here, we provide experimental evidence that connects G4s to alterations in the deposition of chromatin proteins. We have used metabolic labelling and immunoprecipitation of new and parental proteins in hRPE-1 cells to investigate global chromatin deposition dynamics. We identify a reciprocal, local bias in chromatin protein deposition at G4 sites favoring the association of parental proteins with the G4 and new proteins with the C4 DNA strand. The deposition bias at G4 sites does not depend on replication directionality and is strengthened by G4 stabilization. Slowing down replication forks upon hydroxyurea treatment reverses the bias, supposedly affected by decoupling between helicase and polymerase. Interestingly, upon combined G4 stabilization and slowing of the replication forks, new proteins exhibit a redistribution pattern similar to G4 stabilization alone, while parental protein redistribution more resembles one after hydroxyurea treatment, hinting at mechanistic differences between parental and new histone distribution. We also report that the genomic distribution of putative quadruplexes is not random and depends on loop size, where G4s with shorter loops have a preference for the DNA strand replicated by leading and G4s with longer loops by lagging strand replication. These findings provide insight into the mechanisms behind G4 occurrence and its role in epigenetic instability and help to improve our understanding of the factors influencing biases in global chromatin protein redeposition.
Dense local haplotypes can now readily be extracted from long-read or droplet-based sequence data. However, these methods struggle to combine subchromosomal haplotype blocks into global chromosome-length haplotypes. Strand-seq is a single cell sequencing technique that uses read orientation to capture sparse global phase information by sequencing only one of two DNA strands for each parental homolog. In combination with dense local haplotypes from other technologies, Strand-seq data can be used to obtain complete chromosome-length phase information. In this chapter, we run the R package StrandPhaseR to phase SNVs using publicly available sequence data for sample HG005 of the Genome in a Bottle project.
Mutations in the Blm gene can cause Bloom Syndrome, a genetic disorder characterized by genome instability and cancer predisposition. Blm encodes a helicase which was reported to resolve G-quadruplex DNA structures in vitro . The G-quadruplex resolving activity of the BLM helicase has been previously implicated in altering gene expression. However, the exact mechanisms of how G-quadruplex structures may affect gene expression remain to be elucidated. We employed experimentally defined G-quadruplex forming DNA sequences and generated transcriptomes for several Bloom Syndrome patient-derived cell lines and BLM-deficient mouse embryonic stem cells to further investigate the effect of G-quadruplexes on gene expression. Our results do not support the previous findings that G-quadruplexes located within a gene play a major role in altering its expression in BLM-deficient cells. We found concerted large-scale changes in transcript abundance, splicing, nucleosome occupancy, and phasing, that cannot be linked to the local presence of G-quadruplex sequences in either gene bodies or promotors. The investigation of genomic features associated with large-scale differences in nucleosome density highlights the rDNA locus and active enhancers as the most strongly affected regions. We hypothesize that global changes in chromatin structure rather than local G4s might mediate the transcriptome changes in the absence of BLM.
Hundreds of loci in human genomes have alleles that are methylated differentially according to their parent of origin. These imprinted loci generally show little variation across tissues, individuals, and populations. We show that such loci can be used to distinguish the maternal and paternal homologs for all human autosomes without the need for the parental DNA. We integrate methylation-detecting nanopore sequencing with the long-range phase information in Strand-seq data to determine the parent of origin of chromosome-length haplotypes for both DNA sequence and DNA methylation in five trios with diverse genetic backgrounds. The parent of origin was correctly inferred for all autosomes with an average mismatch error rate of 0.31% for SNVs and 1.89% for insertions or deletions (indels). Because our method can determine whether an inherited disease allele originated from the mother or the father, we predict that it will improve the diagnosis and management of many genetic diseases.
How many times does a typical hematopoietic stem cell (HSC) divide to maintain a daily production of over 1011 blood cells over a human lifetime? It has been predicted that relatively few, slowly dividing HSCs occupy the top of the hematopoietic hierarchy. However, tracking HSCs directly is extremely challenging due to their rarity. Here, we utilize previously published data documenting the loss of telomeric DNA repeats in granulocytes, to draw inferences about HSC division rates, the timing of major changes in those rates, as well as lifetime division totals. Our method uses segmented regression to identify the best candidate representations of the telomere length data. Our method predicts that, on average, an HSC divides 56 times over an 85-year lifespan (with lower and upper bounds of 36 and 120, respectively), with half of these divisions during the first 24 years of life.
Mammalian genomes encode over a hundred different helicases, many of which are implicated in the repair of DNA lesions by acting on DNA structures arising during DNA replication, recombination or transcription. Defining the in vivo substrates of such DNA helicases is a major challenge given the large number of helicases in the genome, the breadth of potential substrates in the genome and the degree of genetic pleiotropy among DNA helicases in resolving diverse substrates. Helicases such as WRN, BLM and RECQL5 are implicated in the resolution of error-free recombination events known as sister chromatid exchange events (SCEs). Single cell Strand-seq can be used to map the genomic location of individual SCEs at a resolution that exceeds that of classical cytogenetic techniques by several orders of magnitude. By mapping the genomic locations of SCEs in the absence of different helicases, it should in principle be possible to infer the substrate specificity of specific helicases. Here we describe how the genome can be interrogated for such DNA repair events using single-cell template strand sequencing (Strand-seq) and bioinformatic tools. SCEs and copy-number alterations were mapped to genomic locations at kilobase resolution in haploid KBM7 cells. Strategies, possibilities, and limitations of Strand-seq to study helicase function are illustrated using these cells before and after CRISPR/Cas9 knock out of WRN, BLM and/or RECQL5.
Aneuploidy and chromosomal instability are both commonly found in cancer. Chromosomal instability leads to karyotype heterogeneity in tumors and is associated with therapy resistance, metastasis and poor prognosis. It has been hypothesized that aneuploidy per se is sufficient to drive CIN, however due to limited models and heterogenous results, it has remained controversial which aspects of aneuploidy can drive CIN. In this study we systematically tested the impact of different types of aneuploidies on the induction of CIN. We generated a plethora of isogenic aneuploid clones harboring whole chromosome or segmental aneuploidies in human p53-deficient RPE-1 cells. We observed increased segregation errors in cells harboring trisomies that strongly correlated to the number of gained genes. Strikingly, we found that clones harboring only monosomies do not induce a CIN phenotype. Finally, we found that an initial chromosome breakage event and subsequent fusion can instigate breakage-fusion-bridge cycles. By investigating the impact of monosomies, trisomies and segmental aneuploidies on chromosomal instability we further deciphered the complex relationship between aneuploidy and CIN.