Understanding the genetic basis of gene expression can shed light on the regulatory mechanisms underlying complex traits and diseases. Single-cell resolved measures of RNA levels and single-cell expression quantitative trait loci (sc-eQTLs) have revealed genetic regulation that drives sub-tissue cell states and types across diverse human tissues. Here, we describe the first phase of TenK10K, the largest-to-date dataset of matched whole-genome sequencing (WGS) and single-cell RNA-sequencing (scRNA-seq). We leverage scRNA-seq data from over 5 million cells across 28 immune cell types and matched WGS from 1,925 individuals. This provides power to detect associations between rare and low-frequency genetic variants that have largely been uncharacterised in their impact on cell-specific gene expression. We map the effects of both common and rare variants in a cell type specific manner using SAIGE-QTL. This newly developed method increases power by modelling single cells directly using a Poisson model rather than relying on aggregated 'pseudobulk' counts. We identify putative common regulatory variants for 83% of all 21,404 genes tested and cumulative rare variant signals for 47% of genes. We explore how genetic effects vary across cell type and state spectra, develop a framework to determine the degree to which sc-eQTLs are cell type specific, and show that about half of the effects are observed only in one or a few cell types. By integrating our results with functional annotations and disease information, we further characterise the likely molecular modes of action for many disease-associated variants. Finally, we explore the effects of genetic variants on gene expression across different cell states and functions, as well as effects that directly vary cell state abundance. ### Competing Interest Statement E.B.D., L.C., and K.K.H.F. are employed at Illumina Inc. D.G.M. is a paid advisor to Insitro and GSK, and receives research funding from Google and Microsoft, unrelated to the work described in this manuscript. G.A.F reports grants from National Health and Medical Research Council (Australia), grants from Abbott Diagnostic, Sanofi, Janssen Pharmaceuticals, and NSW Health. G.A.F reports honorarium from CSL, CPC Clinical Research, Sanofi, Boehringer-Ingelheim, Heart Foundation, and Abbott. G.A.F serves as Board Director for the Australian Cardiovascular Alliance (past President), Executive Committee Member for CPC Clinical Research, Founding Director and CMO for Prokardia and Kardiomics, and Executive Committee member for the CAD Frontiers A2D2 Consortium. In addition, G.A.F serves as CMO for the non-profit, CAD Frontiers, with industry partners including, Novartis, Amgen, Siemens Healthineers, ELUCID, Foresite Labs LLC, HeartFlow, Canon, Cleerly, Caristo, Genentech, Artyra, and Bitterroot Bio, Novo Nordisk and Allelica. In addition, G.A.F has the following patents: "Patent Biomarkers and Oxidative Stress" awarded USA May 2017 (US9638699B2) issued to Northern Sydney Local Health District, "Use of P2X7R antagonists in cardiovascular disease" PCT/AU2018/050905 licensed to Prokardia, "Methods for treatment and prevention of vascular disease" PCT/AU2015/000548 issued to The University of Sydney/Northern Sydney Local Health District, "Methods for predicting coronary artery disease" AU202290266 issued to The University of Sydney, and the patent "Novel P2X7 Receptor Antagonists" PCT/AU2022/051400 (23.11.2022), International App No: WO/2023/092175 (01.06.2023), issued to The University of Sydney. The other authors declare no competing interests. ### Funding Statement Data generation and analysis were partly supported by funding from National Health and Medical Research Council (NHMRC) investigator grants (2009982 and 2034556). ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Human Research Ethics Committee of St Vincent's Hospital gave ethical approval for this work. The National Statement on Ethical Conduct in Human Research of the National Health and Medical Research Council gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All gene-level summary statistics of our results are provided as supplementary tables.
Axenfeld-Rieger Syndrome (ARS) is an autosomal dominant condition with both ocular and non-ocular manifestations. ARS is primarily caused by coding variants at the PITX2 or FOXC1 loci, yet many cases still remain undiagnosed. Here we used whole-genome sequencing to identify two non-coding structural variants associated with a typical presentation of PITX2-associated ARS: one with a 450 kb deletion removing a series of conserved enhancer elements distal to PITX2, and the second with a 12.5 Mb inversion displacing the PITX2 gene from these same enhancer elements. Neither variant disrupted the PITX2 gene itself, and therefore both were expected to reduce PITX2 expression by disrupting its proximity or access to enhancer elements. Enhancer-disrupting intergenic inversions therefore represent a unique genetic mechanism for the development of ARS, which should be carefully considered in the context of ARS and other conditions without a conclusive genetic diagnosis.
Collagen VI-related dystrophies manifest with a spectrum of clinical phenotypes, ranging from Ullrich congenital muscular dystrophy (UCMD), presenting with prominent congenital symptoms and characterized by progressive muscle weakness, joint contractures and respiratory insufficiency, to Bethlem muscular dystrophy, with milder symptoms typically recognized later and at times resembling a limb girdle muscular dystrophy, and intermediate phenotypes falling between UCMD and Bethlem muscular dystrophy. Despite clinical and muscle pathology features highly suggestive of collagen VI-related dystrophy, some patients had remained without an identified causative variant in COL6A1, COL6A2 or COL6A3. With combined muscle RNA sequencing and whole-genome sequencing, we uncovered a recurrent, de novo deep intronic variant in intron 11 of COL6A1 (c.930+189C>T) that leads to a dominantly acting in-frame pseudoexon insertion. We subsequently identified and have characterized an international cohort of 44 patients with this COL6A1 intron 11 causative variant, one of the most common recurrent causative variants in the collagen VI genes. Patients manifest a consistently severe phenotype characterized by a paucity of early symptoms followed by an accelerated progression to a severe form of UCMD, except for one patient with somatic mosaicism for this COL6A1 intron 11 variant who manifests a milder phenotype consistent with Bethlem muscular dystrophy. Partial amelioration of the disease phenotype in this individual provides a strong rationale for the development of our pseudoexon skipping therapy to successfully suppress the pseudoexon insertion, resulting in normal COL6A1 transcripts. We have previously shown that splice-modulating antisense oligomers applied in vitro effectively decreased the abundance of the mutant pseudoexon-containing COL6A1 transcripts to levels comparable to the in vivo scenario of the somatic mosaicism shown here, indicating that this therapeutic approach carries significant translational promise for ameliorating the severe form of UCMD caused by this common recurrent COL6A1 variant.
Reanalysis of genomic data in rare disease is highly effective in increasing diagnostic yields but remains limited by manual approaches. Automation and optimization for high specificity will be necessary to ensure scalability, adoption and sustainability of iterative reanalysis. We developed a publicly available automated tool, Talos, and validated its performance using data from 1,089 individuals with rare genetic disease. Trio-based analysis identified 86% of known in-scope diagnoses, returning one variant per case on average. Variant burden reduced to one variant per 200 cases on iterative monthly reanalysis cycles. Application to an unselected cohort of 4,735 undiagnosed individuals identified 248 diagnoses (5.2% yield): 73 (29%) due to new gene-disease relationships, 56 (23%) due to new variant-level evidence, and 119 (48%) due to improved filtering and analysis strategies. Our automated, iterative reanalysis model, applied to thousands of rare disease patients, demonstrates the feasibility of delivering frequent, systematic reanalysis at scale.
Incomplete penetrance, or absence of disease phenotype in an individual with a disease-associated variant, is a major challenge in variant interpretation. Studying individuals with apparent incomplete penetrance can shed light on underlying drivers of altered phenotype penetrance. Here, we investigate clinically relevant variants from ClinVar in 807,162 individuals from the Genome Aggregation Database (gnomAD), demonstrating improved representation in gnomAD version 4. We then conduct a comprehensive case-by-case assessment of 734 predicted loss of function variants in 77 genes associated with severe, early-onset, highly penetrant haploinsufficient disease. Here, we identify explanations for the presumed lack of disease manifestation in 701 of 734 variants (95%). Individuals with unexplained lack of disease manifestation in this set of disorders are rare, underscoring the need and power of deep case-by-case assessment presented here to minimize false assignments of disease risk, particularly in unaffected individuals with higher rates of secondary properties that result in rescue.
Recently, de novo variants in an 18 nucleotide region in the centre of RNU4-2 were shown to cause ReNU syndrome, a syndromic neurodevelopmental disorder (NDD) that is predicted to affect tens of thousands of individuals worldwide 1,2 . RNU4-2 is a non-protein-coding gene that is transcribed into the U4 small nuclear RNA (snRNA) component of the major spliceosome 3 . ReNU syndrome variants disrupt spliceosome function and alter 5’ splice site selection 1,4 . Here, we performed saturation genome editing (SGE) of RNU4-2 to identify the functional and clinical impact of variants across the entire gene. The resulting SGE function scores, derived from variants’ effects on cell fitness, discriminate ReNU syndrome variants from those observed in the population and dramatically outperform in silico variant effect prediction. Using these data, we redefine the ReNU syndrome critical region at single nucleotide resolution, resolve variant pathogenicity for variants of uncertain significance, and show that SGE function scores delineate variants by phenotypic severity. Further, we identify variants impacting function in regions of RNU4-2 that are critical for interactions with other spliceosome components. We show that these variants cause a novel recessive NDD that is clinically distinct from ReNU syndrome. Together, this work defines the landscape of variant function across RNU4-2 , providing critical insights for both diagnosis and therapeutic development.
Genome-wide association studies (GWASs) are often performed on ratios composed of a numerator trait divided by a denominator trait. Examples include body mass index (BMI) and the waist-to-hip ratio, among many others. Explicitly or implicitly, the goal of forming the ratio is typically to adjust for an association between the numerator and denominator. While forming ratios may be clinically expedient, there are several important issues with performing GWAS on ratios. Forming a ratio does not "adjust"for the denominator in the sense of conditioning on it, and it is unclear whether associations with ratios are attributable to the numerator, the denominator, or both. Here we demonstrate that associations arising in ratio GWAS can be entirely denominator driven, implying that at least some associations uncovered by ratio GWAS may be due solely to a putative adjustment variable. In a survey of 10 common ratio traits, we find that the ratio model disagrees with the adjusted model (performing GWAS on the numerator while conditioning on the denominator) at around 1/3 of loci. Using BMI as an example, we show that variants detected by only the ratio model are more strongly associated with the denominator (height), while variants detected by only the adjusted model are more strongly associated with the numerator (weight). Although the adjusted model provides effect sizes with a clearer interpretation, it is susceptible to collider bias. We propose and validate a simple method of correcting for the genetic component of collider bias via leave-one-chromosome-out polygenic scoring.
Genetic variants in RNU4-2, which encodes U4, a key non-coding small nuclear RNA (snRNA) component of the major spliceosome, were recently shown to cause a prevalent neurodevelopmental disorder (NDD) called ReNU syndrome. These variants, which almost exclusively arise de novo, act in a dominant fashion and are clustered within 18 nucleotides (nt) in the centre of RNU4-2. Here we describe a novel recessive NDD associated with homozygous and compound heterozygous variants in RNU4-2. We identified 32 individuals with biallelic variants outside of the 18 nt ReNU syndrome region, that cluster within other functionally important elements of U4, including the Stem II region, the k-turn motif, and the Sm protein binding site. We characterise the clinical phenotype in 27 of these individuals, demonstrating that the recessive disorder is clinically distinct from dominant ReNU syndrome and is associated with distinctive white matter abnormalities, including enlarged perivascular spaces. Together, these findings expand the genotypic and phenotypic spectrum of RNU4-2-associated NDDs. ### Competing Interest Statement N.W. receives research funding from Novo Nordisk and BioMarin Pharmaceutical. D.G.M. is a paid consultant for GlaxoSmithKline, Insitro and Overtone Therapeutics and receives research support from Microsoft. D.P. provides consulting service to Ionis Pharmaceuticals, Acadia Pharmaceuticals and M2DS Therapeutics. S.J.S. receives research funding from BioMarin Pharmaceutical. A.OD.-L. is on the scientific advisory board for Congenica, was a paid consultant for Tome Biosciences, Ono Pharma USA Inc. and at present for Addition Therapeutics, and received reagents from PacBio to support rare disease research. Y.C. has a PhD studentship funded by Novo Nordisk. All other authors declare no competing interests. ### Funding Statement N.W. is supported by a Wellcome Career Development Award (grant no. 305292/Z/23/Z), a Lister Institute research prize, and grant funding from Novo Nordisk. Y.C. is supported by a studentship from Novo Nordisk. The Francis Crick Institute receives its core funding (G.M.F.) from Cancer Research UK (CC2190), the UK Medical Research Council (CC2190), and the Wellcome Trust (CC2190). A.B. is supported by a Wellcome PhD Training Fellowship for Clinicians and the 4Ward North PhD Programme for Health Professionals (223521/Z/21/Z). Analysis was supported by the Centre for Population Genomics (Garvan Institute of Medical Research and Murdoch Childrens Research Institute) and was funded in part by a National Health and Medical Research Council investigator grant (2009982) and the Medical Research Future Fund (MRFF) Genomics Health Futures Mission (2032931). Massimos Mission acknowledges funding support from the Australian Government Department of Health and Aged Care (EPCD000034). O.M. is supported by the Hazem Ben-Gacem Tunisia Medical Fellowship Fund. D.G.C. was supported by the National Institute of Neurological Disorders and Stroke of the National Institutes of Health under award number K12NS098482. This study was supported by the National Institute for Health and Care Research (NIHR) Manchester Biomedical Research Centre (NIHR203308) and funded in part by US National Institutes of Health Genomics Research Elucidates Genetics of Rare, GREGoR, Program (HG011758). Sequencing and analysis of Individual 24 were provided by the Broad Institute Center for Mendelian Genomics (Broad CMG) and were funded by the National Human Genome Research Institute (NHGRI) grants U01HG011755 (GREGoR consortium), R01HG009141, and in part by the Chan Zuckerberg Initiative Donor-Advised Fund at the Silicon Valley Community Foundation (funder DOI 10.13039/100014989) grants 2020-224274, 2022-309464, 2022-316726, and 2022-316726 (https://doi.org/10.37921/236582yuakxy). Research reported in this publication was supported by the National Institute Of Neurological Disorders And Stroke of the National Institutes of Health under Award Number U01NS134358. The content is solely the responsibility of the authors and does not necessarily represent the official views of the funding agencies. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: Informed consent was obtained for all patients included in this study from their parent(s) or legal guardian, with the study approved by the local regulatory authority. The 100,000 Genomes Project Protocol has ethical approval from the HRA Committee East of England Cambridge South (REC Ref 14/EE/1112). I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes Research on the de-identified patient data used in this publication from the Genomics England 100,000 Genomes Project and the NHS GMS dataset can be carried out in the Genomics England Research Environment subject to a collaborative agreement that adheres to patient-led governance. All interested readers will be able to access the data in the same manner that the authors accessed the data. For more information about accessing the data, interested readers may contact research-network{at}genomicsengland.co.uk or access the relevant information on the Genomics England website: https://www.genomicsengland.co.uk/research. Genomic and phenotypic data from the GREGoR consortium, including the RGP cohort, and the UDN are available through the dbGaP accession numbers phs003047.v1.p1 and phs001232.v5.p2, respectively, with at least annual data releases. Access is managed by a data access committee designated by dbGaP and is based on intended use of the requester and allowed use of the data submitter as defined by consent codes.
Genetic kidney disease (GKD) significantly affects the community and is responsible for a notable portion of adult kidney disease cases and about half of cases in paediatric patients. It substantially impacts the quality of life and life expectancy for affected children and adults across all stages of kidney disease. Precise genetic diagnosis in GKD promises to improve patient outcomes, provide access to targeted treatments, and reduce the disease burden for individuals, families, and healthcare systems. Genetic investigations are increasingly used in nephrology practice; however, many patients who undergo testing still lack a definitive diagnosis. The KidGen National Kidney Genomics Study aims to increase diagnostic yield for those with suspected monogenic kidney disease without a diagnosis after standard diagnostic genetic testing. The program will seek to enrol up to 200 families from KidGen Collaborative kidney genetics clinics across Australia who have yet to receive conclusive diagnoses despite prior testing. Participants will undergo a personalised pathway of research genomic investigations. These include re-analysing existing data and/or undergoing advanced genomic testing methods, including short and long-read whole-genome sequencing, RNA sequencing, and functional genomics strategies using mouse modelling or kidney organoids. The KidGen National Kidney Genomics Study is a coordinated, multidisciplinary extension of previous research projects that aims to assess the diagnostic yield of advanced genomic approaches. The study's evidence will drive changes to current diagnostic pathways, including identifying which chronic kidney disease patients are most likely to benefit from a more comprehensive genomic approach to diagnosis.
Tandem repeats (TRs) - highly polymorphic, repetitive sequences dispersed across the human genome - are crucial regulators of gene expression and diverse biological processes, but have remained underexplored relative to other classes of genetic variation due to historical challenges in their accurate calling and analysis. Here, we leverage whole genome and single-cell RNA sequencing from over 5.4 million blood-derived cells from 1,925 individuals to explore the impact of variation in over 1.7 million polymorphic TR loci on blood cell type-specific gene expression. We identify over 62,000 single-cell expression quantitative trait TR loci (sc-eTRs), 16.6% of which are specific to one of 28 distinct immune cell types. Further fine-mapping uncovers 4,283 sc-eTRs as candidate causal drivers of gene expression in 13.6% of genes tested genome-wide. We show through colocalization that TRs are likely mediators of genetic associations with immune-mediated and hematological traits in over 700 genes, and further identify novel TRs warranting investigation in rare disease cohorts. TRs are critical, yet long-overlooked, contributors to cell type-specific gene expression, with implications for understanding rare disease pathogenesis and the genetic architecture of complex traits.
Abstract Inherited bone marrow failure syndromes (IBMFS) are a group of monogenic diseases of diverse pathogenesis manifesting as single or multilineage cytopenia typically due to hypoproliferative dyshematopoiesis. Accurately diagnosing IBMFS is challenging given the overlapping clinicopathological features between individual genetic syndromes as well as with acquired BMFS (e.g. immune aplastic anemia). Accurate genetic diagnosis in IBMFS is critical and most commonly involves targeted panel DNA sequencing or whole exome sequencing. These approaches do not cover the full spectrum of possible genomic abnormalities which in turn may contribute to a significant proportion of patients with IBMFS remaining undiagnosed. We aimed to comprehensively evaluate the unbiased upfront approach of whole genome transcriptome sequencing (WGTS) in patients with suspected IBMFS (the IBMDx study). The IBMDx study aimed to (i) determine the diagnostic rate and clinical impact of upfront WGTS (ii) assess the acceptability of WGTS to patients and physicians through an implementation science framework (iii) evaluate the health-economic impact and cost-effectiveness of the approach and (iv) discover and functionally characterize novel genes and variants in IBMFS. 237 patients were enrolled from March 2022 to March 2025. The median age of the cohort was 30 years (range 6m-78 years; M:F 0.93:1). 70/237 (29.5%) were under 18 years. WGS was performed exclusively on non-hematological DNA (hair follicle DNA [n=178], skin biopsy/cultured skin fibroblasts [n=59]). Trio WGS was performed in 18 families. WGS was performed by a clinically accredited service and results returned to physicians/patients in real time for patient management and segregation/predictive testing as required. A genomic diagnosis for the hematological phenotype was established in 88/237 (37.1%) patients. Diagnoses made included hereditary thrombocytopenia (n=19), telomere biology disorders (TBDs) (n=14), Diamond-Blackfan anemia (n=12), primary red cell disorders (n=6), severe congenital neutropenia (n=6), Shwachman Diamond Syndrome (n=5) and Fanconi anemia (n=4). Targeted sequencing of the hematological compartment demonstrated clonal hematopoiesis in 32/217 (14.7%) patients. In addition, four patients had genetic diagnoses made of symptomatic and clinically significant diseases unrelated to the hematological phenotype (TAP2, IRF2BP2, ACADM and mosaic trisomy 7). Secondary genomic findings requiring further management were detected in 7 patients (TNNI3, MSH6, Monosomy X mosaic, BRCA1, BRIP1, ATM, FBN1). WGTS revealed novel genomic abnormalities not previously established in IBMFS such as retrotransposon-mediated gene disruption, polyadenylation site loss and disruption of novel regulatory regions. Novel genomic abnormalities in genes associated with Diamond-Blackfan anaemia (RPL31), TBDs (TERT/TERC) and hereditary thrombocytopenia (TPM4) underwent in vivo and in vitro functional assessment to inform pathogenicity. Finally, novel associations were identified with primary IBMFS-like presentations (THRA) as well as definitive evidence of gene-disease association were established (MEIS1, TUBB). Using a formal Theoretical Framework of Acceptability, WGTS was found to be highly acceptable to patients and carers as evidenced by positive perceptions of clinicians, alignment with healthcare expectations, and minimal effort required to participate. However, concern regarding equity of access to technology for all patients with suspected IBMFS was a recurrent theme amongst patients and carers interviewed. Health economic analysis showed that the cost of achieving maximum diagnostic yield, incorporating varied diagnostic strategies, to be between $7,800 - $8,200 USD per hematological diagnosis. In summary, we have comprehensively evaluated upfront WGTS in a large cohort of adult and pediatric patients with suspected IBMFS and have found this approach has a high diagnostic rate (37.1%), is highly acceptable to patients, and uncovered multiple genomic abnormalities that would have been missed with more traditional diagnostic approaches. Moreover, we have uncovered new genomic mechanisms of disease and genes associated with IBMFS leading to new areas of research into the underlying biology of these challenging diseases.
Genome-wide association studies (GWAS) have been instrumental in uncovering the genetic basis of complex traits. When integrated with expression quantitative trait loci (eQTL) mapping, they can elucidate how risk loci influence traits through gene regulatory mechanisms. Recent single-cell eQTL (sc-eQTL) studies suggest that genetic effects on gene expression are often cell type- and subtype-specific, but such datasets have so far been underpowered for causal inference. Here, we leverage results from sc-eQTL mapping in the TenK10K project, comprising 154,932 common variant sc-eQTL across 28 immune cell types derived from matched whole-genome sequencing (WGS) and single-cell RNA-sequencing (scRNA-seq) of over 5 million peripheral blood mononuclear cells (PBMCs) from 1,925 individuals. We present a catalogue of cell type-specific causal effects of gene expression on 53 diseases (spanning 58,058 causal associations across 8,672 genes and 28 cell types), and 31 biomarker traits (spanning 681,480 causal associations across 16,085 genes and 28 cell types). By quantifying polygenic enrichment at both the single-cell and cell-type levels, we identify distinct immune cell contributions to both immune-related and systemic conditions. We demonstrate differential polygenic enrichment of Crohn's disease and COVID-19 amongst dendritic cell subtypes, and high activity of B cell interferon II response in SLE. Integration with clinical drug development data reveals that therapeutic compounds targeting gene-trait associations identified in this study are three times more likely to have secured regulatory approval. Using Crohn's disease as a motivating example, we demonstrate how population-based sc-eQTL data can pinpoint risk loci, effector genes and cell types, complementing findings from disease-focused tissue samples. Our findings provide a foundational resource for understanding the cell type-specific genetic architecture of disease and for guiding therapeutic discovery. ### Competing Interest Statement D.G.M. is a paid advisor to Insitro and GSK, and receives research funding from Google and Microsoft, unrelated to the work described in this manuscript. G.A.F reports grants from National Health and Medical Research Council (Australia), Abbott Diagnostic, Sanofi, Janssen Pharmaceuticals, and NSW Health, and honorarium from CSL, CPC Clinical Research, Sanofi, Boehringer-Ingelheim, Heart Foundation, and Abbott. G.A.F serves as Board Director for the Australian Cardiovascular Alliance (past President), Executive Committee Member for CPC Clinical Research, Founding Director and CMO for Prokardia and Kardiomics, and Executive Committee member for the CAD Frontiers A2D2 Consortium. In addition, G.A.F serves as CMO for the non-profit, CAD Frontiers, with industry partners including, Novartis, Amgen, Siemens Healthineers, ELUCID, Foresite Labs LLC, HeartFlow, Canon, Cleerly, Caristo, Genentech, Artyra, and Bitterroot Bio, Novo Nordisk and Allelica. G.A.F also reports the following patents: "Patent Biomarkers and Oxidative Stress" awarded USA May 2017 (US9638699B2) issued to Northern Sydney Local Health District, "Use of P2X7R antagonists in cardiovascular disease" PCT/AU2018/050905 licensed to Prokardia, "Methods for treatment and prevention of vascular disease" PCT/AU2015/000548 issued to The University of Sydney/Northern Sydney Local Health District, "Methods for predicting coronary artery disease" AU202290266 issued to The University of Sydney, and the patent "Novel P2X7 Receptor Antagonists" PCT/AU2022/051400 (23.11.2022), International App No: WO/2023/092175 (01.06.2023), issued to The University of Sydney. ### Funding Statement Data generation and analysis were partly supported by funding from National Health and Medical Research Council (NHMRC) investigator grants (2009982 and 2034556) ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Human Research Ethics Committee of St Vincent's Hospital gave ethical approval for this work. The National Statement on Ethical Conduct in Human Research of the National Health and Medical Research Council gave ethical approval for this work. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes The main single-cell MR association results will be made publicly available via Zenodo website prior to acceptance.
Penetrance is the probability that an individual with a pathogenic genetic variant develops a specific disease. Knowing the penetrance of variants for monogenic disorders is important for counseling of individuals. Until recently, estimates of penetrance have largely relied on affected individuals and their at-risk family members being clinically referred for genetic testing, a 'phenotype-first' approach. This approach substantially overestimates the penetrance of variants because of ascertainment bias. The recent availability of whole-genome sequencing data in individuals from very-large-scale population-based cohorts now allows 'genotype-first' estimates of penetrance for many conditions. Although this type of population-based study can underestimate penetrance owing to recruitment biases, it provides more accurate estimates of penetrance for secondary or incidental findings. Here, we provide guidance for the conduct of penetrance studies to ensure that robust genotypes and phenotypes are used to accurately estimate penetrance of variants and groups of similarly annotated variants from population-based studies.
Collagen VI-related dystrophies (COL6-RDs) manifest with a spectrum of clinical phenotypes, ranging from Ullrich congenital muscular dystrophy (UCMD), presenting with prominent congenital symptoms and characterised by progressive muscle weakness, joint contractures and respiratory insufficiency, to Bethlem muscular dystrophy, with milder symptoms typically recognised later and at times resembling a limb girdle muscular dystrophy, and intermediate phenotypes falling between UCMD and Bethlem muscular dystrophy. Despite clinical and immunohistochemical features highly suggestive of COL6-RD, some patients had remained without an identified causative variant in COL6A1, COL6A2 or COL6A3. With combined muscle RNA-sequencing and whole-genome sequencing we uncovered a recurrent, de novo deep intronic variant in intron 11 of COL6A1 (c.930+189C>T) that leads to a dominantly acting in-frame pseudoexon insertion. We subsequently identified and have characterised an international cohort of forty-four patients with this COL6A1 intron 11 causative variant, one of the most common recurrent causative variants in the collagen VI genes. Patients manifest a consistently severe phenotype characterised by a paucity of early symptoms followed by an accelerated progression to a severe form of UCMD, except for one patient with somatic mosaicism for this COL6A1 intron 11 variant who manifests a milder phenotype consistent with Bethlem muscular dystrophy. Characterisation of this individual provides a robust validation for the development of our pseudoexon skipping therapy. We have previously shown that splice-modulating antisense oligomers applied in vitro effectively decreased the abundance of the mutant pseudoexon-containing COL6A1 transcripts to levels comparable to the in vivo scenario of the somatic mosaicism shown here, indicating that this therapeutic approach carries significant translational promise for ameliorating the severe form of UCMD caused by this common recurrent COL6A1 causative variant to a Bethlem muscular dystrophy phenotype.