OBJECTIVE:Exome sequencing (ES) detects a monogenic disorder in approximately 30% of pregnancies with unexplained nonimmune hydrops fetalis spectrum (NIHFS), defined as the presence of one or more pathologic fetal fluid collection. Since genome sequencing (GS) can identify additional genetic variants not detectable by ES, we sought to determine the proportion of cases with new reportable findings on GS following nondiagnostic standard of care testing, including ES, for NIHFS. STUDY DESIGN:Cohort study of pregnancies with NIHFS that had nondiagnostic results of karyotype and/or chromosomal microarray analysis (CMA) and ES. Eligible fetal fluid collections included at least one of: ascites, pleural or pericardial effusions, skin edema, cystic hygroma, and/or nuchal translucency ≥3.5 mm. Pregnancies with NIHFS that did not receive a clear diagnosis with prior testing underwent additional testing with GS. ES and GS were performed in a Clinical Laboratory Improvement Amendments-approved laboratory, and genetic variants were classified according to the American College of Medical Genetics and Genomics guidelines. Primary outcome was the proportion of new reportable pathogenic or likely pathogenic (P/LP) variants identified on GS and resulting in a new diagnosis. Secondary outcome was the incremental yield of GS for detecting new secondary findings and variants of uncertain significance (VUSs). RESULTS:Overall, 118 fetal samples underwent GS following negative (n=102) or inconclusive (n=16) ES. GS identified four P/LP copy number variants previously not detected by karyotype, CMA, or ES that led to a new genetic diagnosis in three pregnancies, providing a new positive finding yield of 2.5% (3/118). Additionally, one secondary finding and four VUSs were identified by GS in four different pregnancies. In contrast to the primary outcome, these additional findings were largely the result of updated analytical tools and literature rather than GS-specific technology. CONCLUSION:GS has a small but clinically relevant increased diagnostic yield over chromosome studies and ES for NIHFS, particularly for identifying copy number variants. As interpretation of noncoding regions and complex genomic variations improves, the incremental positive finding yield of GS is expected to further increase.
Pathogenic DEGS1 variants have been reported in individuals with autosomal recessive hypomyelinating leukodystrophy 18 (HLD18; MIM# 618404). Here we describe three participants with HLD features and a previously unreported homozygous DEGS1 5′ splice site variant, c.825+4_825 + 5delAGinsTT (NM_003676.4). We used next-generation DNA and transcriptome sequencing, cell-based splicing assays, and tandem mass spectrometry to detect and characterize the variant’s impact on DEGS1 expression. We then performed RNA structure probing and conventional antisense oligonucleotide screening to investigate molecular mechanisms for potential therapeutic intervention. We show that the splice site variant: (1) was sufficient to induce exon two skipping in most detected transcripts; (2) resulted in structural changes to the 5′ and 3′ splice site regions using RNA structure probing; and (3) corresponds to plasma sphingolipid profiles consistent with loss of sphingolipid delta(4)-desaturase activity. Our RNA and lipidomic evidence proved that the DEGS1 variant c.825+4_825 + 5delAGinsTT is pathogenic and suggested a mechanistic model that explains how exon two skipping is induced.
Yield of reported results from genetic testing provides a proximal measure of clinical usefulness. While ACMG/AMP guidelines provide representations of uncertainty for individual genetic variant classification, additional factors are considered when determining whether results explain a patient's presentation. To standardize cross-consortium analysis, a working group of the Clinical Sequencing Evidence-Generating Research (CSER2) consortium iteratively identified factors used when contextualizing variant-level results to case-level interpretation (i.e., interpretation of an individual's genetic data with respect to the indication for testing). Sites independently categorized results; complex cases were discussed collaboratively, leading to revision of classification categories. Our metric incorporates factors beyond classification of reported variants. Analogous to variant-level results, "Definitive Positive" and "Probable Positive" represent certainty that results may be clinically explanatory. The category "Inconclusive" applies when results may or may not fully explain the patient presentation, with subdivision into multiple (non-exclusive) subcategories. Cases falling outside all of the other categories are considered "Negative". The overall diagnostic yield by this metric and use of categories for inconclusive results varied by CSER project, in part paralleling study design differences. This case-level categorization provides a meaningful assessment of diagnostic yield, and for inconclusive cases identifies potentially resolvable factors for case resolution.
Purpose:Diagnostic yield from exome and genome sequencing varies widely across studies. It remains unclear how much of this variation reflects patient-level factors (e.g., sex, clinical features, race/ethnicity, genetic ancestry) versus site-level practices such as sequencing modality or variant interpretation workflows. We aimed to quantify the contributions of these factors to diagnostic outcomes across five U.S. clinical sequencing sites. Methods:We performed a cross-sectional analysis of 3,008 prenatal, neonatal, and pediatric cases from the NHGRI Clinical Sequencing Evidence-Generating Research (CSER) consortium (2017-2023). Clinical indications spanned neurodevelopmental, neurological, immunological, metabolic, craniofacial, skeletal, cardiac, prenatal, and oncologic presentations. Genetic ancestry was inferred from sequencing data, and variants were interpreted using ACMG/AMP guidelines to classify DNA-based diagnoses. Generalized linear mixed models were used to estimate associations between diagnostic yield and fixed effects (sex, prenatal status, isolated cancer, number of clinical indications, sequencing modality, race/ethnicity, and genetic ancestry), while modeling study site as a random effect to quantify between-site variation. Results:The overall diagnostic yield was 19.0%. Multiple clinical indications (OR=1.47, 95% CI 1.20-1.80, p<0.001) were associated with higher diagnostic yield, and male sex (OR=0.80, 95% CI 0.66-0.96, p=0.017) and prenatal status (OR=0.63, 95% CI 0.44-0.90, p=0.012) were associated with lower yield. Sequencing modality, race/ethnicity, genetic ancestry, and isolated cancer were not statistically significantly associated with diagnostic outcomes.. A model without fixed effects attributed ~10% of variance in diagnostic yield to between-site differences. After adjusting for covariates, site-level variance decreased to 5.7%, indicating consistent variation across sites not explained by measured patient factors. Conclusion:Across five sites, patient-level clinical features influenced diagnostic yield, but substantial site-level variation remained even after adjustment. Differences in variant interpretation, or case-classification practices may contribute to this residual variability. Further efforts to increase consistency in exome- and genome-sequencing diagnostic workflows may help reduce inter-site differences.
Background: The 9p21.3 locus was the first genome-wide significant signal for coronary artery disease (CAD) and replicates across multiple non-African populations, yet is absent in African ancestry cohorts. We hypothesized that ancestry-specific linkage disequilibrium (LD) and haplotype structure, rather than allele frequency or power alone, explain this discrepancy. Methods: We analyzed multi-ancestry data from European, East Asian, South Asian, Middle Eastern, African, and Admixed American groups. Within each ancestry, we performed CAD common-variant associations at 9p21.3, rare-variant tests in whole-genome-sequenced Europeans, local-ancestry inference (LAI) stratified associations, conditional and haplotype analyses, and pleiotropy assessments. Results: CAD associations at 9p21.3 were robust in Europeans, East Asians, Middle Eastern, and Admixed Americans ancestry groups, nominal in South Asians, and absent in Africans despite similar allele frequencies for lead-associated SNPs in other ancestry groups. LAI-stratified analyses showed strong signals in African heterozygotes carrying 9p21.3 European chromosomes but not in 9p21.3-homozygous Africans. The association is due to highly common variants; rare variants did not explain the locus signal. LD blocks were extended in non-Africans but fragmented into smaller blocks in Africans, with marked divergence by fixation index. Fine-mapping identified multiple independent signals, including an East Asian-specific haplotype absent in Europeans and Africans. Hamming-distance analyses revealed that risk alleles are dispersed across multiple haplotypes in homozygous African groups, consistent with weaker LD and restricting risk variation at this locus in that group. PheWAS mirrored these ancestry-specific patterns across metabolic traits. ### Competing Interest Statement The authors have declared no competing interest. ### Funding Statement We gratefully acknowledge the Veterans who participated in the VA Million Veteran Program. This research is based on data from the Million Veteran Program, Office of Research and Development, and Veterans Health Administration, with support from MVP000, VA Merit Award #I01-BX003362 (Chang/Tsao) and the Department of Veterans Affairs (VA) Informatics and Computing Infrastructure (VINCI), including data analytics conducted by its Precision Medicine research team, which is funded under the research priority to Put VA Data to Work for Veterans (VA ORD 24-D4V-02). This work was also supported by the American Heart Association Predoctoral Fellowship (Award ID: 24PRE1201199) and the Lamond Fellowship. This publication does not represent the views of the Department of Veterans Affairs or the United States Government. ### Author Declarations I confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained. Yes The details of the IRB/oversight body that provided approval or exemption for the research described are given below: The Million Veteran Program study received ethical and study protocol approval from the VA Central Institutional Review Board. Ethical approval for the use of genomic and clinical data along with consent from participants for UK Biobank was obtained by the National Health Service National Research Ethics Service (ref: 264 11/NW/0382) and data use approved under application 87255. I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals. Yes I understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance). Yes I have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable. Yes All data produced in the present work are contained in the manuscript
Locus 9p21.3 was the first genome-wide significant locus for coronary artery disease (CAD) and replicates across multiple non-African populations, yet it is absent in African-ancestry cohorts. We analyzed multi-ancestry data from European, East Asian, South Asian, Middle Eastern, African, and admixed American groups. Ancestry was inferred using both global and local-ancestry inference (LAI) approaches. Within each ancestry, we performed CAD associations at 9p21.3 with common and rare variants, conditional and haplotype analyses, and pleiotropy assessments. CAD associations at 9p21.3 were robust in European, East Asian, Middle Eastern, and admixed American ancestry groups, nominal in South Asians, and absent in Africans despite similar allele frequencies for lead-associated SNPs. LAI-stratified analyses showed strong signals in African heterozygotes carrying 9p21.3 European chromosomes but not in homozygous Africans. The association is due to highly common variants; rare variants did not explain the locus signal. Linkage disequilibrium (LD) blocks were extended in non-Africans but fragmented into smaller blocks in Africans. Fine-mapping identified multiple independent signals, including an East Asian-specific haplotype absent in Europeans and Africans. Hamming-distance analyses revealed that risk alleles are dispersed across multiple haplotypes in homozygous-African groups. Phenome-wide association study (PheWAS) mirrored these ancestry-specific patterns across metabolic traits. The non-replication of 9p21.3 in African ancestry is not due to absent risk alleles, reduced power, or poorer imputation of risk alleles due to weaker LD; rather, it reflects greater haplotype diversity that disperses risk alleles across configurations, attenuating case-control contrast. These findings illustrate how ancestry-specific LD and multi-causal haplotype architecture modulate association detectability.
Race and ethnicity are demographic constructs used to characterize individuals in biomedical research, and in particular to assess health disparities. Their use in medicine and research has been discussed and challenged, as well as the degree to which they represent strictly social constructs, or ones also with biological meaning. The relationship of race and ethnicity with genetic ancestry has also been described, and how genetic ancestry reflects historical continental isolation, migration, and mating structure. Race and ethnicity are currently most often assessed by self-report in epidemiology and biomedical applications. Here we further interrogate the relationship between how people self-report their race and ethnicity and their genetic ancestry by examining self-report patterns of 97,671 individuals who are participants in the Kaiser Permanente Northern California Genetic Epidemiology Research on Adult Health and Aging (GERA) cohort. Genetic ancestry was determined from a set of 43,988 SNPs from genome-wide genotyping arrays. We observed that rates of self-identification as African American, East Asian and Latino(a) rise dramatically with a modest amount of African, East Asian and Native American genetic ancestry, respectively. By contrast, the rate of self-identification as White rises only when the European/West Asian genetic ancestry is substantial. This indicates that the majority of people who are genetically admixed, even those with primarily European/West Asian genetic ancestry, self-identify with the minority race/ethnicity group. By contrast, self-report as Native American did not increase with Native American genetic ancestry; instead, it was positively correlated with European genetic ancestry, with only a small minority of individuals self-reporting Native American race/ethnicity having Native American genetic ancestry. These results differ dramatically from the other minority race/ethnicity groups. These findings have important implications on how the different self-report race/ethnicity groups are considered in epidemiologic and biomedical research.
INTRODUCTION:We evaluated the independent associations between high-density lipoprotein cholesterol (HDL-C) and triglyceride (TG) levels with Alzheimer's disease and related dementias (ADRD). METHODS:Among 177,680 members of Kaiser Permanente Northern California who completed a survey on health risks, we residualized TGs and HDL-C conditional on age, sex, and body mass index. We included these residuals individually and concurrently in Cox models predicting ADRD incidence. RESULTS:Low (hazard ratio [HR] 1.06, 95% confidence interval [CI] 1.02-1.10) and high quintiles (HR 1.07, 95% CI 1.03-1.12) of HDL-C residuals were associated with an increased risk of ADRD compared to the middle quintile. Additional adjustment for TGs attenuated the association with high HDL-C (HR 1.03, 95% CI 0.99-1.08). Low TG residuals were associated with an increased ADRD risk (HR 1.10, 95% CI 1.06-1.15); high TG residuals were protective (HR 0.92, 95% CI 0.88-0.96). These estimates were unaffected by HDL-C adjustment. DISCUSSION:Low HDL-C and TG levels are independently associated with increased ADRD risk. The correlation with low TG level explains the association of high HDL-C with ADRD. HIGHLIGHTS:Strong correlations between lipid levels are important considerations when investigating lipids as late-life risk factors for Alzheimer's disease and related dementias (ADRD). Low levels of high-density lipoprotein cholesterol (HDL-C) and triglycerides (TGs) were independently associated with an increased risk of ADRD. We found no evidence for an association between high HDL-C and increased ADRD risk after adjustment for TGs. High levels of TGs were consistently associated with a decreased risk of ADRD. There may be interaction between TG and HDL-C levels, where both low HDL-C and TG levels increase the risk of ADRD compared to average levels of both.
INTRODUCTION:Mixed evidence on how statin use affects risk of Alzheimer's disease and related dementias (ADRD) may reflect heterogeneity across sociodemographic factors. Few studies have sufficient power to evaluate effect modifiers. METHODS:Kaiser Permanente Northern California (KPNC) members (n = 705,061; n = 202,937 with sociodemographic surveys) who initiated statins from 2001 to 2010 were matched on age and low-density lipoprotein cholesterol (LDL-C) with non-initiators and followed through 2020 for incident ADRD. Inverse probability-weighted Cox proportional hazards models were used to evaluate effect modification by age, gender, race/ethnicity, education, marital status, income, and immigrant generation. RESULTS:Statin initiation (vs non-initiation) was not associated with ADRD incidence in any of the 32 subgroups (p > .05). Hazard ratios ranged from 0.964 (95% CI: 0.923 to 1.006) among Asian-identified participants to 1.122 (95% CI: 0.995 to 1.265) in the highest income category. DISCUSSION:Sociodemographic heterogeneity appears to have little to no influence on the relationship between statin initiation and dementia. HIGHLIGHTS:The study includes a large and diverse cohort from Kaiser Permanente (N = 705,061). An emulated trial design of statin initiation on dementia incidence was used. Effect modification by sociodemographic factors was assessed. There were no significant Alzheimer's disease and related dementias (ADRD) risk differences in 32 sociodemographic subgroups (p > 0.05).
Mixed evidence on how statins affect dementia risk may reflect variability in model specifications. Alternate specifications are rarely systematically compared. Using an emulated trial design framework, we investigated variation in the estimated effect of statin initiation on dementia across alternative (1) eligibility criteria, (2) confounding variable sets, and (3) outcome definitions. Kaiser Permanente Northern California members’ linked electronic health records from 1996 to 2020 were used to identify statin initiation and dementia diagnoses. Statin initiators were matched on age and low-density lipoprotein cholesterol with up to 5 non-initiators. Possible covariates included clinical (n = 1.4 million); socioeconomic and behavioral (n = 265,224); and genetic (n = 69,573) variables. Using Cox proportional-hazards models, we estimated variation across 1.27 million intent-to-treat estimates for statin initiation varying specification of eligibility, outcome definition, and covariates. Estimated hazard ratios (HRs) for statin initiation on dementia across all specifications ranged from 0.93 to 1.47. The variance of estimates due to model specification differences was 7.6 times larger than the average variance of specific estimates due to finite sample size. Three modeling decisions notably attenuated coefficients [ln(HR)]: requiring a run-in period prior to the emulated trial start date (0.034); adjustment for diabetes (0.030) and cardiovascular disease (0.039); and excluding the first year of follow-up (0.041). HRs from models with all three specifications ranged from 0.99 to 1.15. No specification we evaluated consistently generated protective effects. Estimates of the association between statin initiation and dementia leveraging real world data are sensitive to model specification, especially decisions related to clinical covariates and time-at-risk.
It has been suggested that diagnostic yield (DY) from Exome Sequencing (ES) may be lower among patients with non-European ancestries than those with European ancestry. We examined the association of DY with estimated continental/subcontinental genetic ancestry in a racially/ethnically diverse pediatric and prenatal clinical cohort. Cases (N = 845) with suspected genetic disorders underwent ES for diagnosis. Continental/subcontinental genetic ancestry proportions were estimated from the ES data. We compared the distribution of genetic ancestries in positive, negative, and inconclusive cases by Kolmogorov–Smirnov tests and linear associations of ancestry with DY by Cochran-Armitage trend tests. We observed no reduction in overall DY associated with any genetic ancestry (African, Native American, East Asian, European, Middle Eastern, South Asian). However, we observed a relative increase in proportion of autosomal recessive homozygous inheritance versus other inheritance patterns associated with Middle Eastern and South Asian ancestry, due to consanguinity. In this empirical study of ES for undiagnosed pediatric and prenatal genetic conditions, genetic ancestry was not associated with the likelihood of a positive diagnosis, supporting the equitable use of ES in diagnosis of previously undiagnosed but potentially Mendelian disorders across all ancestral populations.
Journal Article The Inclusion of Race in Prenatal Screening Algorithms Get access Mary E Norton, Mary E Norton Department of Obstetrics, Gynecology, and Reproductive Sciences, University of California, San Francisco, San Francisco, CA, United StatesInstitute of Human Genetics, University of California, San Francisco, San Francisco, CA, United States Address correspondence to this author at: Department of Obstetrics, Gynecology, and Reproductive Sciences, University of California, San Francisco, 490 Illinois St., 10th floor, San Francisco, CA 94143, United States. Tel 415-353-7865; e-mail mary.norton@ucsf.edu. Search for other works by this author on: Oxford Academic Google Scholar Neil Risch Neil Risch Institute of Human Genetics, University of California, San Francisco, San Francisco, CA, United StatesDepartment of Epidemiology and Biostatistics, University of California, San Francisco, San Francisco, CA, United States Search for other works by this author on: Oxford Academic Google Scholar Clinical Chemistry, hvae073, https://doi.org/10.1093/clinchem/hvae073 Published: 06 June 2024 Article history Received: 18 April 2024 Accepted: 24 April 2024 Published: 06 June 2024
Incomplete penetrance, variable expressivity, and phenotype expansion are well-documented phenomena in clinical genetics. However, when rare diseases are analyzed at the population-scale, phenotypes are typically binarized into simple diagnoses (present vs absent), throwing away much of the variability intrinsic to rare diseases. The impact of this information loss on downstream applications, such as penetrance estimation, remains largely unknown. Here, we demonstrate that multivariate symptom models, when combined with automated phenotype expansion, significantly increase population-based estimates of loss-of-function variant penetrance in a cohort of over 450,000 subjects.
BACKGROUND: Inter-individual variation in blood pressure (BP) arises in part from sequence variants within enhancers modulating the expression of causal genes. We propose that these genes, active in tissues relevant to BP physiology, can be identified from tissue-level epigenomic data and genotypes of BP-phenotyped individuals. METHODS: We used chromatin accessibility data from the heart, adrenal, kidney, and artery to identify cis-regulatory elements (CREs) in these tissues and estimate the impact of common human single-nucleotide variants within these CREs on gene expression, using machine learning methods. To identify causal genes, we performed a gene-wise association test. We conducted analyses in 2 separate large-scale cohorts: 77 822 individuals from the Genetic Epidemiology Research on Adult Health and Aging and 315 270 individuals from the UK Biobank. RESULTS: We identified 309, 259, 331, and 367 genes (false discovery rate <0.05) for diastolic BP and 191, 184, 204, and 204 genes for systolic BP in the artery, kidney, heart, and adrenal, respectively, in Genetic Epidemiology Research on Adult Health and Aging; 50% to 70% of these genes were replicated in the UK Biobank, significantly higher than the 12% to 15% expected by chance ( P <0.0001). These results enabled tissue expression prediction of these 988 to 2875 putative BP genes in individuals of both cohorts to construct an expression polygenic score. This score explained ≈27% of the reported single-nucleotide variant heritability, substantially higher than expected from prior studies. CONCLUSIONS: Our work demonstrates the power of tissue-restricted comprehensive CRE analysis, followed by CRE-based expression prediction, for understanding BP regulation in relevant tissues and provides dual-modality supporting evidence, CRE and expression, for the causality genes.
This article is based on the address given by the author at the 2023 meeting of The American Society of Human Genetics (ASHG). A video of the original address can be found at the ASHG website.
Background and ObjectivesThe associations of high-density lipoprotein cholesterol (HDL-C) and low-density lipoprotein cholesterol (LDL-C) with dementia risk in later life may be complex, and few studies have sufficient data to model nonlinearities or adequately adjust for statin use. We evaluated the observational associations of HDL-C and LDL-C with incident dementia in a large and well-characterized cohort with linked survey and electronic health record (EHR) data.MethodsKaiser Permanente Northern California health plan members aged 55 years and older who completed a health behavior survey between 2002 and 2007, had no history of dementia before the survey, and had laboratory measurements of cholesterol within 2 years after survey completion were followed up through December 2020 for incident dementia (Alzheimer disease-related dementia [ADRD]; Alzheimer disease, vascular dementia, and/or nonspecific dementia) based on ICD-9 or ICD-10 codes in EHRs. We used Cox models for incident dementia with follow-up time beginning 2 years postsurvey (after cholesterol measurement) and censoring at end of membership, death, or end of study period. We evaluated nonlinearities using B-splines, adjusted for demographic, clinical, and survey confounders, and tested for effect modification by baseline age or prior statin use.ResultsA total of 184,367 participants [mean age at survey = 69.5 years, mean HDL-C = 53.7 mg/dL (SD = 15.0), mean LDL-C = 108 mg/dL (SD = 30.6)] were included. Higher and lower HDL-C values were associated with elevated ADRD risk compared with the middle quantile: HDL-C in the lowest quintile was associated with an HR of 1.07 (95% CI 1.03-1.11), and HDL-C in the highest quintile was associated with an HR of 1.15 (95% CI 1.11-1.20). LDL-C was not associated with dementia risk overall, but statin use qualitatively modified the association. Higher LDL-C was associated with a slightly greater risk of ADRD for statin users (53% of the sample, HR per 10 mg/dL increase = 1.01, 95% CI 1.01-1.02) and a lower risk for nonusers (HR per 10 mg/dL increase = 0.98; 95% CI 0.97-0.99). There was evidence for effect modification by age with linear HDL-C (p = 0.003) but not LDL-C (p = 0.59).DiscussionBoth low and high levels of HDL-C were associated with elevated dementia risk. The association between LDL-C and dementia risk was modest.
Supplementary Table S1. Descriptive factors for KP study population, broken down by study. Supplementary Table S2. Genome-wide significant SNPs found in our cohort. Supplementary Table S3. Cis-eQTL expression of rs4646284. Supplementary Table S4. Results at the 105 loci previously found to be associated with prostate cancer. Supplementary Table S5. Risk score of and variance explained by the 105 previously reported hits. Supplementary Figure S1. Manhattan and Q-Q plots of each race/ethnicity and meta-analysis. Supplementary Figure S2. Local plots of novel replicated rs4646284 and suggestive rs2659124. Supplementary Figure S3. Cis-eQTL of SLC22A1 and SLC22A3. Supplementary Figure S4. Comparison of ORs of KP to previous reports by race/ethnicity. Supplementary Figure S5. KP AUC estimates.