
We introduce five points for integrating environmental ethics into human genomic data governance: (i) recognizing the ethical imperative to consider environmental impacts of human genomic data; (ii) fostering collective responsibility for environmental harms; (iii) prospectively assessing benefits and harms; (iv) anticipating barriers to integration of environmental ethics into genomic data governance; and (v) meaningfully engaging all interest-holders. These points will be useful to all involved in the genomic data ecosystem.
A single genomic assay that delivers complete information across variant classes remains an aspirational goal. Currently, researchers and clinicians rely on an inefficient, expensive combination of short-read sequencing for single-nucleotide variants (SNVs) and small indels, comparative genomic hybridization (CGH) arrays for copy number variants (CNVs), and optical mapping and long-read sequencing for complex rearrangements, limiting the full potential of genomic discovery. To address these issues, TruPath Genome provides a one-test-for-all solution. By combining PCR-free whole-genome sequencing (WGS) with proximity-mapped read technology, it achieves high-resolution detection of SNVs and indels alongside long-range phasing for CNVs and structural variant (SV) refinement. We applied TruPath Genome on six clinical samples that were previously resolved by conventional methods. Across the cohort, TruPath Genome delivered coverage and variant-calling performance comparable to conventional WGS while achieving superior long-range phasing and enabling precise breakpoint resolution for clinically relevant structural events. This highlights TruPath Genome’s potential to consolidate genomic testing pipelines, accelerate diagnosis, and expand access to advanced genomic insights. Furthermore, its ultra-long-range data facilitates telomere-to-telomere assemblies and pangenome development, advancing our understanding of genome biology at an unprecedented scale.
Genome sequencing enables accurate detection of genetic variants and is transforming rare disease diagnostics. While data generation is scalable, prioritization and clinical interpretation remain challenging, often requiring expert manual classification. AI-driven decision support systems are therefore needed to assist in causal variant identification or to fully automate large-scale re-analysis of unsolved cases. Existing tools often estimate variant impact on protein function, but few integrate genomic, phenotypic, and clinical annotation data for diagnosis. We present aiDIVA, an ensemble-AI combining statistical and machine learning models trained on genomic and phenotypic data to identify causal variants among tens of thousands per patient. aiDIVA applies a random forest model to classify pathogenicity and generates evidence-based scores for dominant and recessive diseases. These predictions are integrated with clinical metadata to prioritize the most likely causal variants. Large language models further refine and explain results. The aiDIVA-meta model consolidates all scores into a ranked list. aiDIVA-meta reported the causal variant among the top-3 candidates in 97.4% of a pre-training collected cohort with prior evidence in ClinVar or HGMD, and in 93.3% of a post-training collected cohort of previously unreported variants.
Clinical genomic profiling of circulating cell-free DNA (cfDNA) from liquid biopsies is routinely used for non-invasive cancer diagnosis and monitoring. However, differential fragmentation of blood cfDNA can lead to systematic loss of genetic information. It remains unclear whether such degradation compromises clinically relevant oncogenic regions across cancer types. To address this, we introduced UPASHAYA (Uncovering Potential Areas of Shadowed Heterochromatin Associated Yielding Aberrations) for investigating the consequence of cfDNA fragmentation across 317 samples spanning pan-cancer and non-cancer cohorts. We identified ‘penumbra regions’, open chromatin domains undergoing elevated nuclease-mediated degradation, creating clinical blind spots for druggable cancer genes. Structurally, these regions feature an abundance of small fragments (<150 bp) and a distinct fragment-end signature, including elevated C-end and C/G frequencies at the second motif position. This chromatin-driven degradation is independent of tumor shedding, showing consistent penumbra enrichment across low- and high-shedding cancers. Furthermore, gene-dense chromosomes, namely 17 and 19, emerged as major depletion hotspots. These penumbra regions are cancer-type-specific, impacting regulatory elements and actionable oncogenic hotspot exons, potentially hindering minimal residual disease detection. Our findings highlight a critical biological vulnerability and call for careful interpretation of clinical genomic data for precision medicine.
FOXP4 is a transcription factor belonging to the FOX subfamily P, which acts as a multifunctional regulator in cardiac morphogenesis, lung development, and gut development. Six heterozygous missense and loss-of-function (LoF) variants in FOXP4 were reported in a few cases to cause neurodevelopmental phenotypes with multiple congenital abnormalities. Larger studies are needed to further confirm and delineate FOXP4-related disease. In this study, exome/genome sequencing was used to identify variants in 13 participants with short stature, dysmorphic features, neurodevelopmental disorders (NDD), and congenital heart disease (CHD). Additionally, we reviewed the data of 12 published cases. Among 25 participants with FOXP4 variants, 23 carried heterozygous variants, and two were homozygous for LoF variants. Most cases presented with short stature, failure-to-thrive, and dysmorphic features. Other common features included NDD, skeletal anomalies, and CHD. All heterozygous variants recruited in this study were absent from controls, confirmed to be de novo or inherited from affected parents, and predicted to undergo nonsense-mediated decay (NMD) or disrupt canonical splicing. The two homozygous LoF variants were predicted to undergo NMD. This study provides clinical and genetic evidence to confirm the autosomal dominant disease and characterizes the recessive form of FOXP4-related disease.
Hemophagocytic lymphohistiocytosis (HLH) is a severe immunological disorder characterized by dysregulated immune activation. Pathogenic variants in HLH-causative genes serve as diagnostic criteria and guide treatment decisions. However, known genes do not fully explain the molecular basis of many cases, and the polygenic contribution to HLH susceptibility remains poorly characterized. Here, we retrospectively analyzed whole-exome sequencing data from 1241 patients with clinically diagnosed or suspected HLH to characterize the HLH genetic landscape. Rare variant association analysis identified two candidate susceptibility genes, IKBKG and DDX3X. Functional experiments showed that IKBKG knockdown impaired NK cell cytotoxicity and degranulation, supporting a contributory role of IKBKG in HLH-related immune dysfunction. Exploratory common variant analysis identified potential susceptibility loci, and an integrated model incorporating rare variant burden and polygenic risk score achieved an AUC of 0.71 in this cohort. These findings expand understanding of HLH genetic architecture and support contributions from both rare and common variants.
Comprehensive genomic analysis, including whole-genome sequencing (WGS), enables iterative data reanalysis, potentially improving diagnostic yield as new gene–disease associations and analytical methods emerge. Although reanalysis has been shown to increase diagnostic yield across disease groups, uncertainty remains regarding its optimal timing and overall impact. We systematically reprocessed and reanalyzed WGS data from a heterogeneous clinical cohort collected over five years to evaluate the effect of reanalysis on diagnostic yield. Clinical WGS cases (~3000) analyzed at the Department of Genomic Medicine, Copenhagen University Hospital, were reviewed. Cases without a causative variant underwent reanalysis. The cohort included singletons and families across major disease groups, including cancer, neurodevelopmental, hematological, immunological, and endocrine disorders. Samples were reprocessed using the latest in-house germline WGS pipeline, including structural variant analysis, and reinterpreted using updated gene panels and latest disease knowledge. Baseline diagnostic yield was 15% in singletons (n = 2209) and 30% in family probands (n = 245). Reanalysis increased yield to 19% and 36%, respectively. Of 1551 curated variants, 79 were clinically reportable, predominantly identified through expanded in silico gene panels. Reanalysis increased overall diagnostic yield from 17% to 21% (22–26% excluding sporadic cancers), with greatest benefit from broad gene panels and semi-automated workflows.
Genomic screening increasingly reveals variants of uncertain significance, forcing clinicians to make high-stakes decisions as million-dollar drug and gene-replacement therapies emerge. Rapid in vivo zebrafish assays can help to resolve variant pathogenicity within weeks, providing organism-level evidence that informs treatment decisions, prevents unnecessary interventions, and strengthens precision medicine.
Comprehensive interpretation of whole-genome sequencing data for phenotypes with complex genetic architecture requires integrating diverse variant classes, but most workflows remain focused on single-nucleotide changes and small indels. We present IMPACT, an open-source, phenotype-configurable pipeline that harmonizes preprocessing, applies variant-specific annotation, and consolidates single-nucleotide variants, indels, structural variants, and copy-number variants within an interactive R Shiny interface. Applied to the UK Biobank cohorts with deafness (n = 126) and epilepsy (n = 41), IMPACT prioritized 565 and 126 variants, respectively. Of these, 512 and 103 were non-incidental, phenotype-relevant candidates, which were subsequently reviewed using ACMG-aligned evidence and classified as pathogenic or candidate variants of uncertain significance where supported. All deafness participants and 87.8% of epilepsy participants carried at least one candidate variant, with configurations ranging from single high-impact variants to compound heterozygotes spanning variant classes. By coupling phenotype-aware filtering with cross-type visualization, IMPACT addresses critical gaps in genome interpretation and offers a scalable foundation for research-driven curation and precision medicine.
Cornelia de Lange Syndrome (CdLS) is a multisystem disorder caused by pathogenic variants in one of the six genes associated with CdLS (NIPBL, SMC3, SMC1A, HDAC8, RAD21, BRD4) and pathogenic variants in phenocopy genes. We hypothesized that individuals with a clinical diagnosis of CdLS and no molecular diagnosis harbor diagnostic variants that are missed by the current standard of care, exome-focused workflows. We performed a re-analysis of the genome sequencing data from a previously published cohort of 173 individuals with a clinical diagnosis of CdLS and expanded the scope of analysis to include noncoding and structural variants. The comprehensive re-analysis revealed molecular etiologies in an additional 37 probands increasing the overall diagnostic yield from 37.5% to 57.8%. The new diagnoses were enriched for variants beyond the exonic SNVs/ indels, including cryptic non-coding variants, copy number variants, balanced rearrangements such as inversions, and variants in genes that phenocopy CdLS. Transcriptome aided re-analysis uncovered cryptic noncoding variants that lacked sufficient computational evidence for aberrant splicing and yet produced aberrantly spliced mRNA. Our results underscore the need for whole genome (and transcriptome) sequencing and a comprehensive, unbiased analytical protocol to exhaustively mine a phenotypically and genetically heterogeneous cohort to maximize its diagnostic yield.
The Clinical Implementation Pilots (CIPs) within Singapore’s National Precision Medicine (NPM) programme are designed to integrate genomic testing into clinical pathways for conditions including hereditary cancers, familial hypercholesterolemia, breast cancer, primary glomerular disease, and pharmacogenomics. These multidisciplinary pilots emphasize collaboration to enhance cost-effectiveness and improve patient outcomes, thereby establishing a sustainable foundation for future scalability in clinical genomics and ultimately fostering a healthier population through personalized healthcare solutions.
Carrier screening aims to identify couples at risk of having children with serious genetic disorders. However, screening panels vary markedly in size, gene content and price, with little guidance to assess clinical utility or value. We analysed 89 carrier screening panels from 30 global providers, modelling test performance using pathogenic variants from gnomAD v4.1.0 and ClinVar across ten ancestry groups and two synthetic populations (representing the United States and Australia). We evaluated how panel size, content, price and clinical utility interrelate. We also compared value per dollar spent to guide clinical decision-making and fair pricing. Clinical utility showed no consistent relationship with panel size or price. Instead, mid-sized, pan-ancestry “Goldilocks” panels delivered the greatest utility, outperforming both smaller and larger panels, while delivering more equitable outcomes. We developed a visual “clinical utility meter” to compare relative test performance and an “efficiency frontier” framework to identify the best value tests for a given price. This data-driven framework and visual tool can support rational test design and selection for value-based laboratory medicine.
Clinical biobanks linking electronic health records (EHRs) with genotype data enable the study of genomic risk factors in real-world populations. However, recall-by-genotype (RbG) of psychiatric risk variants in diverse healthcare-system biobanks remains scarce. Leveraging BioMe, a multi-ancestry biobank within the Mount Sinai Health System, we recalled carriers of rare copy number variants (CNVs) that confer increased risk for neurodevelopmental disorders (NDDs) to establish empirical benchmarks for RbG implementation. We recontacted 892 participants: 335 NDD CNV carriers, 217 individuals with schizophrenia without NDD CNVs, and 340 neurotypical controls without NDD CNVs. Participants completed clinical and cognitive assessments. Overall, 18% of recontacted participants responded to recruitment, and 8% completed the study: 30 NDD CNV carriers, 20 individuals with schizophrenia, and 23 controls. The mean age was 48.8 years, 66% were female, and self-reported ancestry was 37% African, 34% Hispanic, and 26% European. Seventy percent of NDD CNV carriers had at least one neuropsychiatric or developmental condition, including mood or anxiety disorders (40%). Among 22 NDD CNV carriers at loci implicated in impaired cognition, performance was lower than controls on Digit Span Backward (β = -1.76, FDR = 0.04) and Digit Span Sequencing (β = -2.01, FDR = 0.04). NDD CNV carriers also outperformed the schizophrenia group on verbal learning (β = 4.5, FDR = 0.05). Recall of individuals-including those with psychiatric illness-yielded phenotypes not captured in EHRs and provides empirical benchmarks relevant to RbG implementation and precision psychiatry in diverse healthcare systems.
Precision medicine trials may generate new treatment options for children with poor-prognosis cancer. We examined families' experiences of receiving treatment recommendations in the Australian PRISM trial (Australian and New Zealand Clinical Trials Registry: NCT03336931; https://clinicaltrials.gov/study/NCT03336931; registration date 08/11/2017). Parents and patients (12-17 years) completed questionnaires at enrolment (T0; n = 303 and n = 31) and following results and any treatment recommendations (T1; n = 144 and n = 8). Fifty-eight parents completed an interview (T1). At T0, most parents and patients expected to benefit from participation (87%) and receive a treatment recommendation (68%). Of the 70% of parents who received a treatment recommendation, half recalled this information. Parents felt that treatment recommendations offered hope and options; their absence brought disappointment, but also reassurance that all options were explored. Parents reported high involvement (93/100) and satisfaction (95/100) in treatment decisions. Receiving a treatment recommendation was not associated with regret about trial participation (p > 0.05). Our findings underscore precision medicine's value for families in the setting of a poor-prognosis child with cancer.
Growth modeling is central to human genetics, as deviations from typical growth can signal an underlying disorder. In this cohort study, we developed a generalizable framework for generating growth charts across genetic conditions using electronic health records (EHR). Leveraging 22 years of longitudinal EHR data from 452,470 patients across 15 genetic conditions and unaffected individuals, we generated sex- and condition-specific growth charts using Generalized Additive Models for Location, Scale, and Shape, and quantified differences in size, timing, and intensity using SuperImposition by Translation and Rotation (SITAR). SITAR-derived growth parameters showed strong concordance with established annotations in OMIM and Orphanet, and identified previously unreported growth patterns. We stratified cystic fibrosis by CFTR functional class and observed greater growth impairment in individuals with homozygous minimal-function variants compared to those with residual function. This framework provides a generalizable approach for leveraging EHR data to refine genotype-phenotype relationships and enable continuous updating of growth charts across genetic conditions.
Genetic neurological disorders are highly heterogeneous, and many are driven by complex variants that challenge short‑read genome sequencing (srGS). Long‑read genome sequencing (lrGS) has recently shown promise, but head‑to‑head evaluations are limited. A direct comparison study of srGS and lrGS was performed for 310 families with undiagnosed neurological disorders from the Hong Kong Genome Project to assess their diagnostic performance, technical capabilities and costing. Genome sequencing showed an overall diagnostic yield of 22.6% (n = 70/310, 77 variants). lrGS and srGS showed comparable variant detection rates at 92.2% (n = 71/77) and 96.1% (n = 74/77), respectively. lrGS solely confirmed three repeat expansions with phased methylation data and enhanced accuracy in two complex structural variants. srGS detected six variants in homopolymeric regions that were missed by lrGS. Costing analysis revealed a comparable unit cost per sample for srGS and lrGS with USD1,545.97 and USD1,580.00, respectively. This study demonstrated that srGS excels in detecting variants near homopolymers while lrGS offers better resolution for complex variants and enables the simultaneous analysis of methylation and phasing. Although lrGS is not yet capable of fully replacing srGS, we anticipate that the continued technical improvement would position lrGS as the first‑tier genomic test for neurological disorders.
The evolution of transcriptomic technologies requires effective translation between legacy and modern platforms to fully leverage historical data. We present GANomics, a generative adversarial network (GAN) framework that enables bidirectional translation between microarray and RNA-seq data using a small number of samples profiled by both technologies. By integrating paired and unpaired samples through a pair-aware feedback loss, GANomics enforces one-to-one transcript mappings while preserving global gene expression distributions. Applied to a neuroblastoma cohort (n = 498; 10,042 genes) and benchmarked against six alternative methods, GANomics achieved high per-sample correlations ( > 0.96) between real and synthetic data with as few as ten paired profiles. With fifty paired profiles, it accurately recapitulated differential expression, maintained pathway-level rankings, and enabled a cross-platform classifier to transfer comparable to real data. Consistent performance across five additional benchmark datasets confirmed its robustness. By bridging legacy and contemporary transcriptomic data while retaining biological consistency, GANomics facilitates scalable data integration for biomarker reuse and the expansion of transcriptomic resources in clinical applications.
BCL2 upregulation is a key driver of lymphomagenesis and treatment response. By analyzing large-scale whole-genome datasets, we identify a lymphoma-specific noncoding mutational hotspot encompassing the BCL2 promoter that, independent of other factors such as the t(14;18) translocation, strongly associates with allele-specific BCL2 upregulation. These mutations disrupt regulatory protein-binding sites, alter BCL2 isoform balance, and associate with worse survival in diffuse large B-cell lymphoma.
Pathogenic germline variants (PGVs) in dyskerin pseudouridine synthase 1 (DKC1) cause X-linked recessive dyskeratosis congenita (DC), a telomere biology disorder. Females with heterozygous DKC1 PGVs rarely exhibit DC phenotypes due to favorable X chromosome inactivation (XCI). We report a female with classic DC, pigmentary mosaicism and bone marrow failure associated with a novel de novo DKC1 variant (c.190 G > C, p.Val64Leu) with reduced expression of DKC1 and TERC along with downregulated signatures associated with aberrant telomere biology and ribosome function. Markedly skewed XCI was detected with expression of the mutated allele in skin fibroblasts and wild-type DKC1 expression in the bone marrow. We hypothesize that selective pressure in the bone marrow favored wild type expressing cells which acquired a trisomy 9 but were unable to resume normal hematopoiesis. This study demonstrates DKC1 c. 190 G > C as a likely PGV causing classic DC and highlights the molecular and clinical complexities associated with skewed XCI.
Rhabdoid tumour predisposition syndrome (RTPS) is a highly penetrant cancer predisposition syndrome caused by germline variants in SMARCB1 or less frequently in SMARCA4. Genetic testing for this syndrome involves sequence and deletion/duplication analysis of these two genes. Standard clinical testing is limited in detecting structural variants. Here we describe two patients who tested negative on standard clinical germline panel testing for RTPS but were each found to have a mosaic germline insertion of an SVA (SINE-VNTR-Alu) element in the SMARCB1 gene by more advanced comprehensive genomic analysis. These two cases demonstrate the importance of structural variants and broader genomic sequencing for individuals suspected of having an underlying germline cancer predisposition syndrome, such as RTPS.