
Next generation sequencing (NGS) is routinely performed in clinical practice to detect various types of mutations for targeted therapy, diagnosis, and prognosis. Actionable alterations detected by NGS include not only non-synonymous mutations that lead to functional or structural changes of proteins but also copy number variants (CNV) that affect gene dosage, such as gene copy gains, amplifications or deletions. Among tumor-only CNV detection methods, the use of a Panel of Normals (PoN) for relative comparison has become a common practice, largely due to the lack of matched normal samples. It was therefore hypothesized, once established, a PoN and CNV caller may not fully compensate for all experimental variations - such as differences in probe efficiency across reagent lots. To investigate this, 12,104 clinical sequencing datasets from 1,454 sequencing batches were analyzed over a four-year period. This analysis revealed batch-associated fluctuation patterns in gene-level fold changes that could potentially lead to misinterpretation, such as the incorrect classification of gene copy deletions or gains. In this study, a strategy is presented that calculates the median and median absolute deviation of gene-level fold changes across all samples within each sequencing batch and incorporates these measures into the result interpretation. By providing batch-level reference metrics, putative batch-driven artifacts can be identified, reducing false-positive CNV calls and supporting more reliable interpretation in comprehensive genomic profiling.
Next-generation sequencing is routinely performed in clinical practice to detect various types of mutations for targeted therapy, diagnosis, and prognosis. Actionable alterations detected by next-generation sequencing include not only nonsynonymous mutations that lead to functional or structural changes of proteins but also copy number variants (CNVs) that affect gene dosage, such as gene copy gains, amplifications, or deletions. Among tumor-only CNV detection methods, the use of a panel of normals for relative comparison has become a common practice, largely because of the lack of matched normal samples. It was therefore hypothesized, once established, a panel of normals and CNV caller may not fully compensate for all experimental variations, such as differences in probe efficiency across reagent lots. To investigate this, 12,104 clinical sequencing data sets from 1454 sequencing batches were analyzed over a 4-year period. This analysis revealed batch-associated fluctuation patterns in gene-level fold changes that could potentially lead to misinterpretation, such as the incorrect classification of gene copy deletions or gains. In this study, a strategy is presented that calculates the median and median absolute deviation of gene-level fold changes across all samples within each sequencing batch and incorporates these measures into the result interpretation. By providing batch-level reference metrics, putative batch-driven artifacts can be identified, reducing false-positive CNV calls and supporting more reliable interpretation in comprehensive genomic profiling.
Lung adenocarcinoma (LUAD) remains a leading cause of cancer-related mortality. While for Stage IA LUAD surgery is often curative, recurrence rates remain significant. Liquid biopsy enables monitoring residual disease and predicts recurrence, however utility in Stage IA LUAD is not well-established. Tumor-informed whole genome sequencing (WGS) represents a highly sensitive liquid biopsy method for the analysis of circulating-tumor cell-free DNA (ctDNA) without designing and maintaining patient-specific probes. We aimed to determine the prognostic value of genome-wide tumor-informed minimal residual disease (MRD) monitoring in Stage IA NSCLC using WGS of ctDNA. WGS was performed on 42 patients with Stage IA NSCLC. Tumor was sequenced at 40x and germline DNA and plasma derived ctDNA at 20x and patient specific mutational signatures were developed from the tumor-normal comparison and applied to ctDNA WGS at each time point using AI-supported pattern recognition algorithm. ctDNA results were compared with clinical recurrence. WGS ctDNA was able to predict recurrence with 0.75 sensitivity and 0.83 specificity, with median 16.7 months lead time compared to clinical or imaging recurrence. ctDNA WGS was able to distinguish between a second primary and recurrence in histologically or clinically challenging cases. Whole genome sequencing of ctDNA can predict recurrence in the earliest clinical stage of NSCLC, identifying patients who will recur, and providing information to help guide radiological follow-up and adjuvant therapy.
Publicly available RNA-sequencing data provide a cost-effective springboard for biomarker discovery. However, heterogeneity across studies often complicates analysis. This study presents a modular analytics pipeline that combines publicly available data sets with established open-source tools; standardizes quality control, differential expression analysis, and pathway analysis; and leverages competitive machine learning to unify disparate RNA-sequencing data sets for robust biomarker identification. The workflow is demonstrated across three disease contexts, ranging from small pilot data sets to larger, integrated analyses: i) identifying differentially expressed gene signatures associated with coronavirus disease 2019 (COVID-19) severity, ii) combining differential expression and machine learning techniques to analyze multicohort sepsis data sets, resulting in concise biomarker panels, and iii) using both bulk and single-cell data to examine tissue and cell type specificity of N-acyl-phosphatidylethanolamine phospholipase D in atherosclerosis. These applications demonstrate how an adaptable, modular pipeline using open-source tools can repurpose public data to reduce noise, generate new hypotheses, and reveal meaningful biological insights, thereby establishing a foundation for future research and underlining the importance of public data in exploratory biomarker discovery.
Spinal muscular atrophy (SMA) is a common fatal genetic disorder with a high carrier rate. For the prevention program and treatment plan of this disease, a comprehensive assay is practically needed to gain multigenetic information, enabling effective molecular screening and accurate clinical classification. In this study, a novel single-tube matrix-assisted laser desorption/ionization time-of-flight (MALDI-TOF) mass spectrometry assay was developed to simultaneously analyze the copy numbers of survival motor neuron (SMN), neuronal apoptosis-inhibitory protein (NAIP), small EDRK-rich factor 1A (H4F5), and general transcription factor IIH subunit 2 (GTF2H2); identify six common loss-of-function variants in SMN1; and identify hybrid SMN. The results of this assay were 100% consistent with the predetermined values in 158 genotype-known samples. Furthermore, a cohort of 162 clinical samples was double-blind detected to evaluate this assay, and the results showed a 100% overall concordance with comparative methods. Herein, this novel MALDI-TOF mass spectrometry assay was found to be practical and cost-effective and to have multigenetic information available, which is suitable for large-scale molecular screening and routine clinical diagnosis of spinal muscular atrophy.
Sudden unexplained nocturnal death syndrome (SUNDS), a subtype of sudden unexplained death, predominantly affects young, otherwise healthy individuals, with a higher prevalence in male subjects and a geographic concentration in Southeast Asia, particularly Thailand. Despite extensive investigation, the genetic basis of SUNDS remains incompletely understood. Sarcomeric and non-sarcomeric gene variants were investigated in 98 SUNDS cases using whole-exome sequencing. Postmortem cardiac examination and molecular modeling were performed to assess myocardial abnormalities and the structural impact of selected sarcomeric variants. Eleven missense variants in five sarcomeric genes (MYBPC3, MYH7, TNNI3, TNNT2, and TPM1) were identified in 12 cases (12.2%), whereas 37 variants across 18 non-sarcomeric genes were detected in 29 cases (29.6%). The MYH7 variant c.1562T>C (p.Ile521Thr) was identified as likely pathogenic. Cardiac histopathology revealed heterogeneous myocardial remodeling, including myocyte hypertrophy and interstitial fibrosis; some MYH7 variant carriers displayed increased left ventricular wall thickness (>1.5 cm). Molecular modeling of the TPM1 variant c.641A>G (p.Tyr214Cys), located in the hinge region of the tropomyosin-troponin regulatory complex, suggested disruption of a native π-π stacking interaction, potentially affecting thin filament stability. Sarcomeric variants were associated with heterogeneous myocardial remodeling in SUNDS. These findings highlight the genetic and pathologic heterogeneity of SUNDS and suggest that sarcomeric gene variants represent a previously underrecognized contributor to myocardial remodeling and arrhythmogenic risk.
The Australian Reproductive Genetic Carrier Screening Project, "Mackenzie's Mission," performed couple-based screening for 9107 reproductive couples for >1280 genes associated with severe autosomal and X-linked recessive disorders. It identified close to 1:50 participating couples as having a previously unknown increased risk of having children with a condition screened. Here, the processes, successes, and challenges of the laboratory testing undertaken during Mackenzie's Mission are described. Samples were mouth swabs, self-collected and posted to one of three testing laboratories. Two laboratories used exome sequencing; the other used a targeted gene panel. Both were equally effective in identifying increased risk couples for small sequence variants. FMR1 and SMN1 testing were performed separately. A Variant Review Committee met weekly to discuss reportable variants and variants difficult to classify. This promoted consistent reporting. Exclusion of variants previously classified by a laboratory as benign, likely benign, and/or variant of uncertain significance significantly reduced the analysis required for each reproductive couple. The pan-ancestral screening approach was suitable for data from people of various ancestries despite the increased analysis time required for data from couples of African, Middle Eastern, and Asian ancestry. For consanguineous couples, there was a significant increase in the number of variants requiring review and approximately 10-fold more increased risk reports issued. This study demonstrated large-scale, pan-ancestral reproductive genetic carrier screening is feasible across Australia's diverse population.
Cystoscopy is the gold standard for bladder cancer (BC) diagnosis and surveillance but experiences invasiveness and patient discomfort. Existing urine-based assays with single-gene mutation or methylation biomarkers exhibit insufficient diagnostic efficacy, necessitating a reliable noninvasive detection approach. This study developed EpiMut-BC, a high-specificity urinary PCR assay integrating DNA mutation and methylation signatures based on pyrophosphorolysis-activated polymerization real-time quantitative PCR with engineered polymerases and blocked primers. A total of 502 participants from two Chinese tertiary hospitals, including 302 patients with histologically confirmed BC and 200 noncancer controls, were prospectively enrolled and divided into training (n = 93) and validation (n = 409) cohorts. All diagnostic outcomes were verified against cystoscopy and histopathology. Among 489 valid samples, EpiMut-BC achieved an overall sensitivity of 85.9%, a specificity of 95.8%, and an accuracy of 89.8%, with favorable predictive values and consistent specificity across diverse control subgroups. Notably, it yielded a sensitivity of 70.3% for early-stage tumors and 81.7% for superficial tumors. Validated in a large multicenter cohort, this noninvasive assay enables robust early detection of BC even with low-abundance tumor DNA, which can effectively reduce reliance on invasive cystoscopy and unnecessary repeated surgical resection.
Cell-free total nucleic acid (cfTNA)-based liquid biopsy (LBx) offers a minimally invasive alternative to tissue-based next-generation sequencing (NGS). Most NGS-based LBx assays require several days for results. This study validated a rapid NGS-based LBx assay with a 2-day turnaround time (TAT). Patient plasma samples were used to validate the Oncomine Precision Assay for single-nucleotide variants (SNVs), insertion and deletions (indels), and fusions across 50 genes on the Genexus Sequencer. Manual extraction using the QIAamp cfTNA Kit was performed to optimize fusion detection. Analytical sensitivity in synthetic controls (N = 20) using the automated workflow was 99.2% for SNVs and 95% for indels, with 100% specificity (allele frequency ≥0.5%). In clinical samples (N = 107), sensitivity was 99.4% for SNVs and 100% for indels; specificity was 100% for both, compared with orthogonal assays. Manual extraction improved overall performance compared with the automated extraction and was used for the final clinical workflow. Analytical sensitivity of fusions in synthetic controls (N = 22) was 98.9%, with 100% specificity (≥7 copies). In clinical samples (N = 38), sensitivity and specificity for both SNVs and indels were 100% and 99.9%, respectively; sensitivity for fusions was 72.2%; and specificity was 100%. Overall precision was >99%, and average TAT was ≤2 days. This study reports the feasibility and validation of a 2-day TAT NGS-based cfTNA assay that can potentially reduce time to treatment.
Microscopic examination of blood smears is the current gold standard for laboratory confirmation of malaria. However, it is time-consuming and lacks sensitivity while requiring highly skilled laboratory professionals. This study validated a multiplex real-time PCR test for the detection, speciation, and semiquantification of Plasmodium species in whole blood. Analytical results showed the assays were highly specific and sensitive, with a PCR efficiency of 100.21% and 101.99% for Plasmodium falciparum and Plasmodium vivax, respectively. The limit of detection was 0.0001% parasitemia for both P. falciparum and P. vivax with remarkable precision. Accuracy was validated by testing 128 EDTA-whole blood samples from patients suspected of having malaria, with only two discordant results (positive percentage agreement = 98.65%; negative percentage agreement = 98.15%). Moreover, the linearity between cycle threshold value and parasite density (R2 = 0.941) was validated by testing samples serially collected over a 12-day period from a patient positive for P. falciparum, which enabled semiquantitative estimation of parasitemia. This malaria PCR test demonstrates high sensitivity, specificity, reproducibility with minimal hands-on time, and a quick technical turnaround time of <2 hours with capability of providing parasite density estimation to improve diagnosis and monitor treatment of malaria.
Tumor next-generation sequencing (NGS) is widely used to refine diagnosis and identify therapy targets. However, reporting criteria, schemas, and formats vary greatly, which can affect uniformity of clinical cancer care. With the aim of promoting harmonization, the current state of NGS reporting practices was profiled across Genomics Organization for Academic Laboratories members. The assessment included a group landscape analysis to refine topics followed by a survey of member laboratories and post-survey discussions. A total of 28 surveys covering hematology and/or solid tumor panels from 21 academic laboratories were analyzed. Most responses indicated use of one or more variant tiering systems and reporting of all presumed somatic variants of uncertain significance. Most indicated significant manual effort by directors in generating final reports related to variant annotation using multiple external and internal laboratory databases, evaluation for potential germline variants, and correlation with prior NGS studies and clinical context. Report differences by indication related to more frequent inclusion of longitudinal comparisons for hematologic neoplasms and potential therapies for solid tumor reports. On the basis of areas of strong consensus among survey participants, considerations for best practices are presented along with opportunities for future harmonization, which may require improvements in classification schemas or better software tools.
A custom Genexus myeloid assay (CMA) underwent a technical evaluation for detection of variants from both DNA and RNA in a single assay format. The custom assay was initially verified with commercial DNA and RNA controls containing known myeloid variants. Seventy-five patient specimens with various DNA and RNA variants were selected for replicate testing. The CMA generated consistent data on two Genexus sequencers, with occasional samples failing quality metrics randomly. All 22 control DNA variants were detected, with 95% reported as key. Sensitivity and positive predictive value for clinical variants were 95.60% and 97.60%, respectively, increasing to 98.63% and 99.65%, respectively, for allele frequencies ≥5%. FLT3 duplications were consistently detected at lower expected frequencies. Low-level, false-positive key variants were mostly recurrent and filterable. Fusion calling sensitivity was 100% for control and clinical RNA samples, with 97.09% of replicates reporting expected fusions. The CMA reliably detects key DNA and RNA variants in myeloid specimens within 24 hours on the Genexus platform. Variants and fusion transcripts were identified with high sensitivity and minimal nucleic acid input, providing a rapid and precise workflow for reporting clinically relevant myeloid disease variants.
Next-generation sequencing is a first-tier test in molecular oncology, capable of detecting sequence variants and copy number variants (CNVs). Although most amplicon-based gene panels are not designed to detect segmental chromosomal CNVs, they can still be inferred. This study describes a custom analysis for CNV detection using the amplicon-based Oncomine Comprehensive Assay version 3, for critical glioma biomarker 1p/19q codeletion, +7/-10, EGFR amplification, and CDKN2A/B homozygous deletion. The segmental CNV calling algorithm leverages vendor-normalized gene-level CNV data with intersample normalization to control regions. Assay performance was assessed across three cohorts: a retrospective discovery cohort (N = 52) for threshold establishment, a prospective validation cohort (N = 39) for threshold evaluation against parallel fluorescence in situ hybridization (FISH) testing, and a prospective clinical cohort (N = 53) to assess workflow impact. Positive thresholds were defined to achieve 100% specificity, whereas permissive thresholds triggered reflex FISH to ensure 100% sensitivity. Samples with <30% tumor cellularity were excluded. Threshold evaluation in the validation cohort demonstrated 100% concordance with FISH across all biomarkers. Across glioma subtypes, Oncomine Comprehensive Assay version 3 enabled integrated classification in most cases, particularly in oligodendroglioma and glioblastoma, by simultaneously assessing sequence variants and CNVs. Implementation in the prospective clinical cohort reduced required FISH studies by 90%, significantly shortened turnaround times, and decreased molecular testing costs.
Despite the clinical importance of CD4 testing for identifying advanced HIV disease, access to conventional flow cytometry remains limited in many settings. Epigenetic real-time quantitative PCR (qPCR)-based immune cell quantification represents a molecular alternative that may be compatible with simplified sample types such as dried blood spot (DBS) specimens. This study evaluated the analytical performance of an epigenetic CD4+ T-cell qPCR assay using DBS samples. Residual EDTA whole blood specimens (N = 150) from patients with HIV were applied to Whatman 903 DBS cards and tested following manual DNA extraction and qPCR using the i.Mune CD4 assay. CD4 counts derived from DBS specimens were compared with reference flow cytometry using the AQUIOS PanLeucogating platform. Agreement was assessed by using concordance correlation, Bland-Altman analysis, and percentage similarity. Assay repeatability, batch-related variability, and classification performance at clinically relevant CD4 thresholds were also evaluated. DBS specimen-based epigenetic CD4 quantification showed good agreement with flow cytometry (concordance correlation coefficient, 0.91; 95% confidence interval, 0.88-0.93) with a mean bias of -28 cells/μL (-6.9%). Repeatability was acceptable across the measurement range (%CV, 4.0%-11.2%). At a threshold of 200 cells/μL, sensitivity and specificity were 93.8% and 88.9%, respectively. Increased variability was observed in larger manual extraction batches. These findings show the technical feasibility of epigenetic qPCR-based CD4 quantification from DBS samples and support further optimization and validation.