66 Background: Despite population screening efforts, screening rates for colorectal cancer (CRC) remain suboptimal. A non-invasive, blood-based screening test with high sensitivity and specificity in early-stage disease should improve adherence and ultimately reduce mortality; however, tests based only on tumor-derived biomarkers have limited sensitivity. Here we used a multiomic, machine learning platform to discover, refine, and combine tumor- and immune-derived signals to develop a blood test for the detection of early-stage CRC. Methods: Samples from 591 participants enrolled in a prospective study including average-risk screening and case-control cohorts (NCT03688906) were included in this analysis (CRC: n = 43; colonoscopy-confirmed CRC-negative controls: n = 548). Participants with CRC were 60% male with a mean age of 63, and controls were 55% male with a mean age of 60. Stage distribution was 54% early (I/II) and 34% late (III/IV) with 11% unknown. Plasma was analyzed by whole-genome sequencing, bisulfite sequencing, and protein quantification methods. Computational methods were used to assess and infer the performance of individual and combined assays. Results: For colorectal adenocarcinoma, which represents ~95% of all CRCs, our multiomic test achieved a mean sensitivity of 92% in early stage (n = 17) and 84% in late stage (n = 11) at a specificity of 90%. Across all CRC pathological subtypes, our test achieved a mean sensitivity of 80% in early stage (n = 19) and 83% in late stage (n = 12) at a specificity of 90%; the test detected the single squamous cell carcinoma but missed both neuroendocrine tumors. Individual assays achieved a mean sensitivity of 50% in early stage and 66% in late stage at a specificity of 90%. Conclusions: In a prospective cohort, we demonstrated high sensitivity and specificity for early-stage adenocarcinoma by combining tumor- and immune-derived signals from cfDNA, epigenetic, and protein biomarkers. While most CRCs are adenocarcinomas, detection of all pathological subtypes is required to maximize sensitivity in a screening population. Further analysis of molecular and pathological subtypes, as well as the entire ~3000 patient cohort, is underway. Clinical trial information: NCT03688906.
Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to existing screening tests. However, early-stage detection of cancer using tumor-derived cfDNA has proven challenging because of the small proportion of cfDNA derived from tumor tissue in early-stage disease. A machine learning approach to discover signatures in cfDNA, potentially reflective of both tumor and non-tumor contributions, may represent a promising direction for the early detection of cancer. Whole-genome sequencing was performed on cfDNA extracted from plasma samples (N = 546 colorectal cancer and 271 non-cancer controls). Reads aligning to protein-coding gene bodies were extracted, and read counts were normalized. cfDNA tumor fraction was estimated using IchorCNA. Machine learning models were trained using k-fold cross-validation and confounder-based cross-validations to assess generalization performance. In a colorectal cancer cohort heavily weighted towards early-stage cancer (80% stage I/II), we achieved a mean AUC of 0.92 (95% CI 0.91–0.93) with a mean sensitivity of 85% (95% CI 83–86%) at 85% specificity. Sensitivity generally increased with tumor stage and increasing tumor fraction. Stratification by age, sequencing batch, and institution demonstrated the impact of these confounders and provided a more accurate assessment of generalization performance. A machine learning approach using cfDNA achieved high sensitivity and specificity in a large, predominantly early-stage, colorectal cancer cohort. The possibility of systematic technical and institution-specific biases warrants similar confounder analyses in other studies. Prospective validation of this machine learning method and evaluation of a multi-analyte approach are underway.
In this work we explored building automatic speech recognition models for transcribing doctor patient conversation. We collected a large scale dataset of clinical conversations (14,000 hr), designed the task to represent the real word scenario, and explored several alignment approaches to iteratively improve data quality. We explored both CTC and LAS systems for building speech recognition models. The LAS was more resilient to noisy data and CTC required more data clean up. A detailed analysis is provided for understanding the performance for clinical tasks. Our analysis showed the speech recognition models performed well on important medical utterances, while errors occurred in causal conversations. Overall we believe the resulting models can provide reasonable quality in practice.
High-dimensional data acquired from biological experiments such as next generation sequencing are subject to a number of confounding effects. These effects include both technical effects, such as variation across batches from instrument noise or sample processing, or institution-specific differences in sample acquisition and physical handling, as well as biological effects arising from true but irrelevant differences in the biology of each sample, such as age biases in diseases. Prior work has used linear methods to adjust for such batch effects. Here, we apply contrastive metric learning by a non-linear triplet network to optimize the ability to distinguish biologically distinct sample classes in the presence of irrelevant technical and biological variation. Using whole-genome cell-free DNA data from 817 patients, we demonstrate that our approach, METric learning for Confounder Control (METCC), is able to match or exceed the classification performance achieved using a best-in-class linear method (HCP) or no normalization. Critically, results from METCC appear less confounded by irrelevant technical variables like institution and batch than those from other methods even without access to high quality metadata information required by many existing techniques; offering hope for improved generalization.
Statistical learning on biological data can be challenging due to confounding variables in sample collection and processing. Confounders can cause models to generalize poorly and result in inaccurate prediction performance metrics if models are not validated thoroughly. In this paper, we propose methods to control for confounding factors and further improve prediction performance. We introduce OrthoNormal basis construction In cOnfounding factor Normalization (ONION) to remove confounding covariates and use the Domain-Adversarial Neural Network (DANN) to penalize models for encoding confounder information. We apply the proposed methods to simulated and empirical patient data and show significant improvements in generalization.
Introduction: Despite population screening and availability of several stool-based, non-invasive screening methods, over 20% of colorectal cancers (CRC) in the US are metastatic at the time of diagnosis. Blood-based methods using cell-free DNA (cfDNA) are under development as an alternative to stool-based tests. However, early stage detection of cancer using tumor-derived mutations in cfDNA (circulating tumor DNA, or ctDNA) has proven challenging because of the small proportion of cfDNA derived from tumor tissue (tumor fraction, ctDNA/cfDNA ratio) in early stage disease. Using an artificial intelligence-driven approach based on machine learning (ML) to discover signatures in cfDNA potentially reflective of both tumor and immune contributions may represent a promising direction for the early detection of cancer. Methods: De-identified plasma samples (N=1,040) were received from academic clinical studies and commercial biobanks (n=579 CRC patients; n=461 controls). Whole-genome sequencing was performed to >50M reads on cfDNA extracted from plasma. Reads aligning to expressed sequences in the genome were extracted and read counts were normalized to account for variability in read depth, sequencecontent bias, and technical batch effects. cfDNA tumor fraction was estimated using IchorCNA. ML models were trained using 10-fold cross-validation stratified by sequencing batch to mitigate bias from sequencing batch effects. Results: In a cohort heavily weighted towards early stage cancer (82% stage I/II), our method achieves a sensitivity in cross-validation of 81% (Clopper-Pearson 95% confidence interval, 77-84%) at 85% specificity. Sensitivity generally increased by tumor stage. Stratification by sequencing batch was required for reliable generalization. Further analyses revealed susceptibility to additional confounders, including variation in preanalytical and analytical processes such as institution-specific blood collection protocols. Downsampling the dataset to balance with respect to such confounders can reduce sensitivity at 85% specificity by 20-30%. Conclusion: An ML approach using a single analyte was able to achieve high sensitivity and specificity in an predominantly early-stage CRC cohort. The observation of systematic technical and site-specific biases warrants similar confounder analyses in other retrospective studies. Prospective validation of the presented ML method and evaluation of a previously presented multi-analyte approach are underway.