Introduction: The promise of real-world data in advancing clinical research for patients with hematologic malignancies is often undermined by limitations in data collection. While Electronic Health Records (EHR) are widely adopted, the collected information is often unstructured. For instance, data that is fundamental to leukemia clinical research such as bone marrow pathology reports remain largely unoptimized for automated extraction and thus inaccessible to large-scale analyses. Large language models (LLM) might represent a solution to this problem by allowing efficient extraction of large amounts of data. However, its application in clinical research remains unclear. Hence, we sought to use LLM to systematically extract information from bone marrow biopsy pathology reports. Methods: We collected full-text pathology reports from bone marrow biopsies performed to evaluate new onset pancytopenia at Yale-New Haven Hospital. Data extraction was performed on unprocessed text using OpenAI Generative Pre-Trained Transformer 4.1 (gpt-4.1, version 2025-04-14) and gpt-o3 (version 2025-04-16) in a private, HIPAA-compliant environment. Both models were used in their standard version, without any fine-tuning and with a zero-shot prompting strategy. Extracted fields were Medical Record Number (MRN), final diagnosis, biopsy-derived sample quality, cellularity, fibrosis grading, blast percentage; aspirate-derived sample quality, presence of dysplasia, ring sideroblasts and aspirate blast count; and flow cytometry-derived sample quality and blast count. Performances were compared against a dataset manually annotated by expert hematologist review. Accuracy, Agresti-Coull adjusted 95% confidence intervals (CI), hallucination rate and omission rate were computed for each variable. Categorical variables were assessed using Cohen's kappa (k), while numerical variables were evaluated through Spearman's correlation coefficient (r) and Root Mean Square Error (RMSE). Results: This study included 376 pathology reports from unique patients. For gpt-4.1, accuracy was 96.7%, omission rate was 0.6% and hallucination rate was 0.5%. For categorical variables, perfect extraction (accuracy 100%, k=1) was achieved for MRN and presence of fibrosis. Near perfect (>95%) accuracy was achieved for sample quality for trephine biopsy, aspirate and flow, and presence of ring sideroblasts. Accuracy for final diagnosis was 91% (CI=0.88-0.94, k=0.94). Extraction accuracy of continuous variables was near-perfect (99%) for biopsy cellularity (CI=0.98-1, r=0.99, RMSE=2.1), fibrosis grading (CI=0.98-1, r=0.99, RMSE=0) and aspirate blast count (CI=0.98-1, r=1, RMSE=0). Extraction accuracy was slightly lower for biopsy blast count (95%, CI=0.92-0.97, r=0.99, RMSE=1.1) and flow blast count (97%, CI=0.94-0.98, r=0.99, RMSE=2.2) due to higher hallucination rates (3.2% and 0.8%, respectively). Assessment of dysplasia showed lower accuracy for all lineages: erythroid (94%, k=0.96), granulocytic (92%, k=0.93) and megakaryocytic (83%, k=0.67). Accuracy for gpt-o3 was increased at 97.4%, with omission and hallucination rates at 0.7% and 0.2%. Performances in categorical variables extraction were slightly increased for megakaryocytic dysplasia, with accuracy at 86% (k=0.8) and were comparable for the remaining variables. For continuous variables, near-perfect accuracy (99%) was confirmed for cellularity (CI=0.98-1, r=1, RMSE=0), fibrosis grading (CI=0.98-0.99, r=1, RMSE=0) and aspirate blast count (CI=0.98-1, r=1, RMSE=0). A minor improvement was observed in the accuracy for biopsy blast count at 98% due to a lower hallucination rate (0%). Runtime for gpt-o3 on the full dataset was significantly longer at 122 minutes compared with 12 minutes for gpt-4.1. Conclusions: Although prior studies have explored automated extraction from pathology reports using expert systems, rule-based algorithms, and general LLM approaches, this is the first attempt to apply LLM-mediated extraction specifically within malignant hematology, which presents unique challenges compared to solid oncology. Our methodology achieved near perfect accuracy in extracting key bone marrow biopsy datapoints, is scalable for very large datasets and diverse research settings with minimal code adjustments while remaining HIPAA-compliant. The marginal improvements observed with the larger, reasoning gpt-o3 model suggest that smaller, less expensive models can achieve high accuracy with significantly shorter runtimes.
Background: A Quality Improvement (QI) initiative to reduce invasive Staphylococcus aureus (SA) infections in a level IV neonatal intensive care unit (NICU) successfully eliminated Methicillin-resistant (MRSA) but not Methicillin-susceptible (MSSA) infections. A combination of SA whole genome sequencing (WGS) and environmental culturing helped to better understand the epidemiology of MSSA colonization and infection in the NICU and drive new infection prevention interventions. Methods: Environmental surveillance of high-touchpoint surfaces for SA was performed using Dey and Engley neutralizing agar. Selected isolates were confirmed as SA using Columbia Sheep’s Blood agar and Staphaurex testing. Statistical analyses examined correlations between monthly effective cleaning, hand hygiene compliance, and colonization rates. To better understand MSSA spread in the NICU, WGS was performed on a convenience sample of 42 MSSA isolates, sampled one month before and after an invasive MSSA infection. Data extracted from electronic health records were used for retrospective room tracing of colonized patients with related isolates to determine modes of transmission. Results: WGS analysis MSSA isolates revealed four MSSA strains from 29 patients suggesting within unit transmission, while 13 patients were colonized with unique MSSA isolates suggesting external sources. Retrospective room tracing of colonized patients identified three transmission patterns: subsequent room occupant transmission, intra-pod spread, and inter-pod transmission without patient transfer, with evidence that these strains were endemic within the unit for at least 3-12 months. Statistical analyses showed no significant correlation between environmental cleaning or hand hygiene compliance and colonization rates. Conclusions: Persistent MSSA colonization and invasive infections in the NICU result from both within-unit transmission and the introduction of unique isolates. These findings are being used to inform the development of new interventions, including updated below-the-elbow hand hygiene protocols, revised environmental cleaning plans, nurse-parent communication training, and a virtual reality hand hygiene training program for parents and staff. WGS of pathogenic organisms is a useful tool to drive QI initiatives aimed at reducing hospital-acquired infections.
Primary hyperparathyroidism (PHPT) is the third most common endocrine disorder in the United States. Surgical management of primary hyperparathyroidism frequently incorporates intraoperative testing to confirm the presence of parathyroid tissue. At our institution, we utilize ex-vivo aspiration of excised parathyroid tissue, followed by intraoperative parathyroid hormone testing (aIOPTH), which is a rapid and highly sensitive method for distinguishing parathyroid from surrounding tissues. Despite growing adoption, there is limited research on standardized aIOPTH protocols, their relationship with final histopathological results, and their impact on patient outcomes. Various techniques have been described for this purpose, but the literature offers limited guidance on the recommended approach for different surgical scenarios. This study seeks to offer guidance on recommendations for adenoma or multiglandular disease when imploring aIOPTH. This retrospective study was conducted at a tertiary academic medical center and included patients who underwent parathyroidectomy with aIOPTH testing between June 2023 and January 2025. The quantitative aIOPTH values were categorized as either ‘positive’ or ‘negative’ based on absolute thresholds of 500 and 5,000 pg/mL. These categorized results were compared to the final histopathological diagnoses of the excised tissue, and performance metrics were calculated. Separate analyses were conducted for patients with adenoma versus those with multiglandular disease. A total of 161 unique aIOPTH specimens which had a preoperative diagnosis of PHPT were analyzed. With the 500 pg/mL threshold, sensitivity was 97% and PPV was 97%. Using the 5,000 pg/mL threshold, sensitivity decreased to 86%, while PPV increased to 99%. In the subgroup of patients with multiglandular disease, 93 unique aIOPTH specimens were analyzed. In this subgroup, the sensitivity and PPV were 98% and 96% for the 500 pg/mL cutoff, and 84% and 100% for the 5,000 pg/mL cutoff, respectively. At the time of this analysis, 96.6% of patients with who had six-month postoperative calcium testing (n=29) demonstrated durable biochemical cure. The application of absolute cutoffs in aIOPTH testing protocols demonstrates clinical performance consistent with existing literature. Lower PTH cutoffs enhance sensitivity without a substantial reduction in PPV for patients with multiglandular disease. Additionally, aIOPTH shows comparable performance in patients with multiglandular disease. These findings support the use of ex-vivo aspiration for intraoperative parathyroid testing as an effective decision-making tool in parathyroidectomy and lead to guidance for adenoma and multiglandular disease.
BACKGROUND:Intravenous (IV) fluid contamination within clinical specimens causes an operational burden on the laboratory when detected, and potential patient harm when undetected. Even mild contamination is often sufficient to meaningfully alter results across multiple analytes. A recently reported unsupervised learning approach was more sensitive than routine workflows, but still lacked sensitivity to mild but significant contamination. Here, we leverage ensemble learning to more sensitively detect contaminated results using an approach which is explainable and generalizable across institutions. METHODS:An ensemble-based machine learning pipeline of general and fluid-specific models was trained on real-world and simulated contamination and internally and externally validated. Benchmarks for performance assessment were derived from in silico simulations, in vitro experiments, and expert review. Fluid-specific regression models estimated contamination severity. SHapley Additive exPlanation (SHAP) values were calculated to explain specimen-level predictions, and algorithmic fairness was evaluated by comparing flag rates across demographic and clinical subgroups. RESULTS:The sensitivities, specificities, and Matthews correlation coefficients were 0.858, 0.993, and 0.747 for the internal validation set, and 1.00, 0.980, and 0.387 for the external set. SHAP values provided plausible explanations for dextrose- and ketoacidosis-related hyperglycemia. Flag rates from the pipeline were higher than the current workflow, with improved detection of contamination events expected to exceed allowable limits for measurement error and reference change values. CONCLUSIONS:An accurate, generalizable, and explainable ensemble-based machine learning pipeline was developed and validated for sensitively detecting IV fluid contamination. Implementing this pipeline would help identify errors that are poorly detected by current clinical workflows and a previously described unsupervised machine learning-based method.
Background Observable quantitative variations exist between plasma and serum in routine protein measurements, often not reflected in standard reference intervals. In this study, we describe an indirect approach for estimating a combined reference interval (RI) (i.e., serum and plasma), for commonly ordered protein measurands: total protein, albumin, and globulin. Methods We applied an indirect reference interval estimation for protein measurements in serum and plasma using data from July 2018 to February 2024. The data were divided into three Epochs based on a period of plasma separator tube shortage during the COVID-19 pandemic. Bootstrap resampling was used to calculate RIs and corresponding 95% confidence intervals for each month. Results Our results demonstrate notable changes in RI limits for total protein, albumin, and globulin between Epochs, reflecting the influence of changing sample matrix. A combined RI was identified for all components and verified using plasma and serum samples from 20 healthy individuals and retrospective analysis of flagging rates on our outpatient population using new and historical RIs. Conclusion The study demonstrates notable differences in the RIs for total protein, albumin, and globulin when container type changes. In addition, the results demonstrate the effectiveness of big data analytics in deriving RIs and highlights the necessity of continuous RI assessment and adjustment based on the patient population and acceptable specimen types.
Objectives To introduce quantum computing technologies as a tool for biomedical research and highlight future applications within healthcare, focusing on its capabilities, benefits, and limitations.Target Audience Investigators seeking to explore quantum computing and create quantum-based applications for healthcare and biomedical research.Scope Quantum computing requires specialized hardware, known as quantum processing units, that use quantum bits (qubits) instead of classical bits to perform computations. This article will cover (1) proposed applications where quantum computing offers advantages to classical computing in biomedicine; (2) an introduction to how quantum computers operate, tailored for biomedical researchers; (3) recent progress that has expanded access to quantum computing; and (4) challenges, opportunities, and proposed solutions to integrate quantum computing in biomedical applications.
BACKGROUND:Acute kidney injury (AKI) is a serious complication affecting up to 15% of hospitalized patients. Early diagnosis is critical to prevent irreversible kidney damage that could otherwise lead to significant morbidity and mortality. However, AKI is a clinically silent syndrome, and current detection primarily relies on measuring a rise in serum creatinine, an imperfect marker that can be slow to react to developing AKI. Over the past decade, new innovations have emerged in the form of biomarkers and artificial intelligence tools to aid in the early diagnosis and prediction of imminent AKI.CONTENT:This review summarizes and critically evaluates the latest developments in AKI detection and prediction by emerging biomarkers and artificial intelligence. Main guidelines and studies discussed herein include those evaluating clinical utilitiy of alternate filtration markers such as cystatin C and structural injury markers such as neutrophil gelatinase-associated lipocalin and tissue inhibitor of metalloprotease 2 with insulin-like growth factor binding protein 7 and machine learning algorithms for the detection and prediction of AKI in adult and pediatric populations. Recommendations for clinical practices considering the adoption of these new tools are also provided.SUMMARY:The race to detect AKI is heating up. Regulatory approval of select biomarkers for clinical use and the emergence of machine learning algorithms that can predict imminent AKI with high accuracy are all promising developments. But the race is far from being won. Future research focusing on clinical outcome studies that demonstrate the utility and validity of implementing these new tools into clinical practice is needed.
We describe a novel approach to clinical decision support (CDS) for triaging specimens within the clinical laboratory for severe acute respiratory syndrome coronavirus 2 (SARS‑CoV‑2) nucleic acid amplification tests (NAAT). The use of our CDS tool could help clinical laboratories prioritize and process specimens efficiently, especially during times of high demand. There were significant differences in the turnaround time for specimens differentiated by icons on specimen labels. Further studies are needed to evaluate the impact of our CDS tool on overall laboratory efficiency and patient outcomes.
Journal Article Data Analytics in Clinical Laboratories: Advancing Diagnostic Medicine in the Digital Age Get access Anna E Merrill, Anna E Merrill Clinical Associate Professor of Pathology, University of Iowa College of Medicine; Associate Director of Clinical Chemistry, University of Iowa Hospitals and Clinics, Iowa City, IA, United States Address correspondence to this author at: University of Iowa Hospitals and Clinics, Department of Pathology, 200 Hawkins Drive, RCP 6234, Iowa City, IA, 52242, United States. E-mail anna-merrill@uiowa.edu. https://orcid.org/0000-0002-5945-5937 Search for other works by this author on: Oxford Academic Google Scholar Thomas J S Durant, Thomas J S Durant Assistant Professor of Laboratory Medicine, Biomedical Informatics and Data Science, Yale School of Medicine; Medical Director of Chemical Pathology and Laboratory Informatics, Associate Director of ACGME Chemical Pathology Fellowship, Yale-New Haven Hospital, New Haven, CT, United States Search for other works by this author on: Oxford Academic Google Scholar Jason Baron, Jason Baron Clinical Data Scientist, Roche Diagnostics Corporation, Indianapolis, IN, United States Search for other works by this author on: Oxford Academic Google Scholar J Stacey Klutts, J Stacey Klutts Deputy Director, National Pathology and Laboratory Medicine Service, Veterans Health Administration, Washington, DC, United StatesClinical Associate Professor of Pathology, University of Iowa College of Medicine, Iowa City, IA, United States Search for other works by this author on: Oxford Academic Google Scholar Amrom E Obstfeld, Amrom E Obstfeld Associate Chair of Pathology Informatics, Children's Hospital of Philadelphia; Associate Professor of Clinical Pathology and Laboratory Medicine, University of Pennsylvania Perelman School of Medicine, Philadelphia, PA, United States Search for other works by this author on: Oxford Academic Google Scholar David Peaper, David Peaper Associate Professor of Laboratory Medicine, Yale School of Medicine; Medical Director of Clinical Microbiology Laboratory, Yale-New Haven Hospital, New Haven, CT, United States https://orcid.org/0000-0002-9952-5012 Search for other works by this author on: Oxford Academic Google Scholar Michelle Stoffel, Michelle Stoffel Associate Chief Medical Information Officer for Laboratory Medicine & Pathology, Medical Director of Laboratory Medicine & Pathology Informatics, M Health Fairview; Assistant Professor of Laboratory Medicine & Pathology, University of Minnesota, Minneapolis, MN, United States https://orcid.org/0000-0002-4041-748X Search for other works by this author on: Oxford Academic Google Scholar Sarah Wheeler, Sarah Wheeler Associate Professor of Pathology, University of Pittsburgh Medical Center; Associate Medical Director of Clinical Immunopathology, Medical Director of Automated Laboratory UPMC Mercy and Clinical Chemistry UPMC Children's Hospital of Pittsburgh, Pittsburgh, PA, United States https://orcid.org/0000-0002-7851-9836 Search for other works by this author on: Oxford Academic Google Scholar Mark A Zaydman Mark A Zaydman Assistant Professor of Pathology and Immunology, Washington University School of Medicine, St. Louis, MO, United States Search for other works by this author on: Oxford Academic Google Scholar Clinical Chemistry, Volume 69, Issue 12, December 2023, Pages 1333–1341, https://doi.org/10.1093/clinchem/hvad183 Published: 14 November 2023 Article history Received: 06 October 2023 Accepted: 17 October 2023 Published: 14 November 2023
Background Network-connected medical devices have rapidly proliferated in the wake of recent global catalysts, leaving clinical laboratories and healthcare organizations vulnerable to malicious actors seeking to ransom sensitive healthcare information. As organizations become increasingly dependent on integrated systems and data-driven patient care operations, a sudden cyberattack and the associated downtime can have a devastating impact on patient care and the institution as a whole. Cybersecurity, information security, and information assurance principles are, therefore, vital for clinical laboratories to fully prepare for what has now become inevitable, future cyberattacks. Content This review aims to provide a basic understanding of cybersecurity, information security, and information assurance principles as they relate to healthcare and the clinical laboratories. Common cybersecurity risks and threats are defined in addition to current proactive and reactive cybersecurity controls. Information assurance strategies are reviewed, including traditional castle-and-moat and zero-trust security models. Finally, ways in which clinical laboratories can prepare for an eventual cyberattack with extended downtime are discussed. Summary The future of healthcare is intimately tied to technology, interoperability, and data to deliver the highest quality of patient care. Understanding cybersecurity and information assurance is just the first preparative step for clinical laboratories as they ensure the protection of patient data and the continuity of their operations.
The Coronavirus Disease of 2019 (COVID-19) pandemic has been a challenging event for laboratory medicine and diagnostics manufacturers. We have had to confront numerous unique and previously unthinkable issues on a daily basis in order to continue offering diagnostic testing for not only Severe Acute Respiratory Syndrome Coronavirus-2 (SARS-CoV-2), but other testing that was significantly impacted by supply chain and staffing disruptions related to COVID-19. Out of this tremendously stressful and, at times, chaotic environment, decades of innovations and advances in testing methodologies and instrumentation became essential to handle the overwhelming volume of samples with clinically appropriate turn-around-time. Additionally, a number of novel testing approaches and technological innovations emerged to address laboratory and public health needs for widespread testing. In this review we consider both technological advances in infectious diseases testing and other innovations in sample collection, processing, automation, workflow, and testing that have embodied the laboratory response to the COVID-19 pandemic.
Graph data models are an emerging approach to structure clinical and biomedical information. These models offer intriguing opportunities for novel approaches in healthcare, such as disease phenotyping, risk prediction, and personalized precision care. The combination of data and information in a graph model to create knowledge graphs has rapidly expanded in biomedical research, but the integration of real-world data from the electronic health record has been limited. To broadly apply knowledge graphs to EHR and other real-world data, a deeper understanding of how to represent these data in a standardized graph model is needed. We provide an overview of the state-of-the-art research for clinical and biomedical data integration and summarize the potential to accelerate healthcare and precision medicine research through insight generation from integrated knowledge graphs.
Journal Article The Roadmap to Interoperability and Laboratory Data: Current State and Next Steps Get access Tylis Chang, Tylis Chang Northwell Health Labs, Pathology Service Line, New Hyde Park, NY, USA Search for other works by this author on: Oxford Academic Google Scholar Daniel S Herman, Daniel S Herman University of Pennsylvania, Pathology and Laboratory Medicine, Philadelphia, PA, USA https://orcid.org/0000-0003-2873-587X Search for other works by this author on: Oxford Academic Google Scholar David S McClintock, David S McClintock Mayo Clinic, Laboratory Medicine and Pathology, Rochester, MN, USA https://orcid.org/0000-0003-4297-4367 Search for other works by this author on: Oxford Academic Google Scholar Thomas J S Durant Thomas J S Durant Yale School of Medicine, Department of Laboratory Medicine, New Haven, CT, USA Address correspondence to this author at: Department of Laboratory Medicine, 55 Park St. PS345D, New Haven, CT 06511. E-mail thomas.durant@yale.edu. Search for other works by this author on: Oxford Academic Google Scholar The Journal of Applied Laboratory Medicine, Volume 8, Issue 1, January 2023, Pages 226–228, https://doi.org/10.1093/jalm/jfac082 Published: 04 January 2023 Article history Received: 08 July 2022 Accepted: 19 August 2022 Published: 04 January 2023
Background: Urine drug testing (UDT) monitors prescription compliance and/or drug abuse. However, inter-pretation of UDT results obtained by liquid chromatography-tandem mass spectrometry (LC-MS-MS) can be complicated by the presence of drug impurities that are detected by highly sensitive methods. Hydrocodone is a drug impurity that can be found as high as 1% in oxycodone pills.Objectives: We evaluated the frequency and concentration of hydrocodone and its metabolite, hydromorphone, in patients taking oxycodone to check if the ratio of hydrocodone or hydromorphone to oxycodone could distin-guish between oxycodone only use from those consuming additional opiates. Design & methods: We correlated LC-MS/MS results with medication records of 319 patients with positive oxy-codone results over 7 months (4/2021-11/2021).Results: Fifteen of 319 patients with positive oxycodone results were taking oxycodone only. For these 15 pa-tients, the mean ratio of hydrocodone to oxycodone was 0.57% (range 0.05%-3.35%), and the mean ratio of hydromorphone to oxycodone was 0.81% (range 0.18-3.51%).Conclusions: Hydrocodone and/or hydromorphone are detectable in patients taking only oxycodone and can likely be identified as an impurity if their calculated ratio to oxycodone is <1 %. Further validation of the ratios in a larger sample size is recommended.
Artificial intelligence (AI) applications are an area of active investigation in clinical chemistry. Numerous publications have demonstrated the promise of AI across all phases of testing including preanalytic, analytic, and postanalytic phases; this includes novel methods for detecting common specimen collection errors, predicting laboratory results and diagnoses, and enhancing autoverification workflows. Although AI applications pose several ethical and operational challenges, these technologies are expected to transform the practice of the clinical chemistry laboratory in the near future.
BACKGROUND:Laboratorians are left unguided by a paucity of literature on how to configure rules for the detection of intravenous (IV) fluid contamination in blood samples. We designed a study to determine the in vitro effect of increasing blood sample contamination from commonly used crystalloid solutions and how these observations can guide the derivation of multianalyte delta checks to detect such pre-analytical error.METHODS:In this study, we spiked increasing volumes of commonly used IV fluids (normal saline (NS), lactated ringers (LR), and 5% dextrose) into blood samples that were collected from healthy donors. Routine chemistry analytes were measured and compared between neat and contrived samples. From these observations, we derived several permutations of multianalyte delta checks using the basic metabolic panel framework and evaluated rule performance using retrospective data.RESULTS:The wet chemistry experiments showed that increasing the volume of crystalloid solution contamination significantly changed several analytes. Subsequently derived multianalyte delta check procedures were applied to retrospective data. For all IV fluids tested, smaller magnitudes of analyte change resulted in more samples flagged.CONCLUSION:Multianalyte delta checks may be an effective method for the detection of IV fluid contamination.
A 59-year-old male with a past medical history of chronic low back pain, opioid use disorder, and type 2 diabetes mellitus presented to an outpatient clinic for pain contract-related urine drug screen testing.Current medications included metformin, oxycodone, and Suboxone® (buprenorphine and naloxone combination).The patient stated that 5 days prior to the visit, he stopped taking Suboxone due to complaints of nausea and vomiting.Urine toxicology test results are summarized in Table 1.
Flow cytometry (FCM) allows pathologists to accurately immunophenotype hematopoietic cells by detecting the expression of surface proteins with fluorochrome-conjugated antibodies. Accordingly, FCM plays a pivotal role in the diagnostic workup of many hematologic malignancies. However, post-analytic processing and the interpretation of FCM data are primarily manual processes that impact the consistency and limit the throughput of the method. The post-analytic processing workflow is colloquially referred to as ‘gating’ and involves the identification and characterization of immune cell populations by hand through successions of biaxial plots. The gated FCM data is then used by the pathologist to make a clinical interpretation and to then determine potential diagnoses from the patterns and cell frequencies seen in bivariate plots of the data. Since gating and interpretation are manually performed, post-analytic analyses of FCM data are not only laborious from a workflow perspective but remain subjective and prone to variability based on the experience and skill of both the medical laboratory scientist and pathologist. The objective of this study was to develop a computational pipeline that leverages machine learning-based (ML) solutions to automate gating and clinical interpretation of FCM data to increase the throughput and improve the repeatability of FCM analysis. Raw FCS files from clinical samples being evaluated for the presence of T-cell lymphoproliferative disease were exported from the on-instrument database. Automated gating was performed using open-source, supervised-ML packages for flow cytometry data (flowCore and flowDensity; R). These packages implement all processing steps that would typically be done manually (e.g. applying compensation, quality control (QC), and gating). Plots of interest were then generated from the gated data and classified as normal or abnormal using the clinical interpretations that were applied during the normal clinical workflow. These binary labels were used to train an ML-based classifier (VGG-19; Python). To evaluate pipeline performance, we collected FCS files from 1,188 samples that were analyzed by our flow cytometry lab. The automated gating pipeline was used to gate for CD4+CD3+ cells and to create bivariate plots for CD7/CD26 expression. 1150 (96.8%) passed QC and were all gated correctly by visual inspection. Of the 38 (3.2%) samples that failed QC, 13 (1.1%) were lymphopenic and were gated correctly by visual inspection, and the remaining 25 (2.1%) were gated manually. Using the CD7/CD26 plots, the classifier demonstrated promising predictive performance and achieved a precision of 0.85 and recall of 0.83 (weighted average across classes). Our findings represent a novel effort to automate both the gating and interpretation of FCM data using artificial intelligence. These results suggest that ML-based tools have potential utility in aiding the processing and interpretation of FCM data and can augment the efficiency and consistency of workflows in the clinical flow cytometry laboratory.
As the demand for laboratory testing by mass spectrometry increases, so does the need for automated methods for data analysis. Clinical mass spectrometry (MS) data is particularly well-suited for machine learning (ML) methods, which deal nicely with structured and discrete data elements. The alignment of these two fields offers a promising synergy that can be used to optimize workflows, improve result quality, and enhance our understanding of high-dimensional datasets and their inherent relationship with disease. In recent years, there has been an increasing number of publications that examine the capabilities of ML-based software in the context of chromatography and MS. However, given the historically distant nature between the fields of clinical chemistry and computer science, there is an opportunity to improve technological literacy of ML-based software within the clinical laboratory scientist community. To this end, we present a basic overview of ML and a tutorial of an ML-based experiment using a previously published MS dataset. The purpose of this paper is to describe the fundamental principles of supervised ML, outline the steps that are classically involved in an ML-based experiment, and discuss the purpose of good ML practice in the context of a binary MS classification problem.
BACKGROUND: Clinical babesiosis is diagnosed, and parasite burden is determined, by microscopic inspection of a thick or thin Giemsa-stained peripheral blood smear. However, quantitative analysis by manual microscopy is subject to error. As such, methods for the automated measurement of percent parasitemia in digital microscopic images of peripheral blood smears could improve clinical accuracy, relative to the predicate method. METHODS: Individual erythrocyte images were manually labeled as "parasite" or "normal" and were used to train a model for binary image classification. The best model was then used to calculate percent parasitemia from a clinical validation dataset, and values were compared to a clinical reference value. Lastly, model interpretability was examined using an integrated gradient to identify pixels most likely to influence classification decisions. RESULTS: The precision and recall of the model during development testing were 0.92 and 1.00, respectively. In clinical validation, the model returned increasing positive signal with increasing mean reference value. However, there were 2 highly erroneous false positive values returned by the model. Further, the model incorrectly assessed 3 cases well above the clinical threshold of 10%. The integrated gradient suggested potential sources of false positives including rouleaux formations, cell boundaries, and precipitate as deterministic factors in negative erythrocyte images. CONCLUSIONS: While the model demonstrated highly accurate single cell classification and correctly assessed most slides, several false positives were highly incorrect. This project highlights the need for integrated testing of machine learning-based models, even when models in the development phase perform well.