
The main characteristic of point-of-care testing (PoCT) is to provide rapid test results. It remains uncertain whether faster PoCT results truly improve patient management in the emergency department (ED), and how robust the supporting evidence is. This systematic review aimed to summarize the impact of PoCT in ED on time to clinical decision-making and major clinical outcomes (ED LOS, TAT, mortality, hospitalization, discharge, readmission or antibiotic administration). We searched PubMed, Embase and CENTRAL to identify studies evaluating PoCT in ED. The primary outcomes were the time to decision making and reduction in ED length of stay (LOS). We included 55 studies (approximately 142,000 subjects) evaluating molecular test for virus (n=19) or bacteria (n=2) detection, gastrointestinal infection (n=1), biochemical panel (n=12), troponin (n=9), C-reactive protein (n=2), lactate (n=3), human chorionic gonadotropin (HCG, n=2), d-dimer (n=1), monocyte distribution width (MDW, n=1), blood culture (n=1) and urinary tract infection (n=1) tests. Overall, the use PoCT reduced time to decision-making (ranged from -73 min to -96 min) and ED-LOS (MD -29.24 min 95 % CI from -46.11 to -12.38 min) compared to central laboratory testing. No significant differences were observed for major clinical outcomes. Although substantial statistical heterogeneity was identified, PoCT appears to guide faster clinical decision-making without any detrimental effects on clinical outcomes compared to standard laboratory testing. Future research should better define context dependent applications of PoCT translating faster results into meaningful changes in patient pathways and healthcare processes.
The boom of generative artificial intelligence (AI) is empowering laboratory medicine specialists to create clinical software through conversational prompts without formal programming training, a practice referred to as "vibe coding". Although vibe coding provides unprecedented potential to address workflow inefficiencies, it also introduces significant patient safety risks when AI-generated code is incorporated into lab processes without appropriate validation. A key issue is that laboratory specialists using these AI tools often lack the software engineering background to spot critical flaws in the code. This illusion of competence is exacerbated by the fact that individuals with limited expertise most severely overestimate their abilities, especially when AI tools generate syntactically correct code that looks professional. The distinction between "code that runs" and "code that is safe for clinical use" is significant. Current regulatory frameworks demand a broad set of requirements, such as risk management, extensive validation and verification, rigorous documentation, and quality assurance. Unreviewed and unvalidated vibe-coded software cannot meet these standards. Laboratory medicine is a particularly vulnerable discipline, as even minor errors can impact hundreds or thousands of patients due to the high volume of processed specimens. Nevertheless, vibe coding offers genuine value for rapid prototyping and proof-of-concept development when employed responsibly. An effective approach is that laboratory specialists use generative AI to prototype clinical concepts in a sandbox environment, followed by collaboration with qualified software developers and regulatory experts who translate them into production-grade systems. However, addressing the vibe coding risk necessitates institutional governance to fulfill the innovative potential while preserving the quality systems that laboratory medicine has established over decades.
OBJECTIVES:Accurate quantification of M-proteins is important for clinical management of monoclonal gammopathies. Reliable assessment of trueness on M-protein quantification in External Quality Assessment (EQA) schemes is difficult due to the lack of reference methods and materials for target value assignment. We propose a strategy to produce targeted synthetic M-protein EQA materials by spiking known amounts of human(ised) therapeutic monoclonal antibodies (t-mAb) into null serum. This paper validates homogeneity, commutability, and stability of these materials. METHODS:Homogeneity was assessed on two aliquoted synthetic M-protein EQA materials comparing within-vial variation with between-vial variation for sodium, total protein, IgG and M-protein quantification. Commutability was assessed by comparing between-method differences in 51 patient-derived and nine synthetic M-protein EQA materials among various M-protein quantification methods. Stability was checked by reanalysis of two previously distributed synthetic M-protein EQA materials. RESULTS:Within- and between-vial variation were similar in both samples, and therefore homogeneity criteria were met. Commutability was observed for eight out of nine samples, whereas one hypergammaglobulinemic sample was non-commutable. Long-term storage stability at -80 °C was found acceptable in both samples. CONCLUSIONS:Due to their proven homogeneity, commutability, and stability, targeted synthetic M-protein EQA materials are suitable for use in M-protein EQA schemes. This allows for the assessment of laboratory performances based on target values, and opens the way for harmonisation and exchangeability of M-protein diagnostic methods. Additionally, it provides a new source for the production of M-protein EQA materials, alleviating the demand for large volumes of patient material needed to organise M-protein EQA schemes.
OBJECTIVES:Clinical interpretation of ferritin is hindered by between-assay differences, persisting despite the introduction of different generations of WHO standard. To aid harmonisation efforts, this study investigates the commutability of External Quality Assessment (EQA) materials, proof-of-concept reference materials (cRMs), and the newest generation of WHO standard (19/118). METHODS:Four multi-level serum-based cRMs, 50 clinical samples, four EQA materials, and WHO 19/118 were measured with Abbott Alinity, Beckman Coulter Access DxI, Roche Cobas and Siemens Atellica. Commutability was analysed using the prediction interval approach for EQA materials and the difference in bias and calibration effectiveness approaches for cRMs and WHO standard. RESULTS:Differences in assay selectivity did not allow for formal commutability assessment of EQA materials. Nevertheless, all four investigated EQA materials fell within the 95 % prediction interval around clinical samples in all six method comparisons, indicating their suitability for monitoring harmonisation status. Calibration effectiveness demonstrated commutability for cRM1 and cRM3, in contrast to poor commutability for cRM2, cRM4 and WHO 19/118. Recalibration with cRM1 decreased inter-assay CV four-fold from 34 % to 8.2 %. CONCLUSIONS:Improving harmonisation among current ferritin assays is possible, and success of such efforts can be monitored with EQA materials. Although limited by differences in selectivity, the harmonisation potential of the current assays is substantial, and therefore we encourage harmonisation based on a ISO15194-proof material produced like our cRM1 or cRM3. True standardisation will need much more time, requiring prior measurand characterisation, identification of a clinically relevant ferritin measurand, and subsequent assay redesign to eliminate differences in selectivity.
OBJECTIVES:Using HbA1c for diabetes diagnosis directly links analytical uncertainty to patient classification. Since traditional performance metrics provide limited insight into these clinical implications, we developed a probabilistic framework to translate HbA1c concentrations and components of analytical uncertainty into clinically interpretable diagnostic risk. METHODS:The framework was applied to 100,472 HbA1c results obtained from a real-world Laboratory Information System. Analytical bias, analytical imprecision, and biological variation were integrated to estimate the probability that an HbA1c result would be assigned to an alternative diagnostic category based on medical decision limits. Misclassification risk was evaluated across a broad range of analytical performance conditions, allowing the contribution of analytical performance to be distinguished from the background risk arising solely from biological variation. Benchmarking analyses were performed using five analytical platforms, based on External Quality Assessment (EQA) data. RESULTS:Analytical bias increased misclassification risk more than imprecision. Class-specific analyses showed that analytical error redistributes risk asymmetrically rather than uniformly increasing it; thus, a platform's clinical impact depends strongly on the diagnostic composition of the tested population. EQA benchmarking showed that platforms with apparently acceptable conventional performance produced markedly different EAR profiles. Furthermore, probabilistic reporting scenarios revealed substantial divergence from deterministic classification near diagnostic thresholds. CONCLUSIONS:The proposed Clinical Misclassification Risk Framework (CMRF) quantifies the direct contribution of analytical uncertainty and demographics to patient misclassification. Applicable beyond HbA1c, this framework might support a critical transition in laboratory medicine from assessing metrological quality alone to actively managing clinical decision risk.
Lipoprotein(a) [Lp(a)] is a causal, independent risk factor for atherosclerotic cardiovascular disease and calcific aortic valve stenosis. New guidelines recommend testing everyone at least once in a lifetime, and with new drugs specifically targeting Lp(a) on the horizon, accurate measurement remains vitally important. Yet, its measurement has for decades been challenging because of apolipoprotein(a) size heterogeneity, imperfect calibration strategies, and mixed use of mass (mg/dL) vs. molar (nmol/L) units. This review summarizes Lp(a) biology relevant to its measurement, current analytical approaches and their pitfalls, progress toward standardization and harmonization, including emerging mass-spectrometry-based reference measurement procedures.
INTRODUCTION:Laboratory medicine relies on multiple metrological, analytical, interpretative, and clinical frameworks. Although each addresses a specific question, the relationships linking measurement, interpretation, clinical translation, and action remain incompletely articulated. This fragmentation may encourage inferential overreach, whereby evidence is used to support conclusions beyond the question, intended use, population, or context it can legitimately address. CONTENT:This Opinion Paper proposes the Laboratory Medicine Inference Framework (LMIF), which organises laboratory medicine concepts into five interconnected domains: metrological and analytical foundations; clinical traceability and transferability; result interpretation; clinical translation; and clinical action. Concepts are assigned according to their predominant inferential role while recognising cross-domain dependencies. Two complementary pathways connect the domains: an evidence-to-action pathway towards broader and increasingly context-dependent conclusions, and an intended-use-to-requirements pathway through which clinical needs inform interpretive, transferability, and analytical requirements. Examples involving analytical performance specifications, measurement uncertainty, reference intervals, clinical decision limits, method comparability, compatibility, equivalence, interchangeability, clinical validity, and clinical utility illustrate the framework's inferential boundaries. SUMMARY:The LMIF is underpinned by the Principle of Inferential Sufficiency: a conclusion is justified only when the available evidence adequately addresses the relevant inferential question, intended use, population, and context. The framework complements rather than replaces existing models and provides a common language for identifying evidentiary requirements, transferability needs, and unsupported inferential transitions. OUTLOOK:Future work should evaluate whether the LMIF improves education, method implementation, quality management, scientific communication, clinical consultation, and value-based laboratory medicine, and whether its domains and operational questions require refinement across laboratory and clinical settings.
Cerebrospinal fluid biomarkers, and more recently blood-based biomarkers, are playing a pivotal role in reshaping the clinical management of neurodegenerative diseases, supporting early detection, biological diagnosis, patient stratification, prognostic assessment, therapeutic decision-making, and longitudinal disease monitoring. To effectively support clinical decision-making, biomarker measurements must be communicated in a standardized, transparent, and clinically interpretable manner. Despite international quality standards, considerable heterogeneity persists in reporting practices, including differences in terminology, units of measurement, analytical descriptions, reference frameworks, and interpretative comments, limiting comparability across laboratories and potentially affecting clinical decision-making. In this article, we discuss the principles of harmonized reporting for Alzheimer's disease biomarkers and propose a structured framework for standardized clinical neurochemistry reports. We describe the essential components of a harmonized report, including report identification, analytical information, specification of analytical method and platform, standardized units of measurement, appropriate use of reference intervals and clinical decision thresholds, documentation of pre-analytical and quality-related factors, and evidence-based interpretative comments. We further discuss the importance of structured multimarker interpretation and the need to contextualize biomarker findings within the clinical scenario. Finally, we examine the role of harmonized reporting in promoting interoperability across healthcare systems, facilitating longitudinal patient monitoring, supporting electronic health records and standardized terminologies, and enabling artificial intelligence-driven clinical decision support. As neurodegenerative disease diagnostics continue to evolve and novel biomarkers and technologies emerge, harmonized reporting should be regarded as a critical component of laboratory quality, ensuring that analytical advances translate into consistent, clinically meaningful, and interoperable information that ultimately improves patient care.
OBJECTIVES:Piperacillin is a broad-spectrum antibiotic used to treat critically ill patients with severe infections, requiring therapeutic drug monitoring (TDM). We report the development of an isotope dilution-liquid chromatography-tandem mass spectrometry-based candidate reference measurement procedure (RMP) to quantify piperacillin in human plasma and serum. METHODS:Primary reference material was characterized by quantitative nuclear magnetic resonance (qNMR) to ensure traceability to the International System of Units (SI). Piperacillin was analyzed using LC-MS/MS operating in positive electrospray ionization and multiple reaction monitoring mode. Method validation evaluated selectivity, matrix effects, precision, accuracy, and measurement uncertainty (MU) according to GUM guidelines. RESULTS:This RMP allowed quantification of piperacillin within the range of 0.773 µmol/L (0.400 µg/mL) to 464 µmol/L (240 µg/mL), with selectivity, sensitivity and matrix-independence. Intermediate precision was <2.3 % and <1.1 % for spiked analyte in free human serum and native patient pools, respectively. The repeatability CV ranged from 0.7 to 2.0 % and relative mean bias ranged from -1.5 to 2.9 % across matrices and concentrations. Single measurement expanded MU ranged from 2.8 % to 5.0 %. Expanded MU (k=2) for target value assignment ranged from 2.0 % to 3.4 % across the primary therapeutic range, reaching 7.1 % total error at the lower limit of the measuring interval (LLMI). CONCLUSIONS:This candidate RMP enables accurate determination of piperacillin in human serum and plasma, providing a standardized platform to facilitate reliable clinical TDM and routine assay harmonization.
OBJECTIVES:In vitro serum potassium is affected by ambient temperature with peaks in the winter and troughs in the summer. We investigated whether primary care phlebotomy locations with higher winter hyperkalaemia rates exhibited higher summer hypokalaemia and evaluated a simple screening metric to identify high-risk locations. METHODS:Retrospective audit of Indexor-tracked primary care potassium results analysed at a UK hub laboratory between 1 July 2024 and 30 June 2025 was performed. Bivariate multilevel logistic regression modelling assessed whether locations with higher hyperkalaemia rates in January 2025 (coldest month) also showed higher hypokalaemia rates in June 2025 (warmest month). Locations were ranked by January hyperkalaemia frequency and analysed by quartiles. ROC analysis assessed the performance of ΔK (difference in winter and summer mean potassium) in identifying highest seasonal dyskalaemia locations. RESULTS:Among 382,735 included results, serum potassium was inversely associated with ambient temperature. Hyperkalaemia was 5.24 % in January 2025 vs. 0.90 % in June 2025, while hypokalaemia was 0.39 % vs. 2.22 %, respectively. Location-level propensities for January hyperkalaemia and June hypokalaemia correlated positively (ρ=0.35; 95 % CrI 0.02-0.62). The highest-risk quartile locations had greater winter hyperkalaemia (11.44 % vs. 2.97 %; OR 4.22) and summer hypokalaemia (3.12 % vs. 1.51 %; OR 2.10) compared with the lowest quartile. ΔK discriminated top-quartile risk locations (AUC 0.918 for winter hyperkalaemia), with an optimal threshold ΔK≈0.40 mmol/L. CONCLUSIONS:Phlebotomy locations prone to winter hyperkalaemia also exhibit summer hypokalaemia, reflecting location-specific pre-analytical vulnerabilities. A simple ΔK screen can help targeted mitigation for seasonal dyskalaemia such as on-site centrifugation or transport optimisation.
OBJECTIVES:To compare 24-h urinary excretion of aldosterone (24-h UEA) measured by competitive chemiluminescent immunoassay (cCLIA) and sandwich chemiluminescent immunoassay (sCLIA) with liquid chromatography-tandem mass spectrometry (LC-MS/MS), and to evaluate a direct renin concentration (DRC)-stratified interpretive strategy for primary aldosteronism (PA). METHODS:We prospectively enrolled 705 hypertensive inpatients, including 105 with PA and 600 with essential hypertension. 24-h UEA was measured by cCLIA, sCLIA, and LC-MS/MS, and DRC was measured by CLIA. Method agreement was assessed using Passing-Bablok regression and Bland-Altman analysis. Diagnostic performance was evaluated using receiver operating characteristic (ROC), with internal validation by 1,000 outcome-stratified bootstrap resamples. RESULTS:Both CLIAs correlated strongly with LC-MS/MS (Spearman r=0.946 for cCLIA and 0.950 for sCLIA). Overall areas under the curve (AUCs) were 0.809, 0.814, and 0.812 for cCLIA, sCLIA, and LC-MS/MS, respectively, with no statistically significant differences among the methods. A DRC cutoff of 8.1 mU/L was used to define low-DRC (≤8.1 mU/L) and high-DRC (>8.1 mU/L) subgroups. In the low-DRC subgroup, AUCs were 0.880, 0.879, and 0.881; in the high-DRC subgroup, they were 0.871, 0.878, and 0.859, respectively. Optimism-corrected AUCs remained close to the original estimates. Method-specific high-sensitivity and high-specificity thresholds were used to construct an exploratory DRC-stratified interpretive strategy. CONCLUSIONS:CLIA-based 24-h UEA measurements showed strong correlation and small mean bias with LC-MS/MS and similar diagnostic discrimination in this cohort. DRC stratification may support method-specific interpretation of 24-h UEA, but the proposed thresholds require prospective external validation before clinical implementation.
OBJECTIVES:Serum neurofilament light chain (sNfL) reflects neuroaxonal injury in relapsing-remitting multiple sclerosis (RRMS), but its clinical utility is limited by high interindividual variability and confounders such as renal function. Population-based thresholds may therefore be suboptimal for longitudinal monitoring. The reference change value (RCV), an intraindividual approach based on biological variation, remains insufficiently validated in real-world settings. METHODS:In this retrospective cohort study, 967 sNfL determinations from 295 RRMS patients were analyzed (LUMIPULSE G600II). Measurements were matched to clinical relapse and MRI activity within ±90 days. Associations with log-sNfL were assessed using linear mixed-effects models. RCV and population-based threshold (PBT) were compared as factors associated with inflammatory activity using cluster-robust logistic regression and ROC analysis. Correlation with new MRI lesion count was also evaluated. RESULTS:The intraclass correlation coefficient was 0.536. RCV elevation showed a highly significant association with active MRI (OR 4.93, 95 % CI 2.49-9.77) and clinical relapse (OR 6.48, 95 % CI 2.75-15.25), whereas PBT elevation did not reach statistical significance for either outcome (active MRI: OR 1.52, p=0.290; clinical relapse: OR 2.38, p=0.073). Clinical relapse and MRI activity were independently associated with mean sNfL increases of +72.1 % (95 % CI +53.7 %, +92.7 %) and +24.5 % (95 % CI +13.7 %, +36.3 %), respectively. Renal impairment was a major confounder, with sNfL elevations of +16.0 % in eGFR G2 and +67.8 % in G3-G5. Negative predictive values exceeded 91 % across all markers and outcomes. sNfL correlated with new MRI lesion count (Spearman ρ=0.498, p<0.0001), with a Youden-optimal threshold of ≥4 lesions. CONCLUSIONS:sNfL measured on the LUMIPULSE platform is a valid real-world biomarker of inflammatory activity in RRMS. Its high interindividual variability makes RCV-based intraindividual interpretation substantially more informative than population-based thresholding for longitudinal monitoring. The consistently high negative predictive value supports its use as a rule-out tool for active neuroinflammation and provides a rational basis for risk-adapted MRI surveillance in clinically stable patients.
OBJECTIVES:Thalassemia and hemoglobinopathies are significant genetic disorders in Thailand. Routine diagnostic methods including complete blood count (CBC), hemoglobin (Hb) analysis, and targeted PCR, are effective for detecting common mutations but limited sensitivity for rare and unstable variants, and complex genotypes. Long-read nanopore sequencing may overcome these limitations in a single analytical workflow. This study evaluated the feasibility of CycloneSEQ nanopore sequencing and compared its diagnostic performance with routine methods and DNA nanoball (DNB) short-read sequencing. METHODS:150 EDTA blood samples submitted for routine thalassemia diagnosis were randomly selected. Routine diagnosis included CBC, Hb analysis, and targeted PCR for common α- and β-thalassemia mutations. All samples were further analyzed using DNB and CycloneSEQ sequencing. Diagnostic performance, turnaround time, and cost were compared. RESULTS:Among the three methods, 78.7 % of cases showed concordant results, while 20.0 % could not be fully characterized using routine diagnostic panels. Sequencing methods identified the additional mutations including α+-thalassemia (-α3.7) and rare variants such as Hb Dhonburi, Hb Westmead, etc. Discordant results were observed in 1.3 % of cases. CycloneSEQ required approximately 3-5 days for analysis, whereas DNB sequencing required 10-12 days and routine methods required up to 14 days. Although sequencing costs were approximately 1.7-fold higher than conventional testing, comprehensive mutation detection was achieved in a single step. CONCLUSIONS:CycloneSEQ nanopore sequencing demonstrated enhanced diagnostic performance compared with conventional methods by enabling a single analytical workflow, and shorter turnaround time. This is particularly beneficial for thalassemia management in Thailand and other high-prevalence areas.