
Medical history taking serves a critical role in improving diagnostic safety and patient outcomes. Despite advances in technology, a patient's medical history remains the cornerstone of accurate diagnosis, often yielding diagnoses before tests are performed. This article explores how patient-centered communication, contextual awareness, and evidence-based strategies reduce diagnostic errors, which often stem from miscommunication during the clinical encounter. It highlights practical tools for clinicians and patients, addresses systemic challenges, and examines how innovations like AI and digital platforms can augment, not replace, the human connection essential to care. By reinforcing the value of thorough, empathetic history taking, this resource advocates for a future where technology and human expertise work hand-in-hand to achieve diagnostic excellence.
OBJECTIVES:Many, if not most, diagnostic errors go unrecognized by clinicians unless they cause significant patient harm. Reducing the gap between clinicians' perceived and actual diagnostic performance (i.e., improving calibration) may improve patient outcomes. However, existing approaches to measuring calibration were developed for use in experimental settings rather than realistic practice environments. We reviewed the scientific literature and solicited expert field input to improve approaches for assessing calibration in clinical practice. METHODS:We performed a targeted literature review to identify studies that compared clinicians' diagnostic performance or accuracy to perceived accuracy or confidence. Key findings from the review were provided to 10 subject matter experts in clinical reasoning and diagnostic excellence. We then interviewed these subject matter experts in a series of four small-group teleconferences. Transcripts of group discussions were analyzed for thematic content relevant to improving the measurement of calibration in real-world care settings. RESULTS:Discussions with experts highlighted that, while calibration reflects skill and expertise relevant to the task at hand, it is also fluctuating and context-dependent. Experts recommended that measurement of calibration should focus on clinically meaningful decisions that may be upstream of a final, correct diagnosis. Experts also recommended assessment of contextual factors that shape diagnostic decision-making and self-assessment. CONCLUSIONS:Measures of calibration should focus not only on outcomes but also on processes, which are often more actionable and proximal to clinical decision-making. In response to expert input, we propose a measure of real-world calibration that builds upon a previous measure to include several contextual variables.
OBJECTIVES:Discrepancies have been observed in the rerun results of plasma alkaline phosphatases (ALP) on Cobas c503 Roche analysers. The aims were to 1) determine the percentage of discrepancies in ALP reruns and 2) evaluate the effect of an automatic dilution on discrepancies. METHODS:1) Raw plasma ALP results from patients were extracted from Cobas pro c503 analysers over a 2.7-year period. Reruns were triggered for any ALP above 200 U/L with no history or with a previous normal ALP, and those with a variation exceeding the total allowable error (±15 U/L if ≤150 U/L or ±10 % if >150 U/L) were classified as discrepancies. 2) During five weeks, a systematic 1:5 dilution was triggered on the analysers for all pure tests, concomitantly. RESULTS:1) Among 9,574 reruns, 677 (7.1 %) were discordant (5.8 %, 1.2 % and 0.1 % after one, two and three reruns, respectively), and 90 % of which were decreases. 2) 9,062 pairs of concomitant pure and diluted ALP were measured, with respective medians (2.5th-97.5th percentiles) of 94 (43-559) and 90 (41-528) U/L (median of differences -4.4 %). Of these, 274 were retested, of which 16 pure reruns (5.8 %) were discordant, compared to only two diluted reruns (0.7 %). There were greater negative variations in ALP reruns as the leukocyte count increased (p<0.001). CONCLUSIONS:A substantial rate of discrepancies was observed when elevated ALP were retested on Cobas c503 analysers. This issue was almost eliminated by applying a first-line 1:5 dilution. A modification of the assay method is eagerly awaited from the manufacturer.
OBJECTIVES:Diagnostic error is common, harmful, and costly, yet most health systems lack active interventions to improve diagnostic safety. This study aimed to identify scalable care models that advance diagnostic excellence while reducing costs for healthcare systems across diverse delivery and reimbursement environments. METHODS:We employed a multi-method care model development framework that integrated: (1) a literature review of 1,632 sources, (2) 19 semi-structured expert interviews, (3) in-depth analysis of four exemplar programs, and (4) four iterative refinement cycles with a 12-member cross-institutional expert panel. Candidate interventions were evaluated by diagnostic failure points, potential net cost savings, operational feasibility, stakeholder value alignment, and payment-model fit. RESULTS:We identified three highest value care model archetypes. (1) Diagnostic safety nets identify patients with abnormal findings lacking appropriate follow-up and re-engage them before harm escalates. (2) Optimized navigation routes patients to the right level of care through risk stratification and selective specialist input. (3) Decision support strengthens diagnostic reasoning at the point of care through evidence-based diagnostic tools. These archetypes differed in infrastructure requirements and financial attractiveness across payment models, with diagnostic safety nets most broadly attractive across both fee-for-service and risk-bearing environments. CONCLUSIONS:This framework introduces three diagnostic improvement archetypes that health systems can use to select and sequence interventions based on local failure points, operational feasibility, and reimbursement context. By linking intervention choice to real-world implementation conditions, the framework offers a pragmatic approach for advancing diagnostic excellence across diverse care delivery settings.
Research over the past decades has established diagnostic errors to be the foremost patient safety challenge in health care today; the harm related to diagnostic errors is unacceptably high. We now understand where, when, and why these errors arise, and a wide range of interventions to improve diagnostic safety and quality are now being proposed, focusing on both the system-related, and personal, cognitive aspects of the diagnostic process. Stanford University's Clinical Excellence Research Center (CERC) has done a great service to the field by convening an expert panel to provide consensus recommendations on economically-beneficial interventions that would have the greatest clinical impact. The CERC report identified these 3 areas as the top priorities: 1 - Optimizing patient navigation and care coordination; 2 - Providing clinicians with decision support resources; and 3 - Creating diagnostic 'safety nets' to close the loop on abnormal, critical test results. Although the CERC recommendations represent state-of-the-art, authoritative advice, we believe there are 3 other interventions that have comparable financial and clinical impact profiles: 1 - Finding and learning from diagnostic errors; 2 - Improving clinical reasoning; and 3 - Promoting patient engagement. Healthcare organizations have an obligation to begin improving diagnostic safety and quality, and both the CERC recommendations and our own represent excellent options that should be considered for immediate adoption.
Diagnostic errors are a substantial source of patient harm. As artificial intelligence (AI) integrates into clinical workflows, opportunities are emerging to assess their impacts on diagnostic excellence (DxEx). The Coordinating Center for Diagnostic Excellence (CODEX) at the University of California San Francisco established the Action Incubator to translate research advances in DxEx into tangible strategies for improving diagnosis. The September 2025 in-person inaugural Action Incubator convened 30 multidisciplinary stakeholders representing health systems, patient advocacy, industry, and policy groups. Through structured discussions and breakout sessions, participants identified AI scribes as a near-term, scalable use case for evaluating AI's impact on diagnosis not only because of their widespread adoption, but - as supported by cognitive load theory - because of their potential to reduce cognitive burden and allow clinicians to focus more on diagnosis. A modified Delphi process was used to prioritize candidate measures based on feasibility, and impact. Participants generated 17 candidate measures of AI scribe impact on DxEx. Consensus was reached on two priority metrics as functions of AI scribe usage rates by primary care physicians: (1) Timely follow-up of abnormal test results related to breast and colorectal cancer screening and (2) Patient-reported diagnostic experience. Participating health systems will pilot these measures using electronic health record audit logs and patient surveys. The first CODEX Action Incubator developed a pragmatic, consensus-driven framework for measuring the impact of AI on DxEx. Future annual Action Incubators will take up timely, actionable topics related to DxEx.
OBJECTIVES:We describe a case of diagnostic error arising from unquestioned acceptance of a longstanding, inherited diagnosis. CASE PRESENTATION:A patient developed CSF hypotension syndrome following minor trauma. A sacral cyst on imaging was attributed to a Tarlov cyst documented incidentally a decade earlier. Despite progressive symptoms, exhausted conservative therapy, and a radiological report explicitly questioning the diagnosis, the recommended post-myelography CT was twice deferred. The correct diagnosis - a meningomyelocele causing a traumatic CSF leak - was eventually established through external specialist consultation and treated surgically. CONCLUSIONS:Diagnostic momentum, a non-integrated radiological correction signal, and the low cognitive accessibility of rare alternative possibilities combined to delay recognition of the meningomyelocele. Clinicians should not consider prior diagnostic labels as invariably true or well established. Many of these were simply never fully verified or confirmed over time, and these 'inherited' diagnoses are a risk factor for diagnostic error.
INTRODUCTION:Identify knowledge gaps in applying artificial intelligence in clinical settings, using medical imaging as a primary use case to enhance diagnostic efficacy, efficiency, and patient and provider safety. METHODS:We convened a two-day workshop with 18 interdisciplinary experts from three countries. Experts represented quality and patient safety, human factors and systems engineering, radiology and other medical specialties, nursing, medical informatics, cognitive and perceptual psychology, psychometrics and statistics, and machine learning, drawn from academia, industry, health systems, and government. RESULTS:We identified by consensus six major knowledge-gap domains, with specific research questions for each domain: development, validation, integration and sustainability; redesign of existing healthcare systems; human and team augmentation; deployment of adaptive-learning "Foundation Models;" and balancing innovation, standardization, and regulatory oversight. CONCLUSIONS:We recommend employing a multidisciplinary collaborative approach in future research to leverage transformational AI capabilities anticipated in the next 5-7 years for each of the identified knowledge-gap domains, including ensuring that clinical AI supports diagnostic decision-making, integrates into clinical workflows, and mitigates risks related to automation bias, overreliance, fragmentation of care, and unintended consequences.
OBJECTIVES:Accurate diagnosis depends on concordance among clinical findings, histopathologic features, and patient context; when these elements diverge, even well-characterized diseases may evade timely recognition. CASE PRESENTATION:We report a case of Kaposi sarcoma (KS) in a 47-year-old woman who presented with a chronic, unilateral, non-violaceous eruption of the left lower extremity. Serial clinical evaluations and biopsies demonstrated dermal fibrosis, vascular congestion, and hemosiderin deposition without the classic histopathologic features typically associated with KS, leading to repeated interpretation as reactive or stasis-related disease. Despite progressive clinical findings, apparent concordance between benign-appearing clinical and histologic features delayed diagnostic reassessment. Ultimately, expert dermatopathology consultation with immunohistochemical staining confirmed KS in the setting of previously unrecognized HIV infection. CONCLUSIONS:This case illustrates how atypical disease biology, false reassurance from partial concordance, fragmented care, and social determinants of health can disrupt diagnostic synthesis and obscure evolving malignancy. When disease behavior contradicts initial interpretations, deliberate reassessment and longitudinal integration of data across encounters and specialties are essential to avoid delayed diagnosis, particularly in patients facing barriers to longitudinal continuity of care.
Diagnostic error, defined as missed, wrong, or delayed diagnoses or those not communicated to patients, is common, affecting 5-10 % of hospital admissions and clinic visits. Such errors cause patient harm in up to 1 in 100 of such encounters and account for 10 % of all hospital deaths and serious adverse events. About 80 % of diagnostic errors are potentially preventable, most resulting from flaws in clinician reasoning in formulating and testing diagnostic hypotheses. The advent of artificial intelligence (AI), and large language models (LLMs) in particular, has attracted great interest in how these technologies can reduce diagnostic error within the context of bedside or clinic consultations. This narrative review aims to provide practising clinicians with a comprehensible analysis of where AI and LLMs are currently positioned in assisting diagnostic performance in clinician-patient encounters based on contemporary state-of-the-art research. It attempts to answer seven questions relevant to clinician understanding and adoption of AI/LLMs. It concludes that AI tools have matured to the extent that they can improve diagnostic decision-making of clinicians and can assist institutions in increasing diagnostic safety. The rapid development of LLMs and ongoing release of new versions necessitate continuous monitoring of their evolving diagnostic capabilities. Importantly, a balanced approach is required where LLMs work to complement, rather than replace, the nuanced diagnostic reasoning of clinicians.
OBJECTIVES:The diagnostic value of upper endoscopy (EGD) in asymptomatic patients with iron deficiency anemia (IDA) remains uncertain, given the low prevalence of upper gastrointestinal (GI) malignancy. METHODS:We performed a retrospective cohort study comparing asymptomatic patients with laboratory-confirmed IDA to asymptomatic controls referred for pre-bariatric endoscopy. Between-group comparisons assessed demographic, procedural, and endoscopic features. Diagnostic yield was evaluated for clinically significant findings (CSFs) and histologically confirmed malignancy. Diagnostic efficiency was expressed as the number needed to investigate (NNTI). Predictors of CSF and malignancy were identified by multivariable logistic regression, and receiver operating characteristic (ROC) analysis determined optimal age cutoffs. RESULTS:Among 6,185 IDA patients and 2,010 controls, endoscopic pathology was significantly more frequent in the anemia group. Major CSFs were detected in 8.2 % of IDA patients vs. 2.9 % of controls (p<0.0001), while malignancy occurred in 1.3 vs. 0.1 % (p<0.0001). Diagnostic efficiency was superior in IDA (NNTI˜12 for CSF; <100 for malignancy) compared with controls (>30 and >1,000, respectively). In adjusted models, IDA and age ≥50 years independently predicted CSFs and malignancy, whereas female sex was inversely associated. ROC analysis identified optimal age thresholds of approximately 60 years for CSFs and 65 years for malignancy. CONCLUSIONS:Among asymptomatic patients with IDA, EGD demonstrated a substantially higher yield of CSFs and upper gastrointestinal malignancy than in asymptomatic controls, particularly among older individuals. These findings provide a large real-world benchmark for the isolated diagnostic yield of EGD and may inform future risk-stratified approaches to IDA evaluation.
BACKGROUND:Patient narratives are foundational to diagnosis, yet clinicians frequently and unintentionally distort, minimize, or reinterpret these narratives during clinical encounters. These distortions - termed patient narrative distortion - are unmeasured contributors to diagnostic error [G.D. Schiff, O. Hasan, S. Kim, R. Abrams, K. Cosby, B.L. Lambert et al., Diagnostic error in medicine: analysis of 583 physician-reported errors, Arch Intern Med 169 (2009) 1881-1887; P. Croskerry, The importance of cognitive errors in diagnosis and strategies to minimize them, Acad Med 78 (2003) 775-780; H. Singh, A.N.D. Meyer, E.J. Thomas, The frequency of diagnostic errors in outpatient care, BMJ Qual Saf 23 (2014) 727-731; and M.L. Graber, N. Franklin, R. Gordon, Diagnostic error in internal medicine, Arch Intern Med 165 (2005) 1493-1499], emotional harm [J. Conway, F. Federico, K. Stewart, M.J. Campbell, Respectful Management of Serious Clinical Adverse Events, IHI Innovation Series White Paper, IHI, Cambridge, MA, 2011], and inequity [E.N. Chapman, A. Kaatz, M. Carnes, Physicians and implicit bias: how it affects clinical decision making, Acad Med 88 (2013) 354-360; and M. Marmot, R.G. Wilkinson (Eds.), Social Determinants of Health, 2nd ed., Oxford University Press, Oxford, 2005]. No existing safety tool captures the fidelity with which clinicians preserve patient stories [T. Greenhalgh, B. Hurwitz (Eds.), Narrative Based Medicine: Dialogue and Discourse in Clinical Practice, BMJ Books, London, 1998; and W. Levinson, D.L. Roter, J.P. Mullooly, V.T. Dull, R.M. Frankel, Physician-patient communication: the relationship with malpractice claims, JAMA 277 (1997) 553-559]. The stages at which narrative distortion emerges are illustrated in Figure 1. The objective of this article was to define patient narrative distortion as a measurable construct, develop a five-domain taxonomy, propose a scoring system (PNDI), and outline a workflow and validation strategy for clinical use. METHODS:We conducted iterative conceptual modeling and structured synthesis of the diagnostic-safety and narrative-medicine literatures [G.D. Schiff, O. Hasan, S. Kim, R. Abrams, K. Cosby, B.L. Lambert et al., Diagnostic error in medicine: analysis of 583 physician-reported errors, Arch Intern Med 169 (2009) 1881-1887; P. Croskerry, The importance of cognitive errors in diagnosis and strategies to minimize them, Acad Med 78 (2003) 775-780; H. Singh, A.N.D. Meyer, E.J. Thomas, The frequency of diagnostic errors in outpatient care, BMJ Qual Saf 23 (2014) 727-731; and M.L. Graber, N. Franklin, R. Gordon, Diagnostic error in internal medicine, Arch Intern Med 165 (2005) 1493-1499] to identify core distortion modes, develop domain definitions, item-level anchors, and scoring thresholds. We propose a multi-phase validation plan including content validity, inter-rater reliability, construct validity, criterion validity, and responsiveness. RESULTS:The PNDI taxonomy includes five domains: Narrative Completeness Distortion, Meaning Substitution Distortion, Salience Distortion, Context Stripping Distortion, and Bias-Driven Distortion [E.N. Chapman, A. Kaatz, M. Carnes, Physicians and implicit bias: how it affects clinical decision making, Acad Med 88 (2013) 354-360]. Each domain includes 0-3 severity anchors and real-world clinical examples. The total PNDI score ranges from 0-15, with interpretation bands for narrative integrity and safety risk. Encounter-level, unit-level, and organizational-level workflows for implementation are outlined. The five-domain PNDI taxonomy is shown in Figure 2. CONCLUSIONS:PNDI operationalizes narrative integrity as a measurable dimension of diagnostic safety [The Joint Commission, National Patient Safety Goals: Improving Diagnosis in Health Care 2024-2026, The Joint Commission, Oakbrook Terrace, IL, 2024]. It provides clinicians, educators, and safety teams with a practical tool to detect narrative loss, reduce diagnostic error [G.D. Schiff, O. Hasan, S. Kim, R. Abrams, K. Cosby, B.L. Lambert et al., Diagnostic error in medicine: analysis of 583 physician-reported errors, Arch Intern Med 169 (2009) 1881-1887; P. Croskerry, The importance of cognitive errors in diagnosis and strategies to minimize them, Acad Med 78 (2003) 775-780; H. Singh, A.N.D. Meyer, E.J. Thomas, The frequency of diagnostic errors in outpatient care, BMJ Qual Saf 23 (2014) 727-731; and M.L. Graber, N. Franklin, R. Gordon, Diagnostic error in internal medicine, Arch Intern Med 165 (2005) 1493-1499], and strengthen patient trust [W. Levinson, D.L. Roter, J.P. Mullooly, V.T. Dull, R.M. Frankel, Physician-patient communication: the relationship with malpractice claims, JAMA 277 (1997) 553-559].
OBJECTIVES:Diagnostic errors are multifactorial and tend to occur frequently in the emergency department. Pediatric emergency care presents unique challenges, particularly in the evaluation of trauma in preverbal infants who cannot adequately describe their symptoms. Herein, we present the case of an infant in whom irritability following minor head trauma was incorrectly diagnosed. CASE PRESENTATION:A 1-year-6-month-old boy presented to the emergency department with irritability following minor head trauma. The initial evaluation focused on the head trauma, and the child's irritability was attributed to an upper respiratory tract infection and the effect of the unfamiliar examination environment. The patient was discharged, but several hours later, he returned because of right thigh swelling, which revealed a femoral fracture. Root cause analysis using a fishbone diagram identified the contributing factors to the delayed diagnosis as the vague, clinical presentation of pediatric trauma, multiple cognitive biases, and the role of family concerns in clinical reasoning. CONCLUSIONS:In preverbal children, behavioral changes, such as a sudden refusal to walk, may indicate a serious pathology. Clinicians should avoid anchoring, perform a thorough, whole-body assessment, and take the family's concerns into consideration to ensure diagnostic accuracy in cases of pediatric trauma.
OBJECTIVES:Accurate risk assessment is crucial for appropriate management of patients presenting to the emergency department (ED) with chest pain. This study aimed to evaluate how accurately emergency physicians estimate the probability of acute coronary syndrome (ACS). METHODS:A vignette-based on a real ED chest pain patient was presented in four sequential frames with increasing clinical information: history and symptoms, ECG findings, and elevated troponin levels. Evidence-based reference probabilities of ACS were established for each frame. Emergency physicians in the Skåne region of Sweden were invited to estimate the probability of ACS at each stage. The primary outcome was the difference between physicians' estimates and the reference probabilities. RESULTS:Of 707 respondents, 357 physicians with recent ED experience provided complete responses and were included in the analysis. The median age was 34 years (IQR 30-40), 56 % were male, and slightly more than half were specialists or senior residents. Physicians consistently overestimated the probability of ACS across all frames. The largest discrepancy occurred in the initial history/symptoms frame, where the median estimated probability was approximately 70 %, compared with a reference range of 3-16 %. Although estimation accuracy improved as additional diagnostic information was provided, estimates remained higher than reference probabilities throughout. CONCLUSIONS:In this realistic staged case, ED physicians substantially overestimated the risk of ACS in chest pain patients. Such overestimation may contribute to unnecessary diagnostic testing and overtreatment, highlighting the need for improved training and/or decision-support tools in emergency care.
OBJECTIVES:To analyze how diagnostic success was achieved in a challenging case of occult breast cancer using the SIDER protocol, a structured Safety-II framework for reflecting on successful diagnostic processes. CASE PRESENTATION:A 56-year-old woman developed progressive pain extending from the right upper limb to the right periscapular region 8 weeks before presentation, followed by axillary pain, cutaneous changes, and lymphadenopathy. Cervical radiography, cervical magnetic resonance imaging, breast ultrasonography, and computed tomography did not identify a primary breast lesion. At presentation, the clinical team retained breast malignancy in the differential diagnosis because of the patient's sex, progressive symptoms, axillary lymphadenopathy, and unexplained clinical trajectory. Diagnostic uncertainty was explicitly discussed with the patient, supporting continued follow-up. Two weeks later, positron emission tomography-computed tomography showed intense uptake from the right neck to the supraclavicular and axillary regions. Axillary lymph node biopsy demonstrated metastatic carcinoma with immunohistochemical findings consistent with breast origin, establishing the diagnosis of occult breast cancer. CONCLUSIONS:Applying the SIDER protocol showed that diagnostic progress depended on three reproducible processes: maintaining a low-probability but high-impact diagnosis despite nondiagnostic imaging, preserving patient engagement under uncertainty, and enabling diagnostic re-entry through cross-specialty collaboration. This case extends the use of Safety-II reflection by demonstrating how structured analysis of diagnostic success can generate practical lessons for complex, evolving presentations.
Generative artificial intelligence (gen AI) can support and likely enhance clinical reasoning coaching for struggling medical learners when its use is guided by educational theory. We present a conceptual framework, grounded in deliberate practice, self-regulated learning microanalysis, and psychological safety, for integrating gen AI into clinical reasoning coaching. Each theory was selected because it addresses a distinct challenge in coaching struggling learners, from the need for structured repetition with feedback to metacognitive monitoring and the stigma that impedes engagement. We describe how each theory can guide gen AI use in coaching and include suggested prompts.
OBJECTIVES:Diagnostic errors occur frequently and significantly affect patient prognosis and medical safety. Nurses, who are closest to patients, are often the first to detect abnormalities. However, owing to insufficient knowledge of medical diagnoses and diagnostic errors, they tend to hesitate in voicing concerns. This study investigated the hypothesis that clinical reasoning and diagnostic error education for nurses improve their knowledge of diagnostic errors and learning motivation. METHODS:A self-administered questionnaire survey was conducted before and after a lecture among nurses working in the outpatient department of Juntendo University Hospital. Eight domains related to diagnostic errors, including knowledge, confidence, motivation, feelings of guilt, and others, were assessed using a seven-point Likert scale. Pre- and post-lecture responses were compared using paired t-tests. Statistical significance was set at p<0.05. RESULTS:Valid responses were obtained from 35 of the 70 participants in this study (response rate, 50 %), and the mean years of experience among the participants was 15.0 ± 9.0 years. Knowledge of diagnostic errors significantly increased (Pre: 2.1 ± 1.1 to Post: 4.6 ± 1.3), and motivation to learn about diagnostic errors also significantly increased (Pre: 6.1 ± 1.1 to Post: 6.6 ± 0.7). CONCLUSIONS:The above findings suggest that clinical reasoning and diagnostic error education for nurses significantly improve their knowledge, confidence, and motivation to learn about diagnostic errors. Furthermore, diagnostic error education may influence nurses' perceptions and enhance their ability to participate in the diagnostic process.
Since establishment in 2020 of the global consensus-2 diagnostic criteria for mast cell activation syndrome (MCAS), recognition of this prevalent disease, existing alongside rare cutaneous or systemic mastocytosis, has grown significantly. Despite this progress, some have continued using more restrictive criteria, significantly underdiagnosing this complex but treatable disorder. Consensus-2 has brought diagnosis and effective treatment in many MCAS patients previously labeled with unexplained (mostly inflammatory) syndromes or misdiagnosed as somatization or other primary psychiatric disorders. Fears that consensus-2 diagnostic criteria might bring overdiagnosis have not been realized. Appropriate therapy for MCAS can dramatically improve quality of life, sometimes after decades of morbidity and disability despite extensive past unhelpful workups and treatment. Appreciating the breadth and heterogeneity of MCAS, and achieving good management outcomes, requires understanding the complexity of mast cell biology and pathophysiology as well as the range of comorbidities the disease can drive and which can aggravate the disease. The spectrum of diseases thought possibly rooted in variants of MCAS (or to which MCAS is significantly contributing) continues expanding, calling for more research. New therapeutic options have emerged, and clinicians have gained better appreciation for use of existing treatments and mitigation of impediments to treatment. MCAS research remains in early stages, hampered by many factors including limited awareness of the disease and challenges in objective assessment of treatment response. This review examines developments in awareness, education, clinical care, and research since consensus-2 emerged, while discussing challenges and opportunities for accelerating progress.