
BACKGROUND:The increasing use of real-world data in economic evaluation raises concerns about the potential for confounding bias. Methods to control for this bias have been developed, but there is still much to be understood about how confounding influences economic evaluations. OBJECTIVES:To illustrate the impact of unadjusted confounding variables in economic evaluations of observational studies. METHODS:We simulated the costs and effectiveness of 2 treatments across 9 possible confounding effect scenarios. We considered these scenarios in the context in which one treatment is more costly and more effective with mild correlation between confounders and outcomes. All scenarios examined the incremental cost and effectiveness of 400 randomly generated individuals, reflecting sample sizes commonly seen within observational economic evaluations. Results were illustrated with the use of cost-effectiveness planes and cost-effectiveness acceptability curves (CEAC). RESULTS:Our simulations illustrate that confounding bias can have a significant effect on incremental costs and incremental effectiveness estimates. These simulations also illustrate that confounders can affect the evaluation of uncertainty by causing a shift in the CEACs. Such results hint that inadequate consideration of confounding bias can potentially lead to flawed judgments about the cost-effectiveness of a treatment. DISCUSSION:Results of economic evaluations can be influenced by confounding when they are conducted in an observational setting. Efforts must be made to limit their impact to ensure an accurate assessment of the economic value of treatments and to prevent potential losses of population health.
BACKGROUND:Length and quality of life are frequently combined in health technology assessments to derive a single, generic measure of health improvement due to a treatment or medicine. One such measure is given by quality-adjusted life-years (QALYs), typically calculated assuming discrete health states and quality-of-life values. The accurate estimation of QALYs is crucial for informed decision making in health care policy and resource allocation; however, traditional methods often rely on assumptions that are sometimes biologically and statistically inappropriate. This study aims to develop a framework for estimating QALYs in populations where survival data are collected that addresses these limitations. METHODS:A framework for estimating QALYs in continuous time is introduced, based on joint longitudinal-survival models fitted using maximum likelihood; the proposed framework requires patient-level data, including longitudinal health utility values and overall survival. In contrast to the conventional approach, which involves dichotomising health states and separate models, this method allows the estimation of QALYs from a single model while accounting for all the statistical and biological intricacies of the data, providing a more appropriate estimate for cost-effectiveness modelling. The joint modelling approach is validated using Monte Carlo simulation under realistic data-generating mechanisms. RESULTS:Simulations showed that the joint longitudinal-survival modelling approach could recover the true QALYs without bias. Conversely, a comparison method based on calculating QALYs directly from the health utility trajectories, ignoring the survival process, was biased under most scenarios. CONCLUSIONS:A new methodology for estimating QALYs within a unified, flexible and extensible framework has been developed. This approach can improve QALY estimation and their use in cost-effectiveness analyses in practice, enabling more timely and robust health technology assessments. User-friendly Stata software is provided.
BACKGROUND:Best-worst scaling (BWS) has been increasingly used as a values clarification method to help patients make value-concordant decisions. However, evidence on its use remains limited. OBJECTIVE:To examine 1) the association between BWS-derived treatment preference scores and patients' stated and real-world treatment choices and 2) whether BWS-identified best-match treatments were concordant with treatment choices among older adults with end-stage kidney disease (ESKD). METHODS:We conducted a prospective study among patients aged ≥70 y with incident ESKD in Singapore at the time of decision making. During a renal counseling session with a decision aid, participants completed 2 BWS exercises to clarify values related to treatment choices (dialysis vs kidney supportive care [KSC] and hemodialysis vs peritoneal dialysis). BWS responses were scored by assigning preference weights to treatment-related attributes and aggregating them into continuous treatment preference scores for each treatment option. Logistic regression assessed associations between dialysis preference scores and postcounseling preference for dialysis and dialysis initiation at 6 mo. Concordance was assessed by comparing the BWS-derived best-match treatment, postcounseling preferred treatment, and real-world treatment choice at 6 mo using the Stuart-Maxwell test. RESULTS:Twenty-three patients were enrolled (mean age 77.9 ± 4.7 y; 65% male). Higher BWS scores favoring dialysis were associated with greater odds of preferring dialysis after counseling (odds ratio = 1.31; P = 0.018) and initiating dialysis at 6 mo (odds ratio = 1.41; P = 0.012). For the dialysis and vs decision, the Stuart-Maxwell test indicated no significant differences; 70% selected (P = 0.70) and 65% initiated (P = 0.26) a treatment concordant with the BWS-derived best-match option. CONCLUSIONS:This study provides preliminary evidence that BWS-based values clarification is associated with treatment preferences and subsequent real-world treatment initiation among older adults with ESKD.
BACKGROUND:Population-based screening programs are widely promoted as a cornerstone of preventive health care, yet their structural features may subtly influence individual decision making. In the present research, we investigated whether informed decision making may be challenged by population-based screening programs if these are perceived as establishing a default of participation. We hypothesized that screening intentions are higher if the screening is part of a population-based program and that this effect is mediated by perceived default. METHODS:In a preregistered online experiment, 997 German participants aged 18 to 75 years (M = 47.93 y, SD = 15.62 y) were randomly assigned to one of three conditions. Participants read a scenario describing a hypothetical cancer screening offered either (a) as part of a population-based screening program, (b) upon request, or (c) within a research project. The main dependent variable was intention to screen. RESULTS:Screening intentions were significantly higher when the screening was included in a population-based program (M = 4.83) compared with a screening that was available upon request (M = 4.42, P = 0.013, d = 0.22, 95% confidence interval [CI] [0.07, 0.37]) or part of a research project (M = 4.44, P = 0.021, d = 0.21, 95% CI [0.06, 0.36]). A mediation analysis revealed that the status of screening indirectly influenced the intention to participate through its effect on the perceived default of participation. CONCLUSIONS:These findings indicate that program structures alone may shape decisions in ways that challenge the aim of informed choice, particularly for screenings with contested benefit-harm profiles. To ensure that population-based programs support rather than undermine informed decision making, communication and program design must promote transparency and enable active, preference-sensitive choices.
BACKGROUND:Postmastectomy breast reconstruction (PMBR) is a preference-sensitive decision that requires alignment between clinical evidence and patient values. Shared decision making (SDM) supports patient-centered care, yet its systematic implementation in PMBR remains inconsistent. This mixed-methods systematic review synthesizes barriers and facilitators to SDM adoption and identifies strategies to improve decision quality, patient experience, and equity. METHODS:A systematic review was preregistered (OSF: https://osf.io/mu9hv) and conducted in accordance with PRISMA 2020 guidelines. Six databases (PubMed, Embase, Scopus, Web of Science, Cochrane Library, Trip Database) were searched from inception to November 2025. Eligible studies included qualitative, quantitative, or mixed-methods research examining barriers and/or facilitators to SDM implementation in PMBR across micro (individual), meso (organizational), and macro (system) levels. Data were synthesized using a convergent integrated approach and mapped to the Consolidated Framework for Implementation Research. RESULTS:Thirty-one studies (n = 15-485 participants) from North America, Europe, and Asia were included. SDM interventions-particularly decision aids, digital tools, preconsultation education, and culturally tailored strategies-improved patient knowledge (+6%-32%), decisional clarity, and satisfaction. Decisional conflict decreased by 13 to 25 points, and consultation time was reduced by up to 41%. Facilitators were identified across micro (eg, clinician engagement), meso (eg, multidisciplinary collaboration and organizational support), and macro (eg, supportive implementation strategies), whereas barriers included limited SDM literacy, workflow limitations and insufficient documentation, and structural inequities, respectively. Despite demonstrated effectiveness, sustained implementation was limited. CONCLUSIONS:SDM enhances informed and value-concordant decisions in PMBR. Effective implementation requires multilevel strategies, including clinician training, workflow integration, culturally adapted interventions, and organizational support. Future research should prioritize scalable and equitable models to ensure consistent integration of SDM into routine practice.
BACKGROUND:Mental health intervention studies often lack preference-based measures required to estimate quality-adjusted life-years (QALYs) for health economic evaluations. In such circumstances, a mapping study is the second-best alternative to estimate QALYs. This study aimed to develop a mapping algorithm to predict EQ-5D-5L value sets from the Health of the Nation Outcome Scales (HoNOS) among adults with severe mental illness (SMI). METHODS:Trial data from a community mental health intervention conducted in Germany between 2020 and 2023 were assessed over 24 mo. The data included more than 900 adults aged 18 to 82 y living with SMI. Four econometric approaches were used to map HoNOS to EQ-5D-5L: ordinary least square (OLS), generalised linear regression model (GLM), fractional regression model (FRM), and an adjusted limited dependent variable mixture model (ALDVMM). The German EQ-5D-5L value set was used, while keeping the Norwegian value set in the appendix. Model performance was assessed using the root mean squared error (RMSE), mean absolute error (MAE), and the squared correlation between observed and predicted EQ-5D-5L (r2). A 10-fold cross-validation was applied to validate our mapping algorithms. RESULTS:The FRM consistently performed best across all evaluation criteria in both the full sample and cross-validation. In cross-validation, when HoNOS items were used as predictors of the German value set, the FRM yielded RMSE = 0.2204, MAE = 0.1622, and r² = 36.5%. Similar results were observed when HoNOS subscales were used, with slightly higher RMSE (0.2245) and MAE (0.1664) and a lower r² (34.2%). Calibration plots indicated good model fit, with predictions closely aligning with the reference line representing equality between the observed and predicted values. Scatter plots further supported the superior performance of the FRM. A similar pattern was observed for the Norwegian value set. CONCLUSIONS:The preferred mapping algorithm enables the estimation of EQ-5D-5L utilities from HoNOS data, facilitating the calculation of QALYs when only HoNOS information is available. As our sample includes a broad spectrum of mental health patients, further validation within specific diagnostic groups is recommended to improve predictive accuracy.
INTRODUCTION:Decision aids aim to improve the quality of and satisfaction with decision making. Few scales exist that directly evaluate decision aids from the patient's perspective, particularly with respect to information overload. METHODS:We developed a Decision Aid Evaluation Scale with domains assessing acceptability, satisfaction, cognitive load, and helpfulness for decision making. Cognitive interviews and iterative testing were performed. Convergent and discriminant construct validity were assessed by comparing the novel scale with domains from the validated Decisional Conflict Scale (DCS; values clarity, uncertainty, feeling informed subscales). Internal consistency was measured using Cronbach's alpha. The final scale consisted of Likert-type scale items assessing information amount, values clarity, ease of use, cognitive effort required, nervousness, and overall helpfulness. Participants (any sex or gender, aged 40-60 y) completed the scale after viewing a cancer screening decision aid within a hypothetical screening decision scenario. RESULTS:We surveyed 1,249 participants; 778 completed the Decision Aid Evaluation Scale. As hypothesized, a larger proportion of participants who reported low cognitive load had higher DCS values clarity subscale scores than those with high cognitive load (P = 0.003). Participants who found the decision aid helpful had significantly higher DCS values clarity (P < 0.001), informed (P = 0.004), and certainty subscale scores (P < 0.001). Scores on the acceptability questions were higher among those who found the decision aid helpful (P < 0.001), except for questions related to "length," "relatability of photo," and "photo helped make decision." Participants reporting nervousness had lower DCS certainty subscale scores (P < 0.001). Cronbach's alpha for the acceptability and satisfaction domains were 0.74 and 0.75, respectively, indicating internal consistency. CONCLUSIONS:We developed a patient-centered scale that captures users' perceptions of decision aid usability and information overload. The scale demonstrated acceptable construct validity and internal consistency, supporting its use in evaluating decision aids from the patient perspective.
Purpose: To examine the prevalence and patterns of mental health patients’ consideration and engagement in self-initiated psychiatric medication discontinuation, identify their motivations, and assess associated factors including fear of discontinuation, trust in psychiatrists, perceived support, and clinical diagnosis. Design: We conducted a multicenter cross-sectional study including 1,574 adult mental health patients from 12 psychiatric clinics and general hospital psychiatric departments across Serbia, Croatia, Bosnia and Herzegovina, and Montenegro. Participants were recruited from inpatient, outpatient, and day-hospital settings. Participants completed a questionnaire on psychiatric medication use, discontinuation experiences, and trust in psychiatrists. Analyses focused on 773 participants who had considered discontinuing medication, examining prevalence, motivations, and related factors. Results: Among 1,574 participants, 773 (49.1%) had considered discontinuation of psychiatric medication; of these, 361 (46.7%) had actually discontinued, with 273 (75.6%) doing so independently—149 (54.6%) abruptly and 122 (44.7%) gradually. Consideration was most frequent in alcohol use (62.4%) and anxiety disorders (59.9%) and the least in organic disorders (25.0%). Overall, 400 (51.7%) participants discussed discontinuation with their psychiatrist; perceived support was highest in alcohol use (65.9%) and anxiety (58.2%) and lowest in psychotic and bipolar disorders (32.0%). Primary motivations were perceived recovery and regaining autonomy. Fear was generally mild to moderate, and trust was high, with trust and motivations not differing meaningfully between independent and psychiatrist-guided discontinuation. Conclusions: Self-initiated medication changes were highly prevalent and occurred amid a pronounced communication gap, with deprescribing discussions largely patient initiated. Motivations were perceived recovery and desire to regain autonomy rather than treatment dissatisfaction. High trust in psychiatrists provides an opportunity to offer individualized deprescribing pathways through proactive shared decision making, reducing the need for patients to discontinue medication independently.
Threshold models, formalized by Pauker and Kassirer, guide clinical decisions about testing and treatment by partitioning pretest probabilities into three action zones using two threshold pretest probabilities: the test threshold and the test-treatment threshold. This derivation implicitly assumes that diagnostic tests possess fixed sensitivity and specificity, effectively treating them as binary predictors. However, many diagnostic tests yield continuous or ordinal outputs, in which case the optimal sensitivity-specificity pair varies with the pretest probability. I demonstrate that abstracting from this dependency leads to an under-testing bias: the testing threshold is inflated and/or the test-treatment threshold is deflated, resulting in a narrower than optimal testing window. This bias systematically undervalues continuous diagnostic tests and leads to their underuse. Clinical decision-making models and guidelines should therefore recognize that optimal test-score cut-offs depend on pretest probabilities to avoid this systematic underuse of diagnostic testing.
OBJECTIVES:To systematically evaluate how survival estimates from network meta-analyses (NMA) models affect oncology cost-effectiveness analysis (CEA) results, particularly under different proportional hazards (PH) assumptions and extrapolation strategies. METHODS:A total of 19 time-to-event NMA models were evaluated in this analysis, encompassing Cox proportional hazards (Cox-PH), fractional polynomial (FP), Royston-Parmar (RP) models, piecewise exponential (PWE), generalized gamma (Gengamma), and parametric survival models (PSM), with 2 extrapolation strategies for non-PH models: a constant-tail hazard ratio (Fixed-HR) and a parametric extrapolation approach (Varying-HR). NMA model performance was assessed via relative error (RE) by comparing CEA outcomes derived from NMA-based survival estimates with those from head-to-head clinical data. RESULTS:Fifteen CEA models were conducted. The selection of the NMA model exerted a greater influence on CEA outcomes than traditional model parameters (ie, costs or utilities). RP models appeared to perform best, while PWE, second-order FP, and Gengamma models tended to perform comparatively worse. Fixed-HR extrapolation approaches generally yielded more accurate long-term projections than Varying-HR methods did, particularly under non-PH conditions. Violations of the PH assumption substantially increased uncertainty, especially for non-PH models. Notably, the choice of NMA model led to reversals in cost-effectiveness conclusions in up to one-third of the evaluated CEA cases. CONCLUSIONS:NMA model choice substantially influences oncology CEA outcomes. Underutilized but well-performing RP models show strong potential for broader application. Fixed-HR extrapolation was preferred in our study. These findings are context specific, and further validation is needed. When PH assumption is violated, the selection of non-PH models requires greater caution. Our findings may offer insights to support future refinements in CEA methodological guidance.
BACKGROUND:Physicians often fail to communicate important details about key risks and benefits of treatment in prostate cancer consultations, an information gap that may inhibit effective patient participation in care decisions. We developed a framework to assess detail and personalization of physician communication of key concepts, derived from observations of treatment consultations. METHODS:We recorded and transcribed treatment consultations of 50 men with newly diagnosed prostate cancer across 10 multidisciplinary providers. Using open coding, analysts identified statements related to guideline-endorsed SDM concepts and categorized the level of detail for each topic. Concept-specific hierarchies indicating increasing detail were integrated into a unified framework. We reported the highest observed level for each concept across consultations. RESULTS:Across 19,388 statements, sentences addressing cancer severity (1,016; 5.2%), baseline function (186; 1.0%), oncologic endpoints (343; 1.8%), cancer prognosis (318; 1.6%), life expectancy (111; 0.57%), and treatment-related side effects (952; 4.9%) were extracted. Empirically derived frameworks for evaluating the level of detail and personalization for each concept consistently displayed parallel levels of detail. We created a unified framework, with each level providing more precise, patient-specific risk communication: (0) omission, (1) no quantification, (2) binary assessment, (3) imprecise quantification, (4) specific quantification, and (5) patient-specific estimate. The prevalence of the highest observed level of detail and personalization conveyed varied by concept, ranging from 4% to 84%, 0% to 47%, 0% to 42%, 2% to 8%, 0% to 48%, and 0% to 66% for levels 0 to 5, respectively. CONCLUSIONS:Our framework introduces an empirically derived, structured foundation for characterizing the level of detail and personalization of physician risk communication of key concepts in prostate cancer consultations.
BACKGROUND:The vaginal microbiota test predicts the success of in vitro fertilization (IVF), but with no therapy available to improve a low profile, couples must decide whether to proceed or postpone treatment. We aim to examine how couples interpret vaginal microbiome results and make postponement decisions within a shared decision making (SDM) framework. METHODS:Women undergoing IVF or IVF-intracytoplasmic sperm injection (IVF-ICSI) treatment at 2 Dutch hospitals received the ReceptIVFity test™, which classified the vaginal microbiome as high (52.6% chance of conception), medium (23.6%), or low (5.9%) profile based on predicted implantation success after a fresh embryo transfer. Physicians discussed the results with couples using SDM, after which the couples decided whether to proceed or postpone treatment. The primary outcome was the patients' perceived involvement in shared decision making, assessed with the SDM-Q-9 questionnaire. The secondary outcome was the proportion of couples postponing treatment after a low microbiome profile. RESULTS:Between October 2018 and November 2020, 728 women were enrolled. SDM-Q-9 responses showed high perceived involvement overall but lower scores for "exploring options," reflecting limited alternatives when the choice is to proceed or postpone treatment. A low profile was found in 35.4% (258/728). After the SDM consultation, 49.6% (128/258) chose to postpone treatment, with postponement rates increasing to over 80% among couples in later IVF cycles. Decisions were influenced by personal, emotional, and practical considerations, including the Dutch insurance reimbursement system (3 insured IVF or IVF-ICSI cycles regardless of postponement) and the absence of effective treatment to modify a low profile. CONCLUSIONS:These findings demonstrate that couples can understand and use prognostic information when supported by SDM and that the ReceptIVFity test™ facilitated discussion about chances of success, timing of treatment, decisions to proceed or postpone, and personal values.
BACKGROUND:Researchers widely use discrete choice experiments (DCEs) to assess health preferences across subgroups. However, variations in decision consistency, rather than true differences in preferences, can drive observed utility differences. Despite the growing use of DCEs to assess health preference heterogeneity, recent studies highlight a persistent lack of methodological transparency in accounting for unobserved heterogeneity, underscoring the need for technically robust approaches to support credible and actionable comparisons across groups. This study improves health preference research methods by directly addressing scale heterogeneity and reducing bias when comparing subgroups. METHODS:A simulated DCE evaluated hypothetical cancer treatments across 2 imagined groups (patients, caregivers). Each task presented 3 alternatives (including a status quo), varying in months gained, survival rate, side-effect severity, and out-of-pocket cost. Mixed logit models were estimated. Scale heterogeneity was addressed using the Swait-Louviere 2-step procedure. Willingness to pay (WTP) was computed and compared across groups via the Poe et al. (2005) simulation-based test. RESULTS:The Swait-Louviere test confirmed significant scale heterogeneity (P < 0.05) but no meaningful taste differences (P > 0.10). Once scale effects were accounted for, the analysis revealed a shared preference structure across patients and caregivers, with variability driven by inconsistent decision making rather than true preference divergence. Consistent with this, none of the between-group WTP differences were statistically significant, reinforcing the absence of meaningful subgroup contrasts and underscoring the importance of separating scale from taste to avoid biased inference. CONCLUSIONS:Adjusting for scale heterogeneity strengthens DCE validity by reducing bias from decision noise and enabling accurate subgroup comparisons. Using simulated data, this study applied the Swait-Louviere 2-step and scale-invariant WTP contrasts to separate taste from scale; both methods converged, showing that heterogeneity reflected scale rather than true preference differences, with negligible WTP gaps. Routine scale diagnostics, taste (preference) tests under equalized scale, and welfare space reporting are recommended to ensure valid inference. However, as this study used simulated data with no real respondents, its findings are illustrative only and not intended for real-world inference; generalizability and external drivers of scale heterogeneity were not assessed.Key HighlightsThe study enhances methodological rigor by explicitly addressing scale heterogeneity-an often-overlooked bias that improves the validity and real-world relevance of preference-based insights.Applying the Swait-Louviere test and willingness to pay, whenever possible, enables researchers to distinguish true preference differences from response inconsistency across choice datasets.The findings advocate for the routine inclusion of scale diagnostics in stated-preference research to strengthen health decision making and modeling practice.
BACKGROUND:Alternative diagnostic labels for melanoma in situ may better reflect its lower risk (15-y survival of 98%) compared with invasive melanoma (10-y survival ranging from 98% for American Joint Committee on Cancer stage IA to 19% for stage IV). DESIGN:Secondary analysis of an online randomized experiment in Australian adults without melanoma. Participants were randomized to a hypothetical diagnosis of "melanoma in situ (MIS)" (control), "low-risk melanocytic neoplasm," or "low-risk melanocytic neoplasm, in situ" and completed a survey. OUTCOMES:Perceived risk measures were future invasive melanoma and mortality risk (0%-100%), comparative risk, affective risk, and vulnerability (7-point Likert scales). Calculated risk measures were lifetime invasive melanoma risk (from participants' risk factors) and melanoma mortality probability (Australian sex-/age-specific mortality rates). ANALYSIS:An intention-to-treat analysis across randomized groups was performed, unadjusted and adjusted for covariates (linear regression models). RESULTS:In total, 1,668 adults were recruited. Compared with MIS, perceived melanoma mortality risk was lower for low-risk melanocytic neoplasm (-10.4%, 95% confidence interval [CI]: -13.1% to -7.63%, P < 0.001) and for low-risk melanocytic neoplasm, in situ (-7.4%, 95% CI: -10.2% to -4.6%, P < 0.001). Similar patterns were observed for perceived risk of invasive melanoma; comparative, affective risk; and vulnerability. Participants in all groups substantially overestimated their lifetime risk of invasive melanoma (by 48.7%) and of dying from melanoma (by 32.0%) compared with the calculated risk; overestimation was lower in alternative label groups. CONCLUSIONS:Diagnostic labels without the word "melanoma" reduced risk overestimation, supporting MIS relabeling to mitigate overdiagnosis harm by reflecting its largely indolent nature. TRIAL REGISTRATION:ANZCTR: 386943HighlightsAlternative diagnostic labels for melanoma in situ that do not include the word "melanoma" significantly decreased perceived risk compared with melanoma in situ.Participants substantially overestimated their risk; alternative labels reduced this overestimation of perceived risk compared with calculated risk.A new label for melanoma in situ may better communicate the lower risk of adverse outcomes for this lesion compared with invasive melanoma. This may reduce patient anxiety and allow for management decisions that align with their values and preferences.
Background. Simulation calibration is the process of configuring a simulation model's parameters to improve the agreement between the model output and the desired calibration targets (e.g., observed historical data). For most realistic simulation models, this calibration process can be quite computationally expensive, as it requires running the simulation model for each parameter combination. To alleviate this problem, metamodels offer a tradeoff between accuracy and computational efficiency for extensive simulative analysis with a highly complex parameter space. Method. In this study, we examine 4 simulation calibration approaches. Randomly-Simulate (RS) is a simulation-based benchmark widely used in the literature. Optimally-Predict (OP) is an optimization-based approach that we adapt to the simulation calibration setting. Building on these 2 baselines, we introduce 2 hybrid strategies, Predict-then-Simulate (PtS) and Simulate-then-Predict (StP), which combine simulation runs and metamodel-based optimization in complementary orders. We compare all 4 methods in terms of calibration accuracy and computational cost. Results. While the metamodel-based OP approach substantially reduced computational cost relative to RS and identified parameter combinations near the optimal configuration, the hybrid strategies delivered superior calibration performance. In particular, the PtS approach, which combines metamodel-based optimization with targeted simulation refinement, achieved on average a 46% reduction in total actual error compared with the RS benchmark, while maintaining computational efficiency. Conclusions. The study introduces a novel metamodel-based optimization approach to simulation calibration and illustrates its potential benefits for computationally expensive studies. While developed for deterministic targets, the method provides a foundation for future extensions to settings involving stochastic simulation outputs and other forms of model uncertainty. An open-access Python implementation of the proposed framework is provided to facilitate adoption and reproducibility.HighlightsA metamodel-based optimization approach is proposed for calibrating simulation model parameters, which can offer computational advantages particularly in cases in which direct calibration is expensive due to complex or high-dimensional simulation models.A hybrid approach that narrows down the search space via metamodel-based optimization and then uses simulation runs for fine-tuning offers a sweet spot in the tradeoff between computational efficiency and accuracy.
IntroductionThe role of shared decision making (SDM) has become increasingly pivotal, particularly in nuanced choices such as those involving implantable cardioverter-defibrillator (ICD) therapy. This study evaluates the impact of the Dutch ICD Decision Aid on SDM in patients up for ICD implantation or replacement.MethodsA stepped-wedge randomized controlled trial was conducted across 6 Dutch hospitals between February 2018 and September 2019, involving patients eligible for ICD implantation or pulse-generator exchange. SDM experiences of the patients and involved medical professionals were assessed using SDM-Q-9 and SDM-Q-Doc questionnaires, respectively. The Decisional Conflict Scale (DCS) scores measured effective decision making. The intervention group received the decision aid on top of standard care.ResultsA total of 150 patients and 233 health care providers were included in the study. For health care providers, SDM scores did not differ: the SDM-Q-Doc median score was 36 (28-38) in the control phase and 35 (33-40) in the intervention phase (P = 0.81). Patients in both the intervention and control groups demonstrated high SDM scores as well. Decisional conflict scores were low: the median DCS score was 12.5 (4.3-23.4) in the intervention phase and 16.4 (6.25-25.0) in the control phase (P = 0.45). Patients with a higher education provided more correct answers to the theoretical knowledge questions. In addition, patients up for a pulse-generator exchange also had significantly more correct answers.ConclusionsAlthough the Dutch ICD Decision Aid did not result in significant differences in SDM scores or levels of decisional conflict between patient groups, both measures remained consistently favorable overall. The decision aid still holds promise as a valuable resource. Efforts should focus on refining decision-making tools and improving patient knowledge and the quality of patient-centered care.HighlightsA digital decision aid did not significantly increase shared decision-making (SDM) scores for patients and health care providers, as SDM levels were already high across all groups.Despite high SDM scores, patient knowledge about implantable cardioverter-defibrillator (ICD) therapy remained low, highlighting a gap in understanding.Patients with higher education or prior experience with ICDs demonstrated better knowledge retention, indicating the need for tailored educational interventions.The study emphasizes the ongoing challenge of ensuring unbiased, well-informed decision making in ICD therapy, especially during pulse-generator replacements.
ObjectiveTo evaluate the clinical and cost-effectiveness of SHARE TO CARE (S2C), a complex intervention for hospital-wide, systematic implementation of shared decision making.MethodsWe analyzed clinical effectiveness, health care resource utilization, and implementation costs of S2C from the statutory health insurance perspective using a quasi-experimental difference-in-differences approach with evidence from the Department of Neurology. Clinical outcomes included inpatient hospital admissions, emergency department admissions, and rates of standard and advanced imaging procedures. Implementation costs comprised those related to the conception, development, process integration, ongoing support, and auditing of S2C. Health care utilization data covered inpatient and outpatient care, pharmaceuticals, therapeutic services, assistive devices, and nursing care. We conducted sensitivity analyses to account for uncertainties.FindingsS2C was associated with a reduction in inpatient hospital admissions, emergency department admissions, and imaging rates in the intervention group. The cost analyses aligned with these findings, showing reduced total costs and health care resource utilization in the intervention group. Although none of the estimates reached the predefined thresholds for statistical significance, the primary analysis yielded weak evidence (P < 0.1) of a reduction in emergency department admissions in the intervention group. Overall, savings outweighed the costs of implementing S2C, suggesting cost-effectiveness.ConclusionsS2C has the potential to reduce emergency department admissions and overall health care costs from the statutory health insurance perspective. Further research should investigate generalizability, the timing of the treatment effect, and potential biases introduced by the COVID-19 pandemic. The demonstrated effects of shared decision making (SDM) have encouraged statutory health insurances in Germany to offer additional reimbursement for clinics certified under the S2C program. The S2C model illustrates how payers and providers can collaborate to facilitate the nationwide implementation of SDM.HighlightsThe implementation of SHARE TO CARE (S2C) was associated with a statistically nonsignificant reduction in emergency department admissions after 1 y from the statutory health insurance perspective, based on data from the Department of Neurology.The cost savings from reduced health care utilization outweighed the implementation costs, and despite not reaching statistical significance, the results support the potential cost-effectiveness of S2C.S2C has the potential for nationwide implementation as a systematic form of shared decision making.Future research should investigate the generalizability of the results to other health care settings.
ObjectivesThis study qualitatively explored how bolt-ons affect the perception of EQ-5D-5L core dimensions in a valuation context.MethodsSixty Indonesian adults (aged 20-67 y, 50% female) each valued 10 health states using composite time tradeoff (cTTO). States were presented in either forward (EQ-5D-5L, then with 1 bolt-on, then 2) or backward (reversed sequence) order. Participants were assigned to 1 of 3 bolt-on dyads: vision and tiredness, cognition and social relationships, or skin irritation and self-confidence. In semistructured qualitative interviews, respondents described how adding or removing bolt-ons changed the perceived importance of the 5 core dimensions. We classified these changes as related either to measurement or to valuation, with the latter further categorized as a relative or absolute shift in importance.ResultsCognition and vision generated the most shifts in perceived importance in the forward and backward groups, respectively. Regardless of ordering group, most shifts occurred between the EQ-5D-5L alone and the version with 1 bolt-on (first presented in the dyad), with significantly fewer shifts observed between the 1-bolt-on and 2-bolt-on states. In the forward group, most shifts were classified as measurement (56%) or relative preference (29%), while the reverse was true in the backward group: relative preference (53%) or measurement (31%). Absolute preference was least common across both groups.ConclusionsThis is the first study to explore how individuals reason when valuing EQ-5D-5L+bolt-on health states. Our findings suggest that interactions between dimensions are complex and may be influenced by presentation order. Further qualitative research should directly investigate absolute preferential reasoning.HighlightsNo studies have qualitatively explored how individuals value EQ-5D health states with bolt-ons. Understanding how bolt-ons influence reasoning and interact with core dimensions is crucial for informing valuation methods and modeling strategies.Our findings show that bolt-ons can alter how participants perceive the importance of EQ-5D-5L dimensions, although changes are mostly not preference driven. Participants often rely on accessible reasoning, such as conceptual associations between dimensions. Effects vary by the bolt-on used and presentation order.Interactions between bolt-ons and core dimensions complicate efforts to develop robust valuation approaches. Future qualitative studies should aim to capture preference-based reasoning, while quantitative work is needed to disentangle preferential from nonpreferential effects.