
Background The European Union Medical Device Regulation (EU MDR) requires manufacturers to establish quantitative, state of the art parameters for clinical evaluation. Because most available evidence is heterogeneous and uncontrolled, conventional meta-analytic methods provide limited support for deriving device level benchmarks. Methods A methodological framework was developed using synthetic datasets reflecting typical real-world evidence. Study means, variances and standard errors were combined using fixed effect, DerSimonian–Laird random effects, an Appraisal-Weighted (AW) approach, and a Random Appraisal-Weighted (RAW) method. AW incorporates structured methodological appraisal into random-effects weights, while RAW retains the appraisal-weighted point estimate but uses random-effects confidence intervals. Separate pooled estimates and 95% confidence intervals were generated for state of the art and device specific datasets and interpreted using predefined interpretive performance zones. Results Fixed effect pooling produced overly narrow intervals dominated by a single large, lower quality study. Random effects estimation widened intervals but did not address methodological quality. Appraisal-weighting improved point-estimate accuracy when appraisal was informative of bias but produced overly confident confidence intervals. The RAW approach preserved the accuracy gains of appraisal-weighting while maintaining appropriate interval calibration. Comparison of state of the art and device specific RAW intervals yielded intuitive classification of performance. Conclusions The RAW method offers a transparent, reproducible approach for transforming heterogeneous clinical evidence into quantitative benchmarks suitable for benefit–risk assessment under EU MDR. The framework aligns with regulatory expectations, accommodates evolving PMCF evidence, and supports structured, defensible clinical evaluation.
Objective To synthesize the self-reported prevalence of scientific misconduct in health-related research through a systematic review and meta-analysis, including a broader range of unethical practices than previous investigations. Methods A systematic search was conducted across six electronic databases and three sources of gray literature. Observational studies reporting the prevalence of scientific misconduct among researchers, students, or health professionals were included. The methodological quality was assessed using the Joanna Briggs Institute checklist, and a random-effects meta-analysis was performed using the Freeman-Tukey double arcsine transformation. Results Twenty-nine studies were included, totaling 17,919 analyzed responses. The most prevalent form of misconduct was honorary authorship (32%; 95% CI: 21%–44%), followed by plagiarism (28%; 95% CI: 12%–47%), data falsification (12%; 95% CI: 6%–19%), data fabrication (11%; 95% CI: 6%–19%), and methodological manipulation (4%; 95% CI: 2%–7%). High heterogeneity was observed for most outcomes (I 2 ≥ 95%), except for methodological manipulation (I 2 = 62%). No evidence of publication bias was found (p > 0.05). Conclusion Scientific misconduct remains a widespread issue in health research. The findings underscore the need for institutional policies that promote research integrity, transparent authorship criteria, and ethical training. Despite the potential for underreporting due to self-reported data, the results provide an important overview of the extent and types of misconduct present in the current scientific practice.
Background N-of-1 designs involve randomising an individual between treatments and recording outcomes over several timepoints. They are rarely used in mental health research. This research aims to explore the reasons behind this, identifying both opportunities and barriers to their implementation. Methods A cross-sectional qualitative study was conducted to explore healthcare professionals’ and trial app developers’ views on N-of-1 trials in mental health. Data were collected through semi-structured interviews and analysed using framework analysis, informed by the Theoretical Domains Framework. Results We conducted interviews with nine participants, including three psychiatrists, two mental health nurses, two other health professionals, and two people who were involved in developing apps for research in mental health to explore perspectives on N-of-1 trials in mental health. We generated three key themes: (1) N-of-1 trial designs are appropriate for very specific combinations of patient populations, interventions and settings in mental health. (2) If used appropriately, N-of-1 trials have the potential to inform treatment choice, improve outcomes and empower patients. (3) Improved knowledge and understanding of N-of-1 trials is needed among clinicians, patients and the wider research system to facilitate the use of this trial design in mental health. Conclusions Our findings suggest a potential role for N-of-1 trials in mental health, specifically for patients with stable chronic conditions and for interventions that can be safely discontinued and re-started. N-of-1 trials could inform treatment choices, improve outcomes, and empower patients. However, significant barriers to implementing this trial design exist within the current UK health and research system.
Introduction Platform trials have gained prominence, particularly during the COVID-19 pandemic, due to their flexibility and efficiency in evaluating multiple interventions. However, their complex design introduces methodological and regulatory challenges, especially in randomisation when adding new treatment arms. Method This study uses simulations based on a real platform trial to assess the performance of different randomisation methods: simple randomisation, stratified block randomisation, stratified block urn design, and minimisation. We evaluate these methods under various allocation ratios, focusing on covariate balance, allocation predictability, and allocation accuracy, particularly for newly added arms. We also examine the trade-off between covariate balance and allocation predictability. Results This study highlights the inherent trade-off between achieving good covariate balance and maintaining adequate allocation unpredictability. SBUD slightly outperforms SBR for certain stage 2 allocation ratios, though not consistently across all scenarios. Minimisation achieves the best covariate balance among all methods but at the cost of reduced randomness. Allocating more patients exclusively to the newly added arm does not effectively reduce covariate imbalance; however, allocating more patients to both the control and new arms yields better balance. Additionally, incorporating non-concurrent data in the minimisation process improves covariate balance compared to using only concurrent data. Conclusion The findings suggest that when the primary goal is to achieve covariate balance, especially for new arms, trialists should consider using minimisation with non-concurrent data and allocate patients proportionately to control and new arms. Balancing covariate distribution while maintaining sufficient randomness is crucial to ensuring the methodological robustness and reliability of platform trials.
Background Hierarchical Bayesian modelling using groupings of adverse events (AEs) into system organ classes (SOC) are a set of approaches that have been proposed for analysing safety signals in clinical trials. However AEs may be the expression of more than one clinical pathology and the classification of an AE into a single SOC may not always be clear. Further, medical dictionaries may assign AEs which are difficult to classify into a generic disorders SOC. When modelling AE data using SOCs, the misclassification of an AE may lead to either a potential safety signal being missed, or a safety signal being incorrectly flagged. Methods We investigate the use of mixed membership models as one approach to handling this issue. Conclusions Results indicate that this type of approach does have a real effect on model results, and the implications are discussed.
Background Silicone bands offer a non-invasive method for measuring environmental chemicals; however, their feasibility with young children is uncertain. This mixed-methods study examines toddler compliance with silicone bands, identifies predictors of compliance, and offers recommendations for research. Methods Children wore silicone wrist and ankle bands for 1 week; parents completed daily diaries capturing wear time of each band (hours) to indicate compliance (focal outcome). Mothers completed questionnaires capturing family characteristics and child cognitive and behavioral characteristics. Compliance and its associations with family and child characteristics were examined using descriptive, comparative, and associative statistics, and predictive modeling. Comments provided by parents in diaries and research assistants’ feedback were summarized to inform recommendations for future studies. Results Children ( n = 115, 28–39 months) tended to be highly compliant with wearing both bands (56%) or the ankle band only (23%), as reflected by a median of 11–12 hours/day of wear time over the week. Compliance was higher for ankle than wrist bands ( p -value <0.001). Wear time was not associated with child age, sex, associated clothing when wearing the band, or socioeconomic status ( p -values >0.05). Child language and anxiety were positively associated with wristband compliance ( p -values ≤0.038), whereas higher behavioral inhibition and lower effortful control were associated with higher ankle band compliance ( p -values ≤0.016). Qualitative data suggest improving band appeal and resolving sizing issues could improve compliance. Conclusions Results support the feasibility of deploying silicone bands with toddler-aged children. Studies using child-worn bands should account for hourly wear time, as compliance may predict both chemical exposure levels and child characteristics.
Pulling Back the Curtain on the Consensus Process: A case study of a modified-hybrid Delphi process. Gaining consensus among experts is vital for developing guidelines, standards, and curriculum. The Delphi process and Nominal Group Technique (NGT) are commonly employed, with hybrid models emerging. This paper explores group dynamics during a modified-hybrid Delphi process. This study explores the presence, interplay and impact of group dynamics of a modified-hybrid Delphi consensus building process by examining facilitators’ perceptions and observations during the in-person (NGT) phase with expert panelists. A qualitative single-case study design was employed, utilizing observation notes, audio recordings and post-event facilitator reflections. By applying both inductive and deductive coding, a deeper understanding of the rarely examined aspects of group dynamics in the consensus process was achieved. Analysis revealed two themes: “Gaining Momentum” and “Sharing Perspectives.” Under “Gaining Momentum,” three observations included the development of shared understanding, perspective-taking, and a decreasing time to consensus. “Sharing Perspectives” highlighted concerns about hierarchy, the absence of dominant individuals, and the accommodation of diverse perspectives without succumbing to Groupthink. Key insights included thorough preparation to manage potential group dynamics during the NGT including facilitator training and a participant package to foster shared understanding for process and content. Despite minor issues considering hierarchies, the in-person (NGT) phase supported consensus building while effectively mitigating negative group dynamic issues. This study expands our understanding of group dynamics in consensus building during an NGT phase of a modified-hybrid Delphi process, offering valuable insights and practical recommendations preparing for and implementing successful in-person sessions.
Research into disability is often complex and simultaneously involves multiple aspects of patients, health professionals, carers and the situations within which patients live and care is received. Qualitative research into healthcare in general, and more specifically disability, must likewise incorporate the multiple important components of the care situation. In this paper I present the declarative mapping sentence method for conducting qualitative research with a particular emphasis upon its use in undertaking studies into disability and rehabilitation. The declarative mapping sentence is a sentence in ordinary English language that incorporates the important features (facets) of a domain of research interest along with sub-divisions of these facets (called elements). The facets are joined together using ordinary language (connective ontology) with great care so as to suggest the real life relationships between facets. In this paper I explore the components of the declarative mapping sentence in some detail and offer illustrative examples of a declarative mapping sentence developed for use in disability research. I draw attention to and discuss the advantages of using this method in disability and other health related research and suggest that the approach offers a framework within which to understand patients and healthcare professionals within context.
Providing accessible information during informed consent practice helps people make important decisions about research participation. The aim of this review is to critically appraise the current Australian clinical research guidelines to assess their guidance on delivering accessible informed consent processes for people with vision impairment. An archival search of the Overton database was conducted for policy documents and guidelines pertaining to informed consent practice in Australia. This was cross checked with the Australian Research Council’s online catalogue of national codes and guidelines. Content analysis was performed using an implementation science framework to code the actors, actions, context, target and time (AACTT) elements of informed consent practice. The AACTT framework enables a detailed specification of the behaviours targeted for championing or change. Ten policy documents specific to conducting consent procedures for clinical research or healthcare participation in Australia were identified and included in this review. Five of the 10 guidelines provided sufficient detail for researchers to know what actions to perform to enable inclusive consent practice and mutual engagement with people with vision impairment. However, they varied in their content describing how and when to perform them. Two of the key reference documents for clinical research conduct in Australia entirely lacked recommendations for accessible informed consent practice. Australian research policy and practice documents offer some direction in helping researchers to deliver accessible informed consent practice. However, guidelines produced by peak disability bodies and patient advocacy groups offer more direction on how to engage people with vision impairment and offer useful templates for producing accessible print materials.
Understanding long-term patient outcomes (PROs) following surgery requires an efficacious survey methodology. Leveraging a statewide hernia surgery registry to establish a sampling frame, we conducted a 1-year post-operative survey using measures of patient-reported hernia recurrence (the Ventral Hernia Recurrence Inventory), pain (the PROMIS Pain Intensity 3a), and quality of life (the HerQLes scale). Our responsive design approach varied invitation and reminder contact modes and incentive offer across multiple design phases, with the goal of minimizing non-response bias and maximizing cost effectiveness. Outcomes included: contact and response rates (%); item non-response (%); and the association between reminders and incentive offer with response rates, respondent characteristics, and item non-response (%). Differences in demographic and clinical characteristics of respondents and non-respondents were investigated and adjusted using registry data. Of 7062 patients who received hernia surgery between January 2020 and March 2022, 6068 were sampled, 5645 were contacted (contact rate 93.0%), and 1816 responded to the survey (overall response rate 29.9%). Response rates by cohort were 42.3%, 32.5%, 25.2%, and 25.9%, with overall low item non-response. Response rates increased with number of reminders, but with diminishing returns over time; offer of postpaid incentive over no incentive did not significantly improve response rates or influence item non-response. Weighted respondents were comparable to the survey population. We illustrate a strategy to maximize response rate amongst surgical patients and evaluate the representativeness of long-term PROs using a sample-based registry, targeted multi-mode contact methods, and weighting adjustment methods.
Established analysis methods for composite outcomes can be complex to interpret, clinically and statistically. To counter this, composite outcomes are frequently reduced to a binary outcome. This can simplify interpretation but there exists a danger that deeper understanding of patient experience is lost. Rank-based DOOR methodology is emerging as an attractive option, using all components of a composite outcome. DOOR methodology has been employed retrospectively to the TOPPIC trial (ISRCTN89489788) to determine if making fuller use of data could enhance study results. TOPPIC patients were assigned a clinical outcome rank (1 = most desirable to 8 = least desirable). DOOR methodology was applied, testing the null hypothesis of a DOOR probability of 50% (no difference). Additionally, smokers and non-smokers were considered separately. Win ratio and ordinal logistic regression provided supportive analyses. Employing DOOR methodology demonstrated the distribution of ranks differed between treatment groups - DOOR probability of a more desirable outcome with active treatment of 54.6% (95% CI 53.8% - 55.4%, p < 0.0001). Considering smokers and non-smokers separately, this difference was amplified in smokers (DOOR probability 66.1% (95% CI 62.7% - 69.4%, p < 0.0001). Win ratio methodology and ordinal logistic regression showed no difference between treatments. DOOR methodology was applied successfully to a previously published trial and shows potential to discriminate clinical outcomes effectively by deconstructing a binary endpoint of component parts into an ordinal outcome. This approach should be considered when designing trials, either as principal analysis of a rank-based outcome or as a confirmatory sensitivity analysis.
Randomised Controlled Trials (RCTs) are the most robust method to test new or existing interventions. Recruitment and retention are crucial aspects to their success but are often challenging. Methodological changes to RCTs are often tested by Studies Within A Trial (SWATs). To date there is mixed evidence supporting the use of prenotification newsletters to boost retention. This SWAT aimed to assess this intervention within the PROFHER-2 trial. A two-arm parallel group SWAT embedded at the 24-months follow-up for PROFHER-2 – a RCT evaluating treatment methods for a 3- or 4- part proximal humerus fractures in patients over 65 years. Participants were randomised (1:1) to either receive a newsletter 2-to-4 weeks prior to their 24-months questionnaire, or not. The results have been combined in a meta-analysis with existing evidence. There was no evidence of a difference in retention (OR 2.17, 95% CI 0.21-22.02, p = 0.51). Similarly, there was no statistically significant difference in the completion of the primary outcome of the returned questionnaires (OR 0.72, 95% CI 0.04–12.70, p = 0.82), nor the proximity of completion (HR 1.23, 95% CI 0.74 – 2.05, p = 0.43). This result is similar to that seen in two of the four previous evaluations of this SWAT. Sending a newsletter as a prenotification to a questionnaire in a RCT does not improve the retention rates. However, the sample size for this piece of work is small, and the retention rates were high (94%). When combined with the existing evidence, there is no evidence of an effect.
Infant mortality rates (IMR) and other health disparities (e.g., low provider access; higher obesity rates) exist across rural America, especially in the southeastern United States. These disparities coincide with negative socioeconomic factors such as poverty and low household income. Georgia’s 2020 preterm birth rate was seventh highest among the 50 U.S. states, its rate of low birth weight (LBW) babies was fourth among states. This study used a novel, combined spatiotemporal analysis to assess IMR and health trends in Georgia counties ( n = 159). We spatially regressed IMR on 14 biopsychosocial health variables and assessed IMR 2012–2022 longitudinal trends, rural versus urban or by Health Professional Shortage Area (HPSA) status, using a two-level hierarchical linear model. Analyses demonstrated significant associations between IMR and rural counties as well as counties that have high African American populations, high unemployment, and high uninsurance rates. Principal Care Provider Rates showed a negative spatial regression relationship. There was a −0.056 IMR slope for urban counties from 2015 to 2022, whereas the rural county IMR slope was 0.135. Findings concur with previous research suggesting the need for coordinated county-specific socioeconomic development and corresponding increased health programs. Combined geospatial and multilevel models represent a novel approach to inform public health epidemiology.
To evaluate whether machine learning (ML) methods (Elastic Net (EN), eXtreme Gradient Boosting (XGBoost), Feed Forward Neural Net (FNN)) can improve claims-based inpatient quality measurement by Logistic Regression. This retrospective cohort study used German claims data from the years 2015-2021. The study population encompassed inpatient cases of acute myocardial infarction ( n = 165,130) and proximal humerus fracture ( n = 34,912), for which quality related outcomes were assessed. The performances of risk adjustment models based on machine learning methods (EN, XGBoost, FNN) were compared to stepwise backwards Logistic Regression by Receiver Operating Characteristics-Area under the Curve (ROC-AUC), Precision Recall-Area under the Curve (PR-AUC), Brier Score (BS). The institution-specific quality was measured by Standardised Mortality Ratios (SMR) which were used to visualise the impact of the tested methods on quality assessment. For most of the outcomes none or only marginal gains were found for the machine learning methods. Highest gain in model performance showed the FNN in comparison to Logistic Regression with a gain in ROC-AUC of 2.4%, in PR-AUC of 4.5%, and slightly in the BS with a loss of 0.007. The FNN was followed by XGBoost with a gain in ROC-AUC of 2.3%, anyhow this improvement was not reflected in a lower BS. None of the machine learning methods tested is generally superior for creating quality indicators. Marginal gain in model performance should not be the main basis for choosing an adequate method; instead, interpretability should be emphasised, especially when dealing with new datasets with little knowledge of important risk factors.
Models of positionality tend to focus on the researcher’s paradigm and how they interpret the world around them, but less emphasis has been placed on the concept of time as a methodological lens. The following commentary examines what it takes to embrace such an approach—making use of three components of positionality (identity, power and context) as a framework for examining qualitative inquiry. It is proposed that one of the conditions for validity in qualitative inquiry is for the researcher to embrace positionality as a moment in time. This approach will help qualitative health researchers make purposive methodological choices so that they can find an approach that is ‘right for them’ rather than ‘right per se’. Further, a model that outlines the components of methodological self-consciousness is offered, to support the qualitative health researcher in deciding the approach that is ‘right for now’.
Background STROKE OWL is a quasi-experimental study using claims data from statutory health insurances in Germany to calculate the effect of case managers for stroke. Since there is no recruited control group, a suitable procedure needed to be identified to make the intervention effect measurable. Hence, the objective of this paper is to present an approach for comparing matching procedures before final data analyses take place. Methods We followed a four-step approach to identify an appropriate procedure on a partial dataset. First, we conducted a systematic review for identifying potential confounders of the study’s outcome. Afterwards we checked whether a matching procedure was able to balance the dataset with respect to the outcome, under the assumption that all relevant covariates including the intervention variable were balanced. Within the two last steps we checked covariate balances and remaining group sizes. Three matching procedures – coarsened exact matching, optimal full matching and propensity score matching – were tested. Results The coarsened exact matching was able to balance variables perfectly but on average lost >50% of the observations. Although optimal full and propensity score matching revealed some weaknesses concerning the variable balance, considerably more observations remained in the dataset. Based on the described approach and the external framework conditions of STROKE OWL the optimal full matching was chosen as matching procedure most suitable. Conclusion In summary, it was challenging to identify a suitable matching procedure for this study, since the detailed results of the different balance and group size checks varied.
Background There is no consistent approach in assessing swallowing function both in clinical or research settings. However, its significance in ensuring proper nutrition, growth/ development, and overall quality of life is undeniable. We aim to introduce a novel methodology that enhances swallowing function as an outcome measure in clinical and rare disease research and demonstrate its efficacy in Niemann-Pick Type C1 (NPC1). Methods We reviewed commonly implemented qualitative and quantitative swallowing assessments in current clinical practice, including patient/proxy/clinician reports, clinical swallowing evaluation, videofluoroscopic swallow assessment (VFSS), and post-VFSS interpretive measures [American Speech Hearing Association National Outcomes Measures Scale (ASHA-NOMS) and National Institutes of Health penetration and aspiration scale (NIH-PAS)]. Data analysis involved descriptive statistics, longitudinal statistical modeling to account for NPC1-specific covariates, and Kappa weighting correlations to determine inter-rater reliability. Results Our NPC1 cohort ( n = 120) underwent baseline and longitudinal evaluations, ( n = 269 VFSSs). We identified three statistically significant ( p < 0.05) NPC1-specific variables associated with post-VFSS interpretive measures, providing insights into disease progression and treatment effects. Inter-rater reliability correlations for ASHA-NOMS and NIH-PAS demonstrated strong agreement between two speech pathologists [ASHA-NOMS: 0.76 (95% CI 0.70-0.83); NIH-PAS: 0.88 (95% CI 0.82- 0.93)]. VFSS mean radiation dosage was calculated ( n = 129) and fell within the acceptable range declared by the American College of Radiology (263.79 ± 147.44 cGy*cm2). Conclusions This comprehensive methodology successfully documented functional swallowing status with dietary modifications and aspiration risk in a rare disease cohort (NPC1). Additionally, we effectively employed this methodology to support swallowing function as a research endpoint, specifically for phenotyping and developing therapeutic interventions for rare diseases.
Background Members of the public differ regarding their views on the use and sharing of personal health information. Objective This paper describes the methodology employed to develop a telephone-based survey tool to capture views of the public on the collection, use and sharing of personal health information. Method A rigorous methodology comprising multiple stages was undertaken to develop a vignette/scenario-based survey instrument. These steps included a review of instruments used in other jurisdictions, focus groups, engagement meetings with healthcare professionals, cognitive testing and piloting the final instrument. Informed by the findings of each survey development phase, draft scenarios and accompanying questions were developed. Results The following scenarios were developed: ‘Circle of care,’ ‘Use of information beyond your direct care’ and ‘Digital records.’ Conclusion The findings from this survey will inform national policy in relation to health information and will inform the development and implementation of eHealth initiatives. In turn, this should support the delivery of high-quality, effective health and social care. The learnings from the development of this survey will contribute to future health information policy and governance in countries or jurisdictions considering the development of a national electronic health record system. Moreover, this research will support public and population health management by encouraging public engagement to support successful implementation of new health information systems.