
Prediction with right-censored survival outcomes is challenging because incomplete event-time observation degrades estimation prediction, and uncertainty quantification. Cox proportional-hazards and parametric accelerated-failure-time (AFT) models deteriorate when covariate effects are nonlinear, and risk profiles heterogeneous and deterministic Buckley–James imputation under-represents censoring uncertainty in downstream predictions. We propose Bayes–BJ–LGBM, a partially Bayesian (Bayesian–MAP hybrid), censoring-aware framework that recasts Buckley–James latent-time reconstruction as posterior sampling on the log time scale, coupling a regularized LightGBM ensemble to a nonlinear AFT structure. Latent event times, the residual variance and the linear coefficients are drawn by Gibbs sampling, whereas the tree ensemble is optimized by maximum a posteriori (MAP) updates rather than sampled; the reported posterior predictive survival functions and 95
Cellular senescence is a fundamental mechanism of biological ageing that has emerged as a critical target for therapeutic intervention in age related diseases. The coalesce of artificial intelligence and senescence research provides unprecedented opportunities in advancing our knowledge and treatment approaches. This systematic review study addresses the gap across diverse AI model and the heterogeneity in senescence, by conducting the extensive evaluation of the performance outcomes and methodological rigor of AI models using Prediction model Risk of Bias Assessment Tool + Artificial Intelligence (PROBAST + AI) and reporting completeness using Transparent Reporting of a multivariable prediction model for Individual Prognosis or Diagnosis + Artificial Intelligence extension (TRIPOD + AI), providing the insights into AI models robustness and generalizability. For the articles, released between 2019 to 2025, across major databases 18 eligible studies was used for review, following PRISMA guideline. Quality and applicability were assessed by PROBAST + AI (4 domain) and reporting via TRIPOD + AI (27 items). The quantitative synthesis indicates that deep learning architectures, especially Convolutional Neural Networks (CNNs), are dominant which appeared in about 50
Analyzing massive survival data from distributed sources, such as multicenter registries, is hindered by computational bottlenecks and privacy constraints that preclude centralized pooling. This paper addresses these obstacles for the additive hazards model by proposing a distributed inference framework that integrates optimal subsampling with divide-and-conquer aggregation. Each site draws a local subsample and transmits only low-dimensional summary statistics to a central server, which constructs a global estimator. Consistency and asymptotic normality of the resulting estimator are established under some mild regularity conditions. Extensive simulation studies show stable finite-sample performance across a range of covariate distributions, censoring rates, and levels of heterogeneity between sites. Finally, the practical utility of our methodology is illustrated through an application demonstrating the computational gains of the proposed method.
Although Youden's index J has been a widely used summary measure of medical diagnostic accuracy, it suffers from a paradoxical property: the value of J can be very small even if the observed proportional accuracy (proportion of correct test results) is near perfect. The J paradox is defined and discussed, explaining the need for an alternative index. Exploratory analysis of alternative index formulations produced one particular index that overcomes the J paradox and has all the desirable properties expected of a diagnostic accuracy index. An important basis for such an index was the requirement that the index should account for the fact that some component of the diagnostic accuracy could be due to chance results. The analysis was based on analytical and statistical reasoning as well as empirical data, including real medical diagnostic data from various reported studies. The new diagnostic accuracy index A resulting from the analysis is found to have desirable properties, avoids the J paradox, and takes on values that are entirely reasonable. The outlined statistical inference procedure can be used to construct approximate confidence intervals. The index A is also generalized to a weighted equivalent A_w for which different weights or emphasis can be placed on false negative test results and false positive results. The indices A and A_w can also be easily extended to potential situations involving multiple diagnostic categories. The new index A (and its weighted form A_w ) is proposed as a preferred alternative to Youden's J. While J can in some realistic real situations produce results that would seem to be entirely unreasonable, unjustifiable, and misleading, A (and A_w ) appears to take on values that all all realistic and plausible.
Missing data in patient-reported outcome measures (PROMs) can affect estimation accuracy and inferential validity, particularly under missing-at-random (MAR) mechanisms. Multiple imputation (MI) is commonly recommended, but the relative performance of item-level and score-level imputation for longitudinal PROM composite scores remains insufficiently understood. We conducted a simulation study based on longitudinal clinical trial PROM data. Predictive mean matching (PMM) and random forest (RF) imputation were evaluated at both the item and score level under empirically calibrated MAR mechanisms. Simulation scenarios varied sample size, overall missingness, and the proportion of unit nonresponse. Performance was assessed using root mean squared error (RMSE), bias magnitude, relative efficiency compared with complete case analysis (CCA), conditional confidence interval coverage, and variance ratio estimates. RMSE decreased with increasing sample size and increased with higher proportions of missingness across all methods. Both PMM and RF generally outperformed CCA in terms of RMSE and relative efficiency, particularly under higher missingness. Differences between item-level and score-level imputation were generally modest, although score-level imputation tended to yield slightly lower RMSE and bias magnitudes under more severe missingness conditions. Conditional coverage remained close to nominal levels across most MI settings. Variance tended to be overestimated in some scenarios, although sensitivity analyses with larger numbers of imputations substantially reduced this effect. PMM showed comparatively stable performance across simulation settings. Multiple imputation methods generally outperformed complete case analysis for handling longitudinal PROM data under MAR. While score-level imputation sometimes showed slightly more favorable performance than item-level imputation when targeting composite score means, the magnitude of these differences was frequently small in practical terms. PMM provided stable performance across a wide range of settings and represents a reasonable default approach for many applied PROM analyses.
The use of surrogate endpoints can improve feasibility of clinical trials. The results of trial-level analyses are a key factor affecting regulatory policy regarding the uptake of a surrogate endpoint. Trial-level analyses aim to quantify the strength of association between treatment effects on an established clinical endpoint and treatment effects on the surrogate. Unfortunately, there is a well-documented lack of standardization in the meta-regression models used for these analyses. Common models differ in how they account for estimation errors, leading to variation in surrogacy estimands across approaches. This has caused confusion regarding surrogate quality. Moreover, the most used modeling approaches can lead to pessimistic inferences of surrogate quality. We overview common meta-regression models and their corresponding estimands for evaluating trial-level surrogacy, focusing on how each modeling approach accounts for sampling errors. Two broad classes of models can be differentiated. The models in the first class quantify the association between observed, estimated treatment effects, based on unweighted (least squares) or weighted linear regression (weighted least squares with weights proportional to trial sample size). The second class consists of hierarchical meta-regression models, which quantify the association between latent, true treatment effects. We target the trial-level coefficient of determination (R2) in our inferences. We use patient-level meta-analysis of 66 previously conducted chronic kidney disease clinical trials and a small statistical simulation to characterize differences in results between modeling approaches. Across our analyses, use of simple and weighted linear regression produced R2 estimates which were lower than those produced by the hierarchical models. In simulation analyses, use of simple and weighted linear regression resulted in downward bias in R2 when the estimand is defined as the R2 representing the trial-level association of true treatment effects on the clinical and surrogate endpoints. Commonly used methods for evaluating surrogate endpoints can produce unduly pessimistic conclusions of surrogate quality depending on the target of inference. This is because these methods partially or completely ignore estimation error in the analysis. Hierarchical models can be used to overcome such limitations.
Patient-oriented and community-based participatory research emphasize partnership, shared power, and the meaningful inclusion of patient and community knowledge in health research. Group Concept Mapping (GCM) is a structured mixed-methods methodology that combines qualitative idea generation with quantitative analysis and visual mapping. Although GCM is participatory in design, its patient-centric value depends on how ethically the method is implemented, who is involved, and how findings are interpreted and used. This paper critically examines GCM as a methodological tool for integrating patient voices in participatory health research. We describe the six phases of GCM, including preparation, idea generation, structuring, representation, interpretation, and utilization, and analyze how each phase can support or constrain patient-centric, community-based participatory research (CBPR), and patient-oriented research (POR) principles. We also explain ethical and practical considerations related to recruitment, facilitation, accessibility, statement reduction, representation, digital participation, participant burden, and utilization. GCM can support collective sense-making, shared problem definition, prioritization, and action-oriented planning by engaging patients, community members, researchers, clinicians, and decision-makers in a structured process. Its visual maps can make complex patient- and community-generated knowledge accessible for intervention development, program planning, evaluation, and policy discussion. However, GCM is not inherently patient-centric. Its value depends on intentional design, meaningful representation of lived experience, culturally safe facilitation, transparent analytic decisions, shared interpretation, and accountable use of findings. GCM offers a promising methodological pathway for integrating patient and community voices into participatory health research, but it should not be treated as automatically aligned with CBPR or POR principles. Its capacity to advance patient-centric research depends on the ethical and epistemological commitments guiding its implementation, including shared power, reflexivity, inclusion, and follow-through. When implemented with these conditions, GCM can generate culturally grounded, socially just, and actionable insights for healthcare planning, intervention development, and policy.
Reporting studies limitations is essential for transparent interpretation. While reporting guidelines such as CONSORT, STROBE, PRISMA mandate disclosure of limitations, no widely adopted framework standardizes how these are reported in a structured and comparable format. In practice, limitations sections often extend over several paragraphs reiterating predictable design-related constraints. This limits comparability, contributes to inefficiency and may overshadow genuine study-specific considerations. We propose a semi-structured checklist combining standardized limitation domains with a concise narrative component. This approach aims to ensure minimum systematic disclosure while preserving critical interpretation. Structured reporting may improve clarity, comparability, and enable meta-research on limitations across studies. Consensus-based refinement and pilot implementation are warranted to evaluate feasibility and impact.
Propensity scores are an important tool for using observational data to answer causal questions. Machine learning methods for estimating propensity scores have out-performed more commonly used logistic regression methods in studies considering large, synthetic datasets. To inform propensity estimation methods for smaller datasets from real-world observational studies, we describe the implementation of machine learning algorithms using data from a prison-recruited cohort. This study describes a procedure to use logistic regression, gradient-boosting, random forest and single-hidden-layer neural networks to balance confounders between a treatment and control group. We assess balance using the average standardised absolute mean difference (ASAM), where a lower ASAM indicates better balance. We provide an application to an observational cohort of Australian incarcerated men for a causal question for the effect of emotional support in prison on emergency department presentations within 100 days post prison release. There were 328 participants, of whom 231 (70
Dynamic treatment regimes (DTRs) are statistical methods that use patient-specific information to estimate an optimal sequence of treatments, with the end goal of maximizing a desired outcome of interest. DTRs have many useful applications to chronic disease management when patient characteristics are monitored over multiple periods of follow up to determine whether to continue a course of treatment or adopt an alternate strategy. Applications of DTRs to interference networks, or settings where an individual’s outcome can be affected by the treatments received by other individuals in the network, have received relatively little attention in the literature. Previous work has established DTR estimation for couples in households and in networks where general forms of interference are taking place, but the extent to which interference needs to be accounted for to avoid biased estimates of optimal treatment rules remains unknown. In this paper, we demonstrate how one such DTR method, dynamic weighted ordinary least squares regression (dWOLS), can be modified to account for both within- and between-group interference in nested hierarchical networks, and we compare our implementation with other dWOLS approaches with mis-specified exposure mappings. Our findings show that both the overall rate of treatment and the proportion of individuals experiencing interference from assigned neighbours can affect the performance of dWOLS regardless of the specified exposure mapping. Although standard dWOLS may be acceptable in some settings when the overall rate of treatment and proportion of egos assigned alters is expected to be low, our results demonstrate the improved performance of a dWOLS implementation that accounts for both within- and between-group interference, regardless of the overall rate of treatment or proportion of egos assigned alters.
Inverse probability of treatment weighting (IPTW) targets entire study population and is frequently used to adjust for confounding. When exposures are continuous, such as nutrient intake, we need to deal with the problem of large weights caused by a lack of positivity. Generalized overlap weighting (GOW) and generalized matching weighting (GMW) target the populations whose generalized propensity score distributions most overlap among the compared groups. Considering that these two weighting methods mitigate the influence of a lack of positivity, they may estimate exposure effects with less bias and higher precision. The primary and secondary objectives were to compare the performance of IPTW, GOW and GMW in evaluating the associations between nutrient intake and diabetes complications and to determine the favorable number of bins based on performance metrics when the quantile binning approach was applied to nutrient intake. We reanalyzed a dataset from a nutritional epidemiologic cohort study, Japan Diabetes Complications Study (JDCS), including 1,414 patients with type 2 diabetes. The generalized propensity score was estimated using quantile binning approach and used for each weighting method. We assessed covariate balance, weight variability, and the precision of the estimated risk ratios for complications among patients with type 2 diabetes. The results suggested that relatively small number of bins, 2 or 5 bins, might be favorable in quantile binning approach when considering covariate balance and weight variability, making quantile binning approach feasible for applied research. The performance of three methods was similar, with sufficient overlap in the generalized propensity score distributions. However, with reduced overlap, GOW and GMW showed a better covariate balance with 5 bins, and GOW had the smallest weight variability. Our findings suggested that with reduced overlap, GOW might be useful to adjust for confounding in nutritional epidemiology.
The increasing volume of scientific manuscripts has strained traditional peer-review systems, prompting interest in artificial intelligence (AI) as a potential solution. However, the comparative performance of large language models (LLMs) in critically appraising oncological research remains unexplored. We conducted a cross-sectional study evaluating the peer-review capabilities of three LLMs—Claude Sonnet 4.5, Google Gemini 3.0 Thinking, and OpenAI ChatGPT 5.1—using 50 blinded landmark oncology trials identified from ASCO, ESMO, and NCCN guidelines. Manuscripts were anonymized by removing identifying information and submitted to each model with a standardized prompt. Two investigators independently analyzed reviews for final verdicts, trial awareness, journal suitability assessments, and concerns across five domains (biological, methodological, statistical, trial groups, and reporting). Chi-square tests with Bonferroni correction and logistic regression were used for statistical analysis. Significant differences were observed in reviewer verdicts among models (p < 0.001). Claude Sonnet 4.5 was most stringent (74
There has been increasing interest in randomized trials on the discovery and identification of treatment moderators on an outcome. This reflects the stratified medicine paradigm, which moves beyond average treatment effects to focus on patient subgroups with heterogeneous responses. Baseline biomarkers and clinical variables are commonly evaluated as candidate moderators and standard subgroup analyses are typically used to identify single moderators. Recent research has sought to improve the detection of effect modification by optimally combining multiple moderators into a single composite measure. This composite moderator can offer greater power and sensitivity to detect effect modification over any single moderator alone and can be used in medical applications to derive personalized treatment recommendations (PTRs). Kraemer applied this concept to continuous outcomes, and we extended this approach to time-to-event settings using the expected win time against reference (EWTR), an estimand that measures time spent in better or worse clinical states. We compared the composite moderator with single moderators and explored its utility by evaluating whether there were improved treatment recommendations in simulations that varied (1) effect modification scenarios (quantitative vs qualitative), (2) interaction-to-main-effect ratio magnitudes, and (3) configurations of the moderator correlations with the endpoint. We illustrated this approach using randomized trial data investigating two management strategies for participants with atrial fibrillation. Simulations demonstrated that in the presence of strong effect modification, the composite moderator matched or outperformed the best single moderator. However, its advantage was attenuated by weak interactions, correlated moderators, or inclusion of a null endpoint. We extended Kraemer’s composite-moderator framework to time-to-event outcomes using EWTR. The composite moderator may serve as a pragmatic tool for detecting effect modification and informing personalized treatment recommendations, particularly when effect modification is moderate to strong.
Randomized controlled trials are the gold standard for evaluating interventions. For individually randomized trials, clustering can exist in the intervention arm but not the control arm when the intervention is delivered by study personnels and/or in a group setting. This results in a partially nested design, which methodological investigations have so far been limited to fixed trial settings. We examine whether existing frameworks for group sequential design and sample size reestimation (SSR) based on estimated variance parameters can be directly applied to partially nested trials, where a heteroscedastic mixed effect model is implemented. We propose computing stopping boundaries in the standard way and inflating the required sample size by the group sequential design factor. We consider two recruitment strategies that lead to the same amount of data at the interim analysis of a partially nested design. For SSR, we assume half the initially planned number of clusters are enrolled at the start of the trial. We propose updating the required number of clusters — given the same cluster size as initially planned — using variance estimates obtained at the interim analysis. We conduct proof-of-concept simulation studies to evaluate the operating characteristics of the two types of adaptive designs, respectively. Most of the computed group sequential partially nested designs meet power requirement under both recruitment strategies considered. However, enrolling all clusters at the start — resulting in only data from half the cluster size being available at the interim analysis — produces a greater likelihood of stopping early for efficacy compared with recruiting only half the number of clusters initially (which provides data of full cluster size at interim). For SSR, the updated number of clusters is on average close to the required number for achieving the target power, although its median is slightly lower. Both designs lead to Type I error rates that are close to the nominal values in most cases. With careful planning informed by clinical trial simulations, the proposed group sequential and SSR partially nested designs can satisfy the error-rate control requirements, noting that conventional estimation methods may produce biased treatment effect estimates and confidence intervals.
Rare outcomes are common in clinical research and often result in severe class imbalance. Under such conditions, traditional logistic regression tends to favor the majority class and shows limited sensitivity to clinically important rare events. Although complex machine learning methods may improve prediction, their limited interpretability can restrict clinical use. Therefore, an effective and interpretable modeling strategy for highly imbalanced clinical data remains needed. We developed a cost-sensitive logistic regression (CS-LR) model by incorporating different misclassification costs for different classification outcomes into the logistic loss function. The proposed framework was designed to prioritize rare but clinically important outcomes while retaining odds ratio interpretation. Simulation studies were conducted across four covariate scenarios, imbalance ratios from 1
Lifestyle factors, such as smoking, alcohol consumption, and physical activity, are among the most critical modifiable risk factors for morbidity and mortality; however, their representation in health data remains inconsistent and unstructured. Although electronic medical record (EMR) systems are widely adopted in Korea, interoperability remains limited due to heterogeneous formats and unstandardized content. This study aimed to establish a standardized framework for lifestyle data aligned with international standards. Items were collected from national initiatives, tertiary hospital questionnaires, and clinical reports. Overlapping items were prioritized and supplemented with clinically meaningful additions to develop a draft framework. The draft was refined through focus group interviews (FGIs) with healthcare professionals and hospital information system experts to ensure clinical relevance and technical feasibility. Values of the finalized framework were standardized by mapping to internationally recognized terminologies in collaboration with the Systematized Nomenclature of Clinical Terms (SNOMED CT) Korean National Release Center (NRC) at the Korea Health Information Service. The final framework comprised three classes: smoking class (5 elements, 15 values), alcohol consumption class (4 elements, 12 values), and physical activity class (7 elements, 41 values). All components were mapped to internationally recognized terminologies, such as SNOMED CT and Logical Observation Identifiers Names and Codes (LOINC). This proposed standardized lifestyle data framework harmonizes heterogeneous data sources and provides a foundation for integration into Korea’s interoperability roadmap. By enabling structured and interoperable representation of lifestyle factors, it may support patient-centered care, facilitate secondary data use, and strengthen epidemiological and health informatics research.
Meta-analyses of time-to-event (TTE) outcomes, particularly in Health Technology Assessment (HTA), commonly use a hazard ratio (HR) scale. However, non-proportional hazards in included trials create difficulties. Existing methods are either too complex or not easily incorporated into economic decision models for cost-effectiveness assessment. An alternative approach assumes a treatment-log(time) interaction within a Cox proportional hazards model, allowing the log HR to vary linearly with log(time). A bivariate meta-analysis of the resulting treatment and interaction coefficients then yields an overall time-varying HR (TVHR) with appropriate uncertainty. The TVHR approach was applied to a meta-analysis of 20 trials (4,069 patients) comparing chemotherapy to Standard of Care (SoC) for advanced recurrent gastric cancer, with Progression-Free Survival (PFS) as an outcome (median follow-up 1.2 years). It was also applied to a network meta-analysis (NMA) of 13 treatments across 13 Randomised Controlled Trials (RCTs) in previously untreated advanced BRAF-mutated melanoma (3,913 deaths in 6,378 participants) evaluating Overall Survival (OS). Both applications were compared against standard Bayesian meta-analysis assuming proportional hazards. Five trials in the gastric cancer meta-analysis showed non-proportional hazards for PFS. A standard Bayesian random-effects meta-analysis yielded a pooled HR of 0.78 (95
High-quality clinical research relies on the adequate recruitment of trial participants. The success of this underpins Evidence-Based Medicine. Trials participation is recommended however few patients have the opportunity to participate. The purpose of this study was to evaluate completed trials with the same sponsor to identify, define and evaluate various sponsor-led methods that can be applied to improve recruitment. Building upon the Theories of Evidence-Based Medicine and Implementation Research, this study evaluated published data with the same Australian academic trial sponsor to identify, define and evaluate recruitment methods used in completed trials. Methods were ranked using a novel scale. Ethical and patient activation timeliness were considered. Data was extracted and tabulated (grouped, deidentified). Using descriptive statistics, a multi-disciplinary team synthesised and interpreted the data. Between 2008–2020, 56 trials sponsored by the Australian skin cancer network were screened; three studies met the criteria. We identified, described and ranked four recruitment methods (Trial design, Trial promotion, education and training, Site selection and feasibility, Trial resourcing and incentivization, supported by 27 specific initiatives). Timely ethics administration was shown to be important. Despite some methods being more effortful than others, all methods contributed patients. There are a number of study limitations. Clinical trials are complex to run. Insufficient accrual is commonplace, including for Australian skin cancer trials. Four distinct recruitment methods with examples were identified. The results of this study may help trial sponsors and any other researcher to improve trial recruitment and completion, as well as patient outcomes.
Joint models of longitudinal and time-to-event data are essential for dynamic prediction in chronic disease research, but standard formulations typically assume linear biomarker–hazard associations and unidirectional influence from biomarkers to events. These assumptions are often violated in multimorbidity and oncology, where clinical events alter subsequent biomarker trajectories and where biomarker–risk relationships are non-linear or threshold-like. Methods that relax both assumptions simultaneously, and that are accessible to applied researchers, are currently lacking. We extend the joint modeling framework through the JMbdirect package by introducing semi-parametric association structures based on penalized splines and neural additive components, and by embedding explicit bidirectional feedback between events and biomarkers via post-event level shifts and slope changes in the longitudinal trajectory. Estimation is supported via penalized maximum likelihood with Laplace approximation and a Bayesian alternative using Hamiltonian Monte Carlo. Performance was assessed in simulation studies crossing four design factors (form of the association function, presence and magnitude of feedback, single versus competing events, and baseline hazard specification) at sample sizes n ∈{200, 500, 1000} , and in an applied analysis of the primary biliary cirrhosis (PBC) cohort. Incorporating feedback and flexible links improved discrimination and calibration relative to classical linear specifications, particularly under non-linear risk mechanisms. In the representative competing-risks scenario ( n=500 ), the spline-with-feedback estimator attained AUC@5 of 0.85 and an integrated Brier score of 0.134 with near-nominal 95
In clinical trials, the process of informed consent is intended to support prospective participants in making enrollment decisions that are well informed and aligned with their values and preferences. Despite the importance of relationship building during the consent process, there is little structured guidance for researchers on how to effectively facilitate informed consent conversations with prospective participants. This project aimed to develop and implement a relationship-based communication framework for clinical researchers, with the ultimate goal of supporting prospective participants in making values-aligned decisions. The ENgaging with Values to Inform research enrollment deciSIONs (ENVISION) study was launched to improve the process of research enrollment decision-making within the context of a multisite cystic fibrosis clinical trial. The ENVISION team sought input from individuals with cystic fibrosis, researchers, institutional review board leaders, and clinical research leaders to understand key issues in consent communication in the context of cystic fibrosis trials. A working group of the ENVISION team incorporated these viewpoints into the development of a consent communication framework, which was guided by consideration of elements of the Ottawa Decision Support Framework and incorporated established communication strategies. The communication framework and an accompanying training curriculum were finalized through iterative review with the working group and the full team. The development process resulted in a four-component, values-oriented communication framework, the RIVER framework: Relationship (forming a connection with the prospective participant), Information (sharing about the study), Values Exploration (eliciting and examining relevant values), and Resolution (confirming a values-aligned decision). Each component includes considerations to support prospective participants in making values-aligned decisions. RIVER is implemented through a training curriculum designed for research coordinators. The training curriculum consists of two 2-hour sessions that include didactic teaching, group discussion, and role-play scenarios, which provide an opportunity to practice real-world clinical research enrollment conversations. The RIVER framework offers a structured approach to communication during clinical research consent conversations that is designed to enhance voluntary, informed decisions. Ongoing evaluation of the framework and training curriculum will inform opportunities for future revisions, adaptations, and broader dissemination.