A dynamic treatment regime (DTR) is a sequence of treatment decision rules tailored to an individual's evolving status over time. In precision medicine, much focus has been placed on finding an optimal DTR which, if followed by everyone in the population, would yield the best outcome on average; and extensive investigations have been conducted from both methodological and applied standpoints. The purpose of this tutorial is to provide readers who are interested in optimal DTRs with a systematic, detailed, but accessible introduction, including the formal definition and formulation of this topic within the framework of causal inference, identification assumptions required to link the causal quantity of interest to the observed data, existing statistical models and estimation methods for learning the optimal regime from the data, and application of these methods to both simulated and real data.
BACKGROUND:Because confirmatory clinical trials are costly, large-scale endeavors, the choice of their design carries significant weight. While the current methodological landscape offers tools to address residual pre-trial uncertainty through prespecified adaptations to design elements, design choices remain frequently constrained by prevailing orthodoxies. METHODS:We examine how unacknowledged design uncertainties (e.g., sample size calculations based on incorrect effect size or variance estimates) impact the statistical, ethical, and resource-stewardship goals of confirmatory trials. We contrast how fixed and adaptive design strategies handle these uncertainties, detailing the operational safeguards, firewalls, and trade-offs (e.g., complexity, operational bias) required to preserve trial integrity. RESULTS:Strict adherence to default templates prevents stakeholders from matching the design strategy to the specific complexities of the research question. We observe that the gold standard status of fixed designs often obscures their limitations in handling design uncertainties. A rigid adherence to these conventions can ethically hinder a trial's ability to deliver conclusive results. Conversely, while adaptive designs can offer tools to efficiently reduce uncertainty, we acknowledge that adaptive elements are not without cost; their implementation requires rigorous safeguards to manage specific risks to trial integrity and potential operational biases. CONCLUSION:To better align clinical research with its ethical and scientific mandates, we issue two calls to action. First, trial designers, including statisticians, must clearly articulate how target and design uncertainties impact ethical obligations, promoting an open-minded evaluation of both fixed and adaptive methods. Second, stakeholders must foreground ethical considerations during design selection, requiring explicit justification for the use-or non-use-of adaptive elements. The ethical path forward is not to default to the old or blindly adopt the new, but to explicitly justify the chosen design based on its responsiveness to the specific uncertainties of the trial. We posit that trial design must transition from a habit-based process to one where pre-trial uncertainties are openly discussed and addressed.
Abstract Executive function is an essential cognitive domain for typical human behavior which is disrupted in neurodevelopmental and neurodegenerative disorders, but little is known about its underlying molecular basis. To address this, we perform genome-wide association studies (GWAS) using three different measures of executive function in UK Biobank (N = 84,238) and NIHR BioResource’s Genes and Cognition (N = 9932) study participants, followed by a meta-analysis. The trail-making alphanumeric (TMA) measure is the most heritable phenotype (h²=7-26%), associated with 18 independent loci that exhibit a similar direction of effect in both cohorts. Across these loci, in-silico follow-up implicates 178 genes, of which NT5DC2 and RP11-579E24.2 are independently replicated prior to meta-analysis. TMA is linked to pan-cerebral differences in brain structure, with brain-enriched genes showing a biphasic expression profile from early development through to later life. Our data implicate specific cell types, histone modifications and butyrophilin immunoglobulin family proteins as potential targets for promoting cognitive resilience.
Factor analysis (FA) can be used to identify key biomarkers in biological processes by assuming that latent biological pathways (statistically, "latent factors") drive the activity of measurable biomarkers ("observed variables"). However, biological pathways often interact, meaning that the classical FA assumption of independence between factors is questionable. Motivated by sparsely and irregularly collected longitudinal measurements of metabolites in a COVID-19 study, we propose a dynamic factor analysis model that accounts for cross-correlations between pathways via a multi-output Gaussian processes (MOGP) prior on the factor trajectories. To mitigate against overfitting caused by sparsity of longitudinal measurements, we introduce a roughness penalty upon MOGP hyperparameters and allow for non-zero mean functions. We also propose a scalable stochastic expectation maximization (StEM) algorithm that, in simulations, is both 20 times faster and provides more accurate and stable MOGP hyperparameter estimates than a previously-proposed Monte Carlo Expectation Maximization algorithm. In the motivating COVID-19 study, our methodology identifies a kynurenine pathway that affects the clinical severity of patients with COVID-19 disease and uncovers the role of the biomarker taurine. Our R package DFA4SIL implements the proposed method.
There is growing interest in the role of within-individual variability (WIV) in biomarker trajectories for assessing disease risk and progression. A trajectory-based definition that has attracted recent attention characterizes WIV as the curvature-based roughness of the latent biomarker trajectory, and we refer to the resulting measure as trajectory-based biomarker WIV (TB-WIV). To evaluate TB-WIV associations with clinical outcomes and perform dynamic risk prediction, joint models for longitudinal and time-to-event data (JM) are necessary. However, specifying the longitudinal trajectory is critical and poses methodological challenges. In this work, we investigate three Bayesian semiparametric approaches for longitudinal modeling and TB-WIV estimation. We propose two new approaches based on Bayesian penalized splines (P-splines) and functional principal component analysis, and adapt an existing semiparametric approach to the Bayesian JM setting. Through simulation studies, we evaluate parameter estimation, detection of TB-WIV associations, estimation of their raw and standardized magnitudes, and survival prediction, providing insights for method choice and highlighting the need for caution when interpreting raw TB-WIV association magnitudes. We apply the approaches to UK Cystic Fibrosis Registry data, where we identify significant positive associations between lung function TB-WIV and mortality risk in patients with cystic fibrosis, after adjustment for the underlying lung function level, and provide evidence that TB-WIV adds information for survival prediction.
Optimal dynamic treatment regimes (DTRs), as a key part of precision medicine, have progressively gained more attention recently. To inform clinical decision making, interpretable and parsimonious models for contrast functions are preferred, raising concerns about undue misspecification. It is therefore important to properly evaluate the performance of candidate interpretable models and select the one that best approximates the unknown contrast function. Moreover, since a DTR usually involves multiple decision points, an inaccurate approximation at a later decision point affects its estimation at an earlier decision point when a backward induction algorithm is applied. This paper aims to perform model selection for contrast functions in the context of learning optimal DTRs from observed data. Note that the relative performance of candidate models may heavily depend on the sample size when, for example, the comparison is made between parametric and tree-based models. Therefore, instead of investigating the limiting behavior of each candidate model and developing methods to select asymptotically the `correct' one, we focus on the finite sample performance of each model and attempt to perform model selection under a given sample size. To this end, we adopt the counterfactual cross-validation metric and propose a novel method to estimate the variance of the metric. Supplementing the cross-validation metric with its estimated variance allows us to characterize the uncertainty in model selection under a given sample size and facilitates hypothesis testing associated with a preferred model structure.
Monitoring the incidence of new infections during a pandemic is critical for an effective public health response. General population prevalence surveys for SARS-CoV-2 can provide high-quality data to estimate incidence. However, estimation relies on understanding the distribution of the duration that infections remain detectable. This study addresses this need using data from the Coronavirus Infection Survey (CIS), a long-term, longitudinal, general population survey conducted in the UK. Analyzing these data presents unique challenges, such as doubly interval censoring, undetected infections, and false negatives. We propose a Bayesian nonparametric survival analysis approach, estimating a discrete-time distribution of durations and integrating prior information derived from a complementary study. Our methodology is validated through a simulation study, including its resilience to model misspecification, and then applied to the CIS dataset. This results in the first estimate of the full duration distribution in a general population, as well as methodology that could be transferred to new contexts.
Accurate time-to-event prediction is integral to decision-making, informing medical guidelines, hiring decisions, and resource allocation. Survival analysis, the quantitative framework used to model time-to-event data, accounts for patients who do not experience the event of interest during the study period, known as censored patients. However, many patients experience events that prevent the observation of the outcome of interest. These competing risks are often treated as censoring, a practice frequently overlooked due to a limited understanding of its consequences. Our work theoretically demonstrates why treating competing risks as censoring introduces substantial bias in survival estimates, leading to systematic overestimation of risk and, critically, amplifying disparities. First, we formalize the problem of misclassifying competing risks as censoring and quantify the resulting error in survival estimates. Specifically, we develop a framework to estimate this error and demonstrate the associated implications for predictive performance and algorithmic fairness. Furthermore, we examine how differing risk profiles across demographic groups lead to group-specific errors, potentially exacerbating existing disparities. Our findings, supported by an empirical analysis of cardiovascular management, demonstrate that ignoring competing risks disproportionately impacts the individuals most at risk of these events, potentially accentuating inequity. By quantifying the error and highlighting the fairness implications of the common practice of considering competing risks as censoring, our work provides a critical insight into the development of survival models: practitioners must account for competing risks to improve accuracy, reduce disparities in risk assessment, and better inform downstream decisions.
Electronic health records arise from the complex interaction between patients and the healthcare system. This observation process of interactions, referred to as clinical presence, often impacts observed outcomes. When using electronic health records to develop clinical prediction models, it is standard practice to overlook clinical presence, impacting performance and limiting the transportability of models when this interaction evolves. We propose a multi-task recurrent neural network that jointly models the inter-observation time and the missingness processes characterising this interaction in parallel to the survival outcome of interest. Our work formalises the concept of clinical presence shift when the prediction model is deployed in new settings (e.g. different hospitals, regions or countries), and we theoretically justify why the proposed joint modelling can improve transportability under changes in clinical presence. We demonstrate, in a real-world mortality prediction task in the MIMIC-III dataset, how the proposed strategy improves performance and transportability compared to state-of-the-art prediction models that do not incorporate the observation process. These results emphasise the importance of leveraging clinical presence to improve performance and create more transportable clinical prediction models.
Purpose (the aim of the study): Osteoarthritis (OA) exerts a profound impact on quality of life. In the absence of disease-modifying OA drugs (DMOADs), there is a critical need to unravel the molecular drivers of the disease and discern potential variations among individuals. The advent of large-scale, high-throughput omics-related biotechnologies enable us to address this need at an unprecedented scale. The STEpUP OA (Synovial Fluid To Define Molecular Endotypes by Unbiased Proteomics in Osteoarthritis) consortium was established to answer a primary question of whether distinct molecular OA endotypes exist using large-scale proteomics of knee OA synovial fluid (SF). We hypothesised that molecular endotypes may delineate specific patient groups, thereby influencing diagnosis, predicting disease course, and guiding treatment response.
Background Mendelian randomization is a popular method for causal inference with observational data that uses genetic variants as instrumental variables. Similarly to a randomized trial, a standard Mendelian randomization analysis estimates the population-averaged effect of an exposure on an outcome. Dividing the population into subgroups can reveal effect heterogeneity to inform who would most benefit from intervention on the exposure. However, as covariates are measured post-“randomization”, naive stratification typically induces collider bias in stratum-specific estimates. Method We extend a previously proposed stratification method (the “doubly-ranked method”) to form strata based on a single covariate, and introduce a data-adaptive random forest method to calculate stratum-specific estimates that are robust to collider bias based on a high-dimensional covariate set. We also propose measures based on the Q statistic to assess heterogeneity between stratum-specific estimates (to understand whether estimates are more variable than expected due to chance alone) and variable importance (to identify the key drivers of effect heterogeneity). Result We show that the effect of body mass index (BMI) on lung function is heterogeneous, depending most strongly on hip circumference and weight. While for most individuals, the predicted effect of increasing BMI on lung function is negative, it is positive for some individuals and strongly negative for others. Conclusion Our data-adaptive approach allows for the exploration of effect heterogeneity in the relationship between an exposure and an outcome within a Mendelian randomization framework. This can yield valuable insights into disease aetiology and help identify specific groups of individuals who would derive the greatest benefit from targeted interventions on the exposure.
Selection bias poses a substantial challenge to valid statistical inference in nonprobability samples. This study compared estimates of the first-dose COVID-19 vaccination rates among Indian adults in 2021 from a large nonprobability sample, the COVID-19 Trends and Impact Survey (CTIS), and a small probability survey, the Center for Voting Options and Trends in Election Research (CVoter), against national benchmark data from the COVID Vaccine Intelligence Network. Notably, CTIS exhibits a larger estimation error on average (0.37) compared to CVoter (0.14). Additionally, we explored the accuracy (regarding mean squared error) of CTIS in estimating successive differences (over time) and subgroup differences (for females versus males) in mean vaccine uptakes. Compared to the overall vaccination rates, targeting these alternative estimands comparing differences or relative differences in two means increased the effective sample size. These results suggest that the Big Data Paradox can manifest in countries beyond the United States and may not apply equally to every estimand of interest.
Purpose (the aim of the study): The molecular mechanisms of osteoarthritis (OA) remain largely unexplained. STEpUP OA (Synovial Fluid To Define Molecular Endotypes by Unbiased Proteomics in Osteoarthritis) is an international consortium which has used SomaLogic SomaScan technology to measure over 7000 proteins in the synovial fluid (SF) from >1700 individuals with OA or following joint injury. This has generated an unprecedented molecular and clinical data resource with which to interrogate molecules and pathways associated with disease. As part of this work, we set out to compare the SF proteomic profiles of individuals with advanced and non-advanced radiographic knee OA.
A leading explanation for translational failure in neurodegenerative disease is that new drugs are evaluated late in the disease course when clinical features have become irreversible. Here, to address this gap, we cognitively profiled 21,051 people aged 17-85 years as part of the Genes and Cognition cohort within the National Institute for Health and Care Research BioResource across England. We describe the cohort, present cognitive trajectories and show the potential utility. Surprisingly, when studied at scale, the APOE genotype had negligible impact on cognitive performance. Different cognitive domains had distinct genetic architectures, with one indicating brain region-specific activation of microglia and another with glycogen metabolism. Thus, the molecular and cellular mechanisms underpinning cognition are distinct from dementia risk loci, presenting different targets to slow down age-related cognitive decline. Participants can now be recalled stratified by genotype and cognitive phenotype for natural history and interventional studies of neurodegenerative and other disorders.
Background Osteoarthritis (OA) has a lifetime risk of over 40%, imposing a huge societal burden. Clinical variability suggests that it could be more than one disease. Synovial fluid To detect Endotypes by Unbiased Proteomics in OA (STEpUP OA) was established to test the hypothesis that there are detectable distinct molecular endotypes in knee OA. Methods OA knee synovial fluid (SF) samples (N=1361) were from pre-existing OA cohorts with cross-sectional clinical (radiographic and pain) data. Samples were divided into Discovery (N = 708) and Replication (N=653) datasets. Proteomic analysis was performed using SomaScan V4.1 assay (6596 proteins). Unsupervised clustering was performed using k-means, assessed using the f(k) metric, with and without adjustments for potential confounders. Regression analyses were used to assess protein associations with radiographic (Kellgren and Lawrence) and knee pain (WOMAC pain), with and without stratification by body mass index (BMI) or biological sex. Adjustments were made for cohort (random intercept) or intracellular protein, using an intracellular protein score (IPS). Analyses were carried out in R according to a pre-published plan. Results No distinct SF molecular endotypes were identified in OA but two indistinct clusters were defined in non-IPS regressed data which were stable across subgroup analyses. Clustering was lost after IPS regression adjustment. Strong, replicable protein associations were observed with radiographic disease severity, which were retained after adjustment for cohort or IPS. Pathway analysis identified a strong epithelial to mesenchymal transition (EMT) pathway, and weaker associations with angiogenesis, complement and coagulation. The latter were variably lost after adjustment for BMI or biological sex. Associations with patient reported pain were weaker. Conclusion These data support knee OA as a biologically continuous disease in which disease severity is associated with a strong, robust, tissue remodelling signature. Subtle differences were found in pathways after stratification by BMI or sex.### Competing Interest StatementTAP, YD, PH, SL, AS, NKA, AJP, DF, MK, BM, AMV and SK declare no conflicts of interest. FW has received consultancy fees from Pfizer. LSL has received consultancy fees from Arthro Therapeutics AB, and is an advisory board member of AstraZeneca. LJD has received consultancy fees from Nightingale Health PLC. TLV has no conflicts to declare with the exception of grant income for STEpUP OA from industry partners (see above). RAM is a shareholder of AstraZeneca. SB and JM are employees and shareholders of Novartis. CTA has received consultancy fees from Novartis, and has received honoraria for educational purposes also from Novartis. TJW is a shareholder of Chondropeptix BV. DAW has received consultancy fees from GlaxoSmithKline plc, AKL Research & Development Limited, Pfizer Ltd, Eli Lilly and Company, Contura International, and AbbVie Inc, has received honoraria for educational purposes from Pfizer Ltd and AbbVie Inc, is a board member (Director) of UKRI and Versus Arthritis Advanced Pain Discovery Platform. ### Funding StatementThe study was supported by Kennedy Trust for Rheumatology Research (grant number: 171806), Versus Arthritis (grant number: 22473), Centre for Osteoarthritis Pathogenesis Versus Arthritis (grant numbers: 21621, 20205), Galapagos, Biosplice, Novartis, Fidia, UCB, Pfizer (non-consortium member) and Somalogic (in kind contributions). The funders Kennedy Trust for Rheumatology Research, Versus Arthritis and Pfizer had no role in the study design, data collection and analysis, decision to publish or preparation of the manuscript. The funders Galapagos, Biosplice, Novartis, Fidia, UCB and SomaLogic were all active consortium members, attending consortium meetings. As such they made contributions to the study design and support of data collection, decision to publish and review and commenting on the manuscript. In addition, SomaLogic, UCB and Novartis were members of the Data Analysis Group.### Author DeclarationsI confirm all relevant ethical guidelines have been followed, and any necessary IRB and/or ethics committee approvals have been obtained.YesThe details of the IRB/oversight body that provided approval or exemption for the research described are given below:University of Oxford Medical Sciences Central University Research Ethics Committee (CUREC) granted ethical approval for the processing, storage and use of samples and linked data for this project on 1st November 2019 (R67029/RE001).I confirm that all necessary patient/participant consent has been obtained and the appropriate institutional forms have been archived, and that any patient/participant/sample identifiers included were not known to anyone (e.g., hospital staff, patients or participants themselves) outside the research group so cannot be used to identify individuals.YesI understand that all clinical trials and any other prospective interventional studies must be registered with an ICMJE-approved registry, such as ClinicalTrials.gov. I confirm that any such study reported in the manuscript has been registered and the trial registration ID is provided (note: if posting a prospective study registered retrospectively, please provide a statement in the trial ID field explaining why the study was not registered in advance).YesI have followed all appropriate research reporting guidelines, such as any relevant EQUATOR Network research reporting checklist(s) and other pertinent material, if applicable.YesThe minimal datasets upon which this data relies and all R code, including the html vignette, are available at https://github.com/ndorms-tperry/STEpUP-OA-Primary-Manuscript. The full STEpUP OA dataset may be made available by application to the Data Access and Publication Group of STEpUP OA (stepupoa@kennedy.ox.ac.uk) once the primary analysis manuscript is published, in accordance with what is stipulated in our Consortium Agreement. This may attract an access fee to cover administrative processing. Neither the minimal dataset nor the full STEpUP OA dataset include patient identifiable data.
Objectives To develop a protocol for largescale analysis of synovial fluid proteins, for the identification of biological networks associated with subtypes of osteoarthritis. Methods Synovial Fluid To detect molecular Endotypes by Unbiased Proteomics in Osteoarthritis (STEpUP OA) is an international consortium utilising clinical data (capturing pain, radiographic severity and demographic features) and knee synovial fluid from 17 participating cohorts. 1746 samples from 1650 individuals comprising OA, joint injury, healthy and inflammatory arthritis controls, divided into discovery (n = 1045) and replication (n = 701) datasets, were analysed by SomaScan Discovery Plex V4.1 (>7000 SOMAmers/proteins). An optimised approach to standardisation was developed. Technical confounders and batch-effects were identified and adjusted for. Poorly performing SOMAmers and samples were excluded. Variance in the data was determined by principal component (PC) analysis. Results A synovial fluid standardised protocol was optimised that had good reliability (<20% co-efficient of variation for >80% of SOMAmers in pooled samples) and overall good correlation with immunoassay. 1720 samples and >6290 SOMAmers met inclusion criteria. 48% of data variance (PC1) was strongly correlated with individual SOMAmer signal intensities, particularly with low abundance proteins (median correlation coefficient 0.70), and was enriched for nuclear and non-secreted proteins. We concluded that this component was predominantly intracellular proteins, and could be adjusted for using an ‘intracellular protein score’ (IPS). PC2 (7% variance) was attributable to processing batch and was batch-corrected by ComBat. Lesser effects were attributed to other technical confounders. Data visualisation revealed clustering of injury and OA cases in overlapping but distinguishable areas of high-dimensional proteomic space. Conclusions We have developed a robust method for analysing synovial fluid protein, creating a molecular and clinical dataset of unprecedented scale to explore potential patient subtypes and the molecular pathogenesis of OA. Such methodology underpins the development of new approaches to tackle this disease which remains a huge societal challenge.
Introduction Glucagon-like peptide-1 receptor agonists (GLP-1 RAs), currently marketed for type 2 diabetes and obesity, may offer novel mechanisms to delay or prevent neurotoxicity associated with Alzheimer’s disease (AD). The impact of semaglutide in amyloid positivity (ISAP) trial is investigating whether the GLP-1 RA semaglutide reduces accumulation in the brain of cortical tau protein and neuroinflammation in individuals with preclinical/prodromal AD.Methods and analysis ISAP is an investigator-led, randomised, double-blind, superiority trial of oral semaglutide compared with placebo. Up to 88 individuals aged ≥55 years with brain amyloid positivity as assessed by positron emission tomography (PET) or cerebrospinal fluid, and no or mild cognitive impairment, will be randomised. People with the low-affinity binding variant of the rs6971 allele of the Translocator Protein 18 kDa (TSPO) gene, which can interfere with interpreting TSPO PET scans (a measure of neuroinflammation), will be excluded.At baseline, participants undergo tau, TSPO PET and MRI scanning, and provide data on physical activity and cognition. Eligible individuals are randomised in a 1:1 ratio to once-daily oral semaglutide or placebo, starting at 3 mg and up-titrating to 14 mg over 8 weeks. They will attend safety visits and provide blood samples to measure AD biomarkers at weeks 4, 8, 26 and 39. All cognitive assessments are repeated at week 26. The last study visit will be at week 52, when all baseline measurements will be repeated. The primary end point is the 1-year change in tau PET signal.Ethics and dissemination The study was approved by the West Midlands—Edgbaston Research Ethics Committee (22/WM/0013). The results of the study will be disseminated through scientific presentations and peer-reviewed publications.Trial registration number ISRCTN71283871.
The increasing availability of high-dimensional, longitudinal measures of gene expression can facilitate understanding of biological mechanisms, as required for precision medicine. Biological knowledge suggests that it may be best to describe complex diseases at the level of underlying pathways, which may interact with one another. We propose a Bayesian approach that allows for characterizing such correlation among different pathways through dependent Gaussian processes (DGP) and mapping the observed high-dimensional gene expression trajectories into unobserved low-dimensional pathway expression trajectories via Bayesian sparse factor analysis. Our proposal is the first attempt to relax the classical assumption of independent factors for longitudinal data and has demonstrated a superior performance in recovering the shape of pathway expression trajectories, revealing the relationships between genes and pathways, and predicting gene expressions (closer point estimates and narrower predictive intervals), as demonstrated through simulations and real data analysis. To fit the model, we propose a Monte Carlo expectation maximization (MCEM) scheme that can be implemented conveniently by combining a standard Markov Chain Monte Carlo sampler and an R package GPFDA,which returns the maximum likelihood estimates of DGP hyperparameters. The modular structure of MCEM makes it generalizable to other complex models involving the DGP model component. Our R package DGP4LCF that implements the proposed approach is available on the Comprehensive R Archive Network (CRAN).
Identifying patient subgroups with different treatment responses is an important task to inform medical recommendations, guidelines, and the design of future clinical trials. Existing approaches for treatment effect estimation primarily rely on Randomised Controlled Trials (RCTs), which tend to feature more homogeneous patient groups, making them less relevant for uncovering subgroups in the population encountered in real-world clinical practice. Subgroup analyses established for RCTs suffer from significant statistical biases when applied to observational studies, which benefit from larger and more representative populations. Our work introduces a novel, outcome-guided, subgroup analysis strategy for identifying subgroups of treatment response in both RCTs and observational studies alike. It hence positions itself in-between individualised and average treatment effect estimation to uncover patient subgroups with distinct treatment responses, critical for actionable insights that may influence treatment guidelines. In experiments, our approach significantly outperforms the current state-of-the-art method for subgroup analysis in both randomised and observational treatment regimes.