OBJECTIVES:Psychological outcome measures guide research and clinical decision-making, yet many widely used tools were developed with limited psychometric rigour. Although advanced methods (e.g., item response theory, structural equation modelling) are now widely available, their added value in applied research remains uncertain and applied researcher perspective regarding such are unexplored. We aimed to address this gap in knowledge by examining key stakeholder perspective. METHOD:To explore how these methods are perceived, we conducted semi-structured interviews with 21 stakeholders spanning psychometrics, clinical practice, applied research, statistics and academia. Data were analysed using reflexive thematic analysis. RESULTS:Analysis identified three overarching themes: (1) growing recognition that heterogeneity in latent traits challenges assumptions underlying many measures, (2) the enduring use of entrenched but psychometrically weak tools and (3) nuanced views on when advanced methods meaningfully influence research findings. Participants acknowledged gaps in psychometric literacy and emphasised the need for more training and collaboration with psychometricians. CONCLUSION:These findings highlighted persistent limitations in measurement practice, clarify contexts where psychometric methods add genuine value, and point to opportunities for strengthening outcome measurement in psychology and psychiatry research.
Background Antipsychotic trials in schizophrenia quantify treatment efficacy using clinician-rated scales such as the Positive and Negative Syndrome Scale (PANSS) and Brief Psychiatric Rating Scale (BPRS). If these scales include redundant or weakly informative items, measurement noise can blur drug-comparator differences that may dilute real effect sizes. Modern psychometric methods can refine and re-weight symptom scores that may improve precision and potentially increase detectable treatment effects, although it is unclear which approaches might be best. Aims To 1) determine best psychometric models for the PANSS and BPRS, and 2) test whether psychometrically-informed scoring of PANSS/BPRS alters estimated antipsychotic effect sizes compared with conventional total-score analyses, and 3) to explore whether any changes differ across symptom domains and by sex. Method We will conduct secondary analyses of anonymised individual participant data from randomized Phase II–IV schizophrenia trials accessed via secure data-sharing platforms (YODA Project and Vivli). Item-level PANSS/BPRS data will be harmonised at baseline and outcome. Psychometric models will be derived using confirmatory factor analysis, item response theory, and network analysis. Stability and invariance analyses will be conducted, examining covariates such as age and sex. Antipsychotic efficacy will be estimated using mixed-effects participant-level models, expressed as Cohen’s d standardised mean differences (SMDs) for both conventional and psychometrically-informed outcomes. The primary outcome will be the difference in SMDs between conventional and psychometrically-informed effect sizes, summarised across trials. Conclusions This study will determine whether best practice psychometric refinement of existing measures meaningfully changes antipsychotic effect size estimates, informing outcome measurement choices and interpretation of trial results.
Background Antipsychotic trials in schizophrenia quantify treatment efficacy using clinician-rated scales such as the Positive and Negative Syndrome Scale (PANSS) and Brief Psychiatric Rating Scale (BPRS). If these scales include redundant or weakly informative items, measurement noise can blur drug-comparator differences that may dilute real effect sizes. Modern psychometric methods can refine and re-weight symptom scores that may improve precision and potentially increase detectable treatment effects, although it is unclear which approaches might be best. Aims To 1) determine best psychometric models for the PANSS and BPRS, and 2) test whether psychometrically-informed scoring of PANSS/BPRS alters estimated antipsychotic effect sizes compared with conventional total-score analyses, and 3) to explore whether any changes differ across symptom domains and by sex. Method We will conduct secondary analyses of anonymised individual participant data from randomized Phase II–IV schizophrenia trials accessed via secure data-sharing platforms (YODA Project and Vivli). Item-level PANSS/BPRS data will be harmonised at baseline and outcome. Psychometric models will be derived using confirmatory factor analysis, item response theory, and network analysis. Stability and invariance analyses will be conducted, examining covariates such as age and sex. Antipsychotic efficacy will be estimated using mixed-effects participant-level models, expressed as Cohen’s d standardised mean differences (SMDs) for both conventional and psychometrically-informed outcomes. The primary outcome will be the difference in SMDs between conventional and psychometrically-informed effect sizes, summarised across trials. Conclusions This study will determine whether best practice psychometric refinement of existing measures meaningfully changes antipsychotic effect size estimates, informing outcome measurement choices and interpretation of trial results.
If we are to achieve the reorientation in policy and research that is needed, we have to begin by challenging implicit and dominant beliefs in the nature of science and knowledge. This chapter starts by demonstrating the constructed nature of science from which our beliefs about what counts as knowledge are derived. Structuring knowledge in terms of dualisms has forced social science into artificial and unworkable divisions, often inappropriately framed around quantitative versus qualitative methods. We find recognition of this failure of ontology throughout social theory, and yet, hitherto, there has been no adequate alternative that could both rescue a place for the real and incorporate the role of construction. In working with a complex realist ontology, we argue, we can take account of standpoint, agency and the very real ethical role of the social scientist in producing knowledge for policy.
This chapter seeks to apply the complexity frame to understand the interwoven crises in health systems and the issues this raises for research and governance. The impact of beliefs about what is real and how it can be measured and governed, discussed in Part I of the book, become clear. The chapter examines crises in health in terms of conceptual, organisational and research levels. Counterposing health as individually or socially determined identifies very different implications for action as illustrated by the changing definitions applied by the World Health Organization. Conventional research understanding through reduction is compared with the potential offered by more holistic approaches, and the meaning of this for governance is discussed. The failure to recognise health as a complex system is compounded by underpinning neoliberal policies that have brought health systems to crisis. Such action can be compared with public provision to point to drivers that result in very different system trajectories.
The development and evaluation of psychological outcome measures can involve the application of numerous sophisticated psychometric methods. Recent research noted that the use of different methods can lead to a plurality of modelling outcomes for commonly used measures. Furthermore, using these methods to optimise existing measures and derive weighted scores may not change the substantive results of studies when compared to original findings using total sum scores. In a series of semi-structured interviews, we presented these findings to 21 key stakeholders with psychology/psychiatry research backgrounds to elicit their perceptions regarding the use of advanced psychometric methods in applied research. Participants were purposively sampled and included psychometricians, clinicians, applied researchers, statisticians and academics. Using reflexive thematic analysis, we interpreted three themes; (1) Heterogeneity of latent traits, (2) A legacy of poor measurement, and (3) When psychometrics matter. Our findings highlight a mismatch between the generally held perception that latent traits are heterogeneous in nature and the way in which many measures were established as unitary constructs. The quality of commonly used outcome measures was critiqued, and a shortcoming in knowledge of how psychometrics could improve measurement accuracy was acknowledged. Improved training and increased collaboration with psychometricians could augment psychometric skills among applied researchers and enhance critical evaluation of routinely used outcome measures.
This chapter moves on from discussing dualism to consider the implications of rethinking the role of time and place in the construction of social life. In doing so, we return to the significance of agency and studying time and place as the achievement of actors. Rather than background factors, time and place must be understood as constitutive. This becomes clear when, in recognising the role of systems, we encounter the fundamental dynamic of coevolution as crucial to understanding. The chapter seeks to recognise the ethical implications of the shift from objectivity and to highlight the significance of time and place in governance, particularly in the governance of the future. Three examples of future governance are drawn on to illustrate the issues raised.
This chapter demonstrates how global inequality has shifted from being primarily an issue of inequality among countries to being one of inequality within countries. It shows how inequality has increased, with the most affluent drawing away not only from the poor but also from the crucial middle-income households. As the Organisation for Economic Co-operation and Development has noted, middle-income households, the majority in all high and most high middle income countries, achieved considerable ontological security, which is crucial to maintaining political legitimacy and social order. However, younger generations are unable to access this security in terms of day-to-day life, owning housing and establishing secure pensions. The chapter focuses on how inequality of power is the cause of these changes in post-democracies and begins to explore how that inequality might be confronted through new political forms which engage civil society in real decision making. Confronting impending climate catastrophe makes such developments essential.
This book develops a complex realist approach for addressing the interwoven socio-ecological crises of the 21st century, including climate change, health and social care crises, urban crises, fiscal crises and rising inequality. The authors argue that confronting these polycrises requires transforming research itself into an embedded process of participatory action and social change. Central concepts includ7e the complexity frame of reference for understanding systemic crises, the notion of crisis as a phase shift requiring resilience through system transformation, the interweaving of multiple crises and the role of future-oriented social science research as active praxis aimed at co-creating desired socio-ecological futures. The book critiques the restricted modes of science and dualistic thinking that have hampered the formulation of effective policy responses. It advocates synthesising multiple research modes – disciplinary, interdisciplinary, transdisciplinary – into a socio-ecological approach that integrates complexity, governance and political processes. The key is reconceiving time and social research as oriented towards shaping the future possibility spaces of complex systems. Ultimately, the authors argue applied social research using participatory action research and scenarios will be crucial to the transformative praxis required to achieve socially just, sustainable futures.
ObjectiveAs multiple sophisticated techniques are used to evaluate psychometric scales, in theory reducing error and enhancing measurement of patient reported outcomes, we aimed to determine whether applying different psychometric analyses would demonstrate important differences in treatment effects.Study Design and SettingWe conducted secondary analysis of individual participant data from 20 antidepressant treatment trials obtained from Vivli.org (n=6,843). Pooled item-level data from the HRSD-17 were analysed using confirmatory factory analysis (CFA), item response theory (IRT) and network analysis (NA). Multilevel models were used to analyse differences in trial effects at approximately 8 weeks (range 4-12 weeks) post-treatment commencement, with standardised mean differences calculated as Cohen’s d. Effect size outcomes for the original total depression scores were compared with psychometrically-informed outcomes based on abbreviated and weighted depression scores. ResultsSeveral items performing poorly during psychometric analyses and were eliminated, resulting in different models being obtained for each approach. Treatment effects were modified as follows per psychometric approach: 10.4%-14.9% increase for CFA, 0%-2.9% increase for IRT, 14.9%-16.4% reduction for NA. ConclusionPsychometric analyses differentially moderate effect size outcomes depending on the method used. In a 20-trial sample, factor analytic approaches increased treatment effect sizes relative to the original outcomes, NA decreased them, and IRT results reflected original trial outcomes.
The development and evaluation of psychological outcome measures can involve the application of numerous sophisticated psychometric methods. Recent research noted that the use of different methods can lead to a plurality of modelling outcomes for commonly used measures. Furthermore, using these methods to optimise existing measures and derive weighted scores may not change the substantive results of studies when compared to original findings using total sum scores. In a series of semi-structured interviews, we presented these findings to 21 key stakeholders with psychology/psychiatry research backgrounds to elicit their perceptions regarding the use of advanced psychometric methods in applied research. Participants were purposively sampled and included psychometricians, clinicians, applied researchers, statisticians and academics. Using reflexive thematic analysis, we interpreted three themes; (1) Heterogeneity of latent traits, (2) A legacy of poor measurement, and (3) When psychometrics matter. Our findings highlight a mismatch between the generally held perception that latent traits are heterogeneous in nature and the way in which many measures were established as unitary constructs. The quality of commonly used outcome measures was critiqued, and a shortcoming in knowledge of how psychometrics could improve measurement accuracy was acknowledged. Improved training and increased collaboration with psychometricians could augment psychometric skills among applied researchers and enhance critical evaluation of routinely used outcome measures.
This chapter will outline how research should be done as part of explicit goal-oriented social action when that action is the directed transformation of the whole macro-level socio-ecological global order. It will show how research is embedded in the practices of governance at the national and city regional levels, particularly in the crucial processes of planning. In social transformation, research is an essential part of the whole governance process. Research must be participatory/co-production, done with both agencies of governance and the general population in civil society. We will emphasise the role of scenario construction for the creation of socially just futures. We will also show how collaborative action-focused research can develop such scenarios and identify how what is desired might be brought into existence. The chapter will show how engaged action research must be an essential part of the programme for transformation to achieve an equitable and sustainable future.
The purpose of this chapter is to develop and ground the framing of crisis as a state in systems which cannot endure for long and has to result either in a restoration of previous system state or transformation of system state. Crisis must lead to change. This has to be set in relation to the general understanding of how and why complex systems change. It is important to distinguish between changes which alter the system’s parameters without changing its overall character and those which do change that character. The second type of change results in transformation of kind while the first does not. We explore how transformational change in complex systems happens. Transformational cause in complex systems is always complex, multiple and emergent. It involves many things which come together to engender change. A crucial emphasis in the chapter is on the role of human agency in engendering change.
Background Research has suggested that network analysis can be used to identify important pathology symptoms and inform targeted treatment plans that could lead to more efficacious outcomes in clinical trials. However, unless it can be demonstrated that network models are stable, including when accounting for moderating variables, NA-derived treatment plans may not be appropriate to implement. Aim We aim to assess the stability and invariance properties of two commonly used anxiety outcome measures to determine the suitability of NA methods to inform treatment plan design in clinical settings. Method Individual participant data (IPD) for large multi-trial samples will be accessed via Vivli.org. Exploratory graphical analysis will be used to model empirical networks pre- (baseline) and post-treatment (outcome) for the two most commonly used outcome measures in antidepressant clinical trials, namely the Hamilton Rating Scale for Anxiety (HAM-A) and the anxiety subscale of the Hospital Anxiety and Depression Scale (HADS_A). Bootstrapping and permutation techniques will be used to determine the stability and invariance properties of empirical networks in relation to a range of moderating variables, such as age, sex, treatment type and symptom severity. For networks that are unstable or partially invariant, we will examine item redundancy and remove non-performing items to pursue stable/invariant abbreviated models. Discussion This study will determine the suitability of applying NA methods in clinical trials. Findings could inform the way in which clinical trials, and other such research, are conducted. If outcome measures are stable and invariant, then NA methods will have demonstrable utility to inform more efficacious treatment plans. However, if NA is not found to be suitable, its validity as a robust analytical approach will be questionable.
BACKGROUND:Psychometric methods are used to remove underperforming items and reduce error in existing measures, albeit different approaches can produce different results. This study aimed to determine the implications of applying different psychometric methods for clinical trial outcomes. METHODS:Individual participant data from 15 antidepressant treatment trials from Vivli.org were analyzed. Baseline (pretreatment) and 8-week (range 4-12 weeks) outcome data from the Montgomery-Asberg Depression Rating Scale were subjected to best-practice factor analysis (FA), item response theory (IRT), and network analysis (NA) approaches. Trial outcomes for the original summative scores and psychometric-model scores were assessed using multilevel models. Percentage differences in Cohen's d effect sizes for the original summative and psychometrically modeled scores were the effects of interest. RESULTS:Each method produced unidimensional models, but the modified scales varied from 7 to 10 items. Treatment effects (d = 0.072) were unchanged for IRT (10 items), decreased by 1.3%-2.8% (eight-item abbreviated d = 0.070; weighted score d = 0.071) for NA, and increased by 11%-12.5% (seven-item abbreviated model d = 0.081; weighted score d = 0.080) for FA. DISCUSSION:IRT and NA yielded negligible differences in effect outcomes relative to original trials. FA increased effect sizes and may be the most effective method for identifying the items on which placebo and treatment group outcomes differ.
BACKGROUND:The 10-item Montgomery-Åsberg Depression Rating Scale (MADRS) is a commonly used measure of depression in antidepressant clinical trials. Numerous studies have adopted classical test theory perspectives to assess the psychometric properties of this scale, finding generally positive results. However, its network configural structure and stability is unexplored across different time-points and treatment groups. AIMS:To assess the network structure and stability of the MADRS in clinical settings pre- and post-treatment, and to determine a configurally invariant and stable model across time-points and treatment groups (placebo and intervention). METHOD:Individual participant data for 6440 participants from 14 clinical trials of major depressive disorder was obtained from the data repository Vivli.org. Exploratory Graphical Analysis (EGA) was used to identify empirical models pre-treatment (baseline) and post-treatment (8-week outcome). Bootstrapping techniques were applied to obtain optimised configurally invariant models. RESULTS:Empirical models presented with performance issues at baseline and for the placebo group at outcome. An abbreviated 8-item single-community model was found to be stable and configurally invariant across time-points and treatment groups. Symptoms such as low mood and lassitude showed most centrality across all models. LIMITATIONS:Metric invariance could not be explored due to research environment limitations. CONCLUSIONS:An 8-item one-community variant of the MADRS may provide optimal performance when conducting network analyses of antidepressant clinical trial outcomes. Findings suggest that interventions targeting low mood and lassitude might be most efficacious in treating depression among clinical trial participants. Further considerations of the potential impact on trial design and analysis should be explored.
Conventional approaches to causation in the social sciences draw on approaches in the Philosophy of Science in which a causal force acts on cases and generates change in the form of events. This relies on just one of the Aristotelian conceptions of cause - efficient cause - what brings the effect in to being. We should also pay attention to Final Cause - purpose and Formal cause, what makes something what it is and no other.The somethings are complex far from equilbric socio-ecological systems in which human agency has causal powers. This resonates with the understanding of the nature of effect in the complexity frame of reference as the state of the system both in relation to stability and transformation of kind. Effects are systems states. The argument draws on Hegel's and Dewey's understandings of cause / effect relationships as not separable but intimately interwoven. Effects have continuing reciprocal impacts on causes themselves as in positive feedback in systems. This way of thinking about causation allows us to engage with macro social change. The argument will be illustrated by a discussion of the transformation from industrial to post-industrial character across port city regions in high income countries.
Background The 17-item Hamilton Rating Scale for Depression (HRSD-17) is the most popular depression measure in antidepressant clinical trials. Prior evidence indicates poor replicability and inconsistent factorial structure. This has not been studied in pooled randomised trial data, nor has a psychometrically optimal model been developed. Aims To examine the psychometric properties of the HRSD-17 for pre-treatment and post-treatment clinical trial data in a large pooled database of antidepressant randomised controlled trial participants, and to determine an optimal abbreviated version. Method Data for 6843 participants were obtained from the data repository Vivli.org and randomly split into groups for exploratory (n = 3421) and confirmatory (n = 3422) factor analysis. Invariance methods were used to assess potential sex differences. Results The HRSD-17 was psychometrically sub-optimal and non-invariant for all models. High item variances and low variance explained suggested redundancy in each model. EFA failed at baseline and produced four item models for outcome groups (five for placebo-outcome), which were metric but not scalar invariant. Conclusions In antidepressant trial data, the HRSD-17 was psychometrically inadequate and scores were not sex invariant. Neither full nor abbreviated HRSD models are suitable for use in clinical trial settings and the HRSD's status as the gold standard should be reconsidered.
Introduction Patients with heart failure (HF) can have a markedly impaired QoL compared with many chronic diseases and life expectancy in end stage HF is less than one year. Key goals of management are improving symptoms, quality of life (QoL) and prolonging survival. Sacubitril/valsartan is a first-in-class angiotensin-receptor neprilysin inhibitor licenced to treat HF with reduced ejection fraction (HFrEF). Aims Our aim was to synthesise available randomised controlled trial evidence for QoL and safety for sacubitril/valsartan in patients with chronic HF, compared with usual care. QoL outcomes of interest were EuroQol EQ-5D-3L, New York Heart Association (NYHA) class and the Kansas City Cardiomyopathy Questionnaire (KCCQ). Methods This systematic review was registered on PROSPERO (reference CRD42020162031). We searched literature databases from inception to June 2023. Two authors independently screened all studies for eligibility and extracted data. Our meta-analysis used generic inverse variance (GIV) with random effects statistical models. Results We identified 4,483 records and, following screening, included 15 relevant clinical trials in our meta-analysis. Using the Cochrane Risk of Bias tool, most trials were of low risk of bias. Meta-analysis showed a greater likelihood of improvement in NYHA class (RR 1.12; 95% CI 1.01–1.24, p=0.032), a greater improvement in KCCQ score (MD 1.68; 95% CI 0.70–2.66, p=0.001) and greater likelihood of improvement in KCCQ >5 points (RR 1.13; 95% CI 0.80–1.6, p=0.487) with sacubitril/valsartan compared to control. One study reported an estimate for EQ-5D-3L, showing a greater improvement with sacubitril/valsartan (least squares mean of difference 0.92; 95% CI 0.36–1.48; p=0.001). Our results showed a greater risk of symptomatic hypotension (RR 1.59; 95% CI 1.23–2.07, p<0.001) with sacubitril/valsartan. Conclusion In HFrEF, Sacubitril/valsartan confers an improvement in QOL scores, but increases the risk of adverse effects, which needs to be taken into consideration when prescribing in populations that are more vulnerable.