
Evaluation commissioners need more than a verdict on whether a program worked. They need evidence to inform cost-effectiveness decisions, support scaling, and guide complementary program design. Realist evaluation provides the right causal structure to generate that evidence: in this context, the program activates this mechanism, producing this outcome. This article addresses three problems with realist evaluation: the misconceived debate between realist and counterfactual approaches, which are independent decisions not competing alternatives; poor specification of mechanisms and context; and the failure of evaluations in complex systems to capture how the context evolves over time and at scale. The solution integrates realist evaluation with behavioural science and systems thinking and allows for any data collection and analysis methods. Behavioural science operationalises mechanisms and context using a model of behaviour called “COM-B”, and explains why the same program produces different effects for different people through mindsets – making findings precise enough to inform cost-effectiveness decisions and complementary program design. Systems thinking maps feedback loops, leverage points, and emergent outcomes – identifying where to intervene for lasting or transformative change and what contextual conditions must be in place before scaling. The realist behavioural systems evaluation framework is illustrated through a case study and five-step practical guide.
Ethical considerations are central to evaluation practice, both in formal approval processes and in the everyday realities of managing relationships, power, and integrity. Even evaluators adept at meeting procedural ethics requirements and complying with formal expectations such as those expressed in policies and guidelines, must also navigate ongoing informal day-to-day challenges relating to independence, transparency, and stakeholder engagement for which evaluators often report limited guidance or support. Drawing on mixed-methods data from two rounds of the Everyday Ethics in Evaluation study (2024 and 2025), including surveys, polls, and interviews, this article examines the perceived frequency, nature, and severity of ethical challenges faced by evaluators. Thematic analysis identified three recurring areas of concern: conflict of interest mismanagement, non-protection of participant rights, and data and evaluation quality. The findings reveal meaningful variation by role and tenure: emerging evaluators reported higher self-assessed confidence, potentially reflecting overconfidence linked to limited exposure to complex ethical risk; external evaluators reported a higher frequency of conflict-of-interest dilemmas; while public sector evaluators more often described procedural or governance-related challenges. Drawing on the research findings, the article offers two frameworks to help evaluators understand and respond to challenges identified in the findings, and to contribute to the literature on ethical practice in evaluation. These frameworks are intended to support reflective practice in evaluation and better articulate the ethical challenges encountered by evaluators. The article proposes new research directions that could explore the proposed frameworks in different evaluation contexts and to further explore the current level of understanding around ethical practice within the evaluation community.
Randomised controlled trials (RCTs) are increasingly promoted for evaluating publicly funded initiatives, including government-funded student equity programs in higher education. Yet their application in university settings remains uneven and contested. This article critically examines the role of RCTs in evaluating student equity initiatives in Australian universities. Drawing on a review of the literature, two applied RCT case examples, and analysis of the institutional realities of Australian higher education, it identifies recurring methodological tensions that arise when experimental designs are applied to complex, justice-oriented equity interventions. These tensions include weak alignment between problem diagnosis and intervention logic, homogenisation of heterogeneous equity cohorts, reliance on simplified behavioural proxies, delivery and engagement constraints, and broader institutional barriers to experimentation across universities. The article argues that many apparent weaknesses in RCT design reflect deeper tensions between experimental assumptions and the relational, adaptive, and context-dependent nature of equity practice, rather than simply poor implementation. It reframes RCTs not as a default standard for equity evaluation, but as one tool within a plural evaluation ecosystem, to be used selectively and ethically in alignment with equity purpose, institutional capacity, and clearly specified causal questions.
The language we use to describe professional identity in evaluation matters. Prevailing discourse relies on a developmental binary (‘emerging’ versus ‘expert’) that assumes a linear trajectory of competence over time. Yet many practitioners inhabit an under-theorised ‘murky middle’: no longer ‘emerging’, not comfortably ‘experts’, and working across contexts where capability is situational and negotiated. Drawing on complexity and systems thinking, this paper proposes an alternative: the Emergent Evaluator. Rather than a career stage, the Emergent Evaluator is a practice orientation in which identity and competence are treated as dynamic, relational and context-dependent, arising as emergent properties of the evaluator’s interactions with people, organisations, and systems. The article reviews limitations of linear developmental framings and synthesises insights from complex adaptive systems and emergence to reconceptualise evaluator identity. It defines the Emergent Evaluator, outlines core attributes, and explores implications for capability frameworks, professional development, organisational evaluation cultures, and professional associations. It concludes with implications and future directions for embedding an explicit emergent practice orientation across evaluation sectors.
Over the past decade, evaluation practice in Australia has evolved in response to shifting social, political and cultural contexts. Drawing on ten years of Australian Evaluation Society (AES) publications, submissions, thought pieces and key government strategy documents, this article synthesises patterns in evaluation discourse and practice and identifies six influential ideas shaping the field. These ideas are grouped into three themes: People First, Methods and Technology and Embed and Evolve. This special feature article arising from a presentation at the AES Internation Evaluation Conference in Canberra (2025) in which a panel of evaluation leaders, Associate Professor Amy Gullickson, Eleanor Williams, Doyen Radcliffe and Theo Nabben, explored these ideas and tensions within, emphasising the centrality of people, relationships and meaningful change. Through the lens of the Super Six, they contend that evaluation’s true value lies in its capacity to empower communities and improve lives – re-centring people, relationships and values as critical measures of evaluation quality and impact.
This study evaluates the Youth on Track program, launched by the New South Wales Government in 2013 to improve outcomes and reduce recidivism among young people (aged 10–17) who are at risk of offending. Using a quasi-experimental design and Propensity Score Matching, the study compares individuals referred to Youth on Track (Oct 2016–Mar 2020, n = 1,997) with a matched group ( n = 3,994) from the Human Services DataSet, a de-identified dataset of over 7 million New South Wales service records. Analysis includes Youth on Track administrative data and survival analysis to assess outcomes. Youth on Track referrals had similar or higher engagement with services and higher probabilities of adverse outcomes such as reoffending, custody placement, risk of significant harm reports, out-of-home care placement, homelessness support, and mental health service use. Although the findings suggest that Youth on Track referrals did not achieve better outcomes compared to the matched control group, they are prone to detection bias due to factors such as increased supervision and unmeasured traits like peer/sibling involvement in crime or motivation to stop offending. Thus, results should not guide decision-making without further research. Despite limitations, the study highlights the value of rigorous quasi-experimental methods and Human Services DataSet in evaluating complex interventions and informing future program assessments.
Participatory grantmaking (PGM) applies participatory principles directly to funding, involving communities in determining priorities and allocations. Despite its growing use, empirical evidence on PGM implementation remains limited. This process evaluation examined a PGM pilot implemented by Healthy North Coast, a regional primary health commissioning body in New South Wales, Australia. Using a qualitative process evaluation, semi-structured interviews were conducted with implementation and delivery stakeholders involved in the design, facilitation, observation and governance of the pilot. Data were analysed using reflexive thematic analysis, guided by an evaluative framework focussing on feasibility, fairness and legitimacy, equity of participation, and contribution to learning. Five themes were identified: (1) value of participatory design; (2) funding allocation tensions; (3) emotional impact of peer competition; (4) challenges in achieving equity; and (6) capability-building and collaboration opportunities. Findings reveal inherent tensions between participatory ideals and institutional commissioning realities. While the pilot enhanced legitimacy and organisational learning, competitive elements constrained equity and substantive power redistribution. The study contributes empirical evidence on PGM in practice and offers methodological insights for evaluating participatory funding interventions, emphasising the need to assess power, equity, relational impacts and emotional dimensions alongside procedural fidelity.
Despite being cemented in the Joint Committee Standards as a tenet of high-quality evaluation practice, use of meta-evaluation remains limited. One contributing factor to the limited utility of meta-evaluation may be that meta-evaluation is perceived as a burdensome process that necessitates a large degree of time, resources, and expertise to be completed successfully. To remedy this gap and embed meta-evaluation as a regular aspect of evaluation practice, the authors propose a new framework for conducting meta-evaluation. Termed participatory meta-evaluation, this new framework makes program partners active participants in the meta-evaluation process with the goal of enhancing the utility of meta-evaluation findings and the use of meta-evaluation in practice. The three principles of participatory meta-evaluation as well as the perceived benefits of participatory meta-evaluation compared to traditional forms of meta-evaluation are discussed. The authors conclude with a call to action encouraging other evaluators to practice participatory meta-evaluation and document its effects.
Evaluation is essential for assessing whether policy and program outcomes align with their intended goals. A key critique of Western evaluation approaches is their failure to incorporate Indigenous ways of knowing, being and doing. Developing an Indigenist evaluation concept can enhance evaluation quality by recognising Indigenous perspectives. This study draws on Indigenous Lifeworlds as a theoretical basis to explore the idea of ‘collective capability’, asserting Indigenous societies operate at the ‘collective’ level over the individual and focus on ‘capability’ as a way of realising potential through collective effort. We conducted 20 in-depth interviews with Aboriginal and Torres Strait Islander Knowledge holders involved in evaluation to identify attributes of collective capability. Through manifest and latent content analysis, we defined collective capability and its key components. Our definition describes it as people coming together in a relational, collaborative environment to pursue a common goal, where cultural values are prioritised, knowledge is shared and the process is as important as the outcome . The two main elements of collective capability are relationality and knowledge sharing, each with several sub-elements. The next phase of the research will focus on integrating these elements into the context of Indigenist evaluation, developing criteria and processes for their operationalisation.
Community-based programs involve participants whose perspectives are shaped by the social and geographical context in which they operate. Evaluations of these programs should capture this context and incorporate the diverse perspectives of participants in data collection and findings. This is demonstrated in an evaluation of a community-based disaster resilience program in New South Wales, Australia, in which six community case studies were developed from geo-social statistics, site visit observations and in-depth interviews with key community informants. The evaluation used the case studies to identify commonalities and differences in how the program was understood by participants in different communities and by external stakeholders to the program. The findings showed how context informs a more comprehensive understanding of a community-based program’s impact and how this may vary according to the context. This evaluation reinforces the relevance of Realist Evaluation principles which emphasise the importance of context in evaluating programs.
The Australian government spends millions of dollars funding new programs every year. Taxpayers, policy makers, school leaders, teachers, and students need to know whether these programs are good. Legislation ensures they are evaluated, but do those evaluations report what good looks like, how good these programs are, and for whom? This research sought an answer to that question by analysing publicly available educational evaluations using a new conceptual framework that integrated the logic of evaluation and evaluative reasoning. Both are essential to making a credible, valid, and defensible claim about how good something is: the logic of evaluation makes the judgement legitimate, and evaluative reasoning justifies it. We examined 37 reports using our framework using an adapted systematic quantitative analysis method. Only four provided a legitimate and justified evaluative judgement; the rest we categorised as research – not evaluation. Based on our findings, we propose an updated conceptual framework we called the stairways to heaven which clarifies the steps for evaluation in comparison with research. The evaluation stairway clarifies the logic, justifications, and their relationship, integrating current resources for evaluation practice. It can be used by evaluators, evaluation commissioners, and users to clarify when evaluation is needed and get actually evaluative evaluations that connect values and data to decision-making to drive positive social change.
Indigenous evaluation literature is a powerful approach to upholding the aspirations of Indigenous Peoples. As a site of knowledge production, literature shapes the evaluation processes and practices that serve self-determined priorities: This article offers a synthesis of international Indigenous evaluation theory, to move beyond culturally competent evaluation as a framework for prioritising Indigenous values and aspirations. We acknowledge that ‘culture’ and ‘context’ remain necessary and strategic principles in evaluation design, particularly when evaluation of Indigenous programs continues to be strongly bound by funder and settler government priorities. However, we argue that such principles are insufficient, drawing instead on the broader field of Indigenous evaluation theory that centre Indigenous sovereignties, for programs that enact Indigenous aspirations. Insights from our theoretical review indicate Indigenous evaluation frameworks can play an active role in moving beyond neoliberal discourses of evaluation often adopted by governments and other funders, towards relational responsibilities that inform Indigenous approaches to validity. We take up the frameworks of relationships, relevance, and responsibility (3 Rs) as providing an important framing to foreground relational accountabilities and place-specific priorities in evaluative practice for Indigenous-led social change.
This article explores the innovative evaluation methods developed for Starlight Children’s Foundation’s palliative care program, Moments . This approach focuses on capturing feedback from children and families in a flexible, creative manner while minimising evaluation burden. A review of the literature revealed an absence of standardised tools for evaluating programs in palliative care, especially those involving children. However, research shows that participatory and child-friendly evaluation techniques enhance engagement and provide valuable insights into program impact, amplifying children’s voices in a context where they are often underrepresented. In response, Starlight developed novel gamified evaluation activity books and worksheets tailored to various developmental stages, enabling children to share their feedback in a playful, meaningful way that aligns with the program’s nature. Based on our work, we recommend (a) using gamified and participatory methods that actively involve children; (b) creating multiple formats suited to different ages using plain language; (c) applying opt-in recruitment while being mindful of timing; (d) offering clear, flexible participation options such as explainer videos and both digital and physical submissions; and (e) maintaining rigorous ethical standards. These findings add to broader research on program evaluation and highlight the importance of prioritising children’s perspectives, even in sensitive care contexts.
This article shares a novel theory-driven method for efficient qualitative data collection and participatory first-stage analysis suitable for diverse community-facing programs. Evaluations often explore more than one project – for example, a grants program evaluation may consider the activities of each individually funded project to understand the program as a whole. This presents a challenge in how to efficiently collect data that is relevant to each project that can be aggregated across the program, especially when projects are diverse. Efficiency is also a challenge in analysis and interpretation. Well-established evidence shows the value of participatory approaches in exploring impact and implementation. However, these approaches are also generally regarded as time- and resource-intensive. Faced with these challenges in evaluating the NSW Reconstruction Authority’s COVID-19 Community Connection and Wellbeing Program, we developed a novel method – Participatory Analysis Workshops. This method combines data collection at the outcome-level and exploration of implementation barriers and enablers with collaborative participatory data analysis to build a rich qualitative understanding of projects in a low-resource way. This article presents our method and this case example, outlining key features, identifying strengths and weaknesses, and suggesting modifications to enable other practitioners to implement it in their evaluation projects.
This article describes the development of a questionnaire to assess research and evaluation capacity-building support. This support was provided to public health professionals via a capacity-building partnership, the Western Australian Sexual Health and Blood-borne Virus Applied Research and Evaluation Network (SiREN). Evaluating research and evaluation capacity-building initiatives can be challenging due to their complexity. The development of the questionnaire was informed by systems concepts that acknowledged the complex nature of capacity-building, a literature review and consultation and pilot testing with stakeholders. The final questionnaire contains 17 quantitative and seven qualitative questions. Pilot testing found that the questionnaire was easy to understand and acceptable and enabled service users to provide an accurate description of changes that occurred as a result of receiving support. The development of the questionnaire provides insight into how measurement tools to assess capacity-building initiatives can reflect complexity, including the influence of contextual factors, unintended consequences and the non-linear nature of the capacity-building process. The questionnaire could be adapted to evaluate similar capacity-building efforts and the findings used to strengthen capacity-building efforts.
Program logic models are widespread and useful tools for stakeholder engagement, communication, program development and planning, and evaluation purposes. However, a key potential weakness includes the difficulty conveying the complexity associated with many interventions, which is exacerbated by the limited range of visuals used in most program logics. Program logics are almost always represented as text within rectangular boxes arranged in columns or rows, representing inputs, outputs, and outcomes; with arrows representing the linkages between them. The simplicity and consistency of this graphic language is useful but also presents a risk of over-simplifying and misleading non-evaluator users into a direct, causal set of relationships. In addition, rectangles and arrows have cultural and historical meanings that often go un-examined. This article provides a brief review of the symbology of program logics, identifies several issues with current practice, and provides several alternatives for conveying causal relationships and other aspects of program logics. It is hoped that the focus on an expanded visual vocabulary will help evaluators convey the complex reality of programs.