
There is strong evidence that social conditions drive health outcomes and disparities. Consequently, there is scientific and policy interest in understanding the effectiveness of interventions designed to ameliorate such social conditions, with particular attention on place-based interventions and reducing inequalities. However, neither the science nor the policy interest is well served by conventional evaluation frameworks and, furthermore, few evaluations show that impacts have consistently been achieved. In short, current approaches are ill-suited to either proving what has worked or improving what might work. Transferable lessons are often limited. In this context, recent public health literature is helpful in relation to approaches to evaluation framework-building and theory-building. We describe an alternative approach, in contrast to contribution, attribution or developmental approaches, as a ‘social-centric discovery’’ approach. This explores the dynamic relationship between social context and an evolving intervention viewed as a sequence of events within a complex social system.
Background: Behavioral health is a critical issue in the fire service, with firefighters facing high rates of depression, PTSD, and substance misuse due to repeated trauma exposure. Stress First Aid (SFA), a peer-support intervention adapted from the Marine Corps’ Combat and Operational Stress First Aid framework, offers a promising approach to firefighter mental health. Methods: A mixed-methods evaluation of the Enhanced SFA program (ESFA) assessed implementation, effectiveness, and acceptability through in-person pilot sessions, beta tests, and an online program. Surveys were conducted pre- and post-training, and at 3- and 6-month follow-ups to measure usability, satisfaction, skill application, confidence to intervene, and perceived organizational support. Results: Of 1,500 enrollees, 1,199 completed initial assessments. Participants strongly endorsed ESFA’s relevance and format, with 83% intending to apply learned skills. Confidence to intervene improved from 77% at baseline to 90% at 3-month follow-up. Perceived organizational support also increased. While statistical analysis was limited to descriptive methods due to attrition, the results suggest meaningful engagement with training. Conclusion: ESFA demonstrated high usability and potential for strengthening peer support and intervention confidence. Additional research is needed to assess long-term outcomes, but ESFA offers a scalable foundation for firefighter behavioral health initiatives.
This reflective essay revisits a period of institutional reform within a public sector institution in East Africa, where evaluation capacity (ECB) practices emerged through routine, embedded practices in daily operations. Using Preskill and Boyle’s (2008) model as a lens, the essay reinterprets these embedded practices to make the implicit explicit. Three domains are highlighted: creating supportive structures and culture, building communities of practice, and integrating evaluation into planning and decision-making, all reinforced by leadership that modeled and sustained evaluative habits. The reflection also surfaces the tensions between theory and practice, showing how ECB unfolds in complex institutional realities, where tensions can open spaces for learning and adaptation. Ultimately, the essay contributes to ECB scholarship by demonstrating how capacity can develop organically through institutional routines and by bridging practitioner experience with scholarly frameworks.
Background: Subjective wellbeing valuation is a nascent monetization method designed to address existing challenges in valuation practice. Wellbeing valuation is codified in impact measurement evaluation policy in countries such as the UK, New Zealand, and Australia. Purpose: This article shares the development of wellbeing valuations using US-specific data and describes the use of these valuations in a Social Return on Investment (SROI) impact evaluation of an arts-based social enterprise in Appalachia. This case illustrates the promise and challenges of integrating subjective wellbeing into impact measurement to capture outcomes prioritized by stakeholders. Setting: Art-based social enterprise in Appalachian Ohio Intervention: This social enterprise employs adults with developmental differences as artists through the Creative Abundance Model. Research Design: Social Return on Investment (SROI) evaluation; exploratory mixed methods to construct outcome chains and evidence outcomes; subjective wellbeing valuation using US-primary data and US-derived WELLBY to monetize outcomes most important to stakeholders Data Collection and Analysis: Focus groups with stakeholders, ripple effects mapping focus groups, document analysis, stakeholder surveys. Transcripts from focus groups were thematically coded using emergent coding to construct outcome chains. Qualitative data was triangulated with document analysis and quantitative data. Subjective wellbeing valuation data was sourced from MIDUS Refresher I data and Neumman’s (2014) QALY ranges. Findings: The most important outcome to three stakeholders in the Passion Works case study were found to be wellbeing outcomes. The monetization of those three outcomes using subjective wellbeing valuation replicating the WELLBY method, calculated with representative data from the US yielded values that were then discounted using SROI best practice for accounting for value. Overall, these three wellbeing outcomes alone represent between $654,377.10 and $981,139.20 of value.
Background: Narrowly focused pre-employment programs that evaluate impact in terms of finding work have been considered inadequate for meeting the realities and needs of urban Indigenous youth. Purpose: This article focuses on Odabi (“roots” in Anicinabe), a skills and employment program for youth created and offered by an urban Indigenous organization in Quebec Setting: The study is part of a 4-year community-university partnership with [removed for anonymity] Research design: An Indigenous evaluation framework served to examine the evolution and program’s impact from the participants’ perspective. Data collection & analysis: Informing the study was a combination of participant observation, participatory group activities, sharing circles, and individual interview with program participants (25) and staff (11). Findings: The evaluation research shows that reshaping the program involved broadening the definition of success to encompass a comprehensive Indigenous view of well-being that is greater than finding and maintaining a job. Participants tell the story of what counts in terms of the program being rooted in the collective and group aspect, suggesting that the technical aspects of employment should not be at the program’s core, but rather the relational component and the strengthening of the sense of identity and pride.
Background: Common data collection strategies create barriers to authentically elevating youth voice, which is critical to understanding the complexity of youth’s involvement in structured programs. These barriers may be especially profound when asking youth to share about their positive and negative experiences in programs, including reasons why they choose to return or not return. Purpose: The purpose of this paper is to highlight the use of journey mapping as a potentially effective approach to engaging youth in the continuous improvement process regarding a potentially challenging topic. Setting: The evaluation took place with youth participants from a Canadian summer camp organization serving youth from low-income backgrounds. Intervention: The evaluation method employed creative methods, specifically journey mapping, and semi-structured interviews. Research Design: Retrospective and qualitative. Data Collection and Analysis: Qualitative data were interpreted using a six-step thematic analysis approach. Findings: The findings of this case study suggest that creative methods, such as journey mapping, can be an effective approach to engaging youth from marginalized backgrounds in continuous improvement and evaluation.
Evaluation scholars and practitioners have argued that having more people who can think evaluatively is “essential” for successful evaluation practice, organizational improvement, and even healthy democracies. However, empirical evidence backing those claims is scant. Making a case for investing more resources into building people’s capacity to think evaluatively requires a better understanding of its impact and added value. We need a way to measure it. While two scales to measure evaluative thinking exist, both have shortcomings and neither has been used to build a body of evidence concerning the contributions of evaluative thinking to the evaluation field and beyond. A two-stage, mixed-methods study was conducted to create a new, reliable and valid evaluative thinking scale. Stage One identified dimensions of evaluative thinking through literature review and expert interviews. Stage Two generated and refined items to measure these dimensions, confirmed face validity through focus groups, and administered the refined items to 250 random, yet screened, participants. Factor analysis revealed six factors with strong item loadings, and a Cronbach’s alpha test confirmed reliability (α = 0.96). This study enhances our capacity to measure evaluative thinking and lays the groundwork for future research into understanding its impact more deeply.
Oral storytelling methods, such as oral history, are not often incorporated into program evaluations; as evaluators, we tend to develop interview protocols that are more structured. In this paper, I explain how evaluators can use Personal Interwoven Narratives (PIN), a method adapted from oral history, as a more open-ended approach to collect richer, fuller data than more-structured interviews tend to provide. Though not appropriate for all evaluations, PIN can greatly benefit evaluations that aim to honor stories and raise the voices of marginalized or otherwise underheard groups. I will describe how the Personal Interwoven Narratives method developed from the established oral history technique into an oral storytelling method that specifically meets evaluators’ needs, what types of evaluations it most benefits, and best practices and considerations for evaluators who seek to use this new method.
Background: Empowerment evaluation (EE) is a participant-led evaluation model that aligns with restorative justice (RJ) principles by centering the needs of the person or persons impacted. Purpose: This article shares two longitudinal case studies, one using a practical empowerment evaluation approach with six U.S. Catholic institutions of higher education (IHEs), and another using a transformative empowerment evaluation approach with a cohort of 8 K-12 educators. These cases illustrate the benefits and challenges of utilizing empowerment evaluation for restorative justice programs in educational settings. Setting: U.S. institutions of higher education (IHEs) and U.S. K-12 classrooms. Intervention: Empowerment evaluation as a method to evaluate restorative justice implementation in multiple K-16 contexts. Research Design: Two longitudinal qualitative case studies. Data Collection and Analysis: Restorative circle dialogues; individual interviews; data collected during webinars; and an e-gallery walk provided data. Transcripts from circles and individual interviews were reviewed and organized by key themes to illustrate the benefits and challenges of using empowerment evaluation. Findings: Empowerment evaluation can be conceptualized as a practical approach to evaluation, like formative evaluation, which aligns with the values of restorative justice, and focuses on program improvement. Transformative empowerment evaluation emphasizes liberation by empowering participants to step outside of typical roles, traditional structures, and assumed power relations. Due to the variation in these evaluation approaches, the case studies below vary in the emphasis placed on aspects of the evaluation process. Yet, both case studies highlighted participants’ needs for dedicated time for thinking and planning for restorative interventions; peer-to-peer support from colleagues from one’s own school or campus and from other campuses; sharing strategies for planting the “seeds” of restorative justice; and evaluation as an ongoing process of reflection and action.
Evaluation theory is considered integral to good evaluation. There is, however, a lack of clarity on the distinction between prescriptive evaluation theory and evaluation approaches and perspectives. The distinction is further complicated by the central role of program theory in the praxis of evaluation. Notwithstanding, evaluation theory and theorists are popularly codified in the Alkin tree (Alkin & Christie, 2004; 2006; 2008) and presented in introductory evaluation classes in graduate programs. While the Alkin tree has seen several revisions, few female evaluators and even fewer evaluators of color are represented. In this study, we (three Black female evaluation graduate students) develop Critical-PostColonial Theory (CPT) as the analytical framework to conduct an autoethnography, interrogating and reflecting on the teaching and learning of evaluation theory in introductory program evaluation graduate classes. The paper concludes with suggestions for decolonizing the teaching and learning of evaluation theory within graduate evaluation programs and courses.
Every day in the United States, we use language that is oppressive and rooted in our colonial past. For this article, we specifically focus on the term stakeholder regarding evaluation research. Evaluators have defined this term as: “all of those individuals who have an interest (i.e., are somehow vested in) in the program that is to be evaluated." (Alkin & Vo, 2018, p. 51). However, little attention has been given to changing or reshaping this term to avoid perpetuating inequities to make program evaluations more culturally responsive to all parties involved. Historically, this terminology rooted in a colonial perspective is defined as “the person who drove a stake into the land to demarcate the land s/he was occupying/stealing from Indigenous territories” (Phipps, 2022). Even though the term is used frequently, many people may not understand the nature of the roots of this term and its involvement in the oppression of persons of color. Because of these points, it has been suggested that new terms be used in its stead (Sharfstein, 2016). We ask, after so many years of this country being independent of colonial rule, why are we still using terms that are oppressive and hold a connotation of being non-inclusive? This article aims to explore how terminology has been used in research and offers ideas to consider how to move forward with ensuring that the language used in program evaluation and other areas of research is culturally responsive to people and organizations affected by or show an interest in the findings of a given assessment.
This article explores the use of photovoice as an evaluation method in a high school summer bridge program (the Innovators program). Photovoice is a participatory research method where participants take photographs and share narratives in response to prompts that solicit their perspectives and experiences. With the goal of promoting critical dialogue between participants, photovoice method leverages the power of emic knowledge and visual representation to influence programmatic change. We explain how photovoice was implemented in a K-12 mixed methods evaluation and the practical and ethical considerations that evaluators must weigh. Findings indicate that photovoice not only enhances participant engagement but also serves as a valuable tool for program evaluation, yielding rich, qualitative data that traditional methods may overlook. Drawing on a multi-year evaluation of the Innovators program, the article contributes to the growing literature on photovoice by examining what the method means for young participants, in addition to describing the photovoice design and implementation processes.
This paper examines how theory is conveyed to novice evaluators in university coursework, providing a new perspective on the role of theory in evaluation. Using a survey of university evaluation instructors in the United States, the authors examined instructor perspectives on the role of evaluation theory as well as the evaluation content in assigned textbooks. Analysis indicated a similar lack of coherence and consistency in the use of theory to that seen in evaluation practice. The findings support three takeaways for evaluation educators, practitioners, and scholars. First, the field must continue to clarify the role of theory in evaluation. Second, an in-depth analysis of authors’ approaches to discussing theory in textbooks is warranted. Third, textbook authors and publishers could consider the themes that emerged from instructors’ responses to inform book proposals and revisions of textbook editions to better suit the range of instructor needs.
Many evaluators may be conducting and presenting research on evaluation (RoE) without realizing it. Proposals accepted to present at the AEA 2019 conference were analyzed for whether or not they were RoE (i.e., systematic, empirical, and focused on evaluation). RoE proposals were then coded using a framework by Mark (2008) to determine the common trends across those identified as RoE. Non-RoE proposals were coded for why they were not RoE, including that they were not systematic, empirical, and/or focused on evaluation. A total of 15% of the 732 proposals analyzed were coded as RoE; most proposals examined evaluation activities using the descriptive mode of inquiry. Furthermore, most non-RoE proposals were coded as not RoE because they presented an evaluation rather than a research study on evaluation or because they were reflections on evaluation. This study presents multiple opportunities and future directions for RoE.
The Caribbean encompasses an archipelago of islands with stunning beauty, tropical climates, and beautiful beaches. The region faces many developmental challenges caused by small size, undiversified economies, and regular exposure to various climate and nature induced hazards especially hurricanes which continuously impede developmental progress. The COVID-19 pandemic has worsened the region’s economic quandary, and derailed progress towards the United Nations Sustainable Development Goals (SDG) 2030 Agenda. Like many other parts of the world, the region has not yet developed a culture for monitoring and evaluation (M&E). The region’s development progress is also challenged by insufficient evidence-based data, inadequate disaggregated statistics, and dated national statistical systems. In 2019, governments of the region committed and adopted a results-based management (RBM) policy initiated by the CARICOM Secretariat, the most influential regional body. However, four years later, the implementation of this policy is still very much in an embryonic phase. This paper will summarize the current status of M&E in the region, examine the region’s progress towards Agenda 2030, and discuss why transformation thinking is important in the region’s quest for prosperity, resilience, and sustainability.
Background: Evaluators increasingly consider systems- and complexity-informed approaches to evaluation, especially when considering how evaluation might be transformed to evaluate current complex problems. Although Gregory Bateson was an early contributor to systems thinking, there is almost no reference to his work in the evaluation literature. Purpose: To introduce some of Bateson’s core ideas and to pose initial questions intended to spark reflection and discussion, with the intent of contributing to further development of the concept of evaluative thinking. Setting: Global. Intervention: Not applicable. Research Design: Not applicable. Data Collection and Analysis: Not applicable. Findings: Not applicable.
Background: This article is part of a collection recognizing, appreciating, and celebrating the substantial lifetime evaluation contributions of Ray C. Rist. Purpose: This article presents 10 overarching themes from the impressive body of Ray Rist’s published works. Setting: I write as a long-time colleague of and collaborator with Ray C. Rist. Intervention: N/a Research design: N/a Data Collection and Analysis: In the introductory article, the authors offered an inventory of Ray Rist’s published works, some 34 books and 159 articles and chapters. I undertook a systematic qualitative thematic analysis of those publications as an experienced qualitative analyst. I brought to the identification of themes, and their importance, my own knowledge of evaluation as an experienced evaluation practitioner, theorist, and author. Findings: Ten overarching themes from the impressive body of Ray Rist’s published works: (1) focusing evaluation on results not just activities; (2) taking a systems perspective on M&E; (3) engaging in comparative analysis; (4) valuing methodological diversity and rigor; (5) policy evaluation as a distinct and important focus of evaluation; (6) evaluation capacity building; (7) evaluation serving and advancing social justice; (8) editing prowess as a contribution to sharpen communications and enhance evaluation use; (9) working collaboratively; and (10) addressing and synthesizing leading edge issues. This overview concludes with the challenge of transforming evaluation to evaluate transformation. Ray Rist has long been in the forefront with prescient writings spotting, identifying, and naming transformational trends with implications for evaluation.
Background: The book From Studies to Streams was for me an eye-opener when I worked as Director of the Independent Evaluation Office of the Global Environment Facility (GEF). Right from the start in that position I was working on gathering as much evaluative evidence as we could, mix this with knowledge and insight, and look at how to deliver recommendations and insights to the GEF. From Studies to Streams inspired me to work towards a potential mountain of evidence to inform and inspire the replenishment meetings of the GEF. While the book provided an analogy of evaluation insights and evidence streaming down to the ocean, I felt that the knowledge gathered in the Overall Performance Studies would be more moving up, to reach higher levels of decision-making. At the top of the mountain of evidence and insight, the GEF replenishment meetings would decide on the goals of the Fund in the next four years. Purpose: The chapter aims to provide a historically accurate account of how evaluations at different levels of GEF funded activities were used to inform higher-level evaluations, leading to an integrative perspective of achievements. Evaluations incorporating both scientific and national/local perspectives aimed to capture findings and insights at all levels of the GEF and its partners. While this was not the stream downwards to the ocean, it could be likened to steam that gradually wafted up to the pinnacle of GEF decision-making. Setting: The partnership of the GEF with multilateral banks, five UN organisations and the 120 recipient countries of GEF funding. While all agencies and countries had their own evaluation policies, agreement was reached about a minimum number of common elements that would be reported on. All evaluations that touched upon GEF issues were studied to subtract insights relevant for the GEF. This was combined with knowledge generated through the Scientific and Technical Advisory Panel (STAP) of the GEF, and any relevant knowledge available through literature and expertise. Intervention: Not applicable. Research Design: Not applicable. Data Collection and Analysis: Not applicable. Findings: It turned out to be possible to create a flow of evaluative evidence and other insights and knowledge regarding interventions from the many varieties of evaluation that were undertaken in the GEF and its many partners. This led to a veritable mountain of evidence that was presented every four years to the replenishment meetings for the GEF. While the GEF is a relatively unique international funding organisation, it turned out to be possible to make use of the evidence generated at various levels of the partnership and by different actors, along the lines of the Studies to Streams book.
Background: When the covid 19 pandemic started spreading early 2020 Governments responded in various ways. The merits and drawbacks of national responses is not an academic concern, it was – and remains - a question of survival. Purpose: The purpose of this paper is to analyze the evaluation of the Swedish national response to the pandemic and to assess whether the evaluation provided for accountability of the policy measures that were put in place. Setting: The Swedish Government announced early on that its response was to be evaluated. A Parliamentary Committee was established and was given a comprehensive mandate to evaluate the process and the results of the response. The Committee was to start immediately in mid 2020, to deliver interim reports and a final synthesis report in February 2022. Intervention: Not applicable. Research design: We use a case study design based on a desk study of written documentation concerning the covid 19 evaluation. Our study starts with publications early 2020 and up to the final synthesis report of the evaluation and continues with events/debates through the general elections in September 2022 (when the Government responsible for the covid 19 response lost) and the months immediately afterwards. Data collection and analysis: The key sources are the public mandate for the evaluation, its three evaluation reports, records of the debate in daily papers and professional journals, and the autobiographies of leading actors. Findings: The political/administrative system initiated an evaluation that gave a timely, credible and comprehensive assessment of the virtues and mistakes of the Government’s response to the pandemic. Still, the question of accountability remains elusive. Structures that constrained the response were shaped long ago and the actors responsible cannot be held accountable today. Those that can be held to account made mistakes but also took brave, and in retrospect correct measures to reduce the impact of the pandemic. In addition, new information keeps changing the final judgement, for example the impact of business subsidies is better known today. In sum, the information needed to create accountability was – and is – largely available, but to establish accountability remains an elusive task.
Background: This paper reflects on my long-standing collaboration with Ray C. Rist, which began in 1997 and has evolved through decades of shared work in evaluation. Grounded in international development and governance, our collaboration has explored both tangible and intangible dimensions of evaluation to enhance institutional learning and policymaking. Purpose: The study examines key lessons from this collaboration, emphasizing the role of evaluation in organizational transformation, the interplay of measurable and relational factors, and the strategic use of evaluative knowledge in decision-making. Setting: The paper draws on experiences from diverse global contexts, including international development initiatives, university-industry partnerships, and public sector reforms. It situates evaluation within broader theoretical debates on governance and knowledge production. Data Collection: The analysis is informed by direct involvement in evaluation projects, fieldwork across multiple regions, and scholarly contributions, including the Comparative Policy Evaluation series. Insights are drawn from case studies, interviews, and a review of evaluation practices applied in different institutional settings. Findings: The study identifies four critical lessons: (1) the necessity of grounding evaluations in local realities to ensure relevance, (2) the importance of recognizing both tangible and intangible dimensions in program success, (3) the role of evaluative knowledge in fostering reflection and policy change, and (4) the advancement of evaluation scholarship to bridge theory and practice. These findings reinforce the evolving role of evaluation as a tool for navigating complexity, strengthening institutional capacity, and fostering inclusive governance.