
Abstract Recent research assessment reforms have successfully challenged the overreliance on simplistic quantitative indicators and promoted broader approaches to evaluation. This paper argues that these reforms may also generate a new risk: assessment creep, understood as the gradual accumulation of evaluation criteria without sufficient attention to how they are prioritised and combined. While criteria such as open science, societal impact, international collaboration, leadership, and mentoring may each be valuable, their proliferation can reduce transparency and make evaluation systems more difficult to understand and scrutinise. Less transparent systems may also make it harder to distinguish legitimate expert judgement from arbitrary decision-making and to detect favouritism or other forms of unwarranted advantage. Drawing on recent debates surrounding responsible research assessment, the paper suggests that the next challenge is not choosing between metrics and expert judgement, but designing evaluation systems that remain coherent, transparent, and accountable as they become increasingly complex.
Around the world, researchers are increasingly asked, incentivized, or required to demonstrate how their scholarship can have, or has had, a positive impact on society. This reflects the assumption that research should generate value beyond academia, and the growing pressure on academic and research institutions, under political scrutiny, to show their worth. Encouraged to plan and assess the impact of their work, researchers are often guided towards logic models and language borrowed from intervention-focused evaluation and planning programmes. But there is one central, even existential, problem with this approach: research is not a programme or an intervention, and tensions emerge from efforts to plan or evaluate its potential or real impact as though it is one. Here, we focus on a practical aspect of this issue: the language that many governments, funders and institutions use to frame, understand, and train researchers to plan, achieve and assess research impact. We argue that impact planning, 'evidencing', and theories of change can oversimply the relationship between research and society, and do not adequately guide researchers towards meaningful impact. We propose alternative language and frameworks that align with recent developments in the social science literature on the research-practice-policy interface and identify the skills and resources needed to support a wider range of researchers. This shift in language will help clarify to researchers, higher education administrators, and research funders that impact is not a product to be delivered, but a possibility to be cultivated.
The elevation of research culture to a standalone evaluation dimension in REF 2029 reflects increasing recognition that research performance depends not only on outputs but also on the environments supporting research. However, the role of research culture relative to other organizational characteristics in relation to societal impact remains insufficiently understood. This study develops a quantitative approach to measure research culture based on narrative environment statements from REF 2021. Using five core dimensions of positive research culture, relevant textual evidence is identified and scored using a large language model to generate institution-UOA level indicators. We then examine whether these scores are associated with official environment ratings beyond conventional quantitative indicators. Building on these measures, we estimate regression models across 34 UOAs and compare standardized coefficients and their rankings to assess the relative importance of organizational characteristics. The results show clear disciplinary differences. Staff size and economic resources show stronger associations in science-based fields, while degree training is more prominent in social sciences and arts and humanities. Research culture shows a weaker but heterogeneous association across disciplines. Overall, the findings highlight substantial disciplinary variation and suggest the need for discipline-sensitive approaches to research evaluation and management.
Research evaluation systems place increasing emphasis on scientific impact, yet their consequences for researchers themselves remain underexplored. This paper addresses this gap by examining how scientific impact affects researcher happiness. While scientific impact is essential to progress and may be intrinsically rewarding, it also subjects researchers to intense competitive pressures, particularly when career advancement depends on performance metrics. Drawing on Self-Determination Theory, we explain how the pursuit of externally rewarded outcomes may undermine eudaimonic happiness-the deep fulfilment derived from meaningful and self-concordant activity. Using one of the largest two-wave surveys of researchers to date (over 2,400 Spanish academics), we find that scientific impact paradoxically reduces happiness. The analysis addresses potential endogeneity, common method, and desirability biases. We show that work/personal life balance mediates the negative relationship between impact and happiness. Moreover, prosocial motivation and creativity moderate this pathway. Researchers with altruistic aims or creative engagement report greater happiness, even under high-impact pressure. Notably, tenure enables creativity to exert this positive effect, underlining the importance of stable academic contexts. These findings suggest that evaluation systems centred on scientific impact may inadvertently undermine researcher well-being, with potential consequences for motivation, creativity, and the sustainability of the research enterprise.
New forms of research spanning and engaging different sectors and disciplines promise to provide holistic approaches to complex socio-ecological problems. The increasing complexity in terms of determining the composition of research teams, diversity of disciplines and protocols for the communication and translation of science is problematic, and many questions remain open regarding the evaluation of these novel endeavours between interdisciplinary research (IDR) and transdisciplinary research (TDR). This paper conceptualises and tests a participatory and reflexive evaluative approach tailored to TDR projects. The focus is on the development and testing of psychometric scales for the evaluation of TDR. A case study of a transdisciplinary (TD) collaborative science network is presented as an example of this type of research assessment and evaluation. The results of a longitudinal survey series capturing the TD scales are reported: two psychometric scales are tested; IDR values (IV) and TDR evaluation (TDE). One- and three-factor solutions were identified using confirmatory factor analysis. The authors suggest making these scales part of a holistic evaluative approach that accompanies TDR in progress as it allows evaluators to simultaneously monitor the evolution of values and attitudes towards TD alongside a project self-evaluation. The evaluation scales presented in this paper underline the importance of taking a multi-dimensional approach to research evaluation assessing the diverse dimensions of TDR, such as relevance, additionality, learning, and exchange between disciplines. This paper argues in favour of evaluative approaches that harness the learning potential stemming from constant reflection and self-evaluation in research projects.
Writing quality is no longer a reliable indicator of merit in grant review. AI tools make polished proposals common, forcing reviewers to focus on substance, feasibility, and verifiable evidence of applicants’ ideas. Well-designed systems can recognize genuine ideas rather than just carefully crafted applications.
The goal of scientific assessment is to predict which individuals can make optimal use of limited resources within a specific context to make optimal allocation decisions. In academic contexts that pertain to individual-level allocations, this is most relevant for decisions on whom to hire for academic positions, nominate for awards, or whose research projects to fund. The current perspective paper draws upon insights from decades of psychometric research and more recent research on scientific performance to derive a set of five psychometric criteria that should be met for optimal assessment procedures in academia. Although data-driven decision making has gained popularity in most domains, there is increasing resistance against using quantitative measurements in scientific assessment. Recently, several stakeholders have proposed to jettison such measurements and focus instead on qualitative indicators or narratives. We argue that both quantitative and qualitative assessment do not always meet our five criteria, but solely relying on qualitative indicators appears to be a suboptimal strategy. We argue instead that there are smarter ways to use quantitative indicators so that they become more reliable, predictive, and ultimately also more efficient and equitable. We conclude with a set of recommendations for scientific quality assessment that is based on the most recent psychometric and scientific insights. In an appendix, we apply these recommendations to a Dutch case study of how researcher information is considered in the application procedure for a prestigious individual grant.
This study investigates the capabilities of Large Language Models (LLMs) in evaluating Arabic academic research, focusing on the performance of GPT-4 and Claude across multiple assessment dimensions. Through a pilot analysis of 60 strategically selected Arabic academic papers from social sciences and humanities using institutional quality classifications, we examined the models' ability to assess research quality using criteria adapted from the Research Excellence Framework. The study employed various input formats (full text, abstract, and title) and analyzed model performance across three key components: originality, significance, and rigor. Results demonstrate distinct patterns between the models, with Claude achieving higher agreement rates in specific components (91.7% for rigor) despite poor overall correlations with human assessments but exhibiting an upward scoring bias (mean score 3.27 for full text), while ChatGPT displayed more conservative scoring patterns (mean score 2.54) with greater stability across iterations (SD 0.15-022). Both models showed stronger performance in evaluating methodological rigor (human-model correlation 0.27) compared to originality assessment (correlation -0.03). Performance degraded significantly with reduced input length, particularly for title-only evaluations (MAD increasing from 0.49 to 1.09 for ChatGPT). The findings suggest that while current LLMs show promise in supporting Arabic academic evaluation, particularly in structured assessment components, they are better suited as supplementary tools rather than standalone evaluation systems. This study contributes to understanding the potential and limitations of automated research assessment in non-English academic contexts and highlights the importance of developing culturally aware evaluation systems.
Transdisciplinary research (TDR) projects are increasingly designed and funded with the expectation that they contribute to tackling societal challenges and, consequently, generate societal impact. Yet, the relationship between TDR and societal impact remains complex and insufficiently understood. This study uses the concept of impact practices to offer a processual perspective on how societal impact creation processes unfold within TDR projects. Drawing on insights from practice theory, impact scholarship, and the TDR literature, we conceptualise impact practices as comprising three interrelated building blocks of values, strategies, and networks. Based on a real-time process study, we then identify three interrelated processes in TDR projects through which these impact practices develop: (1) emergence of proto-practices, (2) aligning of values, strategies and networks within these proto-impact practices, and (3) orienting the alignments towards transforming or reproducing societal impact. We show how project participants actively engage in these processes and align and orient a heterogeneity of impact practices within a single project, and how some become collective ways of doing for societal impact, or not. This study contributes to societal impact and TDR scholarship by providing a process model and identifying the key dynamics through which new ways of doing for societal impact creation develop in TDR projects. This is relevant for formative evaluation and management, as it makes the heterogeneous ways through which impact creation processes unfold more visible. In turn, this can contribute to stimulating reflections on impact creation in TDR projects and help make these processes more democratic.
Funding agencies that support public research are increasingly prioritizing societal impact, often in conflict with traditional output-driven academic cultures. This paper evaluates the funding process of the Science, Technology, and Innovation (STI) Hub, an initiative of India's Department of Science and Technology (DST), which seeks to foster inclusive development for marginalized communities through technology-driven solutions. Drawing on qualitative methods including interviews and focus groups with program heads and review committee members, we assessed 200 proposals against 19 program-aligned indicators. Our findings suggest a lack of alignment between proposals and program goals, highlighting the challenge of balancing the output-driven academic research culture with the need for social impact and a participatory research approach in proposal design. We observed that proposal design prioritizes perceived need instead of consultation with the concerned community, reflecting a cultural divergence from funders' expectations. The peer review assessment lacked structured criteria and transparency and indicated probable biases towards institutional and PI reputation based on publication records and legacy. We argue from a normative standpoint that the assessment framework must prioritize equity, contextual relevance, and community engagement to foster impactful research. This apparent divergence from Responsible Research Assessment (RRA) stems from what we call theory attenuation, highlighting an urgent need to redefine research evaluation frameworks and foster a culture that encourages societal-impact-driven research. While these issues mirror global critiques of peer review, our contribution lies in extending these insights to India's underexamined STEM funding ecosystem.
Over the past years, the phenomenon of 'predatory publishing' has undergone fundamental changes raising pressing methodological and conceptual challenges for its study, particularly in the context of research evaluation. The complex constellation of commercial, evaluative, and scholarly actors and logics now involved necessitates an interdisciplinary, geographically-diverse, and collaborative approach to studying scholarly - and especially 'predatory' - publishing. In this piece, we outline four key conceptual and methodological dimensions that, we argue, scholars must account for when studying this phenomenon. Firstly, the constantly changing dynamics of who and what constitutes predatory publishers and practices. Secondly, disentangling the complex relationships between evaluation and practice, accounting methodologically for the myriad factors that influence these ties, and recognising that scholarly practices are not a unidirectional effect of evaluations. Thirdly, scholars must recognise that evaluation regimes are embedded in distinct political economies of academia and that the notion of predatoriness is not universal but culturally, methodologically, and institutionally contingent. Finally, the common practice of using quantitative analyses alone to study questionable publishing practices risks reproducing existing biases and overlooking structural dynamics, and thus mixed approaches incorporating qualitative methods are necessary to ensure a nuanced understanding of the topic. We argue that scholars' approach to 'predatory publishing' crucially shapes what empirical dynamics are observed, and consequently call for scholars to take a holistic approach to studying this phenomenon.
The emergence and proliferation of knowledge brokering organisations (KBOs) as a mechanism for bridging the gap between research and policy has led to a range of studies analysing various aspects of their form and activity. However, methodological, conceptual, and practical challenges have until recently prohibited evaluation of the impact of KBOs. This has led to calls for studies to examine the multifaceted impact of KBOs, and for organisations themselves to take an active role in planning for and capturing their own impact. This paper contributes towards this agenda by describing and reflecting on the work of a KBO-the Wales Centre for Public Policy-to plan for and evaluate the impact of its knowledge brokering activities through the introduction of an Embedded Impact Researcher (EIR). A description of the EIR role is provided, followed by three reflections and recommendations from its implementation for similar organisations looking to capture their own impact. First, we argue for the development of better models of the knowledge brokering process. Second, the potential for 'mechanistic' thinking to aid the process of impact planning is explored. Third, the nature of, and tensions associated with, 'embedded' research are discussed. It is anticipated that by sharing these experiences, researchers and practitioners may identify transferable elements that may inform their own work.
Research institutions face increasing pressure to meet research funders' expectations that the research they support creates desired societal changes. Consequently, there is an increasing demand for validated methods to measure and document societal impact at all research levels. While various frameworks and approaches exist, most rely on qualitative data or tangible outcomes like patents. However, there is a lack of reliable quantitative methods to assess societal impact at the institutional and program levels. To address this gap, we developed the Societal Impact Instrument: Occupational Health and Safety Research (SII: OHSR) in 2020. This instrument generates an impact index for societal impact at the programme and institutional level, with subscales measuring reach, usability, and use of research knowledge based on input from key stakeholders. After initial refinement, the instrument has been applied annually for four years. This article validates the revised instrument, documenting that it enables research institutions to evaluate and monitor societal impact over time. Analyses offer insights into how stakeholders use research knowledge. The instrument supports impact development by identifying patterns in knowledge uptake and use. To our knowledge, this is the first instrument to quantitatively measure societal impact at the institutional and programme level based on data gathered from users. Using context-independent questions, it is designed to be easily adaptable to other comparable research institutions. With this article, we hope to encourage other research institutions to adopt the instrument as a complementary tool for evaluating and documenting societal impact, alongside existing approaches.
In the evaluation of research performance of Higher Education Institutions (HEIs), institutional benchmarking has been proposed as better alternative to rankings since the focal institution is compared with HEIs similar for some core characteristics. The literature advocates an hybrid approach to peer selection, where, first, potential peers are identified based on quantitative data and, second, experts make the final selection based on in-depth knowledge of the institution and of the purpose of the comparison. However, for what concerns subject composition, a fundamental characteristic differentiating HEIs, there is no systematic methodology to operationalize this strategy. In this paper, we address this gap by mobilizing two sources of quantitative data on HEIs' subject profiles, i.e. the distribution of students by fields of education and the journal subject distribution of scholarly publications. We compare these two distributions on a large sample of European HEIs derived from the European Tertiary Education Register enriched with data from the SCImago Institutions Ranking. The comparative analyses highlight that the two distributions are largely complementary and that a combined index yields good results to identify potential peers for focal institutions. An in-depth investigation of exemplary HEIs shows that, given the complexity of the HEIs' landscape, quantitative methods can only suggest potential peers in a first step that should be checked by experts based on more specific information in a second step.
The expanding use of Large Language Models (LLMs) in scientific writing raises critical concerns regarding research accountability, integrity, and transparency. This study aimed to develop and content-validate the CAREFUL-AI framework and apply it to evaluate the scientific quality of AI-generated research manuscripts. A three-round Delphi process with 10 multidisciplinary experts was conducted to develop the CAREFUL-AI framework. Content validity was assessed using item-level and scale-level content validity indices. Using standardized, locked prompts, 150 complete scientific research manuscripts across five study designs were generated by six freely accessible LLM platforms ChatGPT (GPT-5.2), Claude (Claude 4.5), Gemini (Gemini 3 Flash), Grok (Grok 4.1), DeepSeek (DeepSeek-V3.2), and Meta AI (Llama 4 Maverick). Manuscripts were evaluated by the respective prompt administrators using the CAREFUL-AI framework. The framework demonstrated strong content validity (I-CVI: 0.83-1.00; S-CVI/Ave: 0.92). Overall, 38.0% of manuscripts were rated high quality, 52.7% moderate quality, and 9.3% low quality. Claude generated the highest proportion of high-quality manuscripts, while Meta AI produced all low-quality outputs. Reproducibility and uncertainty handling consistently received the lowest domain scores across models. The CAREFUL-AI framework provides researchers with a structured tool to critically appraise AI-generated manuscripts, helping safeguard methodological rigor, transparency, and evidence reliability in scientific writing. Substantial variability exists in the quality of AI-generated manuscripts. Although Claude demonstrated comparatively stronger performance, persistent deficiencies in key integrity domains indicate that AI-assisted manuscript generation requires robust human oversight. CAREFUL-AI provides a content-validated framework to support ethical governance and editorial oversight of AI-assisted research.
The rapid integration of large language models (LLMs) into scientific research raises a fundamental question for research evaluation: what is being assessed when core research activities become partially automated? Although prior studies consistently report efficiency gains, existing evaluation practices do not fully capture changes in verification behavior, epistemic reliability, and dependency risk. This paper draws on a systematic review of 90 peer-reviewed articles to examine how LLM-assisted research activities are evaluated across four workflow stages: search, screening, summarizing, and drafting. The results indicate a systematic asymmetry in evaluation practices. Measures of time savings and productivity are widely reported and typically positive, whereas verification practices, trust calibration, and reliability concerns are rarely specified or directly measured, particularly in high-transformation stages such as summarizing and drafting. These patterns suggest that prevailing evaluation approaches tend to emphasize efficiency while leaving epistemic safeguards insufficiently articulated. In response, this study develops a stage-sensitive evaluation framework that distinguishes between low- and high-transformation research activities and specifies proportionate verification checkpoints. By framing LLM-assisted research as a problem of evaluation design rather than merely productivity enhancement, the study contributes to research evaluation theory and practice.
This study investigates the impact of the COVID-19 pandemic on research productivity in Brazil using a scientometrics approach. We analyze data from 134,492 researchers registered on the Lattes platform, focusing on publications, supervisions, research projects, and conference participation between 2010 and 2022. As expected, our findings reveal a significant decrease in productivity in all measures following the onset of the pandemic in 2020. Despite this overall decrease, there was a surge in research on COVID-19-related topics, particularly in the health sciences. We also found that, while gender did not significantly influence productivity changes, researchers with PQ scholarships and those with more than 19 years of experience since completing their PhD were disproportionately affected. These results highlight the need for policies to support these vulnerable groups during global crises.
Funding allocation in academia reflects broader disciplinary hierarchies and systemic inequalities, particularly within the Social Sciences and Humanities (SSH). This study investigates patterns of funding acknowledgments and funding concentration in the SSH at two universities (University of Copenhagen; University of Toronto) over a 20-year period (2002-22). Using Scopus-indexed journal article data, we analyse the proportion of acknowledged funded versus non-funded research, the distribution of funding sources, and the extent to which funding is concentrated towards specific authors and research topics. Our findings reveal significant disparities in funding allocation, with a small subset of authors benefitting from a disproportionate share of funding. We also observe a strong correlation between interdisciplinary research and funding success, suggesting that SSH scholars have enhanced their funding prospects by aligning their work with applied, policy-relevant, or STEM-adjacent domains.
Altmetrics need to be more critically assessed in terms of the extent to which they reflect impact and quality of research compared to popularity or mere attention. Twitter (now rebranded as X) is a popular platform to, among other things, discuss and share scientific articles. Earlier altmetric studies have often focused on investigating whether the number of tweets mentioning scientific articles could be used as an indicator of scientific impact or attention, with results showing weak to moderate correlations with citation counts. But all tweets may not be equal, as original tweets and retweets may reflect different levels of engagement and impact. Using a dataset of over 330,000 PLOS publications, this study explores whether these two forms of Twitter activity correlate differently with traditional citation metrics and how these relationships vary across disciplines. The findings showed the correlation between citations and original tweets was consistently higher than that between citations and retweets and significant weak or moderate, but higher in Social Science and Humanities than in Natural Science, Engineering and Medicine fields. Also, including zero citation counts improved the correlation coefficients for original tweets, but reduced that of retweets. This indicates that original tweets may be more aligned with citation counts as an indicator of scholarly impact, whereas retweets might reflect broader dissemination and popularity. In conclusion, tweets and retweets are different altmetric indicators and should be considered as two different metrics and analysed separately.
Early-career researchers (ECRs) face a critical and uncertain transition from doctoral training to independent scholarship, shaped by disciplinary norms, research infrastructure, and institutional support. As competitive grants increasingly replace faculty appointments as the pivotal route to research independence, this study reconceptualizes time-to-position (TTP) as the interval between PhD completion and the first PI-eligible competitive award. This reframing captures the structural onset of independent research authority within contemporary funding systems. Using data on 1,714 scholars who secured funding from Taiwan's National Science and Technology Council (NSTC), spanning diverse disciplines and institution types, the study examines cross-field differences in research independence, compares funding outcomes between ECRs and non-ECRs, and evaluates how disciplinary, institutional, and productivity factors jointly shape TTP. Results show marked and systematic disparities: institutional conditions exert the strongest influence on TTP, while scholarly productivity plays a secondary but significant role. These findings underscore that research independence is co-produced by individual outputs and multi-layered organizational environments. Although the analysis is limited to funded researchers, TTP provides a policy-relevant diagnostic of how efficiently and equitably research systems integrate emerging scholars. The indicator offers practical value for assessing funding architectures and designing evidence-informed early-career support mechanisms.