
Data availability statements (DAS) are increasingly required by journals to promote transparency and reproducibility. However, it remains unclear whether their growing adoption has led to greater accessibility of research data. This study assessed trends in data accessibility, compliance with DAS, and the relationship between data sharing and citation impact in oceanographic research. A total of 1 400 oceanography articles indexed in the Web of Science Core Collection were randomly selected from 2018–2024 (200 articles per year). Bibliometric characteristics, DAS presence, and actual data accessibility were recorded for each article. Data availability was classified as publicly available, available upon request, partially available, or unavailable. Corresponding authors of articles without publicly available data were contacted to evaluate data availability upon request. Associations between data availability and citation impact were examined using regression analyses. The proportion of articles containing a DAS increased from 11
Heightened scrutiny in peer review of scholarship that challenges dominant perspectives has made it harder for scholars engaged in such work to contribute to advancing knowledge. In this study, we highlight how biases in the peer review process contribute to epistemic exclusion – the devaluing of marginalized scholars as legitimate producers of knowledge as well as the systematic dismissal of non-traditional forms of scholarship including, qualitative, identity-focused, localized, non-positivist, or non-generalizable research. We do so by characterizing some of the most commonly reported forms of epistemic exclusion conveyed in journal article peer reviews. Using a mixed-methods approach, we examined the review experiences of 74 researchers who conduct research with human participants by considering the relationship between manuscript subject matter and author social identity and the relative incidence of different types of exclusionary feedback authors received. Then, using participants’ reports of specific instances of problematic reviewer feedback, we identified assumptions underlying reviewer and editor comments. Significant associations among distinct types of epistemic exclusion were found based on authors' English-speaking status, racial identity, and sexual orientation. A reflexive thematic analysis uncovered an overarching theme, Science is not for you and service is not for me, and four additional themes: (1) I know science better than you, (2) My people are better than yours, (3) I can't be bothered to teach you and (4) What's in it for me? These findings suggest the need for educational programs aimed at unpacking and challenging reviewer and editor assumptions that contribute to biased and negative reviews, and ultimately to epistemic exclusion.
Abstract Background Generative AI is increasingly used in scholarly research, writing, and publication workflows. Many journal and publisher policies ask authors to disclose relevant AI use, but disclosure alone rarely clarifies how AI-assisted work was performed, what information was entered, how outputs were evaluated, or how human responsibility was maintained. This creates a gap between AI-use disclosure as a publication requirement and AI-use documentation as an open science practice. Methods This article develops a conceptual and practical framework for documenting generative AI use in scholarly workflows. The framework was informed by exploratory, non-systematic source and policy mapping, AI-assisted exploratory evidence mapping, manual review of selected recent literature, and development of accompanying Open Science Framework materials. These steps were used to identify recurring documentation expectations and unresolved policy gaps and to translate them into practical documentation domains, with particular attention to task specificity, proportionality, role-specific documentation, privacy-sensitive transparency, and clinically sensitive contexts. Results The framework distinguishes disclosure from documentation and proposes documentation fields for minimal and extended AI use. Minimal documentation is intended for low-risk uses such as limited language polishing, whereas extended documentation is recommended when AI supports literature synthesis, coding, analysis, interpretation, manuscript drafting, peer-review-related work, clinical material, or research procedures. The framework is accompanied by reusable OSF materials, including documentation templates, prompt-log structures, declaration examples, checklists, clinical redaction guidance, and source-tracking materials. Conclusions AI-use disclosure communicates that AI was used; documentation makes the AI-assisted workflow traceable, inspectable, and accountable. A task-specific, proportionate, role-specific, and privacy-sensitive documentation approach can support responsible AI use while protecting confidential, patient-related, peer-review-related, and methodologically sensitive information. The accompanying bilingual materials are openly available on OSF: https://doi.org/10.17605/OSF.IO/A439J .
Breaches of research integrity have sparked interest in the factors that may help explain when research misbehavior is more likely to occur. Often three clusters of factors are distinguished: individual factors, climate factors and publication factors. Our research question is: to what extent can individual, climate and publication factors explain the variance in frequently perceived research misbehaviors? We used validated measurement instruments for these three clusters of factors to survey academic researchers in Amsterdam. Results showed that individual, climate and publication factors combined explain 32% of variance in perceived frequency of research misbehavior. The cluster accounting for the greatest percentage of explained variance was the research climate (23%). The research climate factors included in our study concern perceptions of specific dimensions of the academic organization, such as the existence of research-related norms and socialization activities into responsible research practices within a department, and the quality of resources an institute has available to support researchers in their work. Our results underscore the important role of the research climate in fostering responsible research practices and suggest that the frequency of research misbehaviors might be lowered by putting more emphasis on the socialization into ethical departmental norms and creating an open departmental atmosphere.
Abstract Background Following our previous investigation of 456 articles published by the IHU-Méditerranée Infection (IHU-Mi) institute in Marseille, France, subsequent editorial actions, institutional reviews, and reports from French authorities provided additional evidence relevant to our concerns. We therefore extended our investigation to the broader IHU-Mi bibliography to assess potential breaches of the Declaration of Helsinki, French regulations governing research involving human participants, and the consistency between reported studies and the ethics approvals cited. Methods We assessed 2,674 IHU-MI publications identified through PubMed, Google Scholar, PubPeer, and manual searches. We excluded articles not reporting research involving human participants, articles clearly reporting retrospective studies without ethics flaw, and articles for which editorial investigations had clarified our concerns. We compared the remaining articles with the relevant French legal framework and, where available, with ethics approvals obtained through public transparency procedures or other sources. Results We identified ethical or legal concerns in 853 articles from this institute. These included at least 75 articles with retrospective approvals, 194 describing research conducted abroad without mention of local approvals, 343 not mentioning ethics approval, and 426 describing sample collection from healthy volunteers with concerns. Among 374 articles for which we obtained the cited ethics approval documents, at least 304, or 81.3%, described sample collections that were not covered by the corresponding approval. Since our previous investigation, editorial and institutional responses have led to 64 retractions and 270 expressions of concern, and have clarified or substantiated many of the concerns raised. We also found that the local ethics committee issuing most of the approvals may not have been sufficiently independent from the authors. Conclusion Institutional reports and editorial decisions have substantiated many of the concerns identified in our analysis, although a small subset of cases was attributable to unintentional errors. Our findings indicate widespread and longstanding problems in research ethics compliance within the IHU-Mi bibliography and support the need for continued editorial and institutional review.
The growing use of large language models (LLMs) across research workflows, including literature review, evidence synthesis, peer-review support, citation validation, and research evaluation, has intensified debates about research integrity, accountability, transparency, and the boundary between legitimate assistance and inappropriate delegation. While LLMs may enhance the efficiency and semantic depth of scientometric and scholarly review practices, their use also raises concerns about hallucinated evidence, fabricated references, biased evaluation, unclear authorship, inadequate disclosure, and the erosion of human responsibility. This study conducts a workflow-oriented systematic scoping review with qualitative thematic synthesis to map and critically synthesize the emerging literature on LLMs in scientometrics and scholarly review processes, rather than a comprehensive review of AI and research integrity as a whole. Following PRISMA-informed procedures, searches were conducted in Scopus, the Web of Science Core Collection, and IEEE Xplore for English-language publications published from 2023 to 2025. After eligibility-based screening and backward citation searching, 74 studies were included as an analytical sample. The review maps five conceptual and methodological approaches: human-in-the-loop and hybrid architectures, retrieval-augmented generation, autonomous and multi-agent research systems, semantic and ontological integration, and prompt engineering with structured evaluation frameworks. It also identifies five major application clusters, review automation and intelligent screening, automated review generation and agentic research workflows, quality assessment and peer-review support, data extraction and semantic classification, and citation analysis and source validation, alongside a governance-oriented category addressing disclosure, confidentiality, accountability, and responsible-use frameworks. The scoping synthesis suggests that retrieval-grounded, domain-adapted, and human-supervised systems offer more reliable pathways than unsupported general-purpose generation, especially for citation-sensitive, evaluative, and high-stakes scholarly tasks. Quantitative performance indicators reported in individual studies were treated descriptively as part of mapping evaluation practices, rather than as directly comparable outcomes for pooled estimation. As a normative implication of the scoping synthesis, the review argues that responsible LLM use should be framed as accountable human–machine collaboration grounded in traceable evidence, transparent disclosure, domain-specific validation, fairness auditing, reproducibility, and non-delegable human responsibility.
Transparency and openness are core scientific values that enable independent verification of reported research results. The Transparency and Openness Promotion (TOP) Guidelines, first published in 2015 (TOP 2015), provided a flexible framework for journals to design and implement publication standards promoting these values. Over the past decade, widespread use of TOP 2015 contributed to valuable evidence and feedback for refinement. The TOP Advisory Board undertook a comprehensive update of TOP 2015 that went into effect in 2025 (TOP 2025). The update introduces an explicit conceptual framework clarifying its primary objective: improving the verifiability of empirical research claims. Standards are reorganized into three categories (Research Practices, Verification Practices, and Verification Studies) that resolve conceptual and structural limitations of TOP 2015. Research Practices include seven discrete practices implemented in three ways: Disclosed, Shared and Cited, and Certified. Verification Practices introduce standards for computational reproducibility and results transparency, while Verification Studies enumerate empirical research designs and publication formats that support verification of results. Language now centers on researchers and studies (rather than journals) to facilitate adoption by funders, research organizations, preprint servers, evidence curators, and other interest-holders. These changes promote flexible implementation and contextual tailoring. TOP 2025 retains the core strengths of TOP 2015 while addressing feedback from a decade of implementation and advances in open science. Rather than a universal mandate, it provides a shared, adaptable framework to coordinate the design, implementation, and communication of open policies across interest-holders.
Unsolicited emails from predatory journals offer a quick opportunity to produce a first author publication for a fee. With increasing emphasis on first author publications for applications, interviews and jobs, the opportunity associated with such journals may encourage some to produce work that can be both easily published and damaging to the medical field. Predatory journals therefore have an opportunity to publish low quality literature quickly for a profit. To analyse the quality of peer review, cost and acceptance rate of AI generated papers submitted to predatory journals contacting by unsolicited emails. Chat GPT3 along with an AI image generator were used to produce scientific manuscripts that were submitted to 50 consecutive unsolicited emails from predatory journals. Data on response time and outcome were collected. Generation and submission times, along with article processing charges (APC) were analysed. All papers were formally withdrawn. From 50 submitted papers 15 (30
Cannabinoid-based medicines and conventional pharmaceutical analgesics often demonstrate similarly modest analgesic effects in randomized controlled trials (RCTs), yet meta-analyses of these treatment classes frequently reach different interpretive conclusions. We examined whether this divergence was compatible with interpretive asymmetry, including potential reverse spin bias, defined as narrative framing that is more cautious, negative, or dismissive than expected based on reported efficacy results. We conducted a critical interpretive synthesis of meta-analyses evaluating cannabinoids and pharmaceutical analgesics for pain. Searches of PubMed, Google Scholar, and reference lists were completed through July 2025. Eligible reviews included double-blind RCTs reporting pain-intensity outcomes. Methodological quality was assessed using AMSTAR-2. For each meta-analysis, pooled efficacy estimates, Results-section wording, and conclusion direction were extracted. Reverse spin bias was assessed at the publication level by comparing reported efficacy findings with narrative conclusions. Twenty-four cannabinoid meta-analyses, contributing 27 analyzable interventions, and 15 pharmaceutical analgesic meta-analyses, contributing 47 analyzable interventions, were included. Median pooled pain-relief effect sizes were similar in the cannabinoids and pharmaceutical analgesics literatures, with a small estimated between-class difference of approximately 0.02 effect-size units. At the publication level, positive conclusions were less frequent in cannabinoid publications than in pharmaceutical analgesic publications (n = 3/24 [12.5
Fraud is a well-documented threat to data integrity in online research. However, little research has examined fraud risk in studies that rely on in-person recruitment but collect data online, particularly when preserving participant anonymity is necessary. The purpose of this study was to describe and evaluate fraud prevention and detection strategies in the context of an anonymous, web-based survey that relied on in-person, targeted recruitment. We evaluated fraud in a cross-sectional survey examining the experiences, needs, and barriers faced by families living in hotels in DeKalb County, Georgia. Eligibility criteria included being 18 years of age or older, currently residing in a hotel in DeKalb County, and living with at least one child under 18 years of age. Eligible participants were recruited through door-to-door outreach and received a flyer containing a survey link. Each survey entry was assessed for potential fraud using a standardized protocol comprising ten criteria and were classified as valid, potentially fraudulent, or fraudulent. Descriptive analyses were performed to summarize violations of fraud detection criteria. Among 246 completed survey entries, 11 (4.5
Some questionable research practices (QRPs) are regarded as core contributors to problems of reproducibility and replication, while others compromise the integrity and credibility of research by obscuring responsibility and incentives. However, evidence on the individual-level drivers of QRP engagement is fragmented. This study examines theory-informed associations between social and psychological factors and self-reported QRP engagement. Using data from a cross-sectional survey of researchers across disciplinary fields (18,376 invited; N = 3,050 respondents), this study examines associations between social and psychological factors and self-reported QRP engagement. The sample reflects a self-selected subset of invited respondents. I found that commitment to universalism (β = -0.14, 95
BackgroundResearch reproducibility - the ability of others to independently verify scientific findings by following the same methods and data - is a fundamental aspect of research integrity. When this fails, trust in evidence is undermined and research resources are wasted. Reproducibility measures describe the standards that define reproducible research, and reproducibility interventions are the actions used to improve those standards. There is little agreement on which practices and initiatives should be prioritised for implementation in practice. This study aimed to establish expert consensus on the key measures and interventions that should be prioritised to strengthen research reproducibility.MethodsWe conducted a Delphi consensus study as part of the EU Horizon Europe iRISE project. Experts from five stakeholder groups (researchers, editors, publishers, funders, and policymakers) evaluated reproducibility measures and interventions identified through prior literature mapping across two online survey rounds, followed by a final virtual consensus panel. Items were rated on a 10-point Likert scale, with consensus defined a priori as agreement by at least 70% of panellists assigning high-priority scores, corresponding to scores of eight to ten on the Likert scale.ResultsSeventy-three panellists from 34 countries participated in the first round, with high retention rates in the subsequent round. Consensus was achieved on eight reproducibility measures and six reproducibility interventions in the first round. Prioritised measures included methodological quality, reporting quality, code and data availability and reuse, computational reproducibility, transparency of research plan, reproducible workflow practices, trial registration, and materials availability and reuse. Prioritised interventions included data management training, data quality checks/feedback, statistical training, data sharing policy/guidelines, protocol/trial registration, and reproducible code/analysis training.ConclusionsThe prioritised measures and interventions provide a structured foundation for improving research reproducibility across disciplines. The findings can inform the development of institutional training curricula in data management and statistical methods, policies supporting data sharing and trial registration, and future empirical evaluation of these practices across research contexts.
BACKGROUND:Despite endorsement by medical journals, reporting guidelines have only modestly affected reporting quality in healthcare research. We aimed to identify influences affecting whether authors successfully adhere to reporting guidelines. METHODS:We searched MEDLINE, Embase, PsychINFO, AMED, WHO Global Index Medicus, SciELO, Chinese Biomedical Literature Database, China National Knowledge Infrastructure, Wanfang Data, VIP Chinese Medical Journal Database, OSF, and MiRoR for qualitative research exploring researchers' experiences of reporting guidelines for healthcare research, published after 1996 in English, Chinese, Spanish, or Portuguese. We appraised studies using CASP-Qual. For thematic synthesis, we applied descriptive codes to all text reporting qualitative findings, then aggregated codes inductively into descriptive themes that captured the codes' meaning. We interpreted and contextualised possible influences from these descriptive themes to create analytic themes. RESULTS:From 18 eligible studies, we developed 12 analytic themes: 1) Researchers may not understand guidance as intended or what reporting guidelines are, even if they think they do; 2) Researchers report a variety of reasons for using reporting guidelines, and some are more important than others; 3) Researchers describe using reporting guidelines for different tasks and wanting guidance delivered in ways that better fit their needs; 4) Using reporting guidelines has costs which researchers may feel outweigh benefits; 5) Reporting guidelines may need to be revised and updated; 6) Researchers may not be able to report all items, which can leave them feeling uncertain or worried; 7) Awareness and accessibility may limit reporting guideline usage; 8) Reporting guidelines may be more useful to less experienced researchers, but these researchers may find them harder to use; 9) Researchers want or need design advice, but reporting guidelines may not be the right place to find it; 10) Reporting guidelines can be harder to use if their scope is too broad, too narrow, or poorly defined; 11) Researchers may have to use multiple sets of reporting guidelines, multiplying complexity and costs; 12) Researchers may use checklists but never read the full guidance. DISCUSSION:We identified many influences despite a paucity of evidence. Addressing these influences when developing, refining, and implementing reporting guidelines may improve adherence.
Peer review is fundamental to quality scientific communication, yet reviewer behavior remain underexplored. The impact of emerging large language models (LLMs) on peer review practices is similarly understudied. We aim to characterize behavioral traits of Chinese medical journal peer reviewers and identify evidence-based recommendations to optimize review willingness, efficiency and quality. An online questionnaire survey was distributed to 532 medical researchers in China through the Wenjuanxing platform in February 2025. The questionnaire (38 questions) assessed four domains: basic information, peer review model and efficiency, peer review quality, and reviewer motivations. Statistical analysis included descriptive statistics, Spearman correlations, Kruskal–Wallis tests, etc. The response rate was 51.9
Scholarly publishing is facing unprecedented challenges. Journal scholarly integrity and ethics have historically focused on author conduct and conflicts of interest. We address apparent journal and editor conflicts of interest. We focus on a series of events attending the publication in a peer-reviewed biomedical journal of a research article and four contemporaneously published editorials. A time series of events and questions which they engender follows the events description. We dissect and question behaviors and motivations, and accompanying transparency and accountability, and how these can affect article and journal credibility. The need for journals and scholarly publishing to recognize their own conflicts is discussed, as is the need for the science publishing to be self-correcting, in addition to science articles themselves. Author conflicts of interest and the need for attendant disclosure have become part of the publication process. Journal conflicts can be no less impactful and important, and merit the same disclosure, transparency, and rigor, in order to protect the public trust in biomedical science and publishing. When issues are encountered, science publishing needs to be self-correcting.
Abstract Background Statistical review is essential for research quality and integrity, yet traditional manual review is inefficient. Large language models (LLMs) offer potential support but are unreliable when used without guidance for precise calculations and raise concerns about accountability. This study evaluated whether a structured, rule-based prompt can reliably constrain an LLM to perform statistical review of comparative categorical data, and characterized both its feasibility and its inherent risks from an accountability perspective. Methods This study employed a two-stage design based on the DeepSeekV3.2. In the first stage, a structured prompt was developed through dozens of "test-fail-iterate" cycles using 20 published medical articles. The prompt assigned the LLM the role of a "statistics expert" and provided a closed set of computational rules and a "recognize data-select calculation formula-calculate" workflow for analyzing categorical data, including Pearson's Chi-square test, continuity correction, and McNemar's tests. In the second stage, the performance of the final prompt was evaluated on a test set of 20 independent manuscripts. The model's output was compared against the results calculated by a senior statistician (the gold standard). The primary outcome measures were the performance in statistical method selection and numerical computation, including accuracy, sensitivity (recall), specificity, positive predictive value, negative predictive value, F1 score, and Cohen's Kappa. Secondary measures included reproducibility and efficiency. Results The test set consisted of 15 manuscripts with independent samples and 5 with paired samples. In the assessment of the appropriateness of statistical method selection for 148 analysis items, the model achieved an accuracy of 99.3% (147/148), a sensitivity of 96.2% (25/26) (F1=98.0%, κ=0.976). For the test of computational consistency in 97 independent sample tests, the accuracy for χ 2 value consistency was 94.8% (92/97) (F1=89.3%, κ=0.859), and for P -value consistency, it was 96.9% (94/97) (F1=90.9%, κ=0.891). In the paired-sample analysis, the model's methods and results were in perfect agreement with the manual review, and prompt optimization eliminated discrepancies in degrees-of-freedom calculation rules. Efficiency analysis showed no statistically significant difference in time consumption between the model (407 s) and manual review (374 s) ( P =0.601). In reproducibility tests, the intraclass correlation coefficients for both χ 2 values and P -values exceeded 0.91. However, qualitative analysis revealed 3 typical failure modes in the task workflow: (1) Instability: The model's failure to produce identical outputs across repeated runs, manifesting as inconsistent data extraction or the failure to process all designated tasks (scope neglect). (2) Performance degradation/"lazy" behavior: A decline in execution quality on long or complex tasks, often characterized by the model abandoning its reasoning process to copy author-provided values without verification. (3) Anchoring effect: The model's tendency to over-rely on author-provided statistical values (the "anchor"), causing its verification process to be unduly influenced. Conclusions A structured, rule-based prompt can guide the DeepSeek to achieve high accuracy in standardized statistical review tasks, but its reliability is contingent on operational stability. Inherent failure modes, including performance instability and a strong anchoring effect on author-provided data, persist and can lead to significant errors, particularly when source data are flawed. These findings suggest that the the DeepSeek is not suitable for autonomous auditing. Their most appropriate application is as assistive tools within a human-in-the-loop framework, where rigorous human supervision is essential for risk mitigation and to maintain ultimate accountability.
Prior evidence suggests that journals requiring open data are associated with higher levels of data sharing in the published psychology literature. Data sharing policies are not, however, consistently implemented or enforced. The American Psychological Association (APA), in 2020, signed onto the Transparency and Openness Promotion Guidelines, which promote increasingly stringent data sharing policies. The current study examined self-reported data sharing in APA journals and whether stricter policies are linked to higher levels of self-reported data sharing. We assessed self-reported data sharing practices in 1,250 articles published between 2023 and 2025 in 25 APA journals. Using logistic regression, we examined the association between journal policy level and self-reported data sharing. We then applied post-stratification weighting, based on the actual distribution of policy levels across APA journals and their associated percentages of data sharing, to estimate the overall percentage of self-reported data sharing. We estimated overall self-reported data sharing to be 30.3
BACKGROUND:Research misconduct poses a serious threat to academic integrity, particularly in medical sciences. This study aimed to estimate the prevalence of various forms of research misconduct include plagiarism, data fabrication or falsification, authorship misconduct, salami slicing, and purchasing research work among postgraduate students in Iranian medical universities using the Unmatched Count Technique (UCT). METHODS:A cross-sectional survey was conducted among postgraduate students from multiple Iranian medical universities using a double-list version of the unmatched count technique (UCT). The questionnaire was administered in two sequential waves, with approximately half of participants completing List A and the remaining participants completing List B, ensuring that each respondent received only one list version. For each research misconduct behavior, prevalence was estimated by calculating the mean difference in endorsement counts between lists containing the sensitive item and corresponding control lists with only non-sensitive items. In the double-list design, prevalence estimates were computed separately for List A and List B, with sensitive items counterbalanced across list positions to control for order effects. The final prevalence was calculated as the average of the two list-specific estimates, improving precision and reducing list-order bias. RESULTS:The most commonly reported misconduct was using others' ideas or phrases without proper citation (43%), followed by dishonest result reporting (38%), data fabrication or deletion (34%), and authorship misrepresentation (34%). Salami slicing was reported by 26%, and 20% admitted to purchasing parts or all of a research project. The UCT survey tool demonstrated acceptable reliability, with intraclass correlation coefficients (ICCs) ranging from 0.64 to 0.84. CONCLUSION:The findings indicate a troubling level of research misconduct among postgraduatestudents in Iran's medical sciences universities. This highlights the need for effective ethics training, stronger academic integrity policies, and enforceable institutional mechanisms to promote responsible research conduct and protect the future credibility of medical professionals.
Experts and their opinions can play a pivotal role in research for decision-making, particularly where empirical data are limited or uncertain. However, the methods by which expert input is solicited, synthesised, and applied in research often lack consistency and transparency. The absence of cross-disciplinary guidelines for how expert opinion should be gathered, integrated into research, and utilised in decision-making leaves the process vulnerable to conflicts of interest, interpersonal dynamics, and researcher discretion. Using health technology assessment (HTA) as a case study, we conducted a systematic review of empirical studies examining the use of expert opinion in research for decision-making. Twenty-three studies were included, from which six broad categories of expert consultation methods were identified, ranging from informal individual consultation to structured expert elicitation protocols. Considerable variation was observed in both the methodological rigour and transparency with which expert consultation was conducted and reported. Drawing on the review findings, we propose a preliminary conceptual framework (INTEGRITY) that synthesises psychosocial, methodological and reporting factors to promote transparency, rigour, inclusivity and objectivity when incorporating expert consultation into research. For low-stakes exploratory judgements, informal consultation may be sufficient. For high-stakes policy or resource-allocation decisions that require quantitative estimates, structured and transparent approaches such as formal elicitation protocols should be considered to support inclusive, objective, rigorous, and transparent use of expert judgement in research.
The Integrity Risk Indicators by SCImago (IRIS) framework, developed by the SCImago Research Group, addresses growing concerns about trust in science by shifting research assessment from traditional output and citation-based measures toward the evaluation of institutional exposure to research integrity risks. Rather than focusing on individual misconduct, IRIS recognizes the central role of institutional incentive systems, governance structures, and operational cultures in shaping research integrity. IRIS provides a constructive, data-informed model for identifying structural vulnerabilities associated with practices such as authorship and affiliation patterns, dissemination and citation dynamics, publication venue credibility, and the sustainability and transparency of publication practices. The indicators are derived from bibliometric and publication data and are intended as signals of potential risk, not as diagnoses or judgments. Applied to Higher Education Institutions included in the SCImago Institutions Rankings, IRIS enables meaningful comparison through z-score normalization, placing heterogeneous indicators on a common, interpretable scale centered on the global mean. The use of standardized scores enhances sensitivity to the magnitude of deviations, allowing IRIS to better function as a risk detection system, rather than a competitive ranking, and to support institutions in strengthening responsible and trustworthy research environments.