Context: Large language models (LLMs) are increasingly used to screen literature for systematic reviews (SRs), but the standard confusion-matrix metrics used to evaluate them can mislead under the imbalanced, cost-asymmetric conditions of screening. Objective: We develop and justify LLM4SCREENLIT — practical recommendations for researchers conducting LLM-screening evaluations and for editors and reviewers assessing such studies — differentiated by study type (retrospective benchmarking vs. deployment for a specific SR). Method: Using Delgado-Chaves et al. (2025), an 18-LLM benchmark across three biomedical SRs, as a motivating example, we reviewed 28 additional papers and extracted their reported metrics. We propose a Weighted Matthews Correlation Coefficient (WMCC) that integrates MCC’s chance-correction with asymmetric misclassification costs, and validated it on three software-engineering (SE) reanalyses (Felizardo et al. 2024; Syriani et al. 2024; Huotala et al. 2025), the largest covering 9 LLMs × 24 SE secondary studies (34,528 articles). Results: Across the 29 papers, only 10% reported MCC, only 24% reported full confusion matrices, and none of the five papers claiming workload savings priced false-negative cost. In the largest SE reanalysis, MCC and WMCC disagree on the best LLM in 55% of evaluable studies; in the most striking 9695-article SE study, the Accuracy-best LLM loses 63.3% of relevant evidence (Lost Evidence), the MCC-best 43.9%, but the WMCC-best only 5.8%. Sensitivity analysis (median crossover at w≈2.7, all <7) supports w=10 as a conservative default. Conclusions: SR-screening evaluations should prioritise Lost Evidence and use cost-sensitive WMCC alongside MCC for ranking. Reporting must include the full confusion matrix and treat unclassifiable outputs as positives requiring human review. Designs should be leakage-aware, with non-LLM baselines when the study aims to inform SR practice and labels are available. Editors and reviewers should require these elements as routine. Extension to full-text screening and data extraction is principled but pending empirical validation.
Context: Generative AI (GenAI) and Large Language Models (LLMs) are increasingly used for academic tasks in software engineering and beyond, including systematic literature reviews (SLRs). However, while capable of summarizing text, there is no guarantee they can meet the rigour, reliability, and transparency that SLRs require. Objectives: To support researchers intending to conduct SLRs using GenAI or those conducting empirical studies evaluating how well GenAI supports SLR tasks. Methods: First, we conducted a rapid review to identify studies that propose guidelines for evaluating and using GenAI and LLMs to support SLRs. Second, we drew on thought experiments, relevant guidance from the literature, and our own experience conducting SLRs and evaluating tools to develop recommendations for how to use and assess GenAI in the context of SLRs. Results: We discuss the problems researchers face when evaluating GenAI for SLRs. We identify and explain process issues to consider when planning, conducting, and reporting both SLRs using GenAI and evaluations of GenAI tools. Finally, we summarize our results as a set of process recommendations, which we name GUEST (GenAI Use and Evaluation in SLR Tasks). Conclusion: We argue that GenAI requires human oversight and is not currently capable of unsupervised systematic studies. However, it offers the prospect of cost-effective assistance for some repetitive tasks and for additional validation of some complex tasks. Our GUEST recommendations should help software engineering researchers both to conduct and report trustworthy SLRs using GenAI and to provide rigorous independent evaluation studies.
Software engineering (SE) experiments often have small sample sizes. This can result in data sets with non-normal characteristics, which poses problems as standard parametric meta-analysis, using the standardized mean difference (StdMD) effect size, assumes normally distributed sample data. Small sample sizes and non-normal data set characteristics can also lead to unreliable estimates of parametric effect sizes. Meta-analysis is even more complicated if experiments use complex experimental designs, such as two-group and four-group cross-over designs, which are popular in SE experiments. Our objective was to develop a validated and robust meta-analysis method that can help to address the problems of small sample sizes and complex experimental designs without relying upon data samples being normally distributed. To illustrate the challenges, we used real SE data sets. We built upon previous research and developed a robust meta-analysis method able to deal with challenges typical for SE experiments. We validated our method via simulations comparing StdMD with two robust alternatives: the probability of superiority ( p̂ ) and Cliffs’ d. We confirmed that many SE data sets are small and that small experiments run the risk of exhibiting non-normal properties, which can cause problems for analysing families of experiments. For simulations of individual experiments and meta-analyses of families of experiments, p̂ and Cliff’s d consistently outperformed StdMD in terms of negligible small sample bias. They also had better power for log-normal and Laplace samples, although lower power for normal and gamma samples. Tests based on p̂ always had better or equal power than tests based on Cliff’s d, and across all but one simulation condition, p̂ Type 1 error rates were less biased. Using p̂ is a low-risk option for analysing and meta-analysing data from small sample-size SE randomized experiments. Parametric methods are only preferable if you have prior knowledge of the data distribution.
A few years ago, rapid reviews (RR) were introduced in software engineering (SE) to address the problem that standard systematic reviews take too long and too much effort to be of value to practitioners. Prior to our study, few practice-driven RRs had been reported, and none involved collaboration with practitioners lacking SE research experience. To investigate practitioners’ perspectives on the use of RRs in supporting SE practices, we aimed to validate and build upon the findings of the seminal RR in SE study, specifically considering practitioners without explicit SE research experience. First, we studied previously conducted RRs in SE through a systematic review. Second, we carried out an external replication of the first study that proposed the use of RRs in SE. Specifically, we conducted an RR for an agile software development team looking to improve its knowledge management practices. Most of the software development team’s perceptions about RR results were positive and strongly consistent with previous research. In particular, RR results were considered more reliable than other sources of information and adequate to address the problems detected. Some months later they confirmed using some of the recommendations. The results show that practitioners without explicit SE research experience appreciate the value of evidence and can make use of the results of RRs. However, SE research may need to be translated from broad recommendations to specific process change options. Our research also reveals that SE RRs reporting needs to be substantially improved.
Context : Several tertiary studies have criticized the reporting of software engineering secondary studies. Objective : Our objective is to identify guidelines for reporting software engineering (SE) secondary studies which would address problems observed in the reporting of software engineering systematic reviews (SRs). Method : We review the criticisms of SE secondary studies and identify the major areas of concern. We assess the PRISMA 2020 (Preferred Reporting Items for Systematic Reviews and Meta-Analyses) statement as a possible solution to the need for SR reporting guidelines, based on its status as the reporting guideline recommended by the Cochrane Collaboration whose SR guidelines were a major input to the guidelines developed for SE. We report its advantages and limitations in the context of SE secondary studies. We also assess reporting guidelines for mapping studies and qualitative reviews, and compare their structure and content with that of PRISMA 2020. Results : Previous tertiary studies confirm that reports of secondary studies are of variable quality. However, ad hoc recommendations that amend reporting standards may result in unnecessary duplication of text. We confirm that the PRISMA 2020 statement addresses SE reporting problems, but is mainly oriented to quantitative reviews, mixed-methods reviews and meta-analyses. However, we show that the PRISMA 2020 item definitions can be extended to cover the information needed to report mapping studies and qualitative reviews. Conclusions : In this paper and its Supplementary Material, we present and illustrate an integrated set of guidelines called SEGRESS (Software Engineering Guidelines for REporting Secondary Studies), suitable for quantitative systematic reviews (building upon PRISMA 2020), mapping studies (PRISMA-ScR), and qualitative reviews (ENTREQ and RAMESES), that addresses reporting problems found in current SE SRs.
Context : Recent papers have proposed the use of grey literature (GL) and multivocal reviews. These papers have raised issues about the practices used for systematic reviews (SRs) in software engineering (SE) and suggested that there should be changes to the current SR guidelines. Objective : To investigate whether current SR guidelines need to be changed to support GL and multivocal reviews. Method : We discuss the definitions of GL and the importance of GL and of industry-based field studies in SE SRs. We identify properties of SRs that constrain the material used in SRs: a) the nature of primary studies; b) the requirements of SRs to be auditable, traceable, and reproducible; and explain why these requirements restrict the use of blogs in SRs. Results : SR guidelines have always considered GL as a possible source of primary studies and have never supported exclusion of field studies that incorporate the practitioners' viewpoint. However, the concept of GL, which was meant to refer to documents that were not formally published, is now being extended to information from sources such as blogs/tweets/Q&A posts. Thus, it might seem that SRs do not make full use of GL because they do not include such information. However, the unit of analysis for an SR is the primary study. Thus, it is not the source but the type of information that is important. Any report describing a rigorous empirical evaluation is a candidate primary study. Whether it is actually included in an SR depends on the SR eligibility criteria. However, any study that cannot be guaranteed to be publicly available in the long term should not be used as a primary study in an SR. This does not prevent such information from being aggregated in surveys of social media and used in the context of evidence-based software engineering (EBSE). Conclusions : Current guidelines for SRs do not require extensions, but their scope needs to be better defined. SE researchers require guidelines for analysing social media posts (e.g., blogs, tweets, vlogs), but these should be based on qualitative primary (not secondary) study guidelines. SE researchers can use mixed-methods SRs and/or the fourth step of EBSE to incorporate findings from social media surveys with those from SRs and to develop industry-relevant recommendations.
Context: Evidence-based practice (EBP) has allowed several disciplines to become more mature by emphasizing the use of evidence from well-designed and well-conducted research in decision-making. Its application in SE, Evidence-based software engineering (EBSE) can help to bridge the gap between academia and industry by bringing together academic rigor and research of practical relevance. To achieve this, it seems necessary to improve its adoption.Objective: We sought both to study the attitudes towards EBSE of stakeholders working in a government agency (GA) and to assess whether knowledge of EBSE would impact their working practices. Method: We conducted a multi-stage field investigation in an Uruguayan national GA that is responsible for digital policies. First, we organized an EBSE awareness lecture and we collected and analyzed participants' perceptions of the value and limitations of EBSE. Sixteen months later, in a second stage, we contacted the agency and asked participants whether they had made use of the information about EBSE we presented to them.Results: Initially, participants reported that EBSE seemed useful for tackling challenging problems and, in particular, considered its use appropriate given the agency's responsibilities. Perceived barriers to EBSE adoption were the need for institutional support, the lack of government practice reports, inadequate skills or motivation, the cost of conducting systematic reviews, and the lack of evidence about emerging issues. In the follow-up survey, although the participants were not undertaking systematic reviews themselves, many reported improvements in how they searched for and evaluated information to support their work.Conclusion: Our study presents some insights to better understand EBSE adoption. With the exception of GA-specific issues, perceived value and barriers to adoption were consistent with those reported in software engineering and other disciplines. Our follow-up study confirms the potential value of evidence in the context of IT regulatory and government bodies.
Context: In empirical software engineering, crossover designs are popular for experiments comparing software engineering techniques that must be undertaken by human participants. However, their value depends on the correlation ( $r$ ) between the outcome measures on the same participants. Software engineering theory emphasizes the importance of individual skill differences, so we would expect the values of $r$ to be relatively high. However, few researchers have reported the values of $r$ . Goal: To investigate the values of $r$ found in software engineering experiments. Method: We undertook simulation studies to investigate the theoretical and empirical properties of $r$ . Then we investigated the values of $r$ observed in 35 software engineering crossover experiments. Results: The level of $r$ obtained by analysing our 35 crossover experiments was small. Estimates based on means, medians, and random effect analysis disagreed but were all between 0.2 and 0.3. As expected, our analyses found large variability among the individual $r$ estimates for small sample sizes, but no indication that $r$ estimates were larger for the experiments with larger sample sizes that exhibited smaller variability. Conclusions: Low observed $r$ values cast doubts on the validity of crossover designs for software engineering experiments. However, if the cause of low $r$ values relates to training limitations or toy tasks, this affects all Software Engineering (SE) experiments involving human participants. For all human-intensive SE experiments, we recommend more intensive training and then tracking the improvement of participants as they practice using specific techniques, before formally testing the effectiveness of the techniques.
Context: Evidence-based software engineering (EBSE) can be an effective resource to bridge the gap between academia and industry by balancing research of practical relevance and academic rigor. To achieve this, it seems necessary to investigate EBSE training and its benefits for the practice. Objective: We sought both to develop an EBSE training course for university students and to investigate what effects it has on the attitudes and behaviors of the trainees. Method: We conducted a longitudinal case study to study our EBSE course and its effects. For this, we collect data at the end of each EBSE course (2017, 2018, and 2019), and in two follow-up surveys (one after 7 months of finishing the last course, and a second after 21 months). Results: Our EBSE courses seem to have taught stu-dents adequately and consistently. Half of the respondents to the surveys report making use of the new skills from the course. The most-reported effects in both surveys indicated that EBSE concepts increase awareness of the value of research and evidence and EBSE methods improve information gathering skills. Conclusions: As suggested by research in other areas, training appears to play a key role in the adoption of evidence-based practice. Our results indicate that our training method provides an introduction to EBSE suitable for undergraduates. However, we believe it is necessary to continue investigating EBSE training and its impact on software engineering practice.
Sharing research data from public funding is an important topic, especially now, during times of global emergencies like the COVID-19 pandemic, when we need policies that enable rapid sharing of research data. Our aim is to discuss and review the revised Draft of the OECD Recommendation Concerning Access to Research Data from Public Funding. The Recommendation is based on ethical scientific practice, but in order to be able to apply it in real settings, we suggest several enhancements to make it more actionable. in particular, constant maintenance of provided software stipulated by the Recommendation is virtually impossible even for commercial software. Other major concerns are insufficient clarity regarding how to finance data repositories in joint private-public investments, inconsistencies between data security and user-friendliness of access, little focus on the reproducibility of submitted data, risks related to the mining of large data sets, and sensitive (particularly personal) data protection. In addition, we identify several risks and threats that need to be considered when designing and developing data platforms to implement the Recommendation (e.g., not only the descriptions of the data formats but also the data collection methods should be available). Furthermore, the non-even level of readiness of some countries for the practical implementation of the proposed Recommendation poses a risk of its delayed or incomplete implementation.
Previous studies have raised concerns about the analysis and meta-analysis of crossover experiments and we were aware of several families of experiments that used crossover designs and meta-analysis. To identify families of experiments that used meta-analysis, to investigate their methods for effect size construction and aggregation, and to assess the reproducibility and validity of their results. We performed a systematic review (SR) of papers reporting families of experiments in high quality software engineering journals, that attempted to apply meta-analysis. We attempted to reproduce the reported meta-analysis results using the descriptive statistics and also investigated the validity of the meta-analysis process. Out of 13 identified primary studies, we reproduced only five. Seven studies could not be reproduced. One study which was correctly analyzed could not be reproduced due to rounding errors. When we were unable to reproduce results, we provide revised meta-analysis results. To support reproducibility of analyses presented in our paper, it is complemented by the reproducer R package. Meta-analysis is not well understood by software engineering researchers. To support novice researchers, we present recommendations for reporting and meta-analyzing families of experiments and a detailed example of how to analyze a family of 4-group crossover experiments.
There are inconsistencies between the formulas for the variance of standardized mean difference (SMD) in the Cochrane Handbook for Systematic Reviews and the variance reported in other sources. Instead of the variance appropriate for the SMD of a crossover experiment, the Cochrane Handbook uses the variance appropriate for a pre-test post-test experiment. This means that if there is a non-negligible time period effect, the formula reported by the Handbook will underestimate both the effect size and its variance. In addition, the formula for the standard error of SMD reported in the Cochrane Handbook (in section 23.2.7.2) is inconsistent with the variance derived from the variance of the related t-test. Even if the period effect is negligible, the Cochrane Handbook formula is biased toward underestimates. The difference between the estimates from the two formulas will be small if either the correlation between the repeated measures, or the magnitude of the SMD estimate, is small, or if the sample size is large. However, it can be can be quite substantial in other circumstances.
Background Examples of questionable statistical practice, when published in high quality software engineering (SE) journals, may lead to novice researchers adopting incorrect statistical practices. Objective Our goal is to highlight issues contributing to poor statistical practice in human-centric SE experiments. Method We reviewed the statistical analysis practices used in the 13 papers that reported families of human-centric SE experiments and were published in high quality journals. Results Reviewed papers related to 45 experiments and involved a total of 1303 human participants. We searched for issues that were related to questionable statistical practice that were found in more than one paper. We observed three types of bad practice: incorrect use of terminology, incorrect analysis of repeated measures designs, and post-hoc power testing. We also found two analysis practices (i.e., multiple testing and pre-testing for normality) where statisticians disagree about good practice. Conclusions Identified issues pose a problem because readers may expect the statistical methods used in papers published in top quality, peer-reviewed journals to be correct. We explain why the practices are problematic and provide recommendations for improved practice.
Reproducibility of Empirical Software Engineering (ESE) studies is an essential part for improving their credibility, as it offers the opportunity to the research community to verify, evaluate and improve their research outcomes. We aim to study reproducibility and credibility in ESE with a case study, by investigating how they have been addressed in studies where SZZ, a widely-used algorithm by Śliwerski, Zimmermann and Zeller to detect the origin of a bug, has been applied. We have performed a systematic literature review to evaluate publications that use SZZ. In total, 187 papers have been analyzed for reproducibility, reporting of limitations and use of improved versions of the algorithm. We have found a situation with a lot of room for improvement in ESE as reproducibility is not commonly found; factors that undermine the credibility of results are common. We offer some lessons learned and guidelines for researchers and reviewers to address this problem. Reproducibility and other related aspects that ensure a high quality scientific process should be taken more into consideration by the ESE community in order to increase the credibility of the research results.
Vegas et al. IEEE Trans Softw Eng 42(2):120:135 (2016) raised concerns about the use of AB/BA crossover designs in empirical software engineering studies. This paper addresses issues related to calculating standardized effect sizes and their variances that were not addressed by the Vegas et al.'s paper. In a repeated measures design such as an AB/BA crossover design each participant uses each method. There are two major implication of this that have not been discussed in the software engineering literature. Firstly, there are potentially two different standardized mean difference effect sizes that can be calculated, depending on whether the mean difference is standardized by the pooled within groups variance or the within-participants variance. Secondly, as for any estimated parameters and also for the purposes of undertaking meta-analysis, it is necessary to calculate the variance of the standardized mean difference effect sizes (which is not the same as the variance of the study). We present the model underlying the AB/BA crossover design and provide two examples to demonstrate how to construct the two standardized mean difference effect sizes and their variances, both from standard descriptive statistics and from the outputs of statistical software. Finally, we discuss the implication of these issues for reporting and planning software engineering experiments. In particular we consider how researchers should choose between a crossover design or a between groups design.
Statistics in MedicineVolume 37, Issue 2 p. 320-323 Letter to the Editor Corrections to effect size variances for continuous outcomes of crossover clinical trials Barbara Kitchenham, Barbara Kitchenham School of Computing and Mathematics, Keele University, Keele, Staffordshire ST5 5BG, U.KSearch for more papers by this authorLech Madeyski, Corresponding Author Lech Madeyski lech.madeyski@pwr.edu.pl orcid.org/0000-0003-3907-3357 Faculty of Computer Science and Management, Wrocław University of Science and Technology, Wyb. Wyspianskiego 27, Wroclaw, 50-370 Poland Correspondence to: Lech Madeyski, Faculty of Computer Science and Management, Wrocław University of Science and Technology, Wyb. Wyspianskiego 27, 50-370 Wroclaw, Poland. E-mail: lech.madeyski@pwr.edu.plSearch for more papers by this authorFrançois Curtin, François Curtin Research Center for Statistics, Geneva School of Economics and Management, University of Geneva, Geneva, Switzerland Geneuro SA, Geneva, SwitzerlandSearch for more papers by this author Barbara Kitchenham, Barbara Kitchenham School of Computing and Mathematics, Keele University, Keele, Staffordshire ST5 5BG, U.KSearch for more papers by this authorLech Madeyski, Corresponding Author Lech Madeyski lech.madeyski@pwr.edu.pl orcid.org/0000-0003-3907-3357 Faculty of Computer Science and Management, Wrocław University of Science and Technology, Wyb. Wyspianskiego 27, Wroclaw, 50-370 Poland Correspondence to: Lech Madeyski, Faculty of Computer Science and Management, Wrocław University of Science and Technology, Wyb. Wyspianskiego 27, 50-370 Wroclaw, Poland. E-mail: lech.madeyski@pwr.edu.plSearch for more papers by this authorFrançois Curtin, François Curtin Research Center for Statistics, Geneva School of Economics and Management, University of Geneva, Geneva, Switzerland Geneuro SA, Geneva, SwitzerlandSearch for more papers by this author First published: 18 December 2017 https://doi.org/10.1002/sim.7379Citations: 7Read the full textAboutPDF ToolsRequest permissionExport citationAdd to favoritesTrack citation ShareShare Give accessShare full text accessShare full-text accessPlease review our Terms and Conditions of Use and check box below to share full-text version of article.I have read and accept the Wiley Online Library Terms and Conditions of UseShareable LinkUse the link below to share a full-text version of this article with your friends and colleagues. Learn more.Copy URL Share a linkShare onFacebookTwitterLinkedInRedditWechat Citing Literature Volume37, Issue2Special Issue: Papers from the Workshop on Infectious Diseases Research: Quantitative Methods and Models in the Era of Big Data30 January 2018Pages 320-323 RelatedInformation
Lech Madeyski合作论文数Wroclaw University of Technology12
Mark Turner合作论文数Case Western Reserve University6