Systematic literature reviews (SLRs) are a major tool for supporting evidence-based software engineering. Adapting the procedures involved in such a review to meet the needs of software engineering and its literature remains an ongoing process. As part of this process of refinement, we undertook two case studies which aimed 1) to compare the use of targeted manual searches with broad automated searches and 2) to compare different methods of reaching a consensus on quality. For Case 1, we compared a tertiary study of systematic literature reviews published between January 1, 2004 and June 30, 2007 which used a manual search of selected journals and conferences and a replication of that study based on a broad automated search. We found that broad automated searches find more studies than manual restricted searches, but they may be of poor quality. Researchers undertaking SLRs may be justified in using targeted manual searches if they intend to omit low quality papers, or they are assessing research trends in research methodologies. For Case 2, we analyzed the process used to evaluate the quality of SLRs. We conclude that if quality evaluation of primary studies is a critical component of a specific SLR, assessments should be based on three independent evaluators incorporating at least two rounds of discussion.
Software engineers find that experiments are difficult to perform. Furthermore, no experiment can be considered flawless no matter how well conducted. In this Special Issue the authors have addressed this problem as well as emphasising the need to undertake more empirical studies with discussing practical and methodological issues associated with evaluation. As so, all the papers in this Special Issue address software engineering aspects and evaluate them through various types of empirical studies ranging from experiments, replications of empirical studies, case studies, surveys, observational studies, field studies, systematic reviews.
Background. We have been undertaking a series of systematic literature reviews as part of a research program aimed at assessing evidence-based software engineering. Medical standards for systematic reviews suggest that initial searches of candidate primary studies can rely on the title and the abstract to determine the eligibility of primary studies. Our experience indicates that software engineering abstracts are of such poor quality that it is often impossible to assess the eligibility of a primary study without reading some parts of the study itself. Medicine and psychology recommend the use of structure abstracts to improve the quality of abstracts. There have been several experimental studies of structured abstracts that have confirmed their value in psychology. Objective. This protocol defines a research plan aimed at assessing whether structured abstracts exhibit better structural properties than conventional abstracts for software engineering articles. Method. Conventional abstracts from two published proceedings of the EASE conference will be rewritten as structured abstracts. Metrics such as length, average sentence length, the Flesch readability index and amount and nature of added information will be obtained from different versions of the abstracts. The metrics for conventional and structured abstracts will be compared. Resource requirements. The experiment will require a research supervisor and two student researchers for a period of 8 weeks, and, in addition, support from other experienced research staff from the EBSE project.
BackgroundIn 2004 the concept of evidence-based software engineering (EBSE) was introduced at the ICSE04 conference.AimsThis study assesses the impact of systematic literature reviews (SLRs) which are the recommended EBSE method for aggregating evidence.MethodWe used the standard systematic literature review method employing a manual search of 10 journals and 4 conference proceedings.ResultsOf 20 relevant studies, eight addressed research trends rather than technique evaluation. Seven SLRs addressed cost estimation. The quality of SLRs was fair with only three scoring less than 2 out of 4.ConclusionsCurrently, the topic areas covered by SLRs are limited. European researchers, particularly those at the Simula Laboratory appear to be the leading exponents of systematic literature reviews. The series of cost estimation SLRs demonstrate the potential value of EBSE for synthesising evidence and making it available to practitioners.
For other domains that have adopted the evidence-based paradigm, the impact has included research outcomes having greater influence in terms of informing and influencing practitioners and policy-makers. We examine how evidence-based practices are being adapted for use in software engineering and discuss how decision-making in our own discipline can be liberated from over-reliance on expert judgement. To support our arguments we discuss some outcomes from recent studies and present an example in which performing a systematic literature review demonstrates the unreliability of depending only upon the outcomes of individual studies. Finally, we identify six challenges that need to be addressed in order to provide software engineers with standards and practices that are underpinned by evidence.
This study aims to compare the use of targeted manual searches with broad automated searches, and to assess the importance of grey literature and breadth of search on the outcomes of SLRs. We used a participant-observer multi-case embedded case study. Our two cases were a tertiary study of systematic literature reviews published between January 2004 and June 2007 based on a manual search of selected journals and conferences and a replication of that study based on a broad automated search. Broad searches find more papers than restricted searches, but the papers may be of poor quality. Researchers undertaking SLRs may be justified in using targeted manual searches if they intend to omit low quality papers; if publication bias is not an issue; or if they are assessing research trends in research methodologies.
We report our experiences with adapting the systematic review procedures to help consolidate software engineering knowledge, seen as a step towards being able to employ evidence-based practices to assist with decision making in IT. We describe the different studies performed, and illustrate our procedures through a fuller description of one of our reviews. Our work has demonstrated the value of a study protocol and identified a number of problems. We conclude that it is practical to employ systematic literature reviews in software engineering, but that some cultural changes are required, particularly for reporting of empirical studies.
When conducting a systematic literature review, researchers usually determine the relevance of primary studies on the basis of the title and abstract. However, experience indicates that the abstracts for many software engineering papers are of too poor a quality to be used for this purpose. A solution adopted in other domains is to employ structured abstracts to improve the quality of information provided. This study consists of a formal experiment to investigate whether structured abstracts are more complete and easier to understand than non-structured abstracts for papers that describe software engineering experiments. We constructed structured versions of the abstracts for a random selection of 25 papers describing software engineering experiments. The 64 participants were each presented with one abstract in its original unstructured form and one in a structured form, and for each one were asked to assess its clarity (measured on a scale of 1 to 10) and completeness (measured with a questionnaire that used 18 items). Based on a regression analysis that adjusted for participant, abstract, type of abstract seen first, knowledge of structured abstracts, software engineering role, and preference for conventional or structured abstracts, the use of structured abstracts increased the completeness score by 6.65 (SE 0.37, p < 0.001) and the clarity score by 2.98 (SE 0.23, p < 0.001). 57 participants reported their preferences regarding structured abstracts: 13 (23%) had no preference; 40 (70%) preferred structured abstracts; four preferred conventional abstracts. Many conventional software engineering abstracts omit important information. Our study is consistent with studies from other disciplines and confirms that structured abstracts can improve both information content and readability. Although care must be taken to develop appropriate structures for different types of article, we recommend that Software Engineering journals and conferences adopt structured abstracts.
A recent report on the state of the UK information technology (IT) industry based most of its findings and recommendations on expert opinion. It is surprising that the report was unable to incorporate more empirical evidence. This paper aims to assess whether it is necessary to base IT industry and academic policy on expert opinion rather than on empirical evidence. Current evidence related to the rate of project failure is identified and the methods used to accumulate that evidence discussed. This shows that the report failed to identify relevant evidence and most evidence related to project failure is based on convenience samples. The status of empirical research in the computing disciplines is reviewed showing that empirical evidence covers a restricted range of subjects and seldom addresses the 'Society' level of analysis. Other more robust designs that would address large-scale IT questions are discussed. We recommend adopting a more systematic approach to accumulating and reporting evidence. In addition, we propose using quasi-experimental designs developed and used in the social sciences to improve the methodology used for undertaking large-scale empirical studies in software engineering.
There is little empirical knowledge of the effectiveness of the object-oriented paradigm. To conduct a systematic review of the literature describing empirical studies of this paradigm. We undertook a Mapping Study of the literature. 138 papers have been identified and classified by topic, form of study involved, and source. The majority of empirical studies of OO (object oriented software) concentrate on metrics, relatively few consider effectiveness.
CONTEXT: Systematic literature reviews largely rely upon using the titles and abstracts of primary studies as the basis for determining their relevance.However, our experience indicates that the abstracts for software engineering papers are frequently of such poor quality they cannot be used to determine the relevance of papers.Both medicine and psychology recommend the use of structured abstracts to improve the quality of abstracts.AIM: This study investigates whether structured abstracts are more complete and easier to understand than non-structured abstracts for software engineering papers that describe experiments.METHOD: We constructed structured abstracts for a random selection of 25 papers describing software engineering experiments.The original abstract was assessed for clarity (assessed subjectively on a scale of 1 to 10) and completeness (measured with a questionnaire of 18 items) by the researcher who constructed the structured version.The structured abstract was reviewed for clarity and completeness by another member of the research team.We used a paired 't' test to compare the word length, clarity and completeness of the original and structured abstracts. RESULTS:The structured abstracts were significantly longer than the original abstracts (size difference =106.4 words with 95% confidence interval 78.1 to 134.7).However, the structured abstracts had a higher clarity score (clarity difference= 1.47 with 95% confidence interval 0.47 to 2.41) and were more complete (completeness difference=3.39with 95% confidence intervals 4.76 to 7.56). CONCLUSIONS:The results of this study are consistent with previous research on structured abstracts.However, in this study, the subjective estimates of completeness and clarity were made by the research team.Future work will solicit assessments of the structured and original abstracts from independent sources (students and researchers).
Context: The success of the evidence-based paradigm in other domains, especially medicine, has raised the question of how this might be employed in software engineering.Objectives: To report the research we are doing to evaluate problems associated with adopting the evidence-based paradigm in software engineering and identifying strategies to address these problems.Method: Currently the experimental paradigms used in a selected set of domains are being examined along with the experimental protocols that they employ. Our aim is to identify those domains that have generally similar characteristics to software engineering and to study the strategies that they employ to overcome the lack of rigorous empirical protocols. We are also undertaking a series of systematic literature reviews to identify the factors that may limit their applicability in the software engineering domain.Conclusions: We have identified two domains that experience problems with experimental protocols that are similar to those occurring for software engineering, and will investigate these further to assess whether the approaches used to aggregate evidence in these domains can be adapted for use in software engineering. Our experiences from performing systematic literature reviews are positive, but reveal infrastructure problems caused by poor indexing of the literature.
This workshop is concerned with defining the procedures that are needed to establish a sound empirical foundation for the practices of Software Engineering. Our goal is to begin building a community that will review, analyse, codify and promulgate software engineering experiences as well as to identify the processes and infrastructure that are needed to support these activities.
This paper reports two trials of an evaluation framework intended to evaluate novel software applications. The evaluation framework was originally developed to evaluate a risk-based software bidding model, and our first trial of using the framework was our evaluation of the bidding model. We found that the framework worked well as a validation framework but needed to be extended before it would be appropriate for evaluation. Subsequently, we compared our framework with a recently completed evaluation of a software tool undertaken as part of the Framework V CLARiFi project. In this case, we did not use the framework to guide the evaluation; we used the framework to see whether it would identify any weaknesses in the actual evaluation process. Activities recommended by the framework were not undertaken in the order suggested by the evaluation process and we found problems relating to that oversight surfaced during the tool evaluation activities. Our experiences suggest that the framework has some benefits but it also requires further practical testing.
This paper discusses the issues involved in evaluating a software bidding model. We found it difficult to assess the appropriateness of any model evaluation activities without a baseline or standard against which to assess them. This paper describes our attempt to construct such a baseline. We reviewed evaluation criteria used to assess cost models and an evaluation framework that was intended to assess the quality of requirements models. We developed an extended evaluation framework and an associated evaluation process that will be used to evaluate our bidding model. Furthermore, we suggest the evaluation framework might be suitable for evaluating other models derived from expert-opinion based influence diagrams.
We are often encouraged to follow experimental procedures in undertaking software engineering studies, however we should not do so blindly as often assumptions are made as part of that process that software engineering methods artefacts and processes breach. One such example is the use of crossover designs. We consider the case where there are period by treatment interactions, (i.e where the treatments are non-commutative) and demonstrate the hazards in using a cross-over design in these cases.
We discuss a method of developing a software bidding model that allows users to visualize the uncertainty involved in pricing decisions and make appropriate bid/no bid decisions. We present a generic bidding model developed using the modeling method. The model elements were identified after a review of bidding research in software and other industries. We describe the method we developed to validate our model and report the main results of our model validation, including the results of applying the model to four bidding scenarios.
Mark Turner合作论文数Case Western Reserve University12