ABSTRACT Inspections and testing are two of the most commonly performed software quality assurance processes today. Typically, these processes are applied in isolation, which, however, fails to exploit the benefits of systematically combining and integrating them. In consequence, tests are not focused on the basis of early defect detection data. Expected benefits of such process integration include higher defect detection rates or reduced quality assurance effort. Moreover, when conducting testing without any prior information regarding the system's quality, it is often unclear how to focus testing. A systematic integration of inspection and testing processes requires context‐specific knowledge about the relationships between inspections and testing. This knowledge is typically not available and needs to be empirically identified and validated. Often, context‐specific assumptions can be seen as a starting point for generating such knowledge. On the basis of the integrated inspection and testing approach, In 2 Test, which uses inspection data to focus testing, we present in this article how knowledge about the relationship between inspections and testing can be gained, documented, and evolved in an analytical or empirical manner. In addition, this article gives an overview of related work and highlights future research directions. Copyright © 2012 John Wiley & Sons, Ltd.
Inspections and testing are two of the most commonly performed software quality assurance processes today. Typically, these processes are applied in isolation, which, however, fails to exploit the benefits of systematically combining and integrating them. Expected benefits of such process integration are higher defect detection rates or reduced quality assurance effort. Moreover, when conducting testing without any prior information regarding the system's quality, it is often unclear which parts or which defect types should be prioritized. Existing approaches do not explicitly use information from inspections in a systematical way to focus testing processes. In this article, we present an integrated two-stage approach that routes inspection data to test processes in order to prioritize code classes and defect types. While an initial version of the approach focused on prioritizing code classes, this article focuses on the prioritization of defect types for testing. Results from a case study where the approach was applied on the code level show that those defect types could be prioritized before the testing that afterwards actually showed up most often during the test process. In addition, an overview of related work and an outlook on future research directions are given.
A systematic approach to decision making in software engineering is required, for instance, if an organization aims at achieving CMMI level three. Rational decision making regarding the selection and introduction of SE technologies requires adequate information about their suitability for the intended organizational context. Research is often unable to provide such information, and this could be one reason why promising techniques are sometimes not adopted in practice. From a research point of view, successful technology transfer requires knowing which information decision makers in industry need, and where they actually look for it. With this knowledge, empirical research can strive to produce the needed information in order to increase the likelihood of successful technology adoption. To address these questions, we conducted an online survey among German software industry decision makers. To focus the survey, we used inspections as an exemplary technology. We invited 9653 companies to participate, from which we received 92 fully completed questionnaires. Our main findings are that information regarding the impact of technologies on product quality, cost, and development time, as well as on technology cost-benefit ratio is considered most important among decision makers. The preferred sources of information are colleagues, textbooks, and industry workshops.
Defect measurement plays a crucial role when assessing quality assurance processes such as inspections and testing. To systematically combine these processes in the context of an integrated quality assurance strategy, measurement must provide empirical evidence on how effective these processes are and which types of defects are detected by which quality assurance process. Typically, defect classification schemes, such as ODC or the Hewlett-Packard scheme, are used to measure defects for this purpose. However, we found it difficult to transfer existing schemes to an embedded software context, where specific document- and defect types have to be considered. This paper presents an approach to define, introduce, and validate a customized defect classification scheme that considers the specifics of an industrial environment. The core of the approach is to combine the software engineering know-how of measurement experts and the domain know-how of developers. In addition to the approach, we present the results and experiences of using the approach in an industrial setting. The results indicate that our approach results in a defect classification scheme that allows classifying defects with good reliability, that allows identifying process improvement actions, and that can serve as a baseline for evaluating the impact of process improvements
Seit der Einführung von Use Cases hat deren Bedeutung zur Spezifikation von Anforderungen stetig zugenommen. Die Qualität der Use Cases ist ein entscheidender Faktor für den Erfolg des Entwicklungsprozesses, da die meisten Entwicklungsschritte auf den Use Cases aufbauen. Trotz der extremen Wichtigkeit der Qualität der Use Cases stellen die meisten use-case-basierten Entwicklungsansätze keine oder nur unzureichende integrierte qualitätssichernde Maßnahmen bereit (z.B. ad-hoc Empfehlungen, Erstellungsrichtlinien, einige Checklisten zur Inspektion von Use Cases). Diese Techniken werden in den meisten Fällen unabhängig voneinander eingesetzt, so dass bestimmte Fehlerklassen in den Use Cases durch mehrere Techniken, andere Fehlerklassen überhaupt nicht adressiert werden. In diesem Artikel wird ein integrierter Ansatz vorgestellt, in dem Use Case Erstellungsrichtlinien, Inspektionen und Simulation in systematischer Weise miteinander verknüpft werden. Der Ansatz basiert auf einer Fehlerklassifikation für Use Cases, die als Grundlage dient, die verschiedenen Techniken auf bestimmte Fehlerarten zu fokussieren .
There is a general agreement among software engineering practitioners that software inspections are an important technique to achieve high software quality at a reasonable cost. However, there are many ways to perform such inspections and many factors that affect their cost-effectiveness. It is therefore important to be able to estimate this cost-effectiveness in order to monitor it, improve it, and convince developers and management that the technology and related investments are worth while. This work proposes a rigorous but practical way to do so. In particular, a meaningful model to measure cost-effectiveness is proposed and a method to determine cost-effectiveness by combining project data and expert opinion is described. To demonstrate the feasibility of the proposed approach, the results of a large-scale industrial case study are presented and an initial validation is performed.
Inspections have been shown to be an effective means of detecting defects early on in the software development life cycle. However, they are not always successful or beneficial as they are affected by a number of technical and managerial factors. To make inspections successful, one important aspect is to understand what are the factors that affect inspection effectiveness (the rate of detected defects) in a given environment, based on project data. In this paper we collected data from over 230 code inspections and performed a multivariate statistical analysis in order to look at how management factors, such as the effort assigned and the inspection rate, affect inspection effectiveness. Because the functional form of effectiveness models is a priori unknown, we use a novel exploratory analysis technique: multiple adaptive regression splines (MARS). We compare the MARS model with more classical regression models and show how it can help understand the complex trends and interactions in the data, without requiring the analyst to rely on strong assumptions. Results are reported and discussed in light of existing studies.
One purpose of empirical software engineering is to enable an understanding of factors that influence software development. Surveys are an appropriate empirical strategy to gather data from a large population (e.g., about methods, tools, developers, companies) and to achieve an understanding of that population. Although surveys are quite often performed, for example, in social sciences and marketing research, they are underrepresented in empirical software engineering research, which most often uses controlled experiments and case studies. Consequently, also the methodological support how to perform such studies in software engineering is rather low. However, with the increasing pervasion of the Internet it is possible to perform surveys easily and cost-effectively over Internet pages (i.e., on-line), while at the same time the interest in performing surveys is growing. The purpose of this paper is twofold. First we want to arise the awareness of on-line surveys and discuss methods how to perform these in the context of software engineering. Second, we report our experience in performing on-line surveys in the form of lessons learned and guidelines.
Assessing and controlling software quality is still an immature discipline. One of the reasons for this is that many of the concepts and terms that are used in discussing and describing quality are overloaded with a history from manufacturing quality. We argue in this paper that a quite distinct approach is needed to software quality control as compared with manufacturing quality control. In particular, the emphasis in software quality control is in design to fulfil business needs, rather than replication to agreed standards. We will describe how quality goals can be derived from business needs. Following that, we will introduce an approach to quality control that uses rich causal models, which can take into account human as well as technological influences. A significant concern of developing such models is the limited sample sizes that are available for eliciting model parameters. In the final section of the paper we will show how expert judgement can be reliably used to elicit parameters in the absence of statistical data. In total this provides an agenda for developing a framework for quality control in software engineering that is freed from the shackles of an inappropriate legacy.
Software engineering processes depend on the context they are applied in. Thus, it is risky to determine the best processes for a given project context without empirical data from this context. This report summarizes important issues on the state of the art in empirical studies in order to provide researchers with the background to conduct their own empirical studies and for practitioners to classify existing empirical studies regarding their usefulness. This report gives an overview on the state of the art in empirical studies in software engineering by prescribing a high-level process for empirical studies , which is refined for the three most often used empirical strategies: Controlled experiments, case studies, and surveys. For each strategy this work discusses representative application reports. The appendix lists a glossary of empirical vocabulary, a standard report outline , a bibliography, and experimental material available on-line. 3 Controlled Experiments 19 3.1 Study Definition 21 3.1.1 Determine the goal of the study and an informal list of hypotheses 22 3.1.2 Record the goal in the measurement goal template 23 3.1.3 Determine context of the experiment 25 3.2 Experiment design 26 3.2.1 Operationalize the goal by identifying and quantifying the dependent and independent variables 26 3.2.2 Select an experimental design that satisfies the experiment purpose 27 3.2.3 Consider validity threats 31 3.2.4 Determine the necessary criteria for potential subjects 32 3.2.5 Determine options for data collection to increase the potential benefit from the experiment 33 3.3
Explicit risk management is gaining ground in industrial software development projects. However, there are few empirical studies that investigate the transfer of explicit risk management into industry, the adequacy of the risk management approaches to the constraints of industrial contexts, or their cost-benefit. This paper presents results from a case study that introduced a systematic risk management method, namely the Riskit method, into a large German telecommunication company. The objective of the case study was (1) to analyze the usefulness and adequacy of the Riskit method and (2) to analyze the cost-benefit of the Riskit method in this industrial context. The results of (1) also aimed at improvement and customization of the Riskit method. Moreover, we compare our findings with results of previous case studies to obtain more generalized conclusions on the Riskit method. Our results showed that the Riskit method is practical, adds value to the project, and that its key concepts are understood and usable in practice. Additionally, many lessons learned are reported that are useful for the general audience who wants to transfer risk management into new projects.
The development of high quality software satisfying cost, schedule, and resource requirements is an essential prerequisite for improved competitiveness of life insurance companies. One major difficulty to master this challenge is the inevitability of defects in software products. Since defects are known to be significantly more expensive if detected in later development phases or testing, companies in this marketplace must use cost-effective technologies to detect defects early on in the development process. A particular promising one is software inspection. This paper describes the ESPRIT/ESSI Process Improvement Experiment "High Quality of Software Products by Early Use of Innovative Reading Techniques (HYPER)". The core of this project has been the transfer of innovative software inspection technologies to the Allianz EURO conversion projects. The innovation in the area of software inspection is based on a systematic reading technique, that is, Perspective-based reading (PBR), that tells inspection participants what to look for and more important how to scrutinise a software artefact for defects. Although numerous controlled experiments have shown the PBR technique to be particularly cost-effective, few results have been reported on its use in the context of development projects. The paper presents in a quantitative manner the final results regarding the application of PBR inspections on requirements and design documents in the ESSI PIE. The results are based on 9 requirements and 44 design inspections and demonstrate the benefits to be expected from PBR inspections in an industrial environment.
Software inspection is one of the most effective methods to detect defects. Reinspection repeats the inspection process for software products that are suspected to contain a significant number of undetected defects after an initial inspection. As a reinspection is often believed to be less efficient than an inspection an important question is whether a reinspection justifies its cost.In this paper we propose a cost-benefit model for inspection and reinspection. We discuss the impact of cost and benefit parameters on the net gain of a reinspection with empirical data from an experiment in which 31 student teams inspected and reinspected a requirements document.Main findings of the experiment are: a) For reinspection benefits and net gain were significantly lower than for the initial inspection. Yet, the reinspection yielded a positive net gain for most teams with conservative cost-benefit assumptions. b) Both the estimated benefits and number of major defects are key factors for reinspection net gain, which emphasizes the need for appropriate estimation techniques.
An important requirement to control the inspection of software artifacts is to be able to decide, based on more objective information, whether the inspection can stop or whether it should continue to achieve a suitable level of artifact quality. A prediction of the number of remaining defects in an inspected artifact can be used for decision making. Several studies in software engineering have considered capture-recapture models to make a prediction. However, few studies compare the actual number of remaining defects to the one predicted by a capture-recapture model on real software engineering artifacts. The authors focus on traditional inspections and estimate, based on actual inspections data, the degree of accuracy of relevant state-of-the-art capture-recapture models for which statistical estimators exist. In order to assess their robustness, we look at the impact of the number of inspectors and the number of actual defects on the estimators' accuracy based on actual inspection data. Our results show that models are strongly affected by the number of inspectors, and therefore one must consider this factor before using capture-recapture models. When the number of inspectors is too small, no model is sufficiently accurate and underestimation may be substantial. In addition, some models perform better than others in a large number of conditions and plausible reasons are discussed. Based on our analyses, we recommend using a model taking into account that defects have different probabilities of being detected and the corresponding Jackknife Estimator. Furthermore, we calibrate the prediction models based on their relative error, as previously computed on other inspections. We identified theoretical limitations to this approach which were then confirmed by the data.
Software inspections have established an impressive track record for early defect detection and correction. To increase their benefits, recent research efforts have focused on two different areas: systematic reading techniques and defect content estimation techniques. While reading techniques are to provide guidance for inspection participants on how to scrutinize a software artifact in a systematic manner, defect content estimation techniques aim at controlling and evaluating the inspection process by providing an estimate of the total number of defects in an inspected document. Although several empirical studies have been conducted to evaluate the accuracy of defect content estimation techniques, only few consider the reading approach as an influential factor.In this paper we examine the impact of two specific reading techniques - a scenario-based reading technique and checklist-based reading - on the accuracy of different defect content estimation techniques. The examination is based on data that were collected in a large experiment with students of the Vienna University of Technology. The results suggest that the choice of the reading technique has little impact on the accuracy of defect content estimation techniques. Although more empirical work is necessary to corroborate this finding, it implies that practitioners can use defect content estimation techniques without any consideration of their current reading technique.
Inspections have been shown to be an effective means of detecting defects early on in the software development life cycle. However, they are not always successful or beneficial as they are affected by a number of technical and managerial factors. One important aspect is to understand what are the factors that affect inspection effectiveness (the rate of detected defects) in a given environment, based on project data. In this paper we look at management factors such as the effort assigned, the inspection rate, and so forth. We collected data on a number of analysis and code inspections, and performed a multivariate statistical analysis. Because the functional form of effectiveness models is a priori unknown, we use a novel exploratory analysis technique: Multiple Adaptive Regression Splines (MARS). We compare the MARS model with more classical regression models and show how it can help understand the complex trends and interactions in the data, without requiring the analyst to rely on strong assumptions. Results are reported and discussed in light of existing empirical results. 1. Introduction Inspections have been shown to be an important defect detection technology [9][13]. However, when one is faced with planning inspections, a number of decisions have to be made. For example, the following questions are considered relevant as they are deemed to have an impact on inspection effectiveness, that is the capacity of inspections to uncover defects: • What overall effort to devote to the inspection? • What should be the inspection rate? • How many participants to involve? • How should the material to be inspected be broken down? In order to answer such questions, which will be discussed in further details below, we need to develop models that relate defect detection effectiveness to variables such as effort, number of participants, or the amount of code inspected. To build such effectiveness models, data on inspections need to be collected and multivariate statistical
Talia, D., PK Srimani, and M. Jazayeri. Guest editors’ introduction: Special issues on architecture-independent languages and software tools for parallel processing [part I]; T-SE Mar 00 193-196 Talia, D., PK Srimani, and M. Jazayeri. Guest editors’ introduction: Special issue on architecture-independent languages and software tools for parallel processing [part II]; T-SE April 2000 289-292
Stefan Biffl合作论文数Department of Software Engineering, Institute of Information Systems Engineering, Technische Universitat Wien3
Barbara Paech合作论文数Institut für Informatik, Institute for Computer Science, Heidelberg University1
Sabine Castano合作论文数& Knowledge Management Lab at DICo;The Information Systems1