Static analysis (SA) tools examine code for flaws without executing the code, and produce warnings ("alerts") about possible flaws. A human auditor then evaluates the validity of the purported code flaws. The effort required to manually audit all alerts and repair all confirmed code flaws is often too much for a project's budget and schedule. An alert triaging tool enables strategically prioritizing alerts for examination, and could use classifier confidence. We developed and tested classification models that predict if static analysis alerts are true or false positives, using a novel combination of multiple static analysis tools, features from the alerts, alert fusion, code base metrics, and archived audit determinations. We developed classifiers using a partition of the data, then evaluated the performance of the classifier using standard measurements, including specificity, sensitivity, and accuracy. Test results and overall data analysis show accurate classifiers were developed, and specifically using multiple SA tools increased classifier accuracy, but labeled data for many types of flaws were inadequately represented (if at all) in the archive data, resulting in poor predictive accuracy for many of those flaws.
This causal discovery analysis is intended as an initial step, and is certainly not the final word. For example, one could apply multiple causal discovery algorithms to measure the sensitivity of the learned structures to the use of the PC algorithm. Moreover, software projects exhibit significant dynamics over time, as code is written, refined, refactored, and so forth. We used static datasets that provide snapshots of the projects at particular moments in time. If we collect longitudinal data about similar variables, then we could start to uncover the underlying causal dynamics. One might also suspect that those dynamics could shift over time, as the software practices and philosophies change, as project members enter and leave, etc. Longitudinal data could also enable us to test for this type of causal non-stationarity. The key point that we have established here, however, is the first demonstration of the applicability and usefulness of causal discovery algorithms applied to observational software engineering datasets.
: Coding errors cause the majority of software vulnerabilities. For example, 64% of the nearly 2,500 vulnerabilities in the National Vulnerability Database in 2004 were caused by programming errors. The CERT Division's Source Code Analysis Laboratory (SCALe) offers conformance testing of C language software systems against the CERT C Secure Coding Standard and the CERT Oracle Secure Coding Standard for Java, using various analysis tools available from commercial software vendors. Unfortunately, the current SCALe analysis process and tools do not collect any statistics about the accuracy of the code analysis tools or about the coding violations they flag, such as frequency of occurrence. This paper describes the approach used to add the ability to collect and statistically analyze data regarding coding violations and tool characteristics along with the initial results. The collected data will be used over time to improve the effectiveness of the SCALe analysis.
Abstract : Trust is a key factor in the effectiveness of the Wireless Emergency Alerts (WEA) service. Alert originators (AOs) must trust WEA to deliver alerts to the public in an accurate and timely manner. Members of the public must also trust the WEA service before they will act on the alerts that they receive. This research aimed to develop a trust model to enable the Federal Emergency Management Agency (FEMA) to maximize the effectiveness of WEA and provide guidance for AOs that would support them in using WEA in a manner that maximizes public safety. The research method included Bayesian belief networks to model trust in WEA because they enable reasoning about and modeling of uncertainty. The research approach was to build models that could predict the levels of AO trust and public trust in specific scenarios, validate these models using data collected from AOs and the public, and execute simulations on these models for numerous scenarios to identify recommendations to AOs and FEMA for actions to take that increase trust and actions to avoid that decrease trust. This report describes the process used to develop and validate the trust models and the resulting structure and functionality of the models.
Abstract : The work described in this report, part of a larger SEI research effort on Quantifying Uncertainty in Early Lifecycle Cost Estimation (QUELCE), aims to develop and validate methods for calibrating expert judgment. Reliable expert judgment is crucial across the program acquisition lifecycle for cost estimation, and perhaps most critically for tasks related to risk analysis and program management. This research is based on three field studies that compare and validate training techniques aimed at improving the participants' skills to enable more realistic judgments commensurate with their knowledge. Most of the study participants completed three batteries of software engineering domain-specific test questions. Some participants completed four batteries of questions about a variety of general knowledge topics for purposes of comparison. Results from both sets of questions showed improvement in the participants' recognition of their true uncertainty. The domain-specific training was accompanied by notable improvements in the relative accuracy of the participants' answers when more contextual information to the questions was given along with reference points about similar software systems. Moreover, the additional contextual information in the domain-specific training helped the participants improve the accuracy of their judgments while also reducing their uncertainty in making those judgments.
: Extensive cost overruns in major defense programs are common, and studies have identified poor cost estimation as a main contributor. Research and experience have identified several factors associated with poor cost estimates. These include the following: (1) optimistic expectations about the program's scope and technology such that it can be delivered on schedule and within budget; (2) the enormous amount of unknowns and uncertainty that exist when these estimates are made about large-scale, unprecedented systems that take years to develop and deploy; and (3) the heavy reliance, of necessity, on expert judgment. In this paper, we describe a new, integrative approach for pre-Milestone A cost estimation called quantifying uncertainty in early life cycle cost estimation (QUELCE). QUELCE synthesizes scenario building, Bayesian belief network modeling, and Monte Carlo simulation into an estimation method that quantifies uncertainties, allows subjective inputs, visually depicts influential relationships among change drivers and outputs, and assists with explicit description and documentation underlying an estimate. We use scenario analysis and dependency structure matrix techniques to limit the combinatorial effects of multiple interacting program change drivers to make modeling and analysis more tractable. Finally, we describe results and insights gained from applying the method retrospectively to a major defense program.
: The Source Code Analysis Laboratory (SCALe) is a proof-of-concept demonstration that software systems can be conformance tested against secure coding standards. CERT' secure coding standards provide a detailed enumeration of coding errors that have resulted in vulnerabilities for commonly used software development languages. The SCALe team at the CERT Program, part of Carnegie Mellon University's Software Engineering Institute, analyzes a developer's source code and provides a detailed report of findings to guide the code's repair. After the developer has addressed these findings and the SCALe team determines that the product version conforms to the standard, the CERT Program issues the developer a certificate and lists the system in a registry of conforming systems. This report details the SCALe process and provides an analysis of selected software systems.
: Difficulties with estimating the costs of developing new systems have been well documented, and are compounded by the fact that estimates are now prepared much earlier in the acquisition lifecycle, before there is concrete technical information available on the particular program to be developed. This report describes an innovative synthesis of analytical techniques into a cost estimation method that models and quantifies the uncertainties associated with early lifecycle cost estimation. The method described in this report synthesizes scenario building, Bayesian Belief Network (BBN) modeling and Monte Carlo simulation into an estimation method that quantifies uncertainties, allows subjective inputs, visually depicts influential relationships among program change drivers and outputs, and assists with the explicit description and documentation underlying an estimate. It uses scenario analysis and design structure matrix (DSM) techniques to limit the combinatorial effects of multiple interacting program change drivers to make modeling and analysis more tractable. Representing scenarios as BBNs enables sensitivity analysis, exploration of scenarios, and quantification of uncertainty. The methods link to existing cost estimation methods and tools to leverage their cost estimation relationships and calibration. As a result, cost estimates are embedded within clearly defined confidence intervals and explicitly associated with specific program scenarios or alternate futures.
: There has been a great deal of discussion of late about what it takes for organizations to attain high maturity status and what they can reasonably expect to gain by doing so. Clarification is needed along with good examples of what has worked well and what has not. This may be particularly so with respect to measurement and analysis. This report contains results from a survey of high maturity organizations conducted by the Software Engineering Institute (SEI) in 2008. The questions center on the use of process performance modeling in those organizations and the value added by that use. The results show considerable understanding and use of process performance models among the organizations surveyed; however there is also wide variation in the respondents? answers. The same is true for the survey respondents? judgments about how useful process performance models have been for their organizations. As is true for less mature organizations, there is room for continuous improvement among high maturity organizations. Nevertheless, the respondents? judgments about the value added by doing process performance modeling also vary predictably as a function of the understanding and use of the models in their respective organizations. More widespread adoption and improved understanding of what constitutes a suitable process performance model holds promise to improve CMMI-based performance outcomes considerably.