In this article, we report on an experiment to assess the possibility of rigorous evaluation of interactive question-answering (QA) systems using the cross-evaluation method. This method takes into account the effects of tasks and context, and of the users of the systems. Statistical techniques are used to remove these effects, isolating the effect of the system itself. The results show that this approach yields meaningful measurements of the impact of systems on user task performance, using a surprisingly small number of subjects and without relying on predetermined judgments of the quality, or of the relevance of materials. We conclude that the method is indeed effective for comparing end-to-end QA systems, and for comparing interactive systems with high efficiency.
Report generation is an integral part of many analytical tasks such as intelligence analysis. Thus, an important issue in developing cognitive assistants for analytical tasks is how the cognitive assistant may help a human analyst in generating reports. In this paper, we first describe the task of report generation in intelligence analysis. Then, we describe a scheme for enabling a cognitive assistant to generate self-explanations. The proposed scheme uses introspection over the knowledge, reasoning, and conclusions of the cognitive assistant. Finally, we analyze the introspective scheme for generating self-explanations from the perspective of report generation.
Interest in the environmental factors that affect biometric image quality is increasing as biometric technologies are currently being implemented in various business applications. This study aims to determine, through repeated trials, the effects of various external factors on the image quality and usability of prints collected by an electronic reader. These factors include age and gender but also the absence or presence of immediate feedback. A key factor in biometric systems that will be used daily or routinely is habituation. The user's behavior could potentially change as a result of acclimatization; one's input might increase in quality as one learns how to use the system better, or decrease in quality since comfort with the system could translate into carelessness.
In this paper we discuss the evaluation method-ologies and metrics we have developed for ARDA's Novel Intelligence for Massive Data (NIMD) program. We discuss the requirements for developing methods and metrics in a situation where software components that were to be tested were in very early stages of development and where investigators who might be on the leading edge with respect to their technology were novices with respect to evaluation. Additionally , we discuss how our process of evaluation design is evolving as we gain experience with metrics and measures that are obtainable, yet have some value as indicators of future software performance in the field.
In this paper we present both a broad and deep look at the use of a collaboration tool in the intelligence community. Through an experimental program, intelligence analysts are given the opportunity to explore and use tools to determine if the tools provide sufficient value to be certified and moved into the analytic work environment. The goal of this program is to bring advanced technologies to the intelligence community through research and experimentation. New tools are evaluated using a metrics based assessment. Tools that successfully pass these evaluations are then introduced on an experimental network. Analysts employed by the experimental program work along side analysts in the intelligence community and look for opportunities where the experimental tools could be useful in current analytic processes. These uses are also evaluated to determine the value of the tools in the analytic environment.
Information retrieval has a long history of dealing with printed materials. More recent work has involved the development of experimental visual interfaces to support users’ attempts to access appropriate documents. This research has matured to the point that usability studies and evaluation of approaches to information visualization are needed to guide further development. The reported studies examine the use of alternative document visualizations in tightly controlled settings. Five types of interface representations were defined, including ordered text, ordered icons, a table format, a x-y graph format and a novel spring-based visualization. To assess the relative utility of the various interfaces, we have chosen to apply two kinds of measures: performance on information retrieval tasks and user preference rankings of the interfaces. The results show that performance is strikingly different across the range of interface types with the ordered icon list and text list producing the best results. Users’ preferences, however, indicated that the textual format was the least desirable, while both of the visualization methods, i.e., icon list and spring-based visual, were preferred. We conclude that performance is more easily and accurately measured and that preferences of users can not be used alone to determine the utility of interfaces.
AbstractA number of efforts are being undertaken to integrate usability engineering and software engineering in the software‐development process. The majority of these integration efforts focus on software developers, usability engineers, or defining new processes. In this article, we report on an effort to involve the consumer of software by providing a mechanism, namely the common industry format (CIF), to formally request usability information on the software to be purchased. Copyright © 2004 John Wiley & Sons, Ltd.
Although many different visual information retrieval systems have been proposed, few have been tested, and where testing has been performed, results were often inconclusive. Further, there is very little evidence of benchmarking systems against a common standard. An approach for testing novel interfaces is proposed that uses bottom-up, stepwise testing to allow evaluation of a visualization, itself, rather than restricting evaluation to the system instantiating it. This approach not only makes it easier to control variables, but the tests are also easier to perform. The methodology will be presented through a case study, where a new visualization technique is compared to more traditional ways of presenting data.
The Industry US ability Reporting (IUSR) Project seeks to help potential corporate consumers of software obtain information about the usability of supplier products, to measure the benefit of more usable software, and to increase communication about usability needs between consumers and suppliers. Human factors and software engineers have developed a Common Industry Format (ANSI/NCITS 354–2001) for sharing usability information. Four pilot studies were conducted by industry which verify its usefulness in procurement and assess the costs and benefits of including usability test results in the software purchase process. Use of the Common Industry Format can increase communications across corporate boundaries and help improve the usability of software for consumers. The standard may also be applicable to setting usability requirements, and measuring usability of websites, hardware, and universal access.
We describe the design and use of SMAT (Synchronous Multimedia and Annotation Tool), a tool designed to be part of a scientific collaboratory for use in a robotic, arc-welding research project at the National Institute of Standards and Technology (NIST). The primary functional requirements of SMAT are to provide the capability to capture, synchronize, play back, and annotate multimedia data in a multi-platform, distributed environment. To meet these requirements, SMAT was designed as a control and integration framework that exploits existing tools to render specific media types and control annotation sessions. SMAT defines a component architecture framework where existing tools can be plugged in and controlled using a distributed, event-driven, tool-bus architecture. SMAT's modular architecture enables control inputs to come from anywhere in the distributed collaborative environment, thus allowing for simultaneous remote and local control of the tool, as well as painless interfacing with the existing collaborative environment. SMAT is built on an agent middleware called AGNI (Agents at NIST), also developed at NIST. We give an overview of AGNI that can be used to build failure-resilient, distributed, event-driven applications. In addition to describing SMAT's design, interface and underlying middleware, we present performance information, an initial analysis of welding users' experiences and feedback, related work, and directions for further SMAT development.
Many researchers believe that groupware can only be evaluated by studying real collaborators in their real contexts, a process that tends to be expensive and time-consuming. Others believe that it is more practical to evaluate groupware through usability inspection methods. Deciding between these two approaches is difficult, because it is unclear how they compare in a real evaluation situation. To address this problem, we carried out a dual evaluation of a groupware system, with one evaluation applying user-based techniques, and the other using inspection methods. We compared the results from the two evaluations and concluded that, while the two methods have their own strengths, weaknesses, and trade-offs, they are complementary. Because the two methods found overlapping problems, we expect that they can be used in tandem to good effect, e.g., applying the discount method prior to a field study, with the expectation that the system deployed in the more expensive field study has a better chance of doing well because some pertinent usability problems will have already been addressed.
Previously, we analyzed data from a field study to determine the usability problems in a groupware system (M. Steves et al., 2001). The fact that 75% of the system use was asynchronous, surprised us. Consequently, we suspected that the user centered method that had been employed in that evaluation, might have been insufficient to detect problems related to such heavy asynchronous use. Re-analysis of the data using an artifact-centered approach revealed additional support for our initial findings and some new usability issues. In combination, we believe that user-centered and artifact centered methods can yield superior usability analyses for systems that support both synchronous and asynchronous collaboration, and provide application developers with more appropriate priorities for addressing usability problems and system design