Recently a new class of applications has emerged that uses AJAX and other web programming techniques to provide a rich user experience in a web browser. This class of applications is being called Web 2.0 and includes Google Maps and Google Suggest. To experiment with this approach, we have developed a Web 2.0 thin client collaborative visualization framework called GeoBoost(TM) that uses Scalable Vector Graphics and AJAX to provide a rich user experience built around collaboration. Our framework includes geospatial maps, standard business charts, node and link displays, and custom visual displays. All of our visualization components run in standard web browsers and provide rich interaction and collaboration.
The TRAQ-M (tracking analysis, quantification-mitigation) platform is a computational system that applies Social Science Models and nonparametric statistical methods to understand complex human behavior patterns. The system focuses on understanding patterns in language and is capable of ingesting millions of documents per day and identifying linguistic patterns. In this paper we focus on source modeling, especially determination of which sources pass false information or are otherwise biased We use nonparametric statistical models to compare document content histograms and linguistic pattern analysis to identity disinformation.
Thin-client Web interfaces are rapidly becoming the de facto standard for accessing information. Unlike traditional client-server, browser-based Web applications that provide an intermittent user interface, thin-client Web applications have the added benefit of a direct-manipulation user interface that is typical of desktop applications. These interfaces are greatly preferred by users, but were not possible on the Web because of latency. Technologies that combine the functionality of a desktop application with the wide reach of a browser are called Web 2.0 (see Figure 1). In Web 2.0, JavaScript code in the browser caches and displays information that has been asynchronously downloaded from the server using extended markup language (XML). (The programming paradigm for this Asynchronous JavaScript XML is called AJAX.) Additional JavaScript code handles panning, zooming, scaling, and data validation. Web 2.0 allows users to work with few delays and provides an interface that can be directly manipulated. In our research, we addressed some of the remaining limitations of this approach by creating a thin-client mapping framework. Dynamic web sites such as Google Gmail, Flikr, and Orkut are based on AJAX. Recently, several new Web applications, such as Google Maps, Microsoft’s Virtual Earth, and Google Suggest, combine AJAX, dynamic hypertext markup language (DHTML), and vector graphics. In these applications, manipulations occur almost instantly without reloading pages. For example, in Microsoft’s Virtual Earth, the user clicks on the map and scrolls to the targeted location with the cursor. Google Suggest automatically attempts to complete a search query. Google and Microsoft have added unique technical features to their products. In addition to map image tiles being Figure 1. Web 2.0 eliminates the intermittent nature of traditional Web applications, provides a direct manipulation user interface, and maintains the wide reach of a Web browser.
Real-time traffic management encompasses many aspects of highway traffic, such as the analysis of congestion levels, incident detection and classification, traffic forecasting, and visualization of all the above. The work presented in this paper focuses on real-time detection and visualization of unusual changes in traffic in the Chicago metropolitan area. Our Chicago Alert System (CAS) considers both the spatial and temporal aspects of the data by clustering sensors that report similar traffic flows and by building baselines that capture the seasonality and variation of data over the period of a year. Outlier values of the traffic flow are then detected using the baseline models . Real-time alerts are visually displayed through an online Web Service. We discuss analytic refinements of the system, including continuous updating of baselines and modeling to include weather, holidays, and other exogenous variables.
In this paper we examine the applicability of several well know document visualization programs to the newly arising problem of visualizing the processing of streams of textual datasets. In the military intelligence community, many new forms of information are arriving as textual streams. In addition to the traditional wire services and intelligence summaries available in electronic form, the Internet and 24/7 international news coverage present huge quantities of open source data to the analyst in streaming textual form. When dealing with the modern asymmetric threat, modern analysts are forced to utilize open sources of information more and more. These analysts are hard pressed to handle this new influx of information using traditional relational data base search and aggregate techniques. We identify several view points into a modern textual stream processing flow where the operation and efficiency of the system can be improved by the appropriate visualization technique. We examine some of the physical scalability limits of visualizing large quantities of information. We then put forth several criteria and evaluate some well known text visualization systems for use in this new environment
There is a need within the intelligence communities to analyze massive streams of multilingual unstructured data. Mathematical transformation algorithms have proven effective at interpreting multilingual, unstructured data, but high computational requirements of such algorithms prevent their widespread use. The rate of computation can be vastly increased with field programmable gate array (FPGA) hardware. To experiment with this approach, we developed a system with FPGAs that ingests content over a network at high data rates. The system extracts basewords, counts words, scores documents, and discovers concepts on data that are carried in TCP/IP network flows as packets over a Gigabit Ethernet link or in cells transported over an OC48 link. These algorithms, as implemented in FPGA hardware, introduce certain constraints on the complexity and richness of the semantic processing algorithms. To understand the implications of these constraints and to benchmark the performance of the system, we have performed a series of experiments processing multilingual documents. In these experiments, we compare techniques to generate basewords for our semantic concepts, score documents, and discover concepts across a variety of processing operational scenarios.
Data quality is a critical issue for the success of data-driven enterprises. The challenge for these enterprises is to provide accurate data inputs, correct codings, and accurate processing so that resulting data products are correct, accurate, and timely. Although one might think that the digitization of business and government would lead to better data, if anything, the reverse appears to be true. In our experience business data warehouses and data marts inevitably contain large amounts of poor quality data. Thus there is a need for better tools to help analysts identify and fix data quality problems. To meet this need we have created a data quality visualization tool called DaVis (Data Quality Visualizer). DaVis uses a tabular reduced visual representation to show a dataset, highlights inaccuracies and invalid data, and shows difference between versions of a dataset. Our experience in using DaVis on several consulting projects is that data quality visualization is quite useful in practice and that applying visualization techniques to address data quality problems is a fruitful research direction. CR
This paper describes a graphical method for visualizing reference database searches. The motivation for inventing this technique comes from analyzing the Current Index of Statistics reference database. This database contains 128 thousand references to articles from statistical journals, conference proceedings, and books, published during the last 20 years. The paper traces the evolution of the bootstrap technique, a statistical research breakthrough, in the statistical literature, shows yearly trends, discovers which journal publish articles on bootstrapping, and identifies books on this subject.
Next generation data processing systems must deal with very high data ingest rates and massive volumes of data. Such conditions are typically encountered in the Intelligence Community (IC) where analysts must search through huge volumes of data in order to gather evidence to support or refute their hypotheses. Their effort is made all the more difficult given that the data appears as unstructured text that is written in multiple languages using characters that have different encodings. Human Analysts have not been able to keep pace with reading the data and a large amount of data is discarded even though it might contain key information. The goal of our project is to assess the feasibility of incrementally replacing humans with automation in key areas of information processing. These areas include document ingest, content categorization, language translation, and context-and-temporally-based information retrieval.
This paper describes the design and construction of a new visualization system for collections of heterogeneous information for intelligence analysis. The system has several novel features that taken together provide a highly modular and reusable framework for creating linked visual metaphors. The system leverages modem web technologies such as XML DOM and Schemas to create an expressive and powerful system. Examples of this are the ability to use the structure of the information being visualized (as expressed in an XML schema) to directly generate the object oriented code for manipulating that information, and the use of static and dynamic binding facilities for creating mappings between internal information items and visualization components. The resulting system blurs the distinction between information visualization and the World Wide Web. It is very modular and flexible, and supports rapid iteration and refinement of input data sources as well as interoperability with other information schemas. New linked visual metaphors can easily be added to the framework as required.
This paper describes a system for analyzing the flow of traffic through web-sites. We decomposed the general path analysis problem into a set of distinct subproblems, and created a visual metaphor for analyzing each of them. Our system works off of multiple representations of the clickstream, and exposes the path extraction algorithms and data to the visual metaphors as web services. We have combined the visual metaphors into a web-based "path analysis portal" that lets the user easily switch between the different modes of analysis.
On August 26, 2001 a one-day workshop on Visual Data Mining was held in conjunction with the KDD-2001 conference. About 50 people attended the workshop, mostly from industry with some academics. The audience included both data mining algorithm experts and visualization specialists. During the workshop thirteen peer-reviewed papers were presented that treated both problemspecific issues as well as broader topics.
The explosive growth on-line activity has established the e-channel as a critical component of how institutions interact with their customers. One of the unique aspects of this channel is the rich instrumentation where it is literally possible to capture every visit, click, page view, purchasing decision, and other fine-grained details describing visitor browsing patterns. The problem is that the huge volume of relevant data overwhelms conventional analysis tools. To overcome this problem, we have deve loped a sequence of novel metaphors for visualizing website structure, paths and flow through the site, and website activity. The useful aspect of our tools is that they provide a rich visual interactive workspace for performing ad hoc analysis, discovering patterns, and identifying correlations that are impossible to find with traditional nonvisual tools.
A key problem in software engineering is changing the code. We present a sequence of visualizations and visual metaphors designed to help engineers understand and manage the software change process. The principal metaphors are matrix views, cityscapes, bar and pie charts, data sheets and networks. Linked by selection mechanisms, multiple views are combined to form perspectives that both enable discovery of high-level structure in software change data and allow effective access to details of those data. Use of the views and perspectives is illustrated in two important contexts: understanding software change by exploration of software change data and management of software development. Our approach complements existing visualizations of software structure and software execution.
Audris Mockus合作论文数Min H. Kao Department of Electrical Engineering and Computer Science, Tickle College of Engineering, University of Tennessee5