Identifying specific molecular markers and developing sensitive detection methods are two of the fundamental requirements for detection and differential diagnosis of cancer. Toward this goal, we first performed cDNA array analysis using 65 non-small cell lung cancer and non-involved normal lung tissues. We then used several complementary statistical and analytical methods to examine gene expression profiles generated by us and others from four independent sets of normal and neoplastic lung tissues. We report here that several sets of roughly 20 genes were sufficient to provide a robust distinction between normal and neoplastic tissues of the lung. Next we assessed the predictive ability of these gene sets by using Flow-Thru Chips (FTC) (MetriGenix, Baltimore, MD) containing 20 genes to screen 48 primary lung tumours and normal lung tissues. Gene expression changes detected by FTC distinguished lung cancers from the normal lung tissues using an RNA amount equivalent to that present in as few as 300 cells. We also used an independent set of 24 genes and showed that their expression profile was equally effective when measured by quantitative polymerase chain reaction (Q-PCR). Our results demonstrate that lung cancers can be identified based on the expression patterns of just 20 genes and that this approach is applicable for cancer diagnosis, prognosis, and monitoring using small amount of tumour or biopsy samples.
We describe interactions between kinetic (moving) and static information displays. We have implemented "moxel" kinetic displays in a classic discovery platform with many standard information visualization and analytic tools, and experimented with interactions between them. Moxels, which generalize pixels, are an advanced, moving, form of iconographic display of the kind first developed in static form by Pickett and White (1966). As with the static graphic icons of those early displays, moxels provide a way of mapping together in one image multiple data variables, but with potentially more potency with in-place motion. We show examples of how the two kinds of displays have been integrated, and discuss issues with the integration of dynamic and static visualizations in a single environment. We discuss several interaction paradigms between them including linked brushing, multiple selections, and operations on selected regions.
Although there are a number of visualization systems to choose from when analyzing data, only a few of these allow for the integration of other visualization and analysis techniques. There are even fewer visualization toolkits and frameworks from which one can develop ones own visualization applications. Even within the research community, scientists either use what they can from the available tools or start from scratch to define a program in which they are able to develop new or modified visualization techniques and analysis algorithms. Presented here is a new general-purpose platform for constructing numerous visualization and analysis applications. The focus of this system is the design and experimentation of new techniques, and where the sharing of and integration with other tools becomes second nature. Moreover, this platform supports multiple large data sets, and the recording and visualizing of user sessions. Here we introduce the Universal Visualization Platform (UVP) as a modern data visualization and analysis system.
The challenge in developing advanced techniques for data visualization is to display large amounts of data for perceptual consumption. From early maps and graphs to n-dimensional displays and threedimensional images, graphical data displays are pushing the limits of human understanding. The increasing amount of data for analysis requires more capable displays. We discuss issues in extending visualization techniques to capitalize on the human perceptual system. Drawing on the work from J. J. Gibson and P. Robertson, we introduce the presentation of data via iconographic natural scene generation.
A bstract : Recent technical advances in combinatorial chemistry, genomics, and proteomics have made available large databases of biological and chemical information that have the potential to dramatically improve our understanding of cancer biology at the molecular level. Such an understanding of cancer biology could have a substantial impact on how we detect, diagnose, and manage cancer cases in the clinical setting. One of the biggest challenges facing clinical oncologists is how to extract clinically useful knowledge from the overwhelming amount of raw molecular data that are currently available. In this paper, we discuss how the exploratory data analysis techniques of machine learning and high-dimensional visualization can be applied to extract clinically useful knowledge from a heterogeneous assortment of molecular data. After an introductory overview of machine learning and visualization techniques, we describe two proprietary algorithms (PURS and RadViz™) that we have found to be useful in the exploratory analysis of large biological data sets. We next illustrate, by way of three examples, the applicability of these techniques to cancer detection, diagnosis, and management using three very different types of molecular data. We first discuss the use of our exploratory analysis techniques on proteomic mass spectroscopy data for the detection of ovarian cancer. Next, we discuss the diagnostic use of these techniques on gene expression data to differentiate between squamous and adenocarcinoma of the lung. Finally, we illustrate the use of such techniques in selecting from a database of chemical compounds those most effective in managing patients with melanoma versus leukemia.
A bstract : Recent technical advances in combinatorial chemistry, genomics, and proteomics have made available large databases of biological and chemical information that have the potential to dramatically improve our understanding of cancer biology at the molecular level. Such an understanding of cancer biology could have a substantial impact on how we detect, diagnose, and manage cancer cases in the clinical setting. One of the biggest challenges facing clinical oncologists is how to extract clinically useful knowledge from the overwhelming amount of raw molecular data that are currently available. In this paper, we discuss how the exploratory data analysis techniques of machine learning and high‐dimensional visualization can be applied to extract clinically useful knowledge from a heterogeneous assortment of molecular data. After an introductory overview of machine learning and visualization techniques, we describe two proprietary algorithms (PURS and RadViz™) that we have found to be useful in the exploratory analysis of large biological data sets. We next illustrate, by way of three examples, the applicability of these techniques to cancer detection, diagnosis, and management using three very different types of molecular data. We first discuss the use of our exploratory analysis techniques on proteomic mass spectroscopy data for the detection of ovarian cancer. Next, we discuss the diagnostic use of these techniques on gene expression data to differentiate between squamous and adenocarcinoma of the lung. Finally, we illustrate the use of such techniques in selecting from a database of chemical compounds those most effective in managing patients with melanoma versus leukemia.