2404 BiTE antibodies are emerging as new class of uniquely T cell-engaging antibodies. A CD19/CD3-bispecific BiTE antibody called MT103/MEDI-538 has shown objective complete and partial responses in relapsed non-Hodgkin’s lymphoma (NHL) patients in an ongoing phase 1 study, providing clinical proof-of-concept for this new class of antibody-derived therapeutics. A number of other BiTE antibodies are in clinical and pre-clinical development targeting EpCAM (CD326), CEA and receptor tyrosine kinase EphA2. As a novel approach for treatment of acute myeloid leukemia (AML) and melanoma, we have constructed new BiTE antibodies with similar in-vitro potency and other properties as observed for more advanced clinical and pre-clinical BiTE antibody candidates. One BiTE antibody for treating AML is specific for CD33, a target validated by the immunotoxin gemtuzumab ozogamicin. Even with current treatment regimens, only about 20 percent of AML patients survive five or more years. The high efficacy of MT103 in NHL suggests that also other blood borne malignancies such as AML will be responsive to the therapeutic principle of BiTE antibodies. The other BiTE antibody is specific for melanoma-associated chondroitin sulfate proteoglycan (MCSP), which is among the best characterized melanoma surface antigens. Melanoma occasionally well respond to T cell-based therapies, such as adoptive T cell transfer or vaccination, but are frequently limited by immune escape mechanism, such as loss of MHC class I expression, to which BiTE antibodies are insensitive. CD33- and MCSP-specific BiTE antibodies of human sequence were selected as preclinical lead candidates based on their potency of redirected target cell lysis, and a superior profile of biophysical and pharmaceutical characteristics. CD33- and MCSP-specific BiTE antibodies showed redirected lysis by human peripheral T cells of CD33- and MCSP-expressing cell lines, respectively, at half-maximum concentrations as low as 5 pg/ml (90 femtomolar). Lysis was highly specific because it did not occur with CD33- or MCSP-negative cell lines, or with a BiTE antibody solely sharing the anti-CD3 portion with CD33- and MCSP-specific BiTE antibodies. CD33- and MCSP-specific BiTE antibodies are being investigated for safety in ongoing animal studies. Preclinical candidates for both CD33-specific and MCSP-specific BiTE antibodies may provide highly efficacious approaches for the treatment of AML and melanoma, respectively, both of which are diseases with a high need for new treatment options.
We have carried out supervised machine learning on a subset (8741 compounds) of the public NCI cancer compound library screened for effectiveness against 60 cancer cell lines. Our focus was on identifying quinone compounds and we found these to be over four-fold enriched compared to the entire NCI cancer compound library. Two-class classifications based upon the cell types' tumor tissue origin classes, identified subsets of compounds that were most effective against either melanoma or leukemia cancer cell types. Both of these compound subsets were enriched in quinone compounds.
A bstract : Recent technical advances in combinatorial chemistry, genomics, and proteomics have made available large databases of biological and chemical information that have the potential to dramatically improve our understanding of cancer biology at the molecular level. Such an understanding of cancer biology could have a substantial impact on how we detect, diagnose, and manage cancer cases in the clinical setting. One of the biggest challenges facing clinical oncologists is how to extract clinically useful knowledge from the overwhelming amount of raw molecular data that are currently available. In this paper, we discuss how the exploratory data analysis techniques of machine learning and high‐dimensional visualization can be applied to extract clinically useful knowledge from a heterogeneous assortment of molecular data. After an introductory overview of machine learning and visualization techniques, we describe two proprietary algorithms (PURS and RadViz™) that we have found to be useful in the exploratory analysis of large biological data sets. We next illustrate, by way of three examples, the applicability of these techniques to cancer detection, diagnosis, and management using three very different types of molecular data. We first discuss the use of our exploratory analysis techniques on proteomic mass spectroscopy data for the detection of ovarian cancer. Next, we discuss the diagnostic use of these techniques on gene expression data to differentiate between squamous and adenocarcinoma of the lung. Finally, we illustrate the use of such techniques in selecting from a database of chemical compounds those most effective in managing patients with melanoma versus leukemia.
Using data mining techniques, we have studied a subset (1400) of compounds from the large public National Cancer Institute (NCI) compounds data repository. We first carried out a functional class identity assignment for the 60 NCI cancer testing cell lines via hierarchical clustering of gene expression data. Comprised of nine clinical tissue types, the 60 cell lines were placed into six classes-melanoma, leukemia, renal, lung, and colorectal, and the sixth class was comprised of mixed tissue cell lines not found in any of the other five classes. We then carried out supervised machine learning, using the GI(50) values tested on a panel of 60 NCI cancer cell lines. For separate 3-class and 2-class problem clustering, we successfully carried out clear cell line class separation at high stringency, p < 0.01 (Bonferroni corrected t-statistic), using feature reduction clustering algorithms embedded in RadViz, an integrated high dimensional analytic and visualization tool. We started with the 1400 compound GI(50) values as input and selected only those compounds, or features, significant in carrying out the classification. With this approach, we identified two small sets of compounds that were most effective in carrying out complete class separation of the melanoma, non-melanoma classes and leukemia, non-leukemia classes. To validate these results, we showed that these two compound sets' GI(50) values were highly accurate classifiers using five standard analytical algorithms. One compound set was most effective against the melanoma class cell lines (14 compounds), and the other set was most effective against the leukemia class cell lines (30 compounds). The two compound classes were both significantly enriched in two different types of substituted p-quinones. The melanoma cell line class of 14 compounds was comprised of 11 compounds that were internal substituted p-quinones, and the leukemia cell line class of 30 compounds was comprised of 6 compounds that were external substituted p-quinones. Attempts to subclassify melanoma or leukemia cell lines based upon their clinical cancer subtype met with limited success. For example, using GI(50) values for the 30 compounds we identified as effective against all leukemia cell lines, we could subclassify acute lymphoblastic leukemia (ALL) origin cell lines from non-ALL leukemia origin cell lines without significant overlap from non-leukemia cell lines. Based upon clustering using GI(50) values for the 60 cancer cell lines laid out by the RadViz algorithm, these two compound subsets did not overlap with clusters containing any of the NCI's 92 compounds of known mechanism of action, a few of which are quinones. Given their structural patterns, the two p-quinone subtypes we identified would clearly be expected to possess different redox potentials/substrate specificities for enzymatic reduction in vivo. These two p-quinone subtypes represent valuable information that may be used in the elucidation of pharmacophores for the design of compounds to treat these two cancer tissue types in the clinic.
New sets of powerful data visualization tools have appeared in the marketplace and in the research community.This, combined with readily available computer memory, speed, and graphics capabilities, makes it possible to explore larger and larger data sets.However, it is difficult to judge the effectiveness of the,se tools for supporting lai'ge scale information exploration and knowledge discovery.In this paper, we describe a set of issues critical to benchmarking and evaluation in this domain.We then propose an approach to constructing an evaluation environment and report on initial results from a prototype environment in which we tested five visualization approaches against nine existing data sets. 1 visualizations or the data mining.To remedy the situation, it is becoming increasingly important to develop appropriate data sets and reproducible benchmark tests to identify the current best practices and to steer development of future systems.In this paper, we discuss some of the issues that need to be addressed in order to provide benchmark testing and evaluation to the visualization and data mining communities.We survey evaluation approaches that have been applied in other
We introduce a graphic primitive, called a dimensional anchor (DA), which facilitates the creation of new visualizations and provides insight into the analysis of information visualizations. The DA represents an attempt to provide a unified framework or model for a variety of visualizations, including Parallel Coordinates, scatter plot matrices, Radviz, Survey Plots and Circle Segments A dimensional anchor is constructed by assigning values to parameters associated with various geometric graphic elements that encode the basics of the above visualizations. We define a visualization vector space in which all of the above visualizations and many new ones are represented by vectors. These encodings make it possible to perform a Grand Tour traveling from Parallel Coordinates to Survey Plot, and visiting many other visualizations in between
Describes data exploration techniques designed to classify DNA sequences. Several visualization and data mining techniques were used to validate and attempt to discover new methods for distinguishing coding DNA sequences (exons) from non-coding DNA sequences (introns). The goal of the data mining was to see whether some other, possibly non-linear combination of the fundamental position-dependent DNA nucleotide frequency values could be a better predictor than the AMI (average mutual information). We tried many different classification techniques including rule-based classifiers and neural networks. We also used visualization of both the original data and the results of the data mining to help verify patterns and to understand the distinction between the different types of data and classifications. In particular, the visualization helped us develop refinements to neural network classifiers, which have accuracies as high as any known method. Finally, we discuss the interactions between visualization and data mining and suggest an integrated approach.
Visualizations that can handle flat files, or simple table data are most often used in data mining. In this paper we survey most visualizations that can handle more than three dimensions and fit our definition of Table Visualizations. We define Table Visualizations and some additional terms needed for the Table Visualization descriptions. For a preliminary evaluation of some of these visualizations see “Benchmark Development for the Evaluation of Visualization for Data Mining” also included in this volume. Data Sets Used Most of the datasets for the visualization examples are either the automobile or the Iris flower dataset. Nearly every data mining package comes with at least one of these two datasets. The datasets are available UC Irvine Machine Learning Repository [Uci97]. • Iris Plant Flowers – from Fischer 1936, physical measurements from three types of flowers. • Car (Automobile) – data concerning cars manufactured in America, Japan and Europe from 1970 to 1982 Definition of Table Visualizations A two-dimensional table of data is defined by M rows and N columns. A visualization of this data is termed a Table Visualization. In our definition, we define the columns to be the dimensions or the variates (also called fields or attributes), and the rows to be the data records. The data records are sometimes called ndimensional points, or cases. For a more thorough discussion of the table model, see [Car99]. This very general definition only rules out some structured or hierarchical data. In the most general case, a visualization maps certain dimensions to certain features in the visualization. In geographical, scientific, and imaging visualizations, the spatial dimensions are normally assigned to the appropriate X, Y or Z spatial dimension. In a typical information visualization there is no inherent spatial dimension, but quite often the dimension mapped to height and width on the screen has a dominating effect. For example in a scatter plot of four-dimensional data one could map two features to the Xand Y-axis and the other two features to the color and shape of the plotted points. The dimensions assigned to the Xand Y-axis would dominate many aspects of analysis, such as clustering and outlier detection. Some Table Visualizations such as Parallel Coordinates, Survey Plots, or Radviz, treat all of the data dimensions equally. We call these Regular Table Visualizations (RTVs). The data in a Table Visualizations is discrete. The data can be represented by different types, such as integer, real, categorical, nominal, etc. In most visualizations all data is converted to a real type before rendering the visualization. We are concerned with issues that arise from the various types of data, and use the more general term “Table Visualization.” These visualizations can also be called “Array Visualizations” because all the data are of the same type. Table Visualization data is not hierarchical. It does not explicitly contain internal structure or links. The data has a finite size (N and M are bounded). The data can be viewed as M points having N dimensions or features. The order of the table can sometimes be considered another dimension, which is an ordered sequence of integer values from 1 to M. If the table represents points in some other sequence such as a time series, that information should be represented as another column.