Abstract. We study the problem of how to cover a polygonal region by a sma ll number of axis–parallel ellipses. This question is well mot ivated by a special pattern recognition task where one has to identify ellipse s haped protein spots in 2–dimensional electrophoresis images. We present and di scuss two algorithmic approaches solving this problem: a greedy brute force me thod and a linear programming formulation. Furthermore we discuss related t h oretical questions.
In proteomics 2-dimensional gel electrophoresis (2-DE) is a separation technique for proteins. The resulting protein spots can be identified by either using picking robots and subsequent mass spectrometry or by visual cross inspection of a new gel image with an already analyzed master gel. Difficulties especially arise from inherent noise and irregular geometric distortions in 2-DE images. Aiming at the automated analysis of large series of 2-DE images, or at the even more difficult interlaboratory gel comparisons, the bottleneck is to solve the two most basic algorithmic problems with high quality: Identifying protein spots and computing a matching between two images. For the development of the analysis software CAROL at Freie Universität Berlin we have reconsidered these two problems and obtained new solutions which rely on methods from computational geometry. Their novelties are: 1. Spot detection is also possible for complex regions formed by several “merged” (usually saturated) spots; 2. User-defined landmarks are not necessary for the matching. Furthermore, images for comparison are allowed to represent different parts of the entire protein pattern, which only partially “overlap”. The implementation is done in a client server architecture to allow queries via the Internet. We also discuss and point at related theoretical questions in computational geometry.
We study the problem of covering a polygonal region by a small number of axis{parallel ellipses. This question is well motivated by the special pattern recognition task of identifying ellipse shaped protein spots in 2{dimensional electrophoresis images. We implemented several algorithms solving this problem: a greedy brute force method and a linear programming formulation. Several other algorithms are presented and discussed from a more theoretical point of view. 1 Detecting Spots in 2{dimensional Gel Electrophoresis Images 1.1 Gel Electrophoresis: The Application Background Proteomics is a rapidly growing eld within computational molecular biology. In proteomics 2{ dimensional gel electrophoresis (2DE) is the best known and widely used technique to separate proteins. A 2DE gel is the product of two separations performed sequentially in acrylamide gel media: isoelectric focusing as the rst dimension and a separation by molecular size as the second dimension. A two-dimensional pattern of spots each representing a protein is the result of that process. Eventually, spots are made visible by staining or radiographic methods. By analyzing series of such 2DE images one hopes to identify the proteins that change their expression (size, intensity) and re ect/cause certain biochemical and biomedical conditions of an organism, see [17]. Ideally, in an gel image each spot has the shape of an axis{parallel ellipse, which is a widely accepted modeling assumption, see e.g. [3] or [12]. However, spots that are very close to each other can partially merge (modeled by overlapping ellipses) and form rather complicated regions as depicted in Figure 1. At Freie Universitat Berlin we have started a few years ago to develop the software system CAROL (see [6]) that answers local and global matching queries for gel images. The novelty of its matching tool is the automatization of the process of setting the landmarks, Alon says: Is landmark well de ned ? which is the process that is done by human in other systems. Alon says: Did I say it right ? Instead, our matching between a source and a target image uses the history of the incremental Delaunay triangulation ([13, 11]) of the target spots. The matching tool starts from the assumption that images are already given as spot lists with each spot represented by point coordinates of its center and a real value describing its intensity. However, since the accessible spot detection algorithms did not supply results precise enough for our approach we developed and included a new detection algorithm into the CAROL system, see [15] for details. Research has been partly supported by Deutsche Forschungsgemeinschaft, grant FL 165/4{1. y Department of Computer Science, the University of Arizona, email:alon@cs.arizona.EDU z Institut f ur Informatik, Freie Universitat Berlin, Takustr. 9, D-14195 Berlin, email:ho mann@inf.fu-berlin.de x Institut f ur Informatik, Freie Universitat Berlin, Takustr. 9, D-14195 Berlin, email:kriegel@inf.fu-berlin.de { Institut f ur Informatik, Freie Universitat Berlin, Takustr. 9, D-14195 Berlin, email:schultz@inf.fu-berlin.de
With the growing importance of proteomics in biomedical and pharmaceutical sciences a need has emerged for computing tools that are capable of digitally visualizing and analyzing protein spot patterns within two-dimensional electrophoresis (2-DE) gel. Matching programs need to meet requirements such as interlaboratory comparison and the comparison of samples from different origins. For such research purposes, we have developed the CAROL system that implements new algorithms for spot detection and matching, which enable researchers to take a different approach to protein spot identification and comparison. The present short communication discusses how the system deals with uncertain geometric spot information that arises from streaks and complex spot regions and how this can be amplified for the matching procedure.
We study the problem of how to cover simple polygonal rectilinear regions by a small set of axis–parallel ellipses. This question is well motivated by a special pattern recognition task where one has to identify ellipse shaped protein spots in 2–dimensional electrophoresis images. We present and discuss various algorithmic approaches towards this problem ranging from a brute force method, to a linear programming formulation and an efficient theoretical solution. 1 Detecting Spots in 2–dimensional Gel Electrophoresis Images 1.1 Gel Electrophoresis: The Application Background With the growing importance of proteomics in biomedical and pharmaceutical sciences , see [7], there is an increasing need for efficient, reliable, and robust algorithmic solutions for various tasks that are part of an automatic analysis tool for 2–dimensional electrophoresis (2DE) gel images. With a resolution separation of several thousand proteins in real samples this method is almost two orders of magnitude better than competing Part of a joint research project with Deutsches Herzzentrum Berlin, supported by Deutsche Forschungsgemeinschaft, grant FL 165/4–1. yCS Dept. Stanford University, email:alon@Graphics.Stanford.EDU zInstitut fur Informatik, Freie Universitat Berlin, Takustr. 9, D14195 Berlin, email:name@inf.fu-berlin.de techniques. A 2DE gel is the product of two separations performed sequentially in acrylamide gel media: isoelectric focusing as the first dimension and a separation by molecular size as the second dimension. A two-dimensional pattern of spots each representing a protein is the result of that process. Eventually, spots are made visible by staining or radiographic methods. Ideally, each spot has the shape of an axis–parallel ellipse. However, as outlined below spots that are very close to each other can partially overlap and form rather complex regions. From the application point of view the overall goal is to analyse images from a whole gel series in order to identify those proteins that changes their expression (size, intensity) what could be a hint that reflects/causes certain biochemical and biomedical conditions of an organism. There are several commercial software packages available (like Melanie, PD Quest, Phoretix), which, together with their corresponding hardware, offer complete solutions to the gel analysis task including statistical analysis. However, these also have drawbacks since they do not support the comparison of images drawn from various sources (like databases in the Internet) which may have different size for example and, on the other side, they are too expensive for somebody who wants too evaluate just a few gels once in a while. With this background we have started a few years ago to develop a softare system CAROL (see [1]) that is able to perform local and global matching queries for gel images (say, given in gif format) via the Internet. Its novel algorithmic idea was to avoid setting landmarks by hand. Instead, the matching between a source and a target image uses the history of the incremental Delaunay triangulation, [3], [4], of the target spots. To this end we assumed that images are already given as spot lists with each spot represented by point coordinates of its center and a real value describing its intensity, as provided e. g. by the PDQuest system. Then, matching criteria are both geometric resemblance of locally intensive spot patterns as well as spot neighborhood comparisons. Meanwhile it has become lucid that wrongly detected spots like twin spots, spots within so called streaks or other complex regions are the main obstacle towards a better performance of the matching tool compared to influence of geometric distortions . That is why we decided to develop and include a new spot detection algorithm into the CAROL system to overcome these difficulties. In Figure 1 a part of a gel image is shown and aside the ellipses representing spots as computed by our spot detection algorithm. The main steps of the detection can be summarized as follows, compare with Figure 2 (numbered left to right). The algorithm, see [6] for details, starts smoothing the pixel image by a Gaussian filter (1) and computes a gradient image (2). Applying the watershed transformation to the gradient image we get a segmentation of the original image (3). Ideally, every obtained segment should represent a spot. In reality many more segments than there are spots are produced. By comparing segments with their neighborhood (using neighborhood graphs) it is possible to select those segments which really are part of a spot (4). All these segments should be merged to spots. Typical simple situations are shown in Figure 3. Some of the segments are isolated and form a single spot (top left). It can happen that two or three neighboring segments have to be merged. This case (top right) is solved by covering each subset by axis–parallel ellipses and comparing which covering fits best, i. e. , minimizes the symetric difference. However, there are situations like the bottom one in Figure 3, where the combinatorial complexity does not allow such an exhaustive search. Nevertheless, each such region has to be interpreted as union of ellipses, since it is typically oversaturated (so gray level values do not help here) and, most important, such intensive regions play an important role in the matching procedure. Typically, in complicated 2DE gel images there are up to 10 such complex regions. In fact, not only for very complex regions but also for twin spots and streaks there is an inherent uncertainty in the 2DE gel images due to the electrophoresis process itself which is highly susceptible to faults and distortions. Consequently, in the recent CAROL version we cope with that problem by maintaining a list of proposals how such an ambiguous region could be covered instead of computing only the best covering. In the matching algorithm part we then have the possibility to accept the matching of two complex region if there is a pair of proposed coverings that match. Figure 1 shows only 15 percent of a full 2DE gel image in GIF format. The total size is 811 900 pixel and our spot detection algorithm used about 9 sec to compute a total of 553 spots. 1.2 Modelling the Covering Problem We assume that a simply connected pixel pattern R is given. In the application it is usually a subpattern of a 100 100–square. Let R denote the polygonal curve describing its boundary. Identifying a pixel with its center point p we define the pixel sets R
We study the problem of how to cover a polygonal region by a small number of axis–parallel ellipses. This question is well motivated by a special pattern recognition task where one has to identify ellipse shaped protein spots in 2–dimensional electrophoresis images. We present and discuss two algorithmic approaches solving this problem: a greedy brute force method and a linear programming formulation. Furthermore we discuss related theoretical questions. 1 Detecting Spots in 2–dimensional Gel Electrophoresis Images 1.1 Gel Electrophoresis: The Application Background In proteomics 2–dimensional gel electrophoresis (2DE) is a widely used technique to separate proteins. A 2DE gel is the product of two separations performed sequentially in acrylamide gel media: isoelectric focusing as the first dimension and a separation by molecular size as the second dimension. A two-dimensional pattern of spots each representing a protein is the result of that process. Eventually, spots are made visible by staining or radiographic methods. By analyzing series of such 2DE images one hopes to identify those proteins that change their expression (size, intensity) and reflect/cause certain biochemical and biomedical conditions of an organism, see [15]. Ideally, in an gel image each spot has the shape of an axis–parallel ellipse, which is a widely accepted modeling assumption, see e.g. [3] or [6]. However, spots that are very close to each other can partially merge and form rather complicated regions as indicated in Figure 1. At Freie Universit ät Berlin we have started a few years ago to develop a software system CAROL (see [1]) that is able to perform local and global matching queries for gel images (given in GIF format). Its novel algorithmic idea was to avoid setting landmarks by hand. Instead, the matching between a source and a target image uses the history of the incremental Delaunay triangulation, [7], [8], of the target spots. To this end we assumed that images are already given as spot lists with each spot represented by point coordinates of its center and a real value describing its intensity, as provided e.g. by the commercial PDQuest system. Meanwhile it has become lucid that wrongly detected spots are the main obstacle towards a better performance of the matching tool compared to influence of pure geometric distortions. That is why we decided to develop and include a new spot detection Research has been partly supported by Deutsche Forschungsgemeinschaft, grant FL 165/4–1. email:alon@Graphics.Stanford.EDU email: hoffmann,kriegel,schultz @inf.fu-berlin.de Fig. 1. Twin spots, streaks and complex region algorithm into the CAROL system to overcome these difficulties. This algorithm, see [14] for details, relies on a watershed transformation applied to the gradient image. The most difficult part is how to interpret twin spots, streaks (left side in Fig. 1) and so called complex regions (right side in Fig. 1) as unions of ellipses. In fact, in [14] the latter case was left open and in the implementation the user had to edit these complex regions by hand. To present an algorithmic solution to this question is the subject of this paper and to our knowledge it is the first algorithmic approach that deals with complex regions. In fact, just by looking at the images it is clear that there is an inherent uncertainty in the 2DE gel images due to the electrophoresis process itself which is highly susceptible to faults and geometric distortions, so there is no hope to come up with a perfect spot detection algorithm. Consequently, in the upcoming CAROL version we try to cope with that problem by maintaining a list of proposals how such an ambiguous region could be covered instead of computing only the ‘best’ covering which could be erroneous. In the matching algorithm part we then have the possibility to accept the matching of two ambiguous regions if there is a pair of proposed coverings that match. 1.2 Modeling the Covering Problem The following formalization is a compromise stemming from discussions with practitioners who solve these covering instances by hand. We assume that a connected pixel pattern is given, which is fat in the sense that there are no short cuts consisting of two pixels only. In the application is usually a 1–connected subpattern of a –square. Let denote the rectilinear polygonal curve describing its boundary. We shrink and expand the boundary in both directions as follows. Identifying a pixel with its center point we define the two pixel sets ! #" $ %'& and, analogously, )(* + -, . ' / 0 1 2 3 4 5$ %'& . Now we can formulate the approximative covering problem we are interested in from the application point of view. Fig. 2. Part of a gel image and spots computed For numbers 1 and a natural number find a smallest possible set of axis–parallel ellipses fulfilling the following conditions. 1. (Shape) For each + the ratio of its halfaxes is in the interval $ . 2. (Fitting) Each ellipse + respects ( , i.e., it does not intersect ( . 3. (Intersection) The boundaries of any pair of ellipses intersects in at most 2 points, and area area area & . 4. (Covering) ! "$#&%' covers at least a )( portion of pixels in . We especially emphasize that the somehow strange intersection condition 3 is justified by the application because two spots (ellipses) can only partially merge and do not form a cross. 2 Brute Force Solution vs. Linear Programming In this section we present two algorithmic solutions to our covering problem that have been implemented and tested on 2D gel images. In the implementation we assumed that the regions are 1–connected and rectilinear, however it is straightforward how to extend the algorithms to arbitrary polygonal regions. 2.1 Computing a Brute Force Solution The basic idea behind the naive brute force solution is to generate all ellipses that fulfill condition 1 and 2. A voting scheme is then set up to select the covering. In a first step of the pixel set is approximated by a sample * of size about 5$,+ of the original vertex set. This is done heuristically in a way that almost convex boundary segments have denser samples. (Almost convex is defined via the average slope of lines connecting the point with predecessor points and successor points.) The second step in the brute force approach is to discretize the parameter space of possible ellipses. Recall that an axis–parallel ellipse is formed by all points fulfilling the equation ( (. ' ( 5 with parameters . To this end we restrict the ellipses to have centers which are pixel centers and one of the halfaxes, say , has to have multiple pixel side length. Finally, for a fixed we restrict the second halfaxis to values from the set ( $ & . Now, the algorithm simply computes for each center 2 ' and for each the maximal halfaxis such that the ellipse with parameters 2 ' does not intersect ( . For each ellipse we store which points from * it covers and their number. Eventually, in a greedy fashion we choose the covering. Assume a partial covering is already chosen. The next ellipse is selected among all ellipses satisfying condition 3 (with respect to already chosen ellipses) according to the following criteria, ordered as ranked: 1. covers a maximal number of previously uncovered points from * 2. maximizes the length of longest chain of consecutive covered points from * 3. minimizes maximal intersection area with an already chosen ellipse. After selecting a best (ties are broken randomly) we update the scores of all other remaining ellipses by deleting the points covered by . We stop augmenting ellipses when condition 4 is met. The drawback of the brute force approach is obvious, too many ellipses are tested. Moreover by discretizing ellipse parameters we may miss interesting ellipses like the big one in the right hand solution in Figure 3, see 2.3. for a discussion. Remark: Let be the family of ellipses 2 ' as described above, and let be the minimal number of ellipses of , needed to cover all points of * and avoiding the ones of )( . Then the standard "!$# * 3 –approximation (see [5] of the optimal solution by the greedy approach cannot be guaranteed, at least for arbitrary sample sets * in polygons with holes. This is due to the restrictive intersection property. An example that illustrates this observation consists of a set of rather thin ellipses arranged in a grid like fashion. Besides the very recent paper [11] we are not aware of approximation results for set covers with restrictive intersection properties. 2.2 Using an LP Approach to Generate Ellipses Compared to the brute force approach we do not want to restrict the set of ellipses under consideration for covering a region by an apriori parameter discretization and we want to avoid to generate to many explicite ellipses. Observe that each element in the pixel set ( forms a constraint for each ellipse in the cover, an ellipse must not contain such pixels. Since we do not want arbitrary small ellipses in the cover and we do not want degenerated halfaxes ratios as well, we can sample these ‘outer’ constraints by choosing every 4th point. Let us denote this sample by * ( . On the other side we want at least a few points (say, at least three) from *! to Fig. 3. Ellipse covering computed by brute force method (left) and by LP approach (right) be included in an ellipse. Therefore, we start from a randomly chosen triplet of mutually visible points from * and ask whether there is an axis–parallel ellipse containing these points such that it does not violate an outer constraint. Once having the information that for a given triplet there is a feasible solution one can try to extend the covered sample subset. The idea how to make use of linear programming for testing feasibility is based on the observation that the parameters of an ellipse can be transformed into variables of an LP in such a way that each of the three points which has to be covered by adds a linear constraint to the outer constraints. The cons
Klaus Kriegel合作论文数School of Business and Economics, Free University of Berlin6