Alzheimer’s disease (AD) is not a normal part of aging and does not inevitably happen in the later stages of life. It is a disorder of nerve cells in the brain that impairs memory and often leads to thinking and behavioral disabilities. AD is characterized by the presence of amyloid deposition and neurofibrillary tangles, together with the loss of cortical neurons and synapses [22]. In this work, we propose a neuroimaging technique to facilitate the early diagnosis of Alzheimer’s disease based on longitudinal studies of the affected population. We also discuss the utility of a recently developed information theoretic technique for bias correction in this context. We present the preliminary results and discuss the difficulties in analyzing them.
We describe a methodology for creating new technologies for assisted living in residential environments. The number of eldercare clients is expected to grow dramatically over the next decade as the baby boom generation approaches 65 years of age. The UMass/Smith ASSIST framework aims to alleviate the strain on centralized medical providers and community services as their clientele grow, reduce the delays in service, support independent living, and therefore, improve the quality of life for the up-coming elder population. We propose a closed loop methodology wherein innovative technical systems are field tested in assisted care facilities and analyzed by social scientists to create and refine residential systems for independent living. Our goal is to create technology that is embraced by clients, supports efficient delivery of support services, and facilitates social interactions with family and friends. We introduce a series of technologies that are currently under evaluation based on a distributed sensor network and a unique mobile manipulator (MM) concept. The mobile manipulator provides client services and serves as an embodied interface for remote service providers. As a result, a wide range of cost-effective eldercare applications can be devised, several of which are introduced in this paper. We illustrate tools for social interfaces, interfaces for community service and medical providers, and the capacity for autonomous assistance in the activities of daily living. These projects and others are being considered for field testing in the next cycle of ASSIST technology development.
The term plankton is used to include phytoplankton and zooplankton.While most of the current study on image classification has focused on mesozooplankton, the challenges involved are common to microzooplankton and phytoplankton.
We propose a novel interactive incremental method for pixel classi er construction. It yields very e cient decision tree classi ers and allows training of hierarchical classi ers for recognition of simple structured objects. Two novel concepts made possible by our paradigm are demonstrated: (1) Incremental training of a decision tree classi er on a sequence of images permits incremental, non-iterative improvement by dynamic addition of user-speci ed informative training pixels, i.e. pixels that are currently misclassi ed and will thus change the classi er. In experiments on a realistic terrain classi cation task, the number of training instances involved in building a classi er was reduced by several orders of magnitude, at no perceivable loss of classi cation accuracy. (2) Hierarchical classi cation extends the concept of pixel classi cation from labeling pixels directly with their categories to utilizing these class labels to describe more complex objects. We propose a set of simple and generic feature extractors that characterize spatial relationships between class labels. This makes our system capable of solving object recognition tasks beyond the realm of traditional pixel classi er systems, which is illustrated by a wildlife survey example, a seagull counting problem. KURZFASSUNG Wir prasentieren eine neuartige Methode zur interaktiven, inkrementellen Erstellung von Pixel-Klassi katoren. Sie erm oglicht die schnelle und einfache Erstellung sehr e zienter Entscheidungsbaum-Klassi katoren, sowie die Konstruktion einer Hierarchie von Klassi katoren f ur die Erkennung einfacher strukturierter Objekte. Wir demonstrieren zwei neuartige Konzepte: (1) Ein Entscheidungsbaum-Klassi kator kann durch inkrementelles Training auf Beispiel-Bildern, die in monotoner Folge prasentiert werden, konstruiert und verfeinert werden. Der Nutzer wahlt dabei interaktiv informative Trainingspixel aus, d.h. solche Pixel, die vom aktuellen Klassi kator misklassi ziert werden. Experimente anhand eines realistischen Klassi kationsproblems ergaben eine drastische Reduktion der Anzahl von Trainingspixeln gegen uber einem herkommlichen Trainingsverfahren, ohne erkennbaren Genauigkeitsverlust. (2) Hierarchische Klassi kation erweitert das Anwendungsgebiet von Pixel-Klassi katoren: Das Ergebnis einer Klassi kation wird von einem weiteren Klassi kator verwendet, um komplizierte Klassi kationsprobleme zu losen. Wir prasentieren drei einfache generische Funktionen zur Extraktion topologischer Relationen zwischen klassi zierten Pixeln. Diese konnen von einem hierarchischen Klassi kator genutzt werden, um einfache Objekte zu erkennen, was wir am Beispiel von Seem owen demonstrieren.
A wide variety of data sets produced by individual investigators are now synthesized to address ecological questions that span a range of spatial and temporal scales. It is important to facilitate such syntheses so that "consumers" of data sets can be confident that both input data sets and synthetic products are reliable. Necessary documentation to ensure the reliability and validation of data sets includes both familiar descriptive metadata and formal documentation of the scientific processes used (i.e., process metadata) to produce usable data sets from collections of raw data. Such documentation is complex and difficult to construct, so it is important to help "producers" create reliable data sets and to facilitate their creation of required metadata. We describe a formal representation, an "analytic web," that aids both producers and consumers of data sets by providing complete and precise definitions of scientific processes used to process raw and derived data sets. The formalisms used to define analytic webs are adaptations of those used in software engineering, and they provide a novel and effective support system for both the synthesis and the validation of ecological data sets. We illustrate the utility of an analytic web as an aid to producing synthetic data sets through a worked example: the synthesis of long-term measurements of whole-ecosystem carbon exchange. Analytic webs are also useful validation aids for consumers because they support the concurrent construction of a complete, Internet-accessible audit trail of the analytic processes used in the synthesis of the data sets. Finally we describe our early efforts to evaluate these ideas through the use of a prototype software tool, SciWalker. We indicate how this tool has been used to create analytic webs tailored to specific data-set synthesis and validation activities, and suggest extensions to it that will support additional forms of validation. The process metadata created by SciWalker is readily adapted for inclusion in Ecological Metadata Language (EML) files.
Earth's oceans are a soup of living micro-organisms known as plankton. As the foundation of the food chain for marine life, plankton are also an integral component of the global carbon cycle which regulates the planet's temperature. In this paper, we present a technique for automatic identification of plankton using a variety of features and classification methods including ensembles. The images were obtained in situ by an instrument known as the flow cytometer and microscope (FlowCAM), that detects particles from a stream of water siphoned directly from the ocean. The images are of necessity of limited resolution, making their identification a rather difficult challenge. We expect that upon completion, our system will become a useful tool for marine biologists to assess the health of the world's oceans.
The retrieval of images from a large database of images is an important and emerging area of research. Here, a technique to retrieve images based on appearance that works effectively across large changes of scale is proposed. The database is initially filtered with derivatives of a Gaussian at several scales. A user defined template is then created from an image of an object similar to those being sought. The template is also filtered using Gaussian derivatives. The template is then matched with the filter outputs of the database images and the matches ranked according to the match score. Experiments demonstrate the technique on a number of images in a database. No prior segmentation of the images is required and the technique works with viewpoint changes up to 20 degrees and illumination changes.
Environmental monitoring applications require seamless registration of optical data into large area mosaics that are geographically referenced to the world frame. Using frame-by-frame image registration alone, we can obtain seamless mosaics, but it will not exhibit geographical accuracy due to frame-to-frame error accumulation. On the other hand, the 3D geo-data from GPS, a laser profiler, an INS system provides a globally correct track of the motion without error propagation. However, the inherent (absolute) errors in the instrumentation are large for seamless mosaicing. The paper describes an effective two-track method for combining two different sources of data to achieve a seamless and geo-referenced mosaic, without 3D reconstruction or complex global registration. Experiments with real airborne video images show that the proposed algorithms are practical in important environmental applications.
In this paper, we present a new method for automatically and efficiently generating stereoscopic mosaics by seamless registration of images collected by a video camera mounted on an airborne platform. Using a parallel-perspective representation, a pair of geometrically registered stereo mosaics can be precisely constructed under quite general motion. A novel parallel ray interpolation for stereo mosaicing (PRISM) approach is proposed to make stereo mosaics seamless in the presence of obvious motion parallax and for rather arbitrary scenes. Parallel-perspective stereo mosaics generated with the PRISM method have better depth resolution than perspective stereo due to the adaptive baseline geometry. Moreover, unlike previous results showing that parallel-perspective stereo has a constant depth error, we conclude that the depth estimation error of stereo mosaics is in fact a linear function of the absolute depths of a scene. Experimental results on long video sequences are given.
This paper presents a panoramic virtual stereo vision approach to the problem of detecting and localizing multiple moving objects (e.g., humans) in an indoor scene. Two panoramic cameras, residing on different mobile platforms, compose a virtual stereo sensor with a flexible baseline. A novel "mutual calibration" algorithm is proposed, where panoramic cameras on two cooperative moving platforms are dynamically calibrated by looking at each other. A detailed numerical analysis of the error characteristics of the panoramic virtual stereo vision (mutual calibration error, stereo matching error, and triangulation error) is given to derive rules for optimal view planning. Experimental results are discussed for detecting and localizing multiple humans in motion using two cooperative robot platforms.
An adaptive panoramic stereo approach for two cooperative mobile platform is presented. There are four key features in the approach: 1) omnidirectional stereovision with an appropriate vertical FOV, and a simple camera calibration method; 2) cooperative mobile platforms for mutual dynamic calibration and best view planning; 3) 3D matching after meaningful object (human subject) extraction; and 4) real-time performance. The integration of omnidirectional vision with mutual awareness and dynamic calibration strategies allows intelligent cooperation between visual agents. This provides an effective way to solve the problems of limited resources, view planning, occlusion, and motion detection of movable robotic platforms. Experiments have shown that this approach is quite promising.
An adaptive panoramic stereo vision approach for localizing 3D moving objects has been developed in the Department of Computer Science at the University of Massachusetts at Amherst. This research focuses on cooperative robots involving cameras (residing on different mobile platforms) that can be dynamically composed into a virtual stereo vision system with flexible baseline in order to detect, track, and localize moving human subjects in an unknown indoor environment. This work was carried out under the support of the DARPA ITO Software for Distributed Robotics (SDR) and Autonomous Mobile Robot Software (MARS). This research project is relevant to both civil and military applications with distributed stationary and mobile perceptual agents. Applications are widespread, such as video surveillance and monitoring, security, crowd control, intelligent rooms and office, etc. For example, we are currently looking into applications in metropolitan scenarios such as New York City for solving problems of finding and protecting people in emergency circumstances, for example, during a terrorist attack or a fire in an office building.
Correlation-based stereo matching is very important to the generation of 3D terrain model. One of the difficulties in this stereo matching is the selection of the window size, since there are two competing factors that must be balanced in any stereo reconstruction process – perspective distortion and stereo matching error. This paper presents how the correlation window size affects the accuracy of stereo matching and suggests a tool to tune the window size in a given image domain. To facilitate the analysis proposed in this paper, we use photo-realistic simulation methodology to generate a pair of photo-realistic synthetic images of the terrain from a pre-acquired DEM(Digital Elevation Map) and ortho-image, which can be served as the pseudo ground truth. We performed 3D reconstruction on synthetic images of a natural terrain with Terrest system and carried out the evaluation of the correlation window on DEM accuracy. Experimental results are consistent with our strong expectation about two competing factors and show that our approach can be a useful tool to tune the window size.
A computer vision application can be defined as a sequence of image processing, feature extraction, and interpretation operations that are used to solve a specific task. In most of the computer vision systems this sequence is pre-defined and static. The work presented here shows a dynamic technique for algorithm selection based on both the value of information of each operation, and the computational complexity of the operations, which is modeled in terms of cost. This technique is used in the Ascender II system to select visual operators to perform aerial image interpretation of urban regions. The results show that the cost and information driven method presented here leads to performance gains. Moreover the hierarchical structure of the system simplifies the addition of new visual operations.
Automatic segmentation of stroke lesions in magnetic resonance imagery is a difficult problem because anatomical knowledge is required for the most accurate decisions. Without such knowledge, classification rules seem inconsistent. We propose a hybrid boundary and region based segmentation model built upon nonlinear scalespace and geometric active contours that captures the various segmentation rules necessary to segment lesions. After a user selects a point within damaged tissue and another point within healthy tissue, the image is examined at several levels of detail. At each such scale, the lesion is segmented several times by varying a parameter that models the range of criteria for boundaries between healthy and damaged tissue. These segmentations are collected, and the relative frequency of tissue being labeled lesion is regarded as a measure of confidence in the classification of the tissue as damaged. Experiments compare volumes and segmentations of lesions given by physicians to those given by the automatic method. Performance upper bounds are established by matching automatic segmentation parameters (scale, threshold, and/or confidence) for each image with each physician’s hand segmentation. These results may be compared with results that fix parameters for a particular physician’s segmentation or all physicians’ segmentations. Sensitivity to parameter values and initializations are tested as well. With little initialization, the model achieves zero error on average with a standard deviation near clinically useful bounds. A modest amount of additional input gives zero error on each image.
This paper presents an efficient method for finding salient differential features in images. We argue that the problem of finding salient features among all the possible ones is equivalent to finding outliers in a high-dimensional data set. We apply outlier detection techniques used in data mining to devise a linear time algorithm to extract the salient features. This yields a definition of saliency which rests on a more principled basis and also produces more reliable feature correspondences between images than the more conventional ones.
There have been attempts in a variety of applications to add 3D information into an image-based mosaic representation. Creating stereo mosaics from two rotating cameras was proposed by [Huang & Hung, 1998], and from a single off-center rotating camera by [Ishiguro, et al, 1990], [Peleg & Ben-Ezra, 1999], and by [Shum & Szeliski, 1999]. In these kinds of stereo mosaics, however, viewpoints — therefore the parallax — are limited to images taken from a very small area. Recently our work at UMass ([Zhu, et al, 1999, Zhu, et al, 2001a, Zhu, et al, 2001b]) has been focused on parallel-perspective stereo mosaics from a dominantly translating camera, which is the typical prevalent sensor motion during aerial surveys. A rotating camera can be easily controlled to achieve the desired motion. On the contrary, the translation of a camera over a large distance is much harder to achieve in real vision applications such as robot navigation ([Zheng & Tsuji, 1992]) and environmental monitoring ([Kumar, et al, 1995, Schultz, et al, 2000, Zhu, et al, 2001a]). In an applications to environmental monitoring, we have previously shown ([Zhu, et al, 1999, Zhu, et aI, 2001a, Zhu, et al, 2001b]) that image mosaicing from a translating camera raises a set of different problems from that of circular projections of a rotating camera.
We present a model-based approach to the automatic detection and reconstruction of buildings from aerial imagery. Buildings are first segmented from the scene in an optical image followed by a reconstruction process that makes use of a corresponding digital elevation map (DEM). Initially, each segmented DEM region likely to contain a building rooftop is indexed into a database of parameterized surface models that represent different building shape classes such as peaked, flat, or curved roofs. Given a set of indexed models, each is fit to the elevation data using a robust iterative procedure that determines the precise position and shape of the building rooftop. The indexed model that converges to the data with the lowest residual fit error is then added to the scene by extruding the fit rooftop surfaces to a local ground plane.The approach is based on the observation that a significant amount of rooftop variation can be modeled as the union of a small set of parameterized models and their combinations. By first recognizing the rooftop as one of the several potential rooftop shapes and fitting only these surfaces, the technique remains robust while still capable of reconstructing a wide variety of building types. In contrast to earlier approaches that presuppose a particular class of rooftops to be reconstructed (e.g., flat roofs), the algorithm is capable of reconstructing a variety of building types including peaked, flat, multi-level flat, and curved surfaces. The approach is evaluated on two datasets. Recognition rates for the different building rooftop classes and reconstruction accuracy are reported.
Zhigang Zhu合作论文数Department of Computer Science, The Grove School of Engineering, The City College of New York;CUNY Graduate Center;The CUNY Computational Vision and Convergence Laboratory15