Advances in CMOS fabrication have enabled low-cost camera nodes with limited communication and computation capabilities. By combining these capabilities within a small form-factor device, multi-camera networks can readily be built. However, cameras are high-data-rate devices, and many computer vision algorithms are computationally expensive, while these camera nodes are communication and computation constrained. In this dissertation, we present lightweight techniques for distributed scene analysis in such resource-constrained camera networks. We show that in this setting we can compute global aggregates from distributed local measurements. In particular we use the camera network to count and localize targets. Counting and localizing are useful in many applications in surveillance, security, and monitoring. Counting multiple objects is difficult because objects often occlude one another. A camera network with multiple views can resolve these ambiguities. To satisfy the resource constraints, only a subset of camera nodes can be selected to answer a query, and these nodes must perform lightweight processing and only communicate limited amounts of data. In this work, the local image processing is background subtraction and the communicated data is less than 1/10,000 of the original image size. A two-dimensional visual hull, representing the maximal spatial occupancy, is then inexpensively computed by aggregating this compressed data. The first part of the dissertation describes the counting algorithm which uses the visual hull. Upper and lower bounds for the number of objects are computed and updated under object motion, and an exact count is reached when the bounds converge. The second part describes how to select camera nodes to compute the visual hull. Selecting an optimal subset can be as effective as using all the cameras and both saves resources and increases scalability. The final part analyzes the best selection and placement of camera nodes for optimal target localization. The formulation is based on linear estimation. Uniform placement is shown to be optimal for cameras with identical noise. The analysis leads to an algorithm for camera selection. The performance of the target counting and localization algorithms is demonstrated in simulation and in real camera networks.
The paper studies the optimal placement of multiple cameras and the selection of the best subset of cameras for single target localization in the framework of sensor networks. The cameras are assumed to be aimed horizontally around a room. To conserve both computation and communication energy, each camera reduces its image to a binary "scan-line" by performing simple background subtraction followed by vertical summing and thresholding, and communicates only the center of the detected foreground object. Assuming noisy camera measurements and an object prior, the minimum mean squared error of the best linear estimate of the object location in 2-D is used as a metric for placement and selection. The placement problem is shown to be equivalent to a classical inverse kinematics robotics problem, which can be solved efficiently using gradient descent techniques. The selection problem on the other hand is a combinatorial optimization problem and finding the optimal solution can be too costly to implement in an energy-constrained wireless camera network. A semi-definite programming approximation for the problem is shown to achieve close to optimal solutions with much lower computational burden. Simulation and experimental results are presented.
In this paper, we study how to task camera sensor nodes to reason about the occupancy of the area around them. Occupancy information is valuable because it can be used to answer many other queries such as determining object tracks or the count of the number of people in an area. Camera sensors are challenging to include in a wireless sensor network (WSN) because they are high data rate devices. To save energy and to satisfy the bandwidth constraint, our camera nodes will only send a very limited amount of data and only a limited number of camera nodes will be tasked. Our r st result, from simulation, gives an upper bound on the number of cameras needed for a given accuracy in the occupancy. Given this number of cameras, we then compare several approaches to tasking the most relevant cameras both in simulation and in a real system of 16 camera nodes. Our incremental greedy tasking algorithm performed the best. Finally, we applied this tasking algorithm to a tracking application. We show that the tracker that used tasking outperformed the same tracker without tasking.
Estimating the number of people in a crowded environment is a central task in civilian surveillance. Most vision-based counting techniques depend on detecting individuals in order to count, an unrealistic proposition in crowded settings. We propose an alternative approach that directly estimates the number of people. In our system, groups of image sensors segment foreground objects from the background, aggregate the resulting silhouettes over a network, and compute a planar projection of the scene's visual hull. We introduce a geometric algorithm that calculates bounds on the number of persons in each region of the projection, after phantom regions have been eliminated. The computational requirements scale well with the number of sensors and the number of people, and only limited amounts of data are transmitted over the network. Because of these properties, our system runs in real-time and can be deployed as an untethered wireless sensor network. We describe the major components of our system, and report preliminary experiments with our first prototype implementation.
This paper presents a novel solution for flow-based tracking and 3D reconstruction of deforming objects in monocular image sequences. A non-rigid 3D object undergoing rotation and deformation can be effectively approximated using a linear combination of 3D basis shapes. This puts a bound on the rank of the tracking matrix. The rank constraint is used to achieve robust and precise low-level optical flow estimation without prior knowledge of the 3D shape of the object. The bound on the rank is also exploited to handle occlusion at the tracking level leading to the possibility of recovering the complete trajectories of occluded/disoccluded points. Following the same low-rank principle, the resulting flow matrix can be factored to get the 3D pose, configuration coefficients, and 3D basis shapes. The flow matrix is factored in an iterative manner, looping between solving for pose, configuration, and basis shapes. The flow-based tracking is applied to several video sequences and provides the input to the 3D non-rigid reconstruction task. Additional results on synthetic data and comparisons to ground truth complete the experiments.
We present the results of a numerical study of the 2 {times}L spin- (1) /(2) Heisenberg ladder. Ground state energies and the singlet-triplet energy gaps for 4{le}L{le}14 and J{sub {perpendicular}}/J{sub {vert_bar}{vert_bar}} = 1 were obtained in a Lanczos calculation and checked against earlier calculations by Barnes {ital et al.} (even L{le} 12). A related moments technique is then employed to evaluate spin response functions for L = 12 and a range of J{sub {perpendicular}}/J{sub {vert_bar}{vert_bar}} (0 {endash} 5). We comment on two issues, the need for reorthogonalization and the rate of convergence, that affect the numerical utility of the moments treatment of response functions. {copyright} {ital 1998} {ital The American Physical Society}
Salih Burak Göktürk合作论文数Ojos, Inc.8
Baris Sumengen合作论文数Electrical and Computer Engineering Department
University of California at Santa Barbara4