This paper presents a novel deep regression network to extract geometric information from Light Field (LF) data. Our network builds upon u-shaped network architectures. Those networks involve two symmetric parts, an encoding and a decoding part. In the first part the network encodes relevant information from the given input into a set of high-level feature maps. In the second part the generated feature maps are then decoded to the desired output. To predict reliable and robust depth information the proposed network examines 3D subsets of the 4D LF called Epipolar Plane Image (EPI) volumes. An important aspect of our network is the use of 3D convolutional layers, that allow to propagate information from two spatial dimensions and one directional dimension of the LF. Compared to previous work this allows for an additional spatial regularization, which reduces depth artifacts and simultaneously maintains clear depth discontinuities. Experimental results show that our approach allows to create high-quality reconstruction results, which outperform current state-of-the-art Shape from Light Field (SfLF) techniques. The main advantage of the proposed approach is the ability to provide those high-quality reconstructions at a low computation time.
Medical image segmentation is a key and fundamental operation in medical image analysis. Many studies have been dedicated to this field and have proposed different techniques or variations, but medical image segmentation is still a challenge due to low contrast, noise, complex shape, and so on. This chapter introduces histogram-based level set methods for medical image segmentation, which evolve level set functions by measuring the similarity of histograms. Without preprocessing, we show a comparable segmentation result on brain image using a simple gradient descent or gradient ascend method.
Coins are extensively used in banks,transportation systems,supermarkets,etc.Every day,too many coins would be hoarded and thus need massive sorting,counting and packing work.In recent years,various automatic sorting,counting and packing machines are developed and commercialized.However,most of them are heavy,bigsized,expensive,or of low efficiency.A new integrated coin counting and packing device with a function of wireless communication is proposed and developed in this paper.Several new mechanisms are designed to improve the coin conveying,sorting and counting efficiency and reliability.New coin feeding and heating mechanisms are designed for packing the coins.The device is simple,cheap and efficient.It has high potential and prospect for future applications.
This paper presents a novel technique for Shape from Light Field (SfLF), that utilizes deep learning strategies. Our model is based on a fully convolutional network, that involves two symmetric parts, an encoding and a decoding part, leading to a u-shaped network architecture. By leveraging a recently proposed Light Field (LF) dataset, we are able to effectively train our model using supervised training. To process an entire LF we split the LF data into the corresponding Epipolar Plane Image (EPI) representation and predict each EPI separately. This strategy provides good reconstruction results combined with a fast prediction time. In the experimental section we compare our method to the state of the art. The method performs well in terms of depth accuracy, and is able to outperform competing methods in terms of prediction time by a large margin.
In this contribution, we present a semi-automatic segmentation algorithm for radiofrequency ablation (RFA) zones via optimal s-t-cuts. Our interactive graph-based approach builds upon a polyhedron to construct the graph and was specifically designed for computed tomography (CT) acquisitions from patients that had RFA treatments of Hepatocellular Carcinomas (HCC). For evaluation, we used twelve post-interventional CT datasets from the clinical routine and as evaluation metric we utilized the Dice Similarity Coefficient (DSC), which is commonly accepted for judging computer aided medical segmentation tasks. Compared with pure manual slice-by-slice expert segmentations from interventional radiologists, we were able to achieve a DSC of about eighty percent, which is sufficient for our clinical needs. Moreover, our approach was able to handle images containing (DSC=75.9%) and not containing (78.1%) the RFA needles still in place. Additionally, we found no statistically significant difference (p<;0.423) between the segmentation results of the subgroups for a Mann-Whitney test. Finally, to the best of our knowledge, this is the first time a segmentation approach for CT scans including the RFA needles is reported and we show why another state-of-the-art segmentation method fails for these cases. Intraoperative scans including an RFA probe are very critical in the clinical practice and need a very careful segmentation and inspection to avoid under-treatment, which may result in tumor recurrence (up to 40%). If the decision can be made during the intervention, an additional ablation can be performed without removing the entire needle. This decreases the patient stress and associated risks and costs of a separate intervention at a later date. Ultimately, the segmented ablation zone containing the RFA needle can be used for a precise ablation simulation as the real needle position is known.
For several decades, image restoration remains an active research topic in low-level computer vision and hence new approaches are constantly emerging. However, many recently proposed algorithms achieve state-of-the-art performance only at the expense of very high computation time, which clearly limits their practical relevance. In this work, we propose a simple but effective approach with both high computational efficiency and high restoration quality. We extend conventional nonlinear reaction diffusion models by several parametrized linear filters as well as several parametrized influence functions. We propose to train the parameters of the filters and the influence functions through a loss based approach. Experiments show that our trained nonlinear reaction diffusion models largely benefit from the training of the parameters and finally lead to the best reported performance on common test datasets for image restoration. Due to their structural simplicity, our trained models are highly efficient and are also well-suited for parallel computation on GPUs.
In this paper we present a trained diffusion model for image inpainting based on the structural similarity measure. The proposed diffusion model uses several parametrized linear filters and influence functions. Those parameters are learned in a loss based approach, where we first perform a greedy training before conducting a joint training to further improve the inpainting performance. We provide a detailed comparison to state-of-the-art inpainting algorithms based on the TUM-image inpainting database. The experimental results show that the proposed diffusion model is efficient and achieves superior performance. Moreover, we also demonstrate that the proposed method has a texture preserving property, that makes it stand out from previous PDE based methods.
Using an optimization algorithm to solve a machine learning problem is one of mainstreams in the field of science. In this work, we demonstrate a comprehensive comparison of some state-of-the-art first-order optimization algorithms for convex optimization problems in machine learning. We concentrate on several smooth and non-smooth machine learning problems with a loss function plus a regularizer. The overall experimental results show the superiority of primal-dual algorithms in solving a machine learning problem from the perspectives of the ease to construct, running time and accuracy.
A novel adaptive aircraft detection method based on level set processing and circle-frequency filter is proposed in this paper. First, the SBGFRLS (Selective Binary and Gaussian Filtering Regularized Level Set) method is used twice to find airport region of interest (ROI) and candidate aircraft areas by local segmentation and global segmentation, respectively, so that sizes of those possible target areas can be computed. Then, the circle-frequency (CF) filter method is utilized adaptively to detect target aircrafts in the airport ROI via the mean radius estimated by sizes of those candidate areas obtained before. Experimental results on real remote sensing airport images demonstrate the efficiency and accuracy of the proposed method.
Subject to limited resolution for targets in many satellite images, low-resolution airplane detection is still difficult and challenging, which plays an important role in remote sensing. In this paper, we propose a new method to detect lowresolution airplanes in satellite images. First, the image is preprocessed by combing the unsharp contrast enhancement (UCE) filtered image and the original image. Second, the Local Edge Distribution (LED), which is susceptible to objects owning clustered edges, e.g., airplane, is calculated to acquire the target candidate regions while restraining large background area. Then, a multi-scale fused gradient feature image is computed to characterize the shapes of targets instead of the original image to overcome the influence from the self-shadow and different coating colors of airplanes. After that, a designed airplane shape filter with a modulated item is used to detect and locate real targets, in which the modulated item can effectively measure the degree of coincidence between the patch region and the airplane shape. Finally, coordinates of target centers are computed in the filtered image. Experimental results demonstrate that the proposed algorithm is effective and robust for detecting low-resolution airplanes in satellite images under various complex backgrounds.
This paper presents a novel active contour tracking method, which is used to estimate the non-grid deformation of the motion target and can get the accurate contour of the tracked target. In the proposed method, level set is employed to represent the target region. By using Bhattacharyya similarity as a metric, the proposed method aims at finding a best candidate region in the video frame, whose foreground distribution and background distribution match maximally those of the predefined tracked target. Based on this metric, we derive a level set based object tracking formulation, which estimates iteratively the contour change of the target. Experimental results on the representative video sequences show that the proposed methods perform better than EM-shift and SOAMST algorithm.
This paper presents a novel object tracking framework by joint registration and active contour segmentation (JRACS), which can robustly deal with the non-rigid shape changes of the target. The target region, which includes both foreground and background pixels, is implicitly represented by a level set. A Bhattacharyya similarity based metric is proposed to locate the region whose foreground and background distributions best match those of the tracked target. Based on this metric, a tracking framework that consists of a registration stage and a segmentation stage is then established. The registration step roughly locates the target object by modeling its motion as an affine transformation, and the segmentation step refines the registration result and computes the true contour of the target. The robust tracking performance of the proposed JRACS method is demonstrated by real video sequences where the objects have clear non-rigid shape changes.
In this paper we present a framework to obtain highly efficient implementations for the narrow band level set method on commercial off-the-shelf (COTS) multicore CPU systems with a cache-based memory hierarchy such as Intel Xeon and Atom processors. The narrow-band level set algorithm tracks wave-fronts in discretized volumes (for instance, explosion shock waves), and is computationally very demanding. At the core of our optimization framework is a novel projection-based approach to enhance data locality and enable reuse for sparse surfaces in dense discretized volumes. The method reduces stencil operations on sparse and changing sets of pixels belonging to an evolving surface into dense stencil operations on meta-pixels in a lower-dimensional projection of the pixel space. These meta-pixels are then amenable to standard techniques like time tiling. However, the complexity introduced by ever-changing meta-pixels requires us to revisit and adapt all other necessary optimizations. We apply adapted versions of SIMDization, multi-threading, DAG scheduling for basic tiles, and specialization through code generation to extract maximum performance. The system is implemented as highly parameterized code skeleton that is auto-tuned and uses program generation.We evaluated our framework on a dual-socket 2.8 GHz Xeon 5560 and a 1.6 GHz Atom N270. Our single-core performance reaches 26%-35% of the machine peak on the Xeon, and 12%-20% on the Atom across a range of image sizes. We see up to 6.5x speedup on 8 cores of the dual-socket Xeon. For cache-resident sizes our code outperforms the best available third-party code (C pre-compiled into a DLL) by about 10x and for the largest out-of-cache sizes the speedup approaches around 200x. Experiments fully explain the high speedup numbers.
目标表示方法对跟踪方法的鲁棒性有着重要影响。将对立色局部二值模式(OCLBP)纹理算子作为研究对象引入目标表示。通过分析不同颜色通道之间的相关性和OCLBP的10种纹理模式的表征能力,选择目标候选区域中具有OCLBP的7种主要模式的关键点的纹理直方图作为目标模型。最后将该目标表示方法嵌入到MeanShift框架中,进行目标跟踪。实验结果表明,提出的基于OCLBP主要模式的目标表示方法显著提高了Mean Shift目标跟踪方法的性能。
A C-V model based on level set and prior information was proposed and was applied to segment weed, wheat and apple images. Based on the characteristics of the image, the image was represented by a model which made the image easy to segment at first, and then the data contents of a region of interest in this model were extracted as the prior information. An initial contour by hue was obtained and the proposed model by this contour was initialized, the level set function was iteratively solved. Finally, a stationary contour was obtained. The correct rates of weed, wheat and apple were 0.999, 0.999 and 0.846 respectively and the error rates were 0, 0 and 0.125 respectively.
Shape deformation is a useful tool for shape modeling and animation in computer graphics. In this paper, we propose a novel surface deformation method based on a feature sensitive (FS) metric. Firstly, taking unit normal vectors into account, we derive a FS Laplacian operator, which is more sensitive to featured regions of mesh models than existing operators. Secondly, we use the 1-ring tetrahedron in the dual mesh, a volumetric structure, to encode geometric details. To preserve the shape of the tetrahedron, we introduce linear tetrahedron constraints minimizing both the distortion of the base triangle and the change of the corresponding height. These ensure that geometric details are accurately preserved during deformation. The time complexity of our new method is similar to that of existing linear Laplacian methods. Examples are included to show that our FS deformation method better preserves mesh details, especially features, than existing Laplacian methods. Copyright (C) 2010 John Wiley & Sons, Ltd.
To get better segmentation results, local information and global information should be taken into consideration together. In this paper, we propose a new energy functional which combines a local intensity fitting term and an auxiliary global intensity fitting term, and we also give the method to adjust the weight of auxiliary global fitting term dynamically by using local contrast of the image. The combination of the two terms improves the accuracy of segmentation results obviously while reduces dependence on location of initial contour. The experiment results proved the effectiveness of our method.
Real-time stereo vision is attractive in many applications like robot navigation and 3-D scene reconstruction. Data parallel platforms, e.g., graphics processing unit (GPU), are often used for real-time stereo, because most stereo algorithms involve a large portion of data parallel computations. In this paper, we propose a stereo system on GPU which pushes the Pareto-efficiency frontline in the accuracy and speed tradeoff space. Our system is based on a hardware-aware algorithm design approach. The system consists of new algorithms and code optimization techniques. We emphasize on keeping the highly data parallel structure in the algorithm design process such that the algorithms can be effectively mapped to massively data parallel platforms. We propose two stereo algorithms: namely, exponential step size adaptive weight (ESAW), and exponential step size message propagation (ESMP). ESAW reduces computational complexity without sacrificing disparity accuracy. ESMP is an extension of ESAW, which incorporates the smoothness term to better model non-frontal planes. ESMP offers additional choice in the accuracy and speed tradeoff space. We adopt code optimization methodologies from the performance tuning community, and apply them to this specific application. Such an approach gives higher performance than optimizing the code in an “ad hoc” manner, and helps understanding the code efficiency. Experiment results demonstrate a speedup factor of 2.7-8.5 over state-of-the-art stereo systems at comparable disparity accuracy.