In this study, a photometric stereo method, which can model the three dimensional surface topography of documents, is proposed. With the creation of high quality three dimensional model, it is aimed to better analyze the security features, pen tip strokes and other impressions on paper by document examination experts. Compared to 2D document images, modeling the three-dimensional surface topography of documents provides us more sufficient information for solving the order of writing (if two lines partially crossed), detecting the destruction of writings, analyzing the change of pressure, and author determination. Experimental studies show that the proposed method can make the three-dimensional model of the document surface topography precise and successful at micron levels.
We present a new approach for fine-grained classification of retail products, which learns and exploits statistical context information about likely product arrangements on shelves, incorporates visual hierarchies across brands, and returns recognition results as "confidence sets'' that are guaranteed to contain the true class at a given confidence level. Our system consists of three important components: 1) a nested hierarchy of product classes are automatically constructed based on visual similarities, 2) a confidence set predictor is trained based on class posteriors by using coarse-to-fine binary classifiers to discriminate each nested cluster of the hierarchy from the remainder of classes and a Bayesian network (BN) model that encodes the joint distribution of classifier scores with the fine-level class variable, and 3) n hidden Markov model (HMM) is trained with nested hidden states from the class hierarchy to model spatial transition across the nodes of product class hierarchy and resolve errors in the context-free confidence set results. Novel aspects of the proposed method include 1) combining confidence sets and context information via a HMM, 2) applying this concept to fine grained recognition of products arranged in retail shelves, and 3) presenting experimental results on four large datasets, collected from actual retail stores. We compare our approach with existing confidence set approaches and state-of-the-art convolutional neural networks classifiers including SENet-154, DenseNet-161, B-CNN, and Inception-Resnet-v2. Our approach performs comparably or better than state-of-the-art deep classifiers and exhibits high accuracy for relatively small confidence set sizes.
Classification systems of retail products have recently been gaining more importance. There are many classes of retail products and the resemblance of these products makes the design of product recognition systems, which have many application areas, more challenging. In this paper, we present a comparison of different classification techniques that are widely used in computer vision for image classification on retail product images taken by smart-phones.
Recently, retail product recognition has become an interesting computer vision research topic. The classification of products on shelves is a very challenging classification problem because many product classes are visually similar in terms of shape, color, texture, and metric size. In shelves, same or similar products are more likely to appear adjacent to each other and displayed in certain arrangements rather than at random. The arrangement of the products on the shelves has a spatial continuity both in brand and metric size. By using this context information, the co-occurrence of the products and the adjacency relations between the products can be statistically modeled. In this work, we present a context-aware hybrid classification system for the problem of fine-grained product class recognition. The proposed hybrid approach improves the accuracy of the context-free image classifiers, by combining them with a probabilistic graphical model based on Hidden Markov Models. The fundamental goal of this paper is to use contextual relationships in retail shelves to improve accuracy of the product classifier.
Classification systems of retail products have recently been gaining more importance. There are many classes of retail products and the resemblance of these products makes the design of product recognition systems, which have many application areas, more challenging. In this paper, we present a comparison of different classification techniques that are widely used in computer vision for image classification on retail product images taken by smart-phones.
Recently, retail product recognition has become an interesting computer vision research topic. The classification of products on shelves is a very challenging classification problem because many product classes are visually similar in terms of shape, color, texture, and metric size. In shelves, same or similar products are more likely to appear adjacent to each other and displayed in certain arrangements rather than at random. The arrangement of the products on the shelves has a spatial continuity both in brand and metric size. By using this context information, the co-occurrence of the products and the adjacency relations between the products can be statistically modeled. In this work, we present a context-aware hybrid classification system for the problem of fine-grained product class recognition. The proposed hybrid approach improves the accuracy of the context-free image classifiers, by combining them with a probabilistic graphical model based on Hidden Markov Models. The fundamental goal of this paper is to use contextual relationships in retail shelves to improve accuracy of the product classifier.
We present a context-aware hybrid classification system for the problem of fine-grained product class recognition in computer vision. Recently, retail product recognition has become an interesting computer vision research topic. We focus on the classification of products on shelves in a store. This is a very challenging classification problem because many product classes are visually similar in terms of shape, color, texture, and metric size. In shelves, same or similar products are more likely to appear adjacent to each other and displayed in certain arrangements rather than at random. The arrangement of the products on the shelves has a spatial continuity both in brand and metric size. By using this context information, the co-occurrence of the products and the adjacency relations between the products can be statistically modeled. The proposed hybrid approach improves the accuracy of context-free image classifiers such as Support Vector Machines (SVMs), by combining them with a probabilistic graphical model such as Hidden Markov Models (HMMs) or Conditional Random Fields (CRFs). The fundamental goal of this paper is using contextual relationships in retail shelves to improve the classification accuracy by executing a context-aware approach.
Multiple researchers recently proposed the use of the digital compass embedded in mobile devices for touchless interaction in the 3D space around them. These methods overcome several limits imposed by other interaction techniques and were evaluated for a variety of uses. However, they do not support collaborative settings and are prone to dynamic noise caused by external conditions, as with most other sensor-based interaction techniques. In this paper, we propose the use of frequency-modulated electromagnets as an input medium for magnetic interaction to overcome its various constraints and further enable multi-user and two-handed input. Furthermore, we demonstrated the hardware design specifications of a novel input device, referred to as electromagnetic stylus, which is prototyped to conduct a user-study on the proposed method. Experimental results indicate that gestures performed simultaneously by four electromagnetic styli can accurately be recognized using a single magnetic field sensor, and dynamic noises can be substantially reduced.
This paper proposes a hardware-oriented trinocular adaptive window size disparity estimation (T-AWDE) algorithm and the first real-time trinocular disparity estimation (DE) hardware that targets high-resolution images with high-quality disparity results. The proposed trinocular DE hardware is the enhanced version of the recently published binocular AWDE implementation. The T-AWDE hardware generates a very high-quality depth map by merging two depth maps obtained from the center-left and center-right camera pairs. The T-AWDE hardware enhances disparity results by applying a double checking scheme which solves most of the occlusion problems existing in the AWDE implementation while providing correct disparity results even for objects located at left or right edge of the center image. The proposed T-AWDE hardware architecture enables handling 55 frames per second on a Virtex-7 FPGA at a 1024×768 XGA video resolution for a 128 pixels disparity range.
The computational complexity of disparity estimation algorithms and the need of large size and bandwidth for the external and internal memory make the real-time processing of disparity estimation challenging, especially for High Resolution (HR) images. This paper proposes a hardware-oriented adaptive window size disparity estimation (AWDE) algorithm and its real-time reconfigurable hardware implementation that targets HR video with high quality disparity results. Moreover, an enhanced version of the AWDE implementation that uses iterative refinement (AWDE-IR) is presented. The AWDE and AWDE-IR algorithms dynamically adapt the window size considering the local texture of the image to increase the disparity estimation quality. The proposed reconfigurable hardware architectures of the AWDE and AWDE-IR algorithms enable handling 60 frames per second on a Virtex-5 FPGA at a 1024×768 XGA video resolution for a 128pixel disparity range.
Depth information is used in a variety of 3D based signal processing applications such as autonomous navigation, robot and driving systems, object detection and tracking, 3D television, and disparity-based rendering. In these applications, high accuracy and speed performances are required for depth map estimation. Depth maps can be generated by using disparity estimation methods, which are obtained from stereo matching between the stereo images. The computational complexity of disparity estimation algorithms and the need of large size memory make the real-time processing of disparity estimation challenging, especially for high resolution images. This thesis proposes binocular (AWDE) and trinocular (TAWDE) hardware-oriented adaptive window size disparity estimation algorithms, which target high resolution video with high quality disparity results. Furthermore, in depth map estimation based applications, disparity estimation should be performed in real-time. To reach real-time, the algorithms also should be suitable for hardware implementation. The proposed binocular and trinocular algorithms can be implemented in hardware very efficiently. We first propose a novel binocular disparity estimation algorithm that is a hybrid solution involving the Binary Window Sum of Absolute Differences and the Census cost computation methods to vote and select the best suitable disparity candidates. It utilizes a pixel intensity based refinement step to remove faulty disparity computations. The AWDE algorithm dynamically adapts the window size considering the local texture of the image to increase the disparity estimation quality. We then propose a new trinocular adaptive window size disparity estimation algorithm. Our trinocular algorithm is the extended version of the binocular algorithm. A novel disparity map fusion method is developed. The proposed trinocular algorithm carefully handles the problem in the binocular disparity estimation results by adding a third camera into the system. Finally, the algorithms are evaluated on the different stereo datasets. The results demonstrate that the proposed AWDE and T-AWDE algorithms are suitable for real-time hardware implementation and their reconfigurable hardware can be used in consumer electronics products where high-quality real time disparity estimation is needed for high resolution video.
The computational complexity of disparity estimation algorithms and the need of large size and bandwidth for the external and internal memory make the real-time processing of disparity estimation challenging, especially for High Resolution (HR) images. This paper proposes a hardware-oriented adaptive window size disparity estimation (AWDE) algorithm and its real-time reconfigurable hardware implementation that targets HR video with high quality disparity results. The proposed algorithm is a hybrid solution involving the Sum of Absolute Differences and the Census cost computation methods to vote and select the best suitable disparity candidates. It utilizes a pixel intensity based refinement step to remove faulty disparity computations. The AWDE algorithm dynamically adapts the window size considering the local texture of the image to increase the disparity estimation quality. The proposed reconfigurable hardware of the AWDE algorithm enables handling 60 frames per second on Virtex-5 FPGA at a 1024×768 XGA video resolution for a 120 pixel disparity range. 1