This work proposes a new error backpropagation approach as a systematic way to configure and train the Multi-net System MNOD, a recently proposed algorithm able to segment a class of visual objects from real images. First, a single node of the MNOD is configured in order to best resolve the visual object segmentation problem using the best combination of parameters and features. The problem is then how to add new nodes in order to improve accuracy and avoid overfitting situations. In this scenario, the proposed approach employs backpropagation of error maps to add new nodes with the aim of increasing the overall segmentation performance. Experiments conducted on a standard dataset of real images show that our configuration method, using only simple edges and colors descriptors, leads to configurations that produced comparable results in visual objects segmentation.
We propose a method to describe how a person is dressed, using an innovative way to extract Visual Information exploiting the Human Pose Estimation. Given the lack of algorithms in this field, we aims to pave the way giving a baseline and publishing a detailed dataset for future comparisons. In particular in this study we show how using the Human Pose Estimation, we are able to extract the essential features for the description of the Visual Attributes. Furthermore, the proposed method is able to manage the problems highlighted in literature regarding the extraction of features from images of people due to their articulated poses. For this reason we also propose a formalization of how describe people’s clothing in order to give a starting point and facilitate the analysis and the Visual Attributes extraction phase. Moreover we show how the use of Deformable Structures let us to extract Visual Attributes without the using of segmentation algorithms.
In recent years a great amount of research has focused on algorithms that learn features from unlabeled data. In this work we propose a model based on the Self-Organizing Map (SOM) neural network to learn features useful for the problem of automatic natural images classification. In particular we use the SOM model to learn single-layer features from the extremely challenging CIFAR-10 dataset, containing 60.000 tiny labeled natural images, and subsequently use these features with a pyramidal histogram encoding to train a linear SVM classifier. Despite the large number of images, the proposed feature learning method requires only few minutes on an entry-level system, however we show that a supervised classifier trained with learned features provides significantly better results than using raw pixels values or other handcrafted features designed specifically for image classification. Moreover, exploiting the topological property of the SOM neural network, it is possible to reduce the number of features and speed up the supervised training process combining topologically close neurons, without repeating the feature learning process.
One fundamental issue in today's Online Social Networks (OSNs) is to give users the ability to control the messages posted on their own private space to avoid that unwanted content is displayed. Up to now, OSNs provide little support to this requirement. To fill the gap, in this paper, we propose a system allowing OSN users to have a direct control on the messages posted on their walls. This is achieved through a flexible rule-based system, that allows users to customize the filtering criteria to be applied to their walls, and a Machine Learning-based soft classifier automatically labeling messages in support of content-based filtering.
We present a new approach for automatic gas meter reading from real world images. The gas meter reading is usually done on site by an operator and a picture is taken from a mobile device as proof of reading. Since the reading operation is prone to errors, the proof image is checked offline by another operator to confirm the reading. In this study, we present a method to support the validation process in order to reduce the human effort. Our approach is trained to detect and recognize the text of a particular area of interest. Firstly we detect the region of interest and segment the text contained using a method based on an ensemble of neural models. Then we perform an optical character recognition using a Support Vector Machine. We evaluated every step of our approach, as well as the overall assessment, showing that despite the complexity of the problem our method provide good results also when applied to degraded images and can therefore be used in real applications.
The proposed model aims to extend the MNOD algorithm adding a new type of node specialized in object classification. For each potential object identified by the MNOD, a set of segments are generated using a min-cut based algorithm with different seeds configurations. These segments are classified by a suitable neural model and then the one with higher value is chosen, in agreement with a proper energy function. The proposed method allows to segment and classify each object simultaneously. The results showed in the experiment section highlight the potential and the cost of having unified segmentation and classification in a single model.
In this study we propose a new strategy to perform an object segmentation using a multi neural network approach. We started extending our previously presented object detection method applying a new segment based classification strategy. The result obtained is a segmentation map post processed by a phase that exploits the GrabCut algorithm to obtain a fairly precise and sharp edges of the object of interest in a full automatic way. We tested the new strategy on a clothing commercial dataset obtaining a substantial improvement on the quality of the segmentation results compared with our previous method. The segment classification approach we propose achieves the same improvement on a subset of the Pascal VOC 2011 dataset which is a recent standard segmentation dataset, obtaining a result which is inline with the state of the art.
We propose a local texture descriptor based on a pyramidal composition of Self Organizing Map (SOM). As with the SOM model, our visual descriptor presents two operational steps: a first unsupervised learning phase and a second mapping phase involving a dimensionality reduction of the input data. During the first step a large number of image patches, including different classes of textures, are presented to the model. At the end of the learning process the neural weights on each layer of the SOM pyramid will contain good prototypes of the patches used in training at different level of detail. During the mapping phase a new texture patch is presented to the model and, by using a winner take all principle, a winner neuron is selected and its 2D spatial location is used to describe the input patch. Exploiting the topological order of the SOM, two different texture descriptions can be compared using the common Euclidean distance. In the experimental section we show that a simple clustering algorithm like K-means, applied to the local descriptor responses, is able to segment complex texture mosaics with very good results, even in difficult areas like boundaries which separate two different textures.
In this study we propose a mobile application which interfaces with a Content-Based Image Retrieval engine for online shopping in the fashion domain. Using this application it is possible to take a picture of a garment to retrieve its most similar products. The proposed method is firstly presented as an application in which the user manually select the name of the subject framed by the camera, before sending the request to the server. In the second part we propose an advanced approach which automatically classifies the object of interest, in this way it is possible to minimize the effort required by the user during the query process. In order to evaluate the performance of the proposed method, we have collected three datasets: the first contains clothing images of products taken from different online shops, whereas for the other datasets we have used images and video frames of clothes taken by Internet users. The results show the feasibility in the use of the proposed mobile application in a real scenario.
Given the lack of modern techniques to ensure the digital privacy of individuals, we want to pave the way for a new approach to make pedestrians in cityscape images anonymous. To address these concerns, we propose an automated method to replace any unknown pedestrian with another one which is extracted from a controlled and authorized dataset. The techniques used up to now to make people anonymous are based mainly on the blurring of people's faces, but even so it is possible to trace the identity of the subject starting from his clothing, personal items, hairstyle, the place and time where the photo was taken. The proposed method aims to make the pedestrians completely anonymous, and consists of four phases: firstly we identify the area where the pedestrian is located, we separate the pedestrian from the background, we select the most similar pedestrian from a controlled dataset and subsequently we substitute it. Our case study is Google Street View because it is one of the online services which suffers most from this kind of privacy issues. The experimental results show how this technique can overcome the problems of digital privacy with promising results.
In this study we propose a method for the automatic extraction of Visual Attributes from images. In particular, our case study concerns the processing of images related to commercial offers in the fashion domain and the results show how the use of the proposed method can be successfully applied in a real context. This method is based on a pre-processing phase in which an object detection algorithm identifies the object of interest, subsequently the visual attributes are extracted using a descriptor based on the Pyramid of Histograms of Orientation Gradients. In order to classify these descriptions, we have trained a discriminative model using a manually annotated dataset of commercial offers, which we released for future comparisons. To increase the performance of the visual attributes extraction, the results provided by the previous step have been refined with an a priori probability which models the occurrence of each visual attribute with a specific product type, opportunely estimated on the dataset.
Nowadays an increasing number of people own mobile phones with built-in camera, able to take pictures. Thus, having a fast and fully automatic algorithm of image retrieval is considered a promising way to identify plant leaves on a mobile device. Our solution proposes a Support Vector Machine that provides a multi-class probability estimation with radial basis function kernel based on two descriptors: PHOG and a vari- ant of HAAR. With our method we placed seventh among all the fully automatic methods who participated in the ImageCLEF Plant Identi- cation 2012
We describe a web application that takes advantage of new computer vision techniques to allow the user to make searches based on visual similarity of color and texture related to the object of interest. We use a supervised neural network strategy to segment different classes of objects. A strength of this solution is the high speed in generalization of the trained neural networks, in order to obtain an object segmentation in real time. Information about the segmented object, such as color and texture, are extracted and indexed as text descriptions. Our case study is the online commercial offers domain where each offer is composed by text and images. Many successful experiments were done on real datasets in the fashion field.
This work aims at defining an extension of a competitive method for matching correspondences in stereoscopic image analysis. The method we extended was proposed by Venkatesh. Y.V. et al where the authors extend a Self-Organizing Map by changing the neural weights updating phase in order to solve the correspondence problem within a two-frame area matching approach and producing dense disparity maps. In the present paper we have extended the method mentioned by adding some details that lead to better results. Experimental studies were conducted to evaluate and compare the solution proposed.
This paper proposes a system enforcing content-based message filtering for On-line Social Networks (OSNs). The system allows OSN users to have a direct control on the messages posted on their walls. This is achieved through a flexible rule-based system, that allows a user to customize the filtering criteria to be applied to their walls, and a Machine Learning based soft classifier automatically labelling messages in support of content-based filtering.
Ignazio Gallo合作论文数DiSTA, University of Insubria11