In many reinforcement learning (RL) tasks, the state-action space may be subject to changes over time (e.g., increased number of observable features, changes of representation of actions). Given these changes, the previously learnt policy will likely fail due to the mismatch of input and output features, and another policy must be trained from scratch, which is inefficient in terms of sample complexity. Recent works in transfer learning have succeeded in making RL algorithms more efficient by incorporating knowledge from previous tasks, thus partially alleviating this problem. However, such methods typically must provide an explicit state-action correspondence of one task into the other. An autonomous agent may not have access to such high-level information, but should be able to analyze its experience to identify similarities between tasks. In this paper, we propose a novel method for automatically learning a correspondence of states and actions from one task to another through an agent’s experience. In contrast to previous approaches, our method is based on two key insights: i) only the first state of the trajectories of the two tasks is paired, while the rest are unpaired and randomly collected, and ii) the transition model of the source task is used to predict the dynamics of the target task, thus aligning the unpaired states and actions. Additionally, this paper intentionally decouples the learning of the state-action corresponce from the transfer technique used, making it easy to combine with any transfer method. Our experiments demonstrate that our approach significantly accelerates transfer learning across a diverse set of problems, varying in state/action representation, physics parameters, and morphology, when compared to state-of-the-art algorithms that rely on cycle-consistency.
Quality control of flour is essential to control the quality of bread produced from it. We propose a control method based on the morphological characteristics of the granules of starch. The automation of the identification, segmentation and determination of the average size of the granules of starch of each of the cereals that make up a flour, from microscopy images, is an essential procedure for producers who want to produce bread under the protected geographical indication (PGI) 'Galician Bread'. This identification and counting procedure, if performed manually, is a tedious activity for a trained expert, and is very time consuming. Thus, automating this task would streamline the process, in addition to saving a great deal of time. This paper addresses this problem by using deep learning approaches (Mask R-CNN) to predict the type of the granule of starch and its size for the first time. The trained models are then evaluated with the same raw microscopy images of these granules observed under polarized light, as has been previouly used for manual identification and counting. A dataset comprising 1308 2564 x 1924-pixel images is analysed. The images contain 17000 labelled granules of starch for two types of wheat: commercial wheat flour from 'Castilla' (type 0) and the Galician autochthonous flour 'Caaveiro' (type 1). The number of samples is approximately the same for each class. Instance segmentation with Mask R-CNN (Model II) achieved valid results for unseen images, with a categorical global accuracy of about 88.6% and with a discrepancy with respect to the areas of the granules as estimated by a human expert of less than 4%. The performance achieved by Mask R-CNN produces a strong correlation between the results of an expert and the results of the network, confirming the practical validity of our proposal.
This chapter serves as an introduction to 3D representations of scenes or Structure From Motion (SfM) from straight line segments. Lines are frequently found in captures of man-made environments, and in nature are mixed with more organic shapes. The inclusion of straight lines in 3D representations provide structural information about the captured shapes and their limits, such as the intersection of planar structures. Line based SfM methods are not frequent in the literature due to the difficulty of detecting them reliably, their morphological changes under changes of perspective and the challenges inherent to finding correspondences of segments in images between the different views. Additionally, compared to points, lines add the dimensionalities carried by the line directions and lengths, which prevents the epipolar constraint to be valid along a straight line segment between two different views. This chapter introduces the geometrical relations which have to be exploited for SfM sketch or abstraction based on line segments, the optimization methods for its optimization, and how to compare the experimental results with Ground-Truth measurements.
A benchmark of saliency models performance with a synthetic image dataset is provided. Model performance is evaluated through saliency metrics as well as the influence of model inspiration and consistency with human psychophysics. SID4VAM is composed of 230 synthetic images, with known salient regions. Images were generated with 15 distinct types of low-level features (e.g. orientation, brightness, color, size...) with a target-distractor pop-out type of synthetic patterns. We have used Free-Viewing and Visual Search task instructions and 7 feature contrasts for each feature category. Our study reveals that state-of-the-art Deep Learning saliency models do not perform well with synthetic pattern images, instead, models with Spectral/Fourier inspiration outperform others in saliency metrics and are more consistent with human psychophysical experimentation. This study proposes a new way to evaluate saliency models in the forthcoming literature, accounting for synthetic images with uniquely low-level feature contexts, distinct from previous eye tracking image datasets.
Scene recognition is still a very important topic in many fields, and that is definitely the case in robotics. Nevertheless, this task is view-dependent, which implies the existence of preferable directions when recognizing a particular scene. Both in human and computer vision-based classification, this actually often turns out to be biased. In our case, instead of trying to improve the generalization capability for different view directions, we have opted for the development of a system capable of filtering out noisy or meaningless images while, on the contrary, retaining those views from which is likely feasible that the correct identification of the scene can be made. Our proposal works with a heuristic metric based on the detection of key points in 3D meshes (Harris 3D). This metric is later used to build a model that combines a Minimum Spanning Tree and a Support Vector Machine (SVM). We have performed an extensive number of experiments through which we have addressed (a) the search for efficient visual descriptors, (b) the analysis of the extent to which our heuristic metric resembles the human criteria for relevance and, finally, (c) the experimental validation of our complete proposal. In the experiments, we have used both a public image database and images collected at our research center.
The authors regret that the font style of the variable “S” with subscript “D”, namely “S_D” in Eq. (11) was incorrectly converted as mathcal font. The correct one is “S_D”, not “\mathcal{S_D}”. The authors would like to apologise for any inconvenience caused.
•New method for obtaining a 3D reconstruction of objects and scenes based on lines.•The model does not require previous result from other SfM pipelines.•Resilient to low texture and a low number of images. Aimed for UAVs.
In this study we provide the analysis of eye movement behavior elicited by low-level feature distinctiveness with a dataset of synthetically-generated image patterns. Design of visual stimuli was inspired by the ones used in previous psychophysical experiments, namely in free-viewing and visual searching tasks, to provide a total of 15 types of stimuli, divided according to the task and feature to be analyzed. Our interest is to analyze the influences of low-level feature contrast between a salient region and the rest of distractors, providing fixation localization characteristics and reaction time of landing inside the salient region. Eye-tracking data was collected from 34 participants during the viewing of a 230 images dataset. Results show that saliency is predominantly and distinctively influenced by: 1. feature type, 2. feature contrast, 3. temporality of fixations, 4. task difficulty and 5. center bias. This experimentation proposes a new psychophysical basis for saliency model evaluation using synthetic images.
Finding counterparts for straight lines over multiple images is a fundamental task in image processing, and the base for 3D reconstruction methods using segments. This paper introduces novel insights to improve the state-of-the-art unsupervised line matching over groups of images, aimed to source geometrical relations for 3D reconstruction algorithms. Most of the line-based 3D reconstruction methods published are ballasted as a consequence of sourcing the correspondences from matching methods that are not designed for this purpose. The repetitive line patterns present in many man-made structure turns difficult to came up with an outliers-free set of segment correspondences. The presented approach integrates an outliers detector based on 3D structure into a state of the art line matching algorithm.
Given the amount and variety of saliency models, the knowledge of their pros and cons, the applications they are more suitable for, or which are the more challenging scenes for each of them, would be very useful for the progress in the field. This assessment can be done based on the link between algorithms and public datasets. In one hand, performance scores of algorithms can be used to cluster video samples according to the pattern of difficulties they pose to models. In the other hand, cluster labels can be combined with video annotations to select discriminant attributes for each cluster. In this work we seek this link and try to describe each cluster of videos in a few words.
General dynamic scenes involve multiple rigid and flexible objects, with relative and common motion, camera induced or not. The complexity of the motion events together with their strong spatio-temporal correlations make the estimation of dynamic visual saliency a big computational challenge. In this work, we propose a computational model of saliency based on the assumption that perceptual relevant information is carried by high-order statistical structures. Through whitening, we completely remove the second-order information (correlations and variances) of the data, gaining access to the relevant information. The proposed approach is an analytically tractable and computationally simple framework which we call Dynamic Adaptive Whitening Saliency (AWS-D). For model assessment, the provided saliency maps were used to predict the fixations of human observers over six public video datasets, and also to reproduce the human behavior under certain psychophysical experiments (dynamic pop-out). The results demonstrate that AWS-D beats state-of-the-art dynamic saliency models, and suggest that the model might contain the basis to understand the key mechanisms of visual saliency. Experimental evaluation was performed using an extension to video of the well-known methodology for static images, together with a bootstrap permutation test (random label hypothesis) which yields additional information about temporal evolution of the metrics statistical significance.
A novel approach for line matching is proposed, aimed at achieving good performance with low-textured scenes, under uncontrolled illumination conditions. Line matching is performed by an iterative process that uses structural information collected through the use of different line neighbourhoods, making the set of matched lines grows robustly at each iteration. Results show that this approach is suitable to deal with low-textured scenes, and also robust under a wide variety of image transformations.
A novel approach for line detection and matching is proposed, aimed at achieving good performance with low-textured scenes, under uncontrolled illumination conditions. Line detection is performed by means of phase-based edge detector over Gaussian scale-space, followed by a multi-scale fusion stage which has been proven to be profitable in minimizing the number of fragmented and overlapped segments. Line matching is performed by an iterative process that uses structural information collected through the use of different line neighborhoods, making the set of matched lines grow robustly at each iteration. Results show that this approach is suitable to deal with low-textured scenes, and also robust under a wide variety of image transformations.
Industrial applications of photogrammetry usually involve the detection and recognition of targets placed on the surface of objects. The existing algorithms that allow to cope with the varying conditions of industrial environments are computationally very expensive. This paper proposes an implementation on graphics processing unit (GPU) for the detection of circular targets based on radial symmetry detection and parallel implementation in central processing unit (CPU) for location and recognition of these targets. The tests we have performed show the efficiency of GPU and CPU implementation in terms of execution times and performance.
This paper describes a unified framework for the static and dynamic saliency detection by whitening both the chromatic characteristics and the spatio-temporal frequencies. This approach is grounded in the statistical adaptation to the input data, resembling the human visual system’s early codification. Our approach, AWS-D, outperforms state-of-the-art models in the ability to predict human eye fixations while freely viewing a set of videos from three open access datasets (task free). We used as assessment measure an adaptation of the shuffling-AUC metric to spatio-temporal stimulus, together with a permutation test. Under this criterion, AWS-D not only reaches the highest AUC values, but also holds significant AUC figures for longer periods of time (more frames), over all the videos used in the test. The model also reproduces psychophysical results obtained in pop-out experiments in agreement with human behavior.
This work presents a new measure for radial symmetry and an algorithm for its computation. This measure identifies radially symmetric blobs as locations with contributions from all orientations at some scale. Hence, at a given scale, radial symmetry is computed as the product of the responses of a set of even symmetric feature detectors, with different orientations. This operator presents low sensitivity to shapes lacking radial symmetry, is robust to noise, contrast changes and strong perspective distortions, and shows a narrow point spread function. A multi-resolution measure is provided, computed as the maximum of the symmetry measure evaluated over a set of scales. We have applied this measure in the field of photogrammetry for the detection of circular coded fiducial targets. The detection of local maxima of multi-resolution radial symmetry is combined with a step of false-positive rejection, based on elliptical model fitting. In our experiments, the efficiency of target detection with this method is improved regarding a well-known commercial system, which is expected to improve the performance of bundle adjustment techniques. In order to fulfill all steps previous to bundle adjustment, we have also developed our own method for recognition of coded targets. This is accomplished by a standard procedure of segmentation and decoding of the ring sequence. Nevertheless, we have included a step for the verification of false positives of decoding based on correlation with reference targets. As far as we know, this approach cannot be found in literature.
This paper presents a novel approach to visual saliency that relies on a contextually adapted representation produced through adaptive whitening of color and scale features. Unlike previous models, the proposal is grounded on the specific adaptation of the basis of low level features to the statistical structure of the image. Adaptation is achieved through decorrelation and contrast normalization in several steps in a hierarchical approach, in compliance with coarse features described in biological visual systems. Saliency is simply computed as the square of the vector norm in the resulting representation. The performance of the model is compared with several state-of-the-art approaches, in predicting human fixations using three different eye-tracking datasets. Referring this measure to the performance of human priority maps, the model proves to be the only one able to keep the same behavior through different datasets, showing free of biases. Moreover, it is able to predict a wide set of relevant psychophysical observations, to our knowledge, not reproduced together by any other model before.
A hierarchical definition of optical variability is proposed that links physical magnitudes to visual saliency and yields a more reductionist interpretation than previous approaches. This definition is shown to be grounded on the classical efficient coding hypothesis. Moreover, we propose that a major goal of contextual adaptation mechanisms is to ensure the invariance of the behavior that the contribution of an image point to optical variability elicits in the visual system. This hypothesis and the necessary assumptions are tested through the comparison with human fixations and state-of-the-art approaches to saliency in three open access eye-tracking datasets, including one devoted to images with faces, as well as in a novel experiment using hyperspectral representations of surface reflectance. The results on faces yield a significant reduction of the potential strength of semantic influences compared to previous works. The results on hyperspectral images support the assumptions to estimate optical variability. As well, the proposed approach explains quantitative results related to a visual illusion observed for images of corners, which does not involve eye movements.
In this work we study how we can use a novel model of spatial saliency (visual attention) combined with image features to significantly accelerate a scene recognition application and, at the same time, preserve recognition performance. To do so, we use a mobile robotlike application where scene recognition is carried out through the use of image features to characterize the different scenarios, and the Nearest Neighbor rule to carry out the classification. SIFT and SURF are two recent and competitive alternatives to image local featuring that we compare through extensive experimental work. Results from the experiments show that SIFT features perform significantly better than SURF features achieving important reductions in the size of the database of prototypes without significant losses in recognition performance, and thus, accelerating scene recognition. Also, from the experiments it is concluded that SURF features are less distinctive when using very large databases of interest points, as it occurs in the present case. Visual attention is the process by which the Human Visual System (HVS) is able to select from a given scene regions of interest that contain salient information, and thus, reduce the amount of information to be processed (Treisman, 1980; Koch, 1985). In the last decade, several computational models biologically motivated have been released to implement visual attention in image and video processing (Itti, 2000; Garcia-Diaz, 2008). Visual attention has also been used to improve object recognition and scene analysis (Bonaiuto, 2005; Walther, 2005). In this chapter, we study the utility of using a novel model of spatial saliency to improve a scene recognition application by reducing the amount of prototypes needed to carry out the classification task. The application is based on mobile robot-like video sequences taken in indoor facilities formed by several rooms and halls. The aim is to recognize the different scenarios in order to provide the mobile robot system with general location data. The visual attention approach is a novel model of bottom-up saliency that uses local phase information of the input data where the statistic information of second order is deleted to achieve a Retinoptical map of saliency. The proposed approach joints computational mechanisms of the two hypotheses largely accepted in early vision: first, the efficient coding