Deep Neural Networks lead the state of the art of computer vision tasks. Despite this, Neural Networks are brittle in that small changes in the input can drastically affect their prediction outcome and confidence. Consequently and naturally, research in this area mainly focus on adversarial attacks and defenses. In this paper, we take an alternative stance and introduce the concept of Assistive Signals, which are optimized to improve a model's confidence score regardless if it's under attack or not. We analyse some interesting properties of these assistive perturbations and extend the idea to optimize assistive signals in the 3D space for real-life scenarios simulating different lighting conditions and viewing angles. Experimental evaluations show that the assistive signals generated by our optimization method increase the accuracy and confidence of deep models more than those generated by conventional methods that work in the 2D space. In addition, our Assistive Signals illustrate the intrinsic bias of ML models towards certain patterns in real-life objects. We discuss how we can exploit these insights to re-think, or avoid, some patterns that might contribute to, or degrade, the detectability of objects in the real-world.
Deep Neural Networks are brittle in that small changes in the input can drastically affect their prediction outcome and confidence. Consequently, research in this area mainly focus on adversarial attacks and defenses. In this paper, we take an alternative stance and introduce the concept of Assistive Signals, which are perturbations optimized to improve a model’s confidence score regardless if it’s under attack or not. We analyze some interesting properties of these assistive perturbations and extend the idea to optimize them in the 3D space simulating different lighting conditions and viewing angles. Experimental evaluations show that the assistive signals generated by our optimization method increase the accuracy and confidence of deep models more than those generated by conventional methods that work in the 2D space. ‘Assistive Signals’ also illustrate bias of ML models towards certain patterns in real-life objects.
In remote sensing image-blurring is induced by many sources such as atmospheric scatter, optical aberration, spatial and temporal sensor integration. The natural blurring can be exploited to speed up target search by fast template matching. In this paper, we synthetically induce additional non-uniform blurring to further increase the speed of the matching process. To avoid loss of accuracy, the amount of synthetic blurring is varied spatially over the image according to the underlying content. We extend transitive algorithm for fast template matching by incorporating controlled image blur. To this end we propose an Efficient Group Size (EGS) algorithm which minimizes the number of similarity computations for a particular search image. A larger efficient group size guarantees less computations and more speedup. EGS algorithm is used as a component in our proposed Optimizing auto-correlation (OptA) algorithm. In OptA a search image is iteratively non-uniformly blurred while ensuring no accuracy degradation at any image location. In each iteration efficient group size and overall computations are estimated by using the proposed EGS algorithm. The OptA algorithm stops when the number of computations cannot be further decreased without accuracy degradation. The proposed algorithm is compared with six existing state of the art exhaustive accuracy techniques using correlation coefficient as the similarity measure. Experiments on satellite and aerial image datasets demonstrate the effectiveness of the proposed algorithm.
We present an image set classification algorithm based on unsupervised clustering of labeled training and unlabeled test data where labels are only used in the stopping criterion. The probability distribution of each class over the set of clusters is used to define a true set based similarity measure. To this end, we propose an iterative sparse spectral clustering algorithm. In each iteration, a proximity matrix is efficiently recomputed to better represent the local subspace structure. Initial clusters capture the global data structure and finer clusters at the later stages capture the subtle class differences not visible at the global scale. Image sets are compactly represented with multiple Grassmannian manifolds which are subsequently embedded in Euclidean space with the proposed spectral clustering algorithm. We also propose an efficient eigenvector solver which not only reduces the computational cost of spectral clustering by many folds but also improves the clustering quality and final classification results. Experiments on five standard datasets and comparison with seven existing techniques show the efficacy of our algorithm.
In remote sensing image-blurring is induced by many sources such as atmospheric scatter, optical aberration, spatial and temporal sensor integration. The natural blurring can be exploited to speed up target search by fast template matching. In this paper, we synthetically induce additional non-uniform blurring to further increase the speed of the matching process. To avoid loss of accuracy, the amount of synthetic blurring is varied spatially over the image according to the underlying content. We extend transitive algorithm for fast template matching by incorporating controlled image blur. To this end we propose an Efficient Group Size (EGS) algorithm which minimizes the number of similarity computations for a particular search image. A larger efficient group size guarantees less computations and more speedup. EGS algorithm is used as a component in our proposed Optimizing auto-correlation (OptA) algorithm. In OptA a search image is iteratively non-uniformly blurred while ensuring no accuracy degradation at any image location. In each iteration efficient group size and overall computations are estimated by using the proposed EGS algorithm. The OptA algorithm stops when the number of computations cannot be further decreased without accuracy degradation. The proposed algorithm is compared with six existing state of the art exhaustive accuracy techniques using correlation coefficient as the similarity measure. Experiments on satellite and aerial image datasets demonstrate the effectiveness of the proposed algorithm.
We present automatic extraction of local 3D features (L3DF) from ear and face biometrics and their combination at the feature and score levels for robust identification. To the best of our knowledge, this paper is the first to present feature level fusion of 3D features extracted from ear and frontal face data. Scores from L3DF based matching are also fused with iterative closest point algorithm based matching using a weighted sum rule. We achieve identification and verification (at 0.001 FAR) rates of 99.0% and 99.4%, respectively, with neutral and 96.8% and 97.1% with non-neutral facial expressions on the largest public databases of 3D ear and face.
The complete nucleotide sequences of more than 100 isolates of Potato spindle tuber viroid (PSTVd) collected from locations in the territory of Russia and the former USSR have been determined. These sequences represent 43 individual sequence variants, each containing 1–10 mutations with respect to the “intermediate” or type strain of PSTVd (GenBank Acc. No. V01465). Isolates containing 2–5 mutations were the most common, and 24 sequence variants are described here for the first time. Twenty one isolates contained a mutation found only in Russian and Ukrainian isolates of PSTVd up till now; i.e., replacement of the adenine at position 121 with cytosine (A121C). Many of these isolates contained two mutations—deletion of one of three adenine residues occupying positions 118–120 plus replacement of the adenine at position 121 with either uracil or cytosine (−A120, A121U/C). Both combinations of mutations were phenotypically neutral, i.e. symptom expression in Rutgers tomato was unaffected. Phylogenetic analysis of the sequences of different PSTVd isolates presented in work together with sequences of other naturally-occurring isolates obtained from Internet databases suggesting that known PSTVd isolates may be divided into four groups: (i) a group of isolates from potato and ornamentals where the type strain of PSTVd (PSTVd.018) may be considered to represent the ancestral sequence, (ii) a second group of isolates from potato and ornamentals where PSTVd.123 play the same role as PSTVd.018 for the first group, and iii) a group of potato isolates where PSTVd.125 is a possible ancestral sequence. The fourth and most divergent group of PSTVd isolates differs significantly from these first three groups. The majority of isolates in this group originate from New Zealand and Australia and infect different solanaceous hosts (tomato, pepper, cape gooseberry, potato, and others).
Simple nearest neighbor classification fails to exploit the additional information in image sets. We propose self-regularized nonnegative coding to define between set distance for robust face recognition. Set distance is measured between the nearest set points (samples) that can be approximated from their orthogonal basis vectors as well as from the set samples under the respective constraints of self-regularization and nonnegativity. Self-regularization constrains the orthogonal basis vectors to be similar to the approximated nearest point. The nonnegativity constraint ensures that each nearest point is approximated from a positive linear combination of the set samples. Both constraints are formulated as a single convex optimization problem and the accelerated proximal gradient method with linear-time Euclidean projection is adapted to efficiently find the optimal nearest points between two image sets. Using the nearest points between a query set and all the gallery sets as well as the active samples used to approximate them, we learn a more discriminative Mahalanobis distance for robust face recognition. The proposed algorithm works independently of the chosen features and has been tested on gray pixel values and local binary patterns. Experiments on three standard data sets show that the proposed method consistently outperforms existing state-of-the-art methods.
We propose an efficient and robust solution for image set classification. A joint representation of an image set is proposed which includes the image samples of the set and their affine hull model. The model accounts for unseen appearances in the form of affine combinations of sample images. To calculate the between-set distance, we introduce the Sparse Approximated Nearest Point (SANP). SANPs are the nearest points of two image sets such that each point can be sparsely approximated by the image samples of its respective set. This novel sparse formulation enforces sparsity on the sample coefficients and jointly optimizes the nearest points as well as their sparse approximations. Unlike standard sparse coding, the data to be sparsely approximated are not fixed. A convex formulation is proposed to find the optimal SANPs between two sets and the accelerated proximal gradient method is adapted to efficiently solve this optimization. We also derive the kernel extension of the SANP and propose an algorithm for dynamically tuning the RBF kernel parameter while matching each pair of image sets. Comprehensive experiments on the UCSD/Honda, CMU MoBo, and YouTube Celebrities face datasets show that our method consistently outperforms the state of the art.
Biometric-based human recognition is rapidly gaining popularity due to breaches of traditional security systems and the lowering cost of sensors. The current research trend is to use 3D data and to combine multiple traits to improve accuracy and robustness. This article comprehensively reviews unimodal and multimodal recognition using 3D ear and face data. It covers associated data collection, detection, representation, and matching techniques and focuses on the challenging problem of expression variations. All the approaches are classified according to their methodologies. Through the analysis of the scope and limitations of these techniques, it is concluded that further research should investigate fast and fully automatic ear-face multimodal systems robust to occlusions and deformations.
Classification based on image sets has recently attracted great research interest as it holds more promise than single image based classification. In this paper, we propose an efficient and robust algorithm for image set classification. An image set is represented as a triplet: a number of image samples, their mean and an affine hull model. The affine hull model is used to account for unseen appearances in the form of affine combinations of sample images. We introduce a novel between-set distance called Sparse Approximated Nearest Point (SANP) distance. Unlike existing methods, the dissimilarity of two sets is measured as the distance between their nearest points, which can be sparsely approximated from the image samples of their respective set. Different from standard sparse modeling of a single image, this novel sparse formulation for the image set enforces sparsity on the sample coefficients rather than the model coefficients and jointly optimizes the nearest points as well as their sparse approximations. A convex formulation for searching the optimal SANP between two sets is proposed and the accelerated proximal gradient method is adapted to efficiently solve this optimization. Experimental evaluation was performed on the Honda, MoBo and Youtube datasets. Comparison with existing techniques shows that our method consistently achieves better results.
Image based human pose recovery has many applications in different industries such as games, entertainment, physiological rehabilitation and biometrics. This paper presents a new pose estimation algorithm from monocular images based on a nonlinear mapping of human silhouettes, coded using a collection of local image moments, to the pose space using a mixture of Neural Networks (NN) regressors. All parameters are estimated automatically. Experiments and comparative results show a superior performance of the proposed method.
The potential of viroid infection to dwarf citrus growing in intensive plantings is well established.How viroids exert this dwarfing effect is not known, but one possibility involves limiting the size of the root system.As part of an ongoing effort to develop Citrus viroid III (CVd-III) for use with rootstocks other than trifoliate orange or its hybrids, we have studied the effects of viroid infection on root development under greenhouse conditions using three rootstock/scion combinations; i.e., rooted Etrog citron cuttings, trifoliate orange seedlings, and young Valencia orange/trifoliate orange grafted trees.Groups of 10 young Etrog cuttings growing under greenhouse conditions were slash-inoculated with CVd-IIIb RNA transcripts and then observed for up to 12 mo with periodic cutbacks.Three months post-inoculation, the viroid-infected Etrog plants were significantly shorter than the uninoculated controls.By 6 mo post-inoculation, the inhibitory effect of CVd-IIIb on root dry weight had also become statistically significant.Between 6 and 12 mo, effects on root weight and development continued to intensify.Graft inoculation of trifoliate orange seedlings or Valencia scions growing on trifoliate orange rootstocks with either CVd-IIIa or CVd-IIIb resulted in a similar (though not statistically significant) inhibition of root dry weight accumulation over an 18 mo period.
Forty PSTVd isolates collected from five regions of Russia (North-western, Central, Volga region, Northern Caucasus and the Far East) were sequenced during 2006-2008. All isolates lacked the adenine residue present at position 123 of the type strain; i.e., PSTVd-intermediate (GenBank V01465). Nineteen Russian isolates also contained an adenosine --> cytosine substitution at position 120. Twenty two additional unique changes were also observed in one or more of the isolates sequenced.
The use of biometrics in human recognition is rapidly gaining in popularity. Considering the limitations of biometric systems with a single biometric trait, a current research trend is to combine multiple traits to improve performance. Among the biometric traits, the face is considered to be the most non-intrusive and the ear is the most promising candidate to be combined with the face. In this survey, existing approaches involving these two biometric traits are summarized and their scopes and limitations are discussed Some challenges to be addressed are discussed and few research directions are outlined.
Gareth Loy合作论文数Royal Institute of Technology3