Computer Vision and Biometrics benefit from the recent advances in Pattern Recognition and Artificial Intelligence, which tends to make model-based face recognition more efficient. Also, deep learning combined with data augmentation tends to enrich the training sets used for learning tasks. Nevertheless, face recognition still is challenging, especially because of imaging issues that occur in practice, such as changes in lighting, appearance, head posture and facial expression. In order to increase the reliability of face recognition, we propose a novel supervised appearance-based face recognition method which creates a low-dimensional orthogonal subspace that enforces the face class separability. The proposed approach uses data augmentation to mitigate the problem of training sample scarcity. Unlike most face recognition approaches, the proposed approach is capable of handling efficiently grayscale and color face images, as well as low and high-resolution face images. Moreover, proposed supervised method presents better class structure preservation than typical unsupervised approaches, and also provides better data preservation than typical supervised approaches as it obtains an orthogonal discriminating subspace that is not affected by the singularity problem that is common in such cases. Furthermore, a soft margins Support Vector Machine classifier is learnt in the low-dimensional subspace and tends to be robust to noise and outliers commonly found in practical face recognition. To validate the proposed method, an extensive set of face identification experiments was conducted on three challenging public face databases, comparing the proposed method with methods representative of the state-of-the-art. The proposed method tends to present higher recognition rates in all databases. In addition, the experiments suggest that data augmentation also plays an essential role in the appearance-based face recognition, and that the CIELAB color space (L*a*b) is generally more efficient than RGB for face recognition as it attenuates lighting variations.
O melanoma é o tipo mais letal de câncer de pele, uma vez que é mais propenso à metástase. Especificamente, a taxa de pacientes que sobrevivem pelo menos cinco anos após o diagnóstico dessa doença no estágio inicial é superior a 99%. No entanto, essa taxa diminui para cerca de 25% se a detecção ocorre somente no último estágio. Nesse contexto, sistemas que auxiliem no diagnóstico precoce do melanoma podem desempenhar um papel de extrema importância, especialmente em regiões nas quais o acesso a dermatologistas é precário. Contudo, diferenciar um melanoma de lesões melanocíticas benignas pode ser uma tarefa desafiadora, mesmo para especialistas experientes. Para lidar com esse problema, nesta tese, propõe-se um sistema automático para detectação de melanoma a partir de uma simples fotografia digital, o qual baseia-se em modelos de representações esparsas. Os resultados apresentados pelo sistema proposto são promissores e sugerem que o sistema proposto pode potencialmente superar alternativas estado-da-arte e até mesmo médicos treinados.
Melanoma is the most lethal type of skin cancer, since it is most prone to metastasis. Specifically, the rate of patients who survive at least five years after early stage diagnosis of this disease is over 99%. However, this rate decreases to about 25% if detection occurs only at the last stage. In this context, systems that assist in the early diagnosis of melanoma can play an extremely important role, especially in regions where access to dermatologists is poor. However, differentiating melanoma from benign melanocytic lesions can be a challenging task, even for experienced specialists. To address this problem, in this thesis, an automatic system is proposed for melanoma detection from a simple digital photograph, which is based on sparse representation models. The results presented by the proposed system are promising and suggest that it can potentially outperform state-of-the-art alternatives and even trained dermatologists.
Despite the difficulties imposed by the worldwide spread of COVID-19, some recent advances have been made by the Technical Committee 17 (TC-17) in topics regarding measurements and materials.
Object tracking is challenging, and recently, correlation filters methods have been proposed for this task. Most of these methods focus on the central portion of the target and are negatively affected by changes in the target size and shape. This work proposes a collaborative scheme using several local correlation filters combined with a global correlation filter for improving the performance of object tracking methods based on correlation filters. The proposed correlation filter used in this scheme is based on features extracted from multiple layers of deep convolutional neural networks, and a strategy to identify when these models should also be updated is presented. Experiments show that the proposed scheme tends to be consistent and to achieve better results than other comparative tracking approaches. The proposed collaborative approach can be applied to other correlation filters, which tends to further improve the tracker performance.
Object tracking remains a challenging problem in computer vision. Recently, methods based on correlation filters have achieved good performance in benchmarks. Usually, most of these visual trackers rely on tracking the object as a whole, not being able to handle the object variations. This paper proposes a scheme using several local correlation filters combined with a global correlation filter for improving the performance of object tracking methods based on correlation filters. We integrated this scheme into the traditional Kernelized Correlation Filter (KCF) method to evaluate our proposed approach. Experiments show that the proposed scheme is consistent and achieves better results compared to the baseline.
Melanoma is one of the deadliest types of pigmented skin lesions, and if identified in the earlier stages can increase the patient survival rate. The use of digital cameras as an alternative to other devices, such as dermatoscope, is gaining space in skin lesion prescreening with e-health systems used in the macroscopic diagnose of pigmented skin lesion images. The traditional framework used to classify macroscopic pigmented skin lesion (MPSL) images consists of a preprocessing step to remove hair and shading effects, followed by the lesion area detection and segmentation. Next, techniques are used to extract a set of features from the obtained region, and these attributes make it possible to distinguish between malignant and benign cases. Usually, the features are extracted from data labeled by a specialist, and are used to train a machine learning algorithm, which is then used to suggest a diagnosis for an undiagnosed skin lesion image. In this work, we present a review of some of the most recent advances in MPSL segmentation and classification.
Face biometry is a popular user authentication scheme that is easy to use and tends to be less invasive than other user authentication approaches. Despite the success achieved by face biometrics, face spoofing attacks (or presentation attacks) still pose a challenge to researchers. In practice, fraudsters may deceive a face authentication system by displaying fake copies of an authorised user face, such as photos or videos, and gain unauthorised access to the system. This work proposes a method for detecting unauthorised access attempts using misrepresentations of the identity of an authorised user. The proposed methodology for presentation attack detection relies on the observation of imaging and liveness attributes, such as the detection of liveness using the face deformation energy, and imaging attributes usually found in authentic accesses such as facial and background textures, and steganalysis features. Based on the experimental results, the proposed approach potentially can detect face spoofing attacks at each frame of video sequences with error rates of half total error rate (HTER)= {6.51, 4.93}%, and also in full video sequences with HTER = {5.55, 0}%, for the CASIA and Nanjing University of Aeronautics and Astronautics databases, respectively.
Biometry-based authentication systems are potential candidates to replace traditional username and password-based access schemes. Facial recognition is becoming widely popular, and many existing devices already include embedded cameras, making this technology easy to use. Nevertheless, facial recognition systems are prone to security breaches, such as facial spoofing attacks, where a impostor tries to gain access to the system by disguising him/herself as a genuine user. The goal of this paper is to propose a countermeasure capable of detecting unauthorized access attempts in facial recognition systems. Most authors uses only the face to detect facial spoofing attacks. However, we argue that more information available in the training data found on Presentation Attack Detection (PAD) datasets should be used, specially when adopting deep learning schemes. We show that using the full frame captured by the camera is more reliable than using only the face since the environment presents rich information that is useful to differentiate a genuine access from an impostor. We present a deep learning method that uses the entire frame instead of just the face to detect presentation attacks. The preliminary experimental results are encouraging, and based on a GoogLeNet architecture, the detection of such attacks potentially can be obtained in more than 96% of the test cases.
Geodesic distance is a natural dissimilarity measure between probability distributions of a specific type, and can be used to discriminate texture in image-based measurements. Furthermore, since there is no known closed-form solution for the geodesic distance between general multivariate normal distributions, we propose two efficient approximations to be used as texture dissimilarity metrics in the context of face recognition. A novel face recognition approach based on texture discrimination in high-resolution color face images is proposed, unlike the typical appearance-based approach that relies on low-resolution grayscale face images. In our face recognition approach, sparse facial features are extracted using predefined landmark topologies, that identify discriminative image locations on the face images. Given this landmark topology, the dissimilarity between distinct face images are scored in terms of the dissimilarities between their corresponding face landmarks, and the texture in each one of these landmarks is represented by multivariate normal distributions, expressing the color distribution in the vicinity of each landmark location. The classification of new face image samples occurs by determining the face image sample in the training set which minimizes the dissimilarity score, using the nearest neighbor rule. The proposed face recognition method was compared to methods representative of the state-of-the-art, using color or grayscale face images, and presented higher recognition rates. Moreover, the proposed texture dissimilarity metric also is efficient in general texture discrimination (e.g. texture recognition of material images), as our experiments suggest.
In this work, a new method based on a multi-model dictionary is proposed for face tracking. A reconstruction and a classification dictionary are combined, and each dictionary is learned from positive and negative examples. This scheme tends to enhance the discrimination between a tracked target face and the background. Also, an efficient scheme that collects data during face tracking is proposed to update the dictionaries in an incremental learning scheme, allowing to track faces even when the face appearance changes (e.g. under different face expressions). The preliminary experimental results suggest that the proposed method tends to perform better than comparative methods, which are representative of the state-of-the-art.
•A scalable graph compression algorithm for image segmentation proposed.•The input image is represented by a region graph model.•Texton dictionaries capture the local texture features in decoupled sub-graphs.•A graph compression algorithm reduces the graph size and segments the image.•Local graph decoupling and recoupling operations lead to an efficient method.
A camera-based scheme is proposed for detecting vehicles at user-defined virtual loops, simulating the operation of inductive loops. False vehicular detections are minimized by a combination of efficient edge detection and color information. The experimental results suggest that the proposed scheme potentially can detect and count vehicles at user-defined virtual loops accurately (with more than 98% correct detection rate, in average), besides being more robust to cast shadows and sudden illumination changes than comparable methods that represent the state of the art. (C) 2018 SPIE and IS&T
•A very high spatial resolution images land-use classification scheme is proposed.•The proposed method relies on dictionaries of deep features.•These dictionaries are very discriminative and compact.•Likelihoods are linked to the sparse representation approach.•The proposed method can be competitive in a comparison with the state-of-the-art.
Superpixels have many applications in visual information processing, and can be used to reduce redundant information of an image, as well as the computational complexity of other expensive tasks (e.g., image segmentation). In this work, an iterative hierarchical stochastic graph contraction (IHSGC) method for multi-scale superpixels generation is proposed. A stochastic strategy is used to generate multi-scale superpixels, and each superpixel is represented by a hierarchical tree and describes an image patch at fine and coarse scales simultaneously. The proposed method consists of two main steps. The first step initializes the method based on a multi-channel unsupervised stochastic over-segmentation at the pixel level. The proposed over-segmentation scheme actually performs hierarchical stochastic clustering of visual features (i.e. pixels, image patches, and potentially can be applied to other visual features as well), while preserving the local spatial relationships across different scales. The second step consists of an iterative hierarchical stochastic graph contraction method. Coarser scales are generated by graph contractions until the desired number of superpixels is obtained. The experimental results based on the popular Berkeley segmentation databases BSDS300 and BSDS500 suggest that the proposed approach potentially can perform better than comparative state-of-the-art methods in terms of boundary recall and under-segmentation error. (C) 2018 Elsevier Ltd. All rights reserved.
Faces carry a lot of information to distinguish different individuals. In this context, biometrics-based verification systems play a major role in terms of recognizing (or confirming) an individual identity, relying on physiological and/or behavioral characteristics among a set of individual biometric traits. In particular, facial recognition is important because it has a relatively low cost (i.e., it can be carried out using standard cameras) and is one of the least intrusive biometric modalities available, since it does not require physical contact like fingerprint recognition or retina scanning.
Diabetic macular edema (DME) affects the retina and reduces the visual acuity of patients with severe diabetic retinopathy. Its conventional treatment involves laser photocoagulation combined with infrared (IR) imaging. However, the laser beam may hit healthy retinal areas and cause unintentional retinal damages if retinal motion occurs. We propose a method for retinal motion detection and compensation that relies on phase correlations in the image and in the log-polar domains, and is robust to affine retinal deformations (e.g., rotations, scales, and translations) in IR videos. The proposed method is also robust to the background noise and illumination changes commonly occurring in this retinal imaging modality. The proposed method can be used to estimate the retinal affine motion parameters and compensate for small retinal motions (nearly 50 μm). The critical method parameters are selected and adjusted optimally, which improves the robustness of the method. The experimental results suggest that the proposed method potentially can be more robust for detecting retinal motion and for estimating the parameters of affine retinal deformations than comparable methods that currently represent the state of the art, which helps to improve the reliability of laser treatments for DME.
Discriminating shadows from the objects casting them often is challenging in practice, since the moving targets and their shadows tend to present similar motion patterns, and foreground detection methods often confuse cast shadows with foreground objects. To overcome these shadow detection difficulties, we propose a new stochastic shadow detection approach. In the proposed method, chromatic and gradient information are integrated with image hypergraph segmentation using a cascade of shadow/non-shadow classifiers, and a stochastic majority voting scheme is used to detect the shadow regions. The proposed method receives as input the segmented foreground objects and their cast shadows (mask), and outputs the shadows detected in the foreground mask. The experimental results were obtained with seven well known datasets, and suggest that the proposed shadow detection scheme can be more robust to different video acquisition conditions than other shadow detection methods, that are representative of the state-of-the-art.
We propose a novel generative approach for face recognition, in which sparse facial features are extracted from high resolution color face images using predefined landmark topologies which mark discriminative locations on face images, unlike the appearance-based approach, in which low resolution grayscale face images are used, reducing the computational complexity. By adopting a common landmark topology, the dissimilarity between distinct face images can be scored in terms of the dissimilarities between their corresponding landmarks, which are obtained by proposed geodesic distance approximations between multivariate normal distributions which represent the color intensities in the vicinities of each landmark location. The classification process of new face samples occurs by the determination of the face image sample present in the training set which minimizes the dissimilarity score. The proposed method was compared with representative current state-of-the-art methods using color or grayscale face images and presented the higher recognition rates. Moreover, these results also support a trend in which color information is relevant in face recognition.
Eliza Yingzi Du合作论文数Department of Electrical and Computer Engineering, School of Engineering and Technology, Indiana University-Purdue University2