Face expression recognition (FER) has been extensively explored by the research community, with comparisons between different FER models typically relying on accuracy metrics. To make model decisions more interpretable, evaluation can be complemented with the usage of (model-agnostic) explainability tools, leading to explainable FER models (XFER). However, the effectiveness of explainability tools in the context of XFER is seldom evaluated. This paper proposes a framework, entitled "face expression recognition explainability evaluation framework, based on action units" (XFER-AU), for evaluating the quality of explanations provided by different explainability tools in the context of XFER. To establish a comparison term, XFER-AU automatically generates a FER explanation ground truth, based on facial action units (AU). The proposed framework thus enables the comparison of explainability tools in terms of their ability to explain the decision made by an FER model. To perform the comparison, an evaluation metric called weighted explainability score (WES) is proposed, which takes into account the number of ground truth AUs covered and the precision of the explainability map produced. The proposed framework is used to compare several explainability tools, notably Local Interpretable Model-agnostic Explanation (LIME), Gradient-weighted Class Activation Mapping (Grad-CAM), SHapley Additive exPlanation (SHAP), Average Removal/ Aggregation (AVG), Minus and PLUS explanation (MinPLUS) and Randomised Input Sampling for Explanation (RISE). The proposed XFER-AU results report that, for the tested datasets, LIME explanations tend to better align with the ground truth. In more challenging scenarios, the FER models struggle to effectively point to the facial AUs relevant to explaining the observed expression, thereby reducing the model’s decision confidence. Results also show that even when an expression is correctly classified, older FER models, such as those based on VGG-16, support their decision on a smaller number of the relevant AUs when compared to more recent models, such as EfficientNet-B0.
This study introduces a framework applicable to videos of habitual behaviors of multiple bird species to automatically assess head angular velocities and frequencies during various behaviors in natural habitats. The process involves detecting birds, identifying key points on their heads, and tracking changes in their positions over time. Bird detection and key point extraction were trained on publicly available datasets, including Animal Kingdom, NABirds, Birdsnap, CUB-200-2011, and eBird, featuring videos and images of diverse bird species in uncontrolled settings. Initial challenges arose due to the complexity of video backgrounds, leading to misidentifications and inaccurate key point estimates. These issues were addressed through validation, refinement, filtering, and smoothing steps. Head angular velocities and rotation frequencies were computed from the refined key points. The algorithm performed well at moderate speeds but was limited by the 30 Hz frame rate of most eBird videos, which constrained measurable angular velocities and frequencies and caused motion blur, affecting key point detection. Our findings suggest that the framework may provide plausible estimates of head motion but also emphasize the importance of high frame rate videos in future research, including extensive comparisons against ground truth data, to fully characterize bird head movements. ### Competing Interest Statement The authors have declared no competing interest.
Artificial intelligence-based face recognition solutions are becoming increasingly popular. Therefore, it is crucial to fully understand and explain how these technologies work in order to make them more effective and acceptable to society. This is the goal of the CHIST-ERA project XAIface, the final results of which are reported in this article: a framework and toolkit for improving AI decision explainability, in the context of automated face recognition, through several novel methods are presented. These methods are integrated into an end-to-end face recognition demonstrator system, which facilitates studying the impact of various influencing factors and system processes on recognition performance. By doing so, we can visually explain the decisions made by the face verification pipeline for specific instances in our test set using heatmaps and locally interpretable features. Furthermore, we offer a comprehensive explanation of the end-to-end model by examining the relationship between verification failures and misclassifications of soft biometric facial traits.
The integration of Face Verification (FV) systems into multiple critical moments of daily life has become increasingly prevalent, raising concerns regarding the transparency and reliability of these systems. Consequently, there is a growing need for FV explainability tools to provide insights into the behavior of these systems. FV explainability tools that generate visual explanations, e.g., saliency maps, heatmaps, contour-based visualization maps, and face segmentation maps show promise in enhancing FV transparency by highlighting the contributions of different face regions to the FV decision-making process. However, evaluating the performance of such explainability tools remains challenging due to the lack of standardized assessment metrics and protocols. In this context, this paper proposes a subjective performance assessment protocol for evaluating the explainability performance of visual explanation-based FV explainability tools through pairwise comparisons of their explanation outputs. The proposed protocol encompasses a set of key specifications designed to efficiently collect the subjects’ preferences and estimate explainability performance scores, facilitating the relative assessment of the explainability tools. This protocol aims to address the current gap in evaluating the effectiveness of visual explanation-based FV explainability tools, providing a structured approach for assessing their performance and comparing with alternative tools. The proposed protocol is exercised and validated through an experiment conducted using two distinct heatmap-based FV explainability tools, notably FV-RISE and CorrRISE, taken as examples of visual explanation-based explainability tools, considering the various types of FV decisions, i.e., True Acceptance (TA), False Acceptance (FA), True Rejection (TR), and False Rejection (FR). A group of subjects with variety in age, gender, and ethnicity was tasked to express their preferences regarding the heatmap-based explanations generated by the two selected explainability tools. The subject preferences were collected and statistically processed to derive quantifiable scores, expressing the relative explainability performance of the assessed tools. The experimental results revealed that both assessed explainability tools exhibit comparable explainability performance for FA, TR, and FR decisions with CorrRISE performing slightly better than FV-RISE for TA decisions.
Every day millions of people travel on highways for work- or leisure-related purposes. Ensuring road safety is thus of paramount importance, and maintaining good-quality road pavements is essential, requiring an effective maintenance policy. The automation of some road pavement maintenance tasks can reduce the time and effort required from experts. This paper proposes a simple system to help speed up road pavement surface inspection and its analysis towards making maintenance decisions. A low-cost video camera mounted on a vehicle was used to capture pavement imagery, which was fed to an automatic crack detection and classification system based on deep neural networks. The system provided two types of output: (i) a cracking percentage per road segment, providing an alert to areas that require attention from the experts; (ii) a segmentation map highlighting which areas of the road pavement surface are affected by cracking. With this data, it became possible to select which maintenance or rehabilitation processes the road pavement required. The system achieved promising results in the analysis of highway pavements, and being automated and having a low processing time, the system is expected to be an effective aid for experts dealing with road pavement maintenance.
Explainable Face Recognition (XFR) is a critical technology to support the large deployment of learning-based face recognition solutions. This paper aims at contributing to the more transparent usage of Vision Transformers (ViTs) for face verification (FV) tasks, by proposing a novel approach for generating FV explainability heatmaps, for both positive and negative decisions. The proposed solution leverages on the attention maps generated by a ViT and employs masking techniques to create masks based on the highlighted regions in the attention maps. These masks are applied to the pair of faces, and the masking technique with most impact on the decision is selected to be used to generate heatmaps for the probe-gallery pair of faces. These heatmaps offer valuable insights into the decision- making process, shedding light on the most important face regions for the verification outcome. The key novelty of this paper lies in the proposed approach for generating explainability heatmaps tailored for verification pairs in the context of ViT models, which combines the ViT attention maps regions of the probe-gallery pair to create masks that allow evaluating those region's impact on the verification decision for both positive and negative decisions.
Many different cultures and countries have fish as a central piece in their diet, particularly in coastal countries such as Portugal, with the fishery and aquaculture sectors playing an increasingly important role in the provision of food and nutrition. As a consequence, fish-freshness evaluation is very important, although so far it has relied on human judgement, which may not be the most reliable at times. This paper proposes an automated non-invasive system for fish-freshness classification, which takes fish images as input, as well as a seabream fish image dataset. The dataset will be made publicly available for academic and scientific purposes with the publication of this paper. The dataset includes metadata, such as manually generated segmentation masks corresponding to the fish eye and body regions, as well as the time since capture. For fish-freshness classification four freshness levels are considered: very-fresh, fresh, not-fresh and spoiled. The proposed system starts with an image segmentation stage, with the goal of automatically segmenting the fish eye region, followed by freshness classification based on the eye characteristics. The system employs transformers, for the first time in fish-freshness classification, both in the segmentation process with the Segformer and in feature extraction and freshness classification, using the Vision Transformer (ViT). Encouraging results have been obtained, with the automatic fish eye region segmentation reaching a detection rate of 98.77%, an accuracy of 96.28% and a value of the Intersection over Union (IoU) metric of 85.7%. The adopted ViT classification model, using a 5-fold cross-validation strategy, achieved a final classification accuracy of 80.8% and an F1 score of 81.0%, despite the relatively small dataset available for training purposes.
Heat Map (HM)-based explainable Face Verification (FV) has the goal to visually interpret the decision-making of black-box FV models. Despite the impressive results, state-of-the-art FV explainability methods based on HMs mainly address genuine verification by generating visual explanations that reveal the similar face regions which most contributed for acceptance decisions. However, the similar face regions may not be the unique critical regions for the model decision, notably when rejection decisions are performed. To address this issue, this paper proposes a more complete FV explainability method, providing meaningful HM-based explanations for both genuine and impostor verification and associated acceptance and rejection decisions. The proposed method adapts the RISE algorithm for FV to generate Similarity Heat Maps (S- HMs) and Dissimilarity Heat Maps (D-HMs) which offer reliable explanations to all types of FV decisions. Qualitative and quantitative experimental results show the effectiveness of the proposed FV explainability method beyond state-of-the-art benchmarks.
A family of cellular automata arising from perturbations of a basic cellular automata rule, represented by 3E6IGS58S, in base 32, is studied. These rules can be seen as modeling idealized fluids in non-equilibrium, subject to interaction on distinct phases. Using adaptive techniques such as assembly and singular perturbation of cellular automata, we present several simulations showing the increase of complexity in the perturbed systems behavior, in particular, showing the increasing number of distinct spatial-temporal patterns exhibited.
This special issue of IET Biometrics, “BIOSIG 2021 Special Issue on Efficient, Reliable, and Privacy-Friendly Biometrics”, has as starting point the 2021 edition of the Biometric Special Interest Group (BIOSIG) conference. This special issue gathers works focussing on topics of biometric recognition put under the new light of fostering the efficiency, reliability and privacy of biometrics systems and methods. The “BIOSIG 2021 Special Issue on Efficient, Reliable, and Privacy-Friendly Biometrics” issue contains 12 papers, several of them being extended versions of papers presented at the BIOSIG 2021 conference, dealing with concrete research areas within biometrics such as Presentation Attack Detection for Face and Iris, Biometric Template Protection Schemes and Deep Learning techniques for Biometrics. Paper “Face Morphing Attacks and Face Image Quality: The Effect of Morphing and the Attack Detectability by Quality” was authored by Biying Fu and Naser Damer. This paper addresses the effect of morphing processes both on the perceptual image quality and the image utility in face recognition (FR) when compared to bona fide samples. This work provides an extensive analysis of the effect of morphing on face image quality, including both general image quality measures and face image utility measures, analysing six different morphing techniques and five different data sources using 10 different quality measures. The consistent separability between the quality scores of morphing attack and bona fide samples measured by certain quality measures sustains the proposal of performing unsupervised morphing attack detection (MAD) based on quality scores. The study looks into intra- and inter-dataset detectability to evaluate the generalisability of such a detection concept on different morphing techniques and bona fide sources. The results obtained point out that a set of quality measures, such as MagFace and CNNNIQA, can be used to perform unsupervised and generalised MAD with a correct classification accuracy of over 70%. Paper “Pixel-Wise Supervision for Presentation Attack Detection on ID Cards” was authored by Raghavendra Mudgalgundurao, Patrick Schuch, Kiran Raja, Raghavendra Ramachandra, and Naser Damer. This paper addresses the problem of detection of fake ID cards that are printed and then digitally presented for biometric authentication purposes in unsupervised settings. The authors propose a method based on pixel-wise supervision, using DenseNet, to leverage minute cues on various artefacts such as moiré patterns and artefacts left by the printers. To test the proposed system, a new database was obtained from an operational system, consisting of 886 users with 433 bona fide, 67 print and 366 display attacks (not publicly available due to GPDR regulations). The proposed approach achieves better performance compared to handcrafted features and deep learning models, with an Equal Error Rate (EER) of 2.22% and Bona fide Presentation Classification Error Rate (BPCER) of 1.83% and 1.67% @ Attack Presentation Classification Error Rate (APCER) of 5% and 10%, respectively. Paper “Deep Patch-Wise Supervision for Presentation Attack Detection” was authored by Alperen Kantarcı, Hasan Dertli, and Hazım Ekenel. This paper addresses the generalisation problem in face presentation attack detection (PAD). Specifically, convolutional neural networks (CNN)-based systems have gained significant popularity recently due to their high performance on intra-dataset experiments. However, these systems often fail to generalise to the datasets that they have not been trained on. This indicates that they tend to memorise dataset-specific spoof traces. To mitigate this problem, the authors propose a new presentation attack detection (PAD) approach that combines pixel-wise binary supervision with patch-based CNN. The presented experiments show that the proposed patch-based method forces the model not to memorise the background information or dataset-specific traces. The proposed method was tested on widely used PAD datasets—Replay-Mobile, OULU-NPU— and on a real-world dataset that has been collected for real-world PAD use cases. The results presented show that the proposed approach is found to be superior on challenging experimental setups. Namely, it achieves higher performance on OULU-NPU protocol 3, 4 and on inter-dataset real-world experiments. Paper “Transferability Analysis of Adversarial Attacks on Gender Classification to Face Recognition: Fixed and Variable Attack Perturbation” was authored by Zohra Rezgui, Amina Bassit, and Raymond Veldhuis. This paper focusses on the challenge of transferability of adversarial attacks. This work is motivated by the fact that it was proved in the literature that these attacks, targeting a specific model, are transferable among models performing the same task, however, the transferability scenarios are not considered in the literature for models performing different tasks but sharing the same input space and model architecture. In this paper, the authors study the above mentioned challenge regarding VGG16-based and ResNet50-based biometric classifiers. The impact of two white-box attacks on a gender classifier is investigated and then their robustness to defence methods is assessed by applying a feature-guided denoising method. Once the effectiveness of these attacks was established in fooling the gender classifier, we tested their transferability from the gender classification task to the facial recognition task with similar architectures in a black-box manner. Two verification comparison settings are employed, in which the authors compare images perturbed with the same and different magnitude of the perturbation. The presented results indicate transferability in the fixed perturbation setting for a Fast Gradient Sign Method (FGSM) attack and non-transferability in a Projected Gradient Descent (PGD) attack setting. The interpretation of this non-transferability can support the use of fast and train-free adversarial attacks targeting soft biometric classifiers as means to achieve soft biometric privacy protection while maintaining facial identity as utility. Paper “Combining 2D Texture and 3D Geometry Features for Reliable Iris Presentation Attack Detection using Light Field Focal Stack” was authored by Zhengquan Luo, Yunlong Wang, Nianfeng Liu, and Zilei Wang. In this paper, the authors leverage the merits of both light field (LF) imaging and deep learning (DL) to combine 2D texture and 3D geometry features for iris presentation attack detection (PAD). The proposed study explores off-the-shelf deep features of planar-oriented and sequence-oriented deep neural networks (DNNs) on the rendered focal stack. The proposed framework excavates the differences in 3D geometric structure and 2D spatial texture between bona fide and spoofing irises captured by LF cameras. A group of pre-trained DL models are adopted as feature extractor and the parameters of SVM classifiers are optimised on a limited number of samples. Moreover, two branch feature fusion further strengthens the framework's robustness and reliability against severe motion blur, noise, and other degradation factors. The results indicate that variants of the proposed framework significantly surpass the PAD methods that take 2D planar images or LF focal stack as input, even recent state-of-the-art methods fined-tuned on the adopted database. The results of multi-class attack detection experiments also verify the good generalisation ability of the proposed framework on unseen presentation attacks. Paper “Hybrid Biometric Template Protection: Resolving the Agony of Choice between Bloom Filters and Homomorphic Encryption” was authored by Amina Bassit, Florian Hahn, Chris Zeinstra, Raymond Veldhuis and Andreas Peter. This paper addresses the development of biometric template protection (BTP) schemes investigating the strengths and weaknesses of Bloom filters (BFs) and homomorphic encryption (HE). The paper notes that the pros and cons of BF-based and HE-based BTPs are not well studied in the literature and these two approaches both seem promising from a theoretical viewpoint. Thus, this work presents a comparative study of the existing BF-based BTPs and HE-based BTPs by examining their advantages and disadvantages from a theoretical standpoint. This comparison was applied to iris recognition as a study case, where the biometric and runtime performances of the BTP approaches were tested on the same setting, dataset, and implementation language. As a synthesis of this study, the authors propose a hybrid BTP scheme that combines the good properties of BFs and HE, ensuring unlinkability and high recognition accuracy, while being about 7 times faster than the traditional HE-based approach. The evaluation of the proposed scheme confirmed its biometric accuracy (an EER of 0:17% over the IITD iris database) and runtime efficiency (104:35 ms, 155:15 ms and 171:70 ms for 128,192, and 256 bits security level, respectively). Paper “Locality Preserving Binary Face Representations Using Auto-encoders” was authored by Mohamed Amine HMANI, Dijana Petrovska-Delacrétaz and Bernadette Dorizzi. This paper focusses on template protection schemes for face biometrics and introduces a novel approach to binarising biometric data using Deep Neural Networks (DNN) applied to facial data. The authors propose the use of DNN to extract binary embeddings from face images directly. The proposed binary embeddings give a state-of-the-art performance on two well-known databases (MOBIO and the Labelled Faces in the Wild (LFW)) with almost negligible degradation compared to the baseline. Further, as an application, the paper proposes a cancellable system based on the binary embeddings using a shuffling transformation with a randomisation key as a second factor. The cancellable system is analysed according to the ISO/IEC 24745:2011 standardised metrics. The templates generated by the cancellable system are unlinkable without the disclosure of the second factor. Paper “Reliable Detection of Doppelgängers based on Deep Face Representations” was submitted by Christian Rathgeb, Daniel Fischer, Pawel Drozdowski and Christoph Busch. This paper assesses the impact of doppelgängers (people that look alike) on the HDA Doppelgänger and Disguised Faces in The Wild databases using a state-of-the-art face recognition system, confirming that the existence of doppelgängers significantly increases false match rates. The paper then presents a method able to distinguish doppelgängers from mated comparison trials, by analysing differences in deep representations obtained from face image pairs. The proposed detection system achieves a state-of-the-art detection equal error rate of approximately 2.7% for the task of separating mated authentication attempts from doppelgängers in the mentioned databases. Paper “Benchmarking Human Face Similarity Using Identical Twins” was authored by Shoaib Meraj Sami, John McCauley, Sobhan Soleymani, Nasser Nasrabadi, and Jeremy Dawson. This paper addresses the problem of distinguishing identical twins and non-twin look-alikes in automated facial recognition (FR) applications. This work makes use of one of the largest twin datasets compiled to date to address two FR challenges: 1) determining a baseline measure of facial similarity between identical twins and 2) applying this similarity measure to determine the impact of doppelgangers, or look-alikes, on FR performance for large face datasets. The methodology proposed for facial similarity measure is based on a deep convolutional neural network trained on a tailored verification task designed to encourage the network to group together highly similar face pairs in the embedding space and achieves a test AUC of 0.9799. The proposed network provides a quantitative similarity score for any two given faces and has been applied to large-scale face datasets to identify similar face pairs. An additional analysis which correlates the comparison score returned by a facial recognition tool and the similarity score returned by the proposed network has also been performed. Paper “Discriminative Training of Spiking Neural Networks Organised in Columns for Stream-based Biometric Authentication” was authored by Enrique Argones Rúa, Tim Van hamme, Davy Preuveneers, and Wouter Joosen. In this paper, the authors address stream-based biometric authentication using a novel approach based on spiking neural networks (SNNs). SNNs have proven advantages regarding energy consumption and they are a perfect match with some proposed neuromorphic hardware chips, which can lead to a broader adoption of user device applications of artificial intelligence technologies. One of the challenges when using SNNs is the discriminative training of the network, since it is not straightforward to apply the well-known error backpropagation (EBP), massively used in traditional artificial neural networks (ANNs). Thus, the authors propose to use a network structure based on neuron columns, resembling cortical columns in the human cortex, and a new derivation of error backpropagation for the spiking neural networks that integrates the lateral inhibition in these structures. In the experiments presented, the potential of the proposed approach is tested in the task of inertial gait authentication, where gait is quantified as signals from Inertial Measurement Units (IMU). The proposed approach is compared to state-of-the-art ANNs being shown that SNNs provide competitive results, obtaining a difference of around 1% in Half Total Error Rate when compared to state-of-the-art ANNs in the context of IMU-based gait authentication. Paper “Towards Understanding the Character of Quality Sampling in Deep Learning Face Recognition” was authored by Iurii Medvedev, João Tremoço, Luís Espírito Santo, Beatriz Mano, and Nuno Gonçalves. This paper addresses the problem of the inconsistency between the training data and the deployment scenario in face-based biometric systems, which are developed specifically for dealing with ID document compliant images. This inconsistency is often caused by the choice of unconstrained face images of celebrities for training, motivated by its public availability opposed to the fact that existing document compliant face image collections are hardly accessible due to security and privacy issues. To mitigate the addressed problem, the authors propose to regularise the training of the deep face recognition network with a specific sample mining strategy, which penalises the samples by their estimated quality. This deep learning strategy is expanded to seek for the penalty (sampling character) that better satisfies the purpose of adapting deep learning face recognition for images of ID and travel documents. The presented experiments demonstrate the efficiency of the approach for ID document compliant face images. Paper “Masked Face Recognition: Human versus Machine” was authored by Naser Damer, Fadi Boutros, Marius Süßmilch, Meiling Fang, Florian Kirchbuchner and Arjan Kuijper. This paper focusses on the assessment of the effect of wearing a mask on face recognition (FR) in a collaborative environment. This work provides a joint evaluation and in-depth analyses of the face verification performance of human experts in comparison to state-of-the-art automatic FR solutions. In this paper, an extensive evaluation by human experts is presented along with four automatic recognition solutions. An analysis was made of the correlations between the verification behaviours of human experts and automatic FR solutions under different settings, such as involved unmasked pairs, masked probes and unmasked references, and masked pairs, with real and synthetic masks. The study concludes with a set of take-home messages on different aspects of the correlation between the verification behaviour of humans and machines. The Guest Editorial Board would like to thank all of the authors for their contributions to this Special Issue: “BIOSIG 2021 Special Issue on Efficient, Reliable, and Privacy-Friendly Biometrics”. We would also like to express our appreciation to the reviewers, who have provided insightful comments and suggestions contributing to improving the quality of the manuscripts. Data sharing is not applicable to this article as no new data were created or analysed in this study. Ana F. Sequeira holds a PhD in Electrical and Computer Engineering and a Degree and a Master in Mathematics. Sequeira's research is focussed on fundamental computer vision and machine learning topics and comprises anti-spoofing techniques (for iris, face and fingerprint); biometric recognition for border control; as well as facial analysis topics, such as emotion recognition, image compliance with standardisation requirements; and more recently, the study of interpretability of AI for biometrics. In particular, her research has focussed on biometric applications in challenging scenarios in use cases such as border control or transactions on mobile devices. Currently, Sequeira is a researcher at INESC TEC and, in the past, was a postdoctoral research assistant at the University of Reading, UK, collaborating in two European projects on biometrics for border control (FASTPASS—FP7 312583 and PROTECT—H2020 700259). In addition, Sequeira collaborated with the company IrisGuard UK to evaluate the iris recognition-based technology for monetary transactions on mobile devices—EyePay Technology. Sequeira led the construction of several biometric databases; managed biometric competitions focussing on iris spoofing, multimodal recognition and iris/periocular cross-spectral recognition and has co-authored several research publications recognised by the peers with citations. Marta Gomez-Barrero is a Professor for IT-Security and technical data privacy at the Hochschule Ansbach, in Germany. Between 2016 and 2020, she was a postdoctoral researcher at the National Research Center for Applied Cybersecurity (ATHENE)—Hochschule Darmstadt, Germany. Before that, she received her MSc degree in Computer Science and Mathematics (2011), and her PhD degree in Electrical Engineering (2016), all from Universidad Autonoma de Madrid, Spain. Her current research focusses on security and privacy evaluations of biometric systems, Presentation Attack Detection (PAD) methodologies, and biometric template protection (BTP) schemes. She has co-authored more than 70 publications, chaired special sessions and competitions at international conferences; she is associate editor for the EURASIP Journal on Information Security and represents the German Institute for Standardisation (DIN) in ISO/IEC SC37 JTC1 SC37 on biometrics. Naser Damer is a senior researcher at the Fraunhofer IGD, performing research management, applied research, scientific consulting, and system evaluation. He received his Ph.D. in computer science from TU Darmstadt (2018). His main research interests lie in the fields of biometrics, machine learning, and information fusion. Naser is a research area co-coordinator and a principal investigator at the National Research Center for Applied Cybersecurity ATHENE, Germany. He lectures on Human and Identity-centric Machine Learning, as well as on Ambient Intelligence at TU Darmstadt. Naser is a member of the organising teams of several conferences, workshops, and special sessions, including being a program co-chair of BIOSIG. He serves as an associate editor for Pattern Recognition (Elsevier) and the Visual Computer (Springer). He represents the German Institute for Standardization (DIN) in the ISO/IEC SC37 international biometrics standardization committee. He is a member of the IEEE Biometrics Council serving on its Technical Activities Committee. Paulo Lobato Correia is Associate Professor at the Department of Electrical and Computer Engineering, Instituto Superior Técnico, Universidade de Lisboa, Portugal. He is a Senior Researcher of the Multimedia Signal Processing research group of Instituto de Telecomunicações. He is Senior Member of the IEEE. Paulo Correia coordinated the participation in several national and international research projects, dealing with image and video analysis and processing. He is Editor in Chief of IET Biometrics for the term 2020–2022. He was Subject Editor (for Multimedia papers) of the Elsevier Signal Processing Journal (2018–2020). He was Associate Editor of the IEEE Transactions on Circuits and Systems for Video Technology (2006–2014), of the Elsevier Signal Processing Journal (2005–2017), and of IET Biometrics (2013–2019). He has been Guest Editor of several special issues for scientific journals and cooperated in many conference organising committees. He is a founding member of the Advisory Board of the European Signal Processing Association (EURASIP) and was the elected chairman of EURASIP's Technical Area Committee on “Biometrics, Data Forensics and Security” for the term 2018–2020. He has co-authored more than 140 journal and conference papers. The main research interests are about video analysis and processing, with emphasis on biometrical signal analysis, targeting recognition, sports, medical and forensic applications.
AI-based face recognition has become increasingly appealing for daily life applications due to its high performance. At the same time, image coding is very commonly used and, thus, many applications perform face recognition with decoded images. However, using decoded images, which may suffer from compression artifacts, may impact the final decision-making process of AI-based face recognition systems and, thus, its overall recognition performance. This paper studies the impact of image coding on the overall face verification performance of a popular and high performing face recognition solution, ArcFace. Face recognition using both original and decoded images, with several compression rates and qualities, is considered. Tests were performed using the Labeled Faces in the Wild (LFW) face dataset, with its images coded using conventional image coding standards, notably JPEG, JPEG 2000, and JPEG XL, as well as three emerging AI-based image codecs. As expected, the experimental results show that coding can have a significant impact on face recognition performance, with its impact becoming increasingly relevant as the coding rate is reduced. It is also observed that the recent AI-based image codecs appear to offer slightly better recognition performance for the same coding rates as a consequence of their better RD performance.
Human motion analysis provides useful information for the diagnosis and recovery assessment of people suffering from pathologies, such as those affecting the way of walking, i.e., gait. With recent developments in deep learning, state-of-the-art performance can now be achieved using a single 2D-RGB-camera-based gait analysis system, offering an objective assessment of gait-related pathologies. Such systems provide a valuable complement/alternative to the current standard practice of subjective assessment. Most 2D-RGB-camera-based gait analysis approaches rely on compact gait representations, such as the gait energy image, which summarize the characteristics of a walking sequence into one single image. However, such compact representations do not fully capture the temporal information and dependencies between successive gait movements. This limitation is addressed by proposing a spatiotemporal deep learning approach that uses a selection of key frames to represent a gait cycle. Convolutional and recurrent deep neural networks were combined, processing each gait cycle as a collection of silhouette key frames, allowing the system to learn temporal patterns among the spatial features extracted at individual time instants. Trained with gait sequences from the GAIT-IT dataset, the proposed system is able to improve gait pathology classification accuracy, outperforming state-of-the-art solutions and achieving improved generalization on cross-dataset tests.
Long Short-Term Memory (LSTM) is a prominent recurrent neural network for extracting dependencies from sequential data such as time-series and multi-view data, having achieved impressive results for different visual recognition tasks. A conventional LSTM network, hereafter referred only as LSTM network, can learn a model to posteriorly extract information from one input sequence. However, if two or more dependent sequences of data are simultaneously acquired, the LSTM networks may only process those sequences consecutively, not taking benefit of the information carried out by their mutual dependencies. In this context, this paper proposes two novel LSTM cell architectures that are able to jointly learn from multiple sequences simultaneously acquired, targeting to create richer and more effective models for recognition tasks. The efficacy of the novel LSTM cell architectures is assessed by integrating them into deep learning-based methods for face recognition with multi-view, light field images. The new cell architectures jointly learn the scene horizontal and vertical parallaxes available in a light field image, to capture richer spatio-angular information from both directions. A comprehensive evaluation, with the IST-EURECOM LFFD dataset using three challenging evaluation protocols, shows the advantage of using the novel LSTM cell architectures for face recognition over the state-of-the-art light field-based methods. These results highlight the added value of the novel cell architectures when learning from correlated input sequences.
Several pathologies can alter the way people walk, i.e., their gait. Gait analysis can be used to detect such alterations and, therefore, help diagnose certain pathologies or assess people's health and recovery. Simple vision-based systems have a considerable potential in this area, as they allow the capture of gait in unconstrained environments, such as at home or in a clinic, while the required computations can be done remotely. State-of-the-art vision-based systems for gait analysis use deep learning strategies, thus requiring a large amount of data for training. However, to the best of our knowledge, the largest publicly available pathological gait dataset contains only 10 subjects, simulating five types of gait. This paper presents a new dataset, GAIT-IT, captured from 21 subjects simulating five types of gait, at two severity levels. The dataset is recorded in a professional studio, making the sequences free of background camouflage, variations in illumination and other visual artifacts. The dataset is used to train a novel automatic gait analysis system. Compared to the state-of-the-art, the proposed system achieves a drastic reduction in the number of trainable parameters, memory requirements and execution times, while the classification accuracy is on par with the state-of-the-art. Recognizing the importance of remote healthcare, the proposed automatic gait analysis system is integrated with a prototype web application. This prototype is presently hosted in a private network, and after further tests and development it will allow people to upload a video of them walking and execute a web service that classifies their gait. The web application has a user-friendly interface usable by healthcare professionals or by laypersons. The application also makes an association between the identified type of gait and potential gait pathologies that exhibit the identified characteristics.
We present a novel LSTM cell architecture capable of learning both intra- and inter-perspective relationships available in visual sequences captured from multiple perspectives. Our architecture adopts a novel recurrent joint learning strategy that uses additional gates and memories at the cell level. We demonstrate that by using the proposed cell to create a network, more effective and richer visual representations are learned for recognition tasks. We validate the performance of our proposed architecture in the context of two multi-perspective visual recognition tasks namely lip reading and face recognition. Three relevant datasets are considered and the results are compared against fusion strategies, other existing multi-input LSTM architectures, and alternative recognition solutions. The experiments show the superior performance of our solution over the considered benchmarks, both in terms of recognition accuracy and complexity. We make our code publicly available at https://github.com/arsm/MPLSTM.
Light field (LF) cameras provide rich spatio-angular visual representations by sensing the visual scene from multiple perspectives and have recently emerged as a promising technology to boost the performance of human-machine systems such as biometrics and affective computing. Despite the significant success of LF representation for constrained facial image analysis, this technology has never been used for face and expression recognition in the wild. In this context, this paper proposes a new deep face and expression recognition solution, called CapsField, based on a convolutional neural network and an additional capsule network that utilizes dynamic routing to learn hierarchical relations between capsules. CapsField extracts the spatial features from facial images and learns the angular part-whole relations for a selected set of 2D sub-aperture images rendered from each LF image. To analyze the performance of the proposed solution in the wild, the first in the wild LF face dataset, along with a new complementary constrained face dataset captured from the same subjects recorded earlier have been captured and are made available. A subset of the in the wild dataset contains facial images with different expressions, annotated for usage in the context of face expression recognition tests. An extensive performance assessment study using the new datasets has been conducted for the proposed and relevant prior solutions, showing that the CapsField proposed solution achieves superior performance for both face and expression recognition tasks when compared to the state-of-the-art.
This is the first editorial written by the new Editors-in-Chief (EiC) of IET Biometrics. The previous role of Prof. Michael Fairhurst is now shared between Paulo Lobato Correia (IST-UL, Instituto de Telecomunicacoes / Instituto Superior Tecnico, Universidade de Lisboa, Portugal) and Xudong Jiang (NTU, Singapore). The recent Editorial by the Honorary Editor-in-Chief, Prof. James Wayman, was devoted to thanking Prof. Fairhurst for his vision and contribution as founder and EiC of IET Biometrics for the past eight years. We want to thank him again for all the work and his constant enthusiasm for this endeavour. We are committed to work closely with a team of outstanding Associate Editors (AEs), who are key elements to help handling submissions. Together, we aim for a faster article review process, while keeping and enhancing constructive and quality peer-review and unbiased editorial decisions. The journal continues to rely on its most valuable assets, the Authors and the Readers. IET Biometrics keeps looking for scientifically sound original research articles, advancing the state-of-the-art, and well-organized review papers on relevant topics, providing a good starting point and guidance for whoever is studying a specific topic, by providing a structured description of the available approaches, along with a critical discussion and highlighting directions for the future. The scope of IET Biometrics has always been relatively wide, to cover not only core technological issues but also those that cross traditional disciplinary boundaries. Besides topics directly related to ‘the automated recognition of individuals based on their behavioural and biological characteristics’, the scope of IET Biometrics also includes ‘any topics where it can be shown that a paper can increase our understanding of biometric systems, signal future developments and applications for biometrics, or promote greater practical uptake for relevant technologies’. To make this even clearer, in 2020 IET Biometrics started accepting other paper types, notably case study papers, reporting on specific instances of an interesting or particular phenomenon or system related to research areas covered by the journal. You are invited to read the first case study paper by Galdi et al., published in the current issue of the journal: ‘PROTECT: Pervasive and useR fOcused biomeTrics bordEr projeCT - a case study’, presenting the multibiometric verification systems developed within the EU project PROTECT. More changes are ahead, as from 2021 IET journals will adopt the Open Access model. This means that published papers will be freely accessible online. And the content published since 2013 will also be made free-to-view online, contributing to an increased visibility of IET Biometrics papers. To conclude this brief editorial, and on behalf of the editorial team, we would like to invite the talented researchers within our community to continue to support IET Biometrics by submitting solid research for publication and playing an active role in the review process. This will ensure that IET Biometrics keeps its identity in this time of many changes and challenges.
In a world where security issues have been gaining growing importance, face recognition systems have attracted increasing attention in multiple application areas, ranging from forensics and surveillance to commerce and entertainment. To help understanding the landscape and abstraction levels relevant for face recognition systems, face recognition taxonomies allow a deeper dissection and comparison of the existing solutions. This paper proposes a new, more encompassing and richer multi-level face recognition taxonomy, facilitating the organization and categorization of available and emerging face recognition solutions; this taxonomy may also guide researchers in the development of more efficient face recognition solutions. The proposed multi-level taxonomy considers levels related to the face structure, feature support and feature extraction approach. Following the proposed taxonomy, a comprehensive survey of representative face recognition solutions is presented. The paper concludes with a discussion on current algorithmic and application related challenges which may define future research directions for face recognition.
Light field cameras are able to capture the intensity of light rays coming from multiple directions, thus representing the visual scene from multiple viewpoints. This paper exploits the rich spatio-angular information available in light field images for facial emotion recognition. In this context, a new deep network is proposed that first extracts spatial features using a VGG16 convolutional neural network. Then, a Bidirectional Long Short-Term Memory (Bi-LSTM) recurrent neural network is used to learn spatio-angular features from viewpoint feature sequences, exploring both forward and backward angular relationships. Additionally, an attention mechanism allows our model to selectively focus on the most important spatio-angular features, thus enabling a more effective learning outcome. Finally, a fusion scheme is adopted to obtain the emotion recognition classification results. Comprehensive experiments have been conducted on the IST-EURECOM Light Field Face database using two challenging evaluation protocols, showing the superiority of our method over the state-of-the-art.
Face recognition has attracted increasing attention due to its wide range of applications, but it is still challenging when facing large variations in the biometric data characteristics. Lenslet light field cameras have recently come into prominence to capture rich spatio-angular information, thus offering new possibilities for advanced biometric recognition systems. This paper proposes a double-deep spatio-angular learning framework for light field-based face recognition, which is able to model both the intra-view/spatial and inter-view/angular information using two deep networks in sequence. This is a novel recognition framework that has never been proposed in the literature for face recognition or any other visual recognition task. The proposed double-deep learning framework includes a long short-term memory (LSTM) recurrent network, whose inputs are VGG-Face descriptions, computed using a VGG-16 convolutional neural network (CNN). The VGG-Face spatial descriptions are extracted from a selected set of 2D sub-aperture (SA) images rendered from the light field image, corresponding to different observation angles. A sequence of the VGG-Face spatial descriptions is then analyzed by the LSTM network. A comprehensive set of experiments has been conducted using the IST-EURECOM light field face database, addressing varied and challenging recognition tasks. The results show that the proposed framework achieves superior face recognition performance when compared to the state of the art.