Ordinal scores occur commonly in medical imaging studies and more recently in black-box studies on forensic identification accuracy. To assess the accuracy of radiologists in medical imaging studies or the accuracy of forensic examiners in biometric studies, one needs to estimate the accuracy measures such as the receiver operating characteristic (ROC) curves and also account for the covariates related to the radiologists or forensic examiners. The novelty of the paper is twofold. First, we propose a new covariate-adjusted homogeneity test for ordinal ROC curves to determine differences in accuracy among multiple rater groups. Second, since the covariance structure among the ROC regression estimators is not available, we obtained the asymptotic covariance matrix of the ROC estimators and derived theoretical results of the proposed test. We conducted extensive simulation studies to evaluate the finite sample performance of the proposed test. The simulation results show that estimated ROC curves are consistent and the empirical coverage of the confidence intervals is close to the nominal level. Our proposed test is applied to a large-scale face recognition study in which participants include facial examiners, facial reviewers, super-recognizers, fingerprint examiners, and students. The results show differences in accuracy among five rater groups. Ad-hoc pairwise comparison tests are then conducted by establishing confidence bands of differences among ROC curves. Those pairwise tests identify statistically significant differences in ROC curves among five participant groups.
Deep learning models trained for facial recognition now surpass the highest performing human participants. Recent evidence suggests that they also model some qualitative aspects of face processing in humans. This review compares the current understanding of deep learning models with psychological models of the face processing system. Psychological models consist of two components that operate on the information encoded when people perceive a face, which we refer to here as 'face codes'. The first component, the core system, extracts face codes from retinal input that encode invariant and changeable properties. The second component, the extended system, links face codes to personal information about a person and their social context. Studies of face codes in existing deep learning models reveal some surprising results. For example, face codes in networks designed for identity recognition also encode expression information, which contrasts with psychological models that separate invariant and changeable properties. Deep learning can also be used to implement candidate models of the face processing system, for example to compare alternative cognitive architectures and codes that might support interchange between core and extended face processing systems. We conclude by summarizing seven key lessons from this research and outlining three open questions for future study.
Human review of consequential decisions by face recognition algorithms creates a collaborative human-machine system. We establish the circumstances under which combining human and machine face identification decisions improves accuracy. Using data from expert and non-expert face identifiers, we show that the benefits of human-human and human-machine collaborations increase as the difference in baseline accuracy between collaborators decreases. This rule holds across a wide range of baseline abilities, from novices to professional forensic face examiners. An important consequence of the rule is that people who are substantially less accurate than the machine, can actually improve decision accuracy when they collaborate with the machine. In a group of individual people collaborating with a machine, "intelligent human-machine fusion" was implemented by selecting people with the potential to increase collaborative accuracy. Performance with intelligent human-machine fusion was more accurate than either the machine operating alone or fusing all humans with the machine. Eliminating the machine from consideration yielded less predictable results, with average performance at or below intelligent human-machine collaboration. However, intelligent human-machine fusion was consistently more effective at minimizing the impact of low-performing humans on accuracy. The results demonstrate a meaningful role for both humans and machines in assuring accurate face identification.
Humans and machines vary in the accuracy with which they recognize faces of different races. This can impact the fairness of face identification in security and forensic settings. We introduce a protocol for designing a cross-race face identification test for evaluating people (e.g., forensic facial examiners, super-recognizers) and machines with superior face-identification ability. We followed this protocol to create a cross-race test and report the test's benchmarks on untrained human participants and two state-of-the-art face recognition algorithms. The goal of the protocol is to select a relatively small number of challenging test items (facial image comparisons) of two races, with approximately equally challenging items of both races. Item selection consisted of pre-screening with an open-source face recognition algorithm, followed by a second round of prescreening using the performance of untrained human participants. We sampled face-images (Black and White identities) from a large biometric data set and applied the protocol to assemble face comparisons. The protocol yielded a cross-race test with 20 comparison pairs portraying Black and White identities (10 same-identity; 10 different-identity). Untrained participants (54 Black; 51 White) judged whether face-image pairs showed the same or different identities using a 7-point scale. By design, the test proved challenging for untrained participants, with performance comparable across Black and White image pairs for both Black and White participants. Two top-performing face recognition systems from the Face Recognition Vendor Test-ongoing [6] scored perfectly (no errors) on both Black and White face-image pairs from the Cross-Race Test. The human and machine benchmarks established here make this test ideal for evaluating cross-race face recognition bias in people with high levels of skill and training.
Deep learning networks trained for facial recognition achieve accuracy equivalent to the highest-performing human participants. Recent evidence shows that they may also model some aspects of face processing in humans. This review compares the current understanding of deep learning networks with psychological models of the face processing system. Psychological models consist of two components: (i) a core system that extracts ‘face codes’ from retinal input encoding invariant and changeable properties of faces; (ii) an extended system linking these codes to personal information about a person and their social context. Studies of face codes in existing deep learning models have revealed some surprising results. For example, face codes in networks designed for identity recognition also encode expression information, which contrasts with psychological models that separate invariant and changeable properties. Deep learning can also be used to implement candidate models of the face processing system, such as comparing alternative cognitive architectures and information codes that might support interchange between core and extended face processing systems. After reviewing the current state of this emerging research topic, we summarise the implications for understanding the face processing system in seven lessons, and pose three questions for future research to address.
Forensic facial professionals have been shown in previous studies to identify people from frontal face images more accurately than untrained participants when given 30 s per face pair. We tested whether this superiority holds in more challenging conditions. Two groups of forensic facial professionals (examiners, reviewers) and untrained participants were tested in three lab-based tasks: other-race face identification, disguised face identification, and face memory. For other-race face identification, on same-race faces, examiners were superior to controls; on different-race identification, examiners and controls performed comparably. Examiners were superior to controls for impersonation disguise, but not consistently superior for evasion disguise. Examiners' performance on the Cambridge Face Memory Test (CFMT+) was marginally better than reviewers and controls. We conclude that under laboratory-style conditions, professional examiners' identification superiority does not generalize completely to other-race and disguised faces. Future work should administer other-race and disguise face identification tests that allow forensic professionals to follow methods and procedures they typically use in casework.
Face recognition algorithms perform more accurately than humans in some cases, though humans and machines both show race-based accuracy differences. As algorithms continue to improve, it is important to continually assess their race bias relative to humans. We constructed a challenging test of 'cross-race' face verification and used it to compare humans and two state-of-the-art face recognition systems. Pairs of same- and different-identity faces of White and Black individuals were selected to be difficult for humans and an open-source implementation of the ArcFace face recognition algorithm from 2019 (5). Human participants (54 Black; 51 White) judged whether face pairs showed the same identity or different identities on a 7-point Likert-type scale. Two top-performing face recognition systems from the Face Recognition Vendor Test-ongoing performed the same test (7). By design, the test proved challenging for humans as a group, who performed above chance, but far less than perfect. Both state-of-the-art face recognition systems scored perfectly (no errors), consequently with equal accuracy for both races. We conclude that state-of-the-art systems for identity verification between two frontal face images of Black and White individuals can surpass the general population. Whether this result generalizes to challenging in-the-wild images is a pressing concern for deploying face recognition systems in unconstrained environments.
Measures of face-identification proficiency are essential to ensure accurate and consistent performance by professional forensic face examiners and others who perform face-identification tasks in applied scenarios. Current proficiency tests rely on static sets of stimulus items and so cannot be administered validly to the same individual multiple times. To create a proficiency test, a large number of items of "known" difficulty must be assembled. Multiple tests of equal difficulty can be constructed then using subsets of items. We introduce the Triad Identity Matching (TIM) test and evaluate it using item response theory (IRT). Participants view face-image "triads" (N = 225) (two images of one identity, one image of a different identity) and select the different identity. In Experiment 3, university students (N = 197) showed wide-ranging accuracy on the TIM test, and IRT modeling demonstrated that the TIM items span various difficulty levels. In Experiment 3, we used IRT-based item metrics to partition the test into subsets of specific difficulties. Simulations showed that subsets of the TIM items yielded reliable estimates of subject ability. In Experiments 3a and b, we found that the student-derived IRT model reliably evaluated the ability of non-student participants and that ability generalized across different test sessions. In Experiment 3c, we show that TIM test performance correlates with other common face-recognition tests. In summary, the TIM test provides a starting point for developing a framework that is flexible and calibrated to measure proficiency across various ability levels (e.g., professionals or populations with face-processing deficits).
Ordinal scores occur commonly in medical imaging studies and in black-box forensic studies \citep{Phillips:2018}. To assess the accuracy of raters in the studies, one needs to estimate the receiver operating characteristic (ROC) curve while accounting for covariates of raters. In this paper, we propose a covariate-adjusted homogeneity test to determine differences in accuracy among multiple rater groups. We derived the theoretical results of the proposed test and conducted extensive simulation studies to evaluate the finite sample performance of the proposed test. Our proposed test is applied to a face recognition study to identify statistically significant differences among five participant groups.
Forensic face examiners outperform untrained participants in face identity matching (Phillips et al., 2018), though it is unclear whether this superiority generalizes to other-race faces. We developed a challenging test that can be performed with the limited time available to professional examiners. To select the most difficult image pairs from a set of Black (n= 3,102) and White (n= 122,728) faces (self-identified race when images were collected), we employed a deep convolutional neural network (DCNN) (Deng et al., 2019) and an experiment with untrained participants. Image pairs (n= 36 per race) were assembled using a DCNN “perceptual” similarity measure. Same-identity (different-identity) image pairs with the lowest (highest) similarity scores were selected from all possible pairs. Untrained participants (White: n= 26, Black: n= 11) judged whether the images showed the same identity or different identities. Ranking by perceptual difficulty, we created a set of 10 Black and 10 White face pairs (half same-identity pairs). “Difficulty” was measured by tallying the number of participants who incorrectly indicated same-identity pair as different identities, and vice versa. Next, we benchmarked the test by computing participants’ accuracy (area under the ROC curve) on the subset of pairs. The test proved challenging for untrained participants [Black participants: (faces: Black= 0.66, White= 0.53); White participants: (faces: Black= 0.56, White= 0.49)]. Participant race, face race, and the interaction did not affect accuracy (p > 0.05). Notably, additional DCNNs performed more accurately on the White face pairs than Black face pairs (Szegedy et al., 2017: Black= 0.72 , White= 1.0; Ranjan et al., 2017: Black= 0.5, White= 0.92). Given that the pattern of performance across race differed for humans and the DCNNs, we conclude that untrained human benchmarks are critical in building a challenging and balanced cross-race test for experts.
We evaluated the detailed, behavioral properties of face matching performance in two specialist groups: forensic facial examiners and super-recognizers. Both groups compare faces to determine identity with high accuracy and outperform the general population. Typically, facial examiners are highly trained; super-recognizers rely on natural ability. We found distinct behaviors between these two groups. Facial examiners took advantage of the full 7-point identity judgment scale; super-recognizers’ judgments clustered toward highly confident decisions. Facial examiners’ identity judgments for same-identities and different-identities mirrored each other; those from super-recognizers did not. Facial examiners showed higher identity judgment agreement than super-recognizers. Despite these qualitative differences, both groups showed insight into their own accuracy: more confident people and those who rated the task to be easier tended to be more accurate. These findings show that to better understand and interpret judgments according to the nature of someone’s facial expertise, evaluations should assess more than accuracy.
Automatic recognition of human faces is a significant problem in the development and application of pattern recognition. In this paper, we introduce a simple technique for identification of human faces in cluttered scenes based on neural nets. In detection phase, neural nets are used to test whether a window of 18x27 pixels contains a face or not. A major difficulty in learning process comes from the large database required for face / nonface images. We solve this problem by dividing these data into two groups. Such division results in reduction of computational complexity and thus decreasing the time and memory needed during the test of an image. The proposed face recognition technique consists of three parts; preprocessing, feature extraction, and recognition steps. Gradient Vector method is used for facial feature extraction. A face recognition system based on recent method which concerned with both representation and recognition using artificial neural networks is presented. It then evaluates the performance of the system by applying two photometric normalization techniques: histogram equalization and homomorphic filtering. The system produces promising results for face verification and face recognition.
Recent years have seen considerable advances in biometric recognition techniques leading to a wide-spread deployment of biometric technology across a number of application domains, ranging from security, border control, and criminal investigations to entertainment, social media, autonomous driving and even health services. To highlight some of these advancements and present the latest research ach...
Face recognition networks generally demonstrate bias with respect to sensitive attributes like gender, skintone etc. For gender and skintone, we observe that the regions of the face that a network attends to vary by the category of an attribute. This might contribute to bias. Building on this intuition, we propose a novel distillation-based approach called Distill and De-bias (D&D) to enforce a network to attend to similar face regions, irrespective of the attribute category. In D&D, we train a teacher network on images from one category of an attribute; e.g. light skintone. Then distilling information from the teacher, we train a student network on images of the remaining category; e.g., dark skintone. A feature-level distillation loss constrains the student network to generate teacher-like representations. This allows the student network to attend to similar face regions for all attribute categories and enables it to reduce bias. We also propose a second distillation step on top of D&D, called D&D++. Here, we distill the `un-biasedness' of the D&D network into a new student network, the D&D++ network, while training this new network on all attribute categories; e.g., both light and dark skintones. This helps us train a network that is less biased for an attribute, while obtaining higher face verification performance than D&D. We show that D&D++ outperforms existing baselines in reducing gender and skintone bias on the IJB-C dataset, while obtaining higher face verification performance than existing adversarial de-biasing methods. We evaluate the effectiveness of our proposed methods on two state-of-the-art face recognition networks: ArcFace and Crystalface.
Previous generations of face recognition algorithms differ in accuracy for images of different races (race bias). Here, we present the possible underlying factors (data-driven and scenario modeling) and methodological considerations for assessing race bias in algorithms. We discuss data-driven factors (e.g., image quality, image population statistics, and algorithm architecture), and scenario modeling factors that consider the role of the “user” of the algorithm (e.g., threshold decisions and demographic constraints). To illustrate how these issues apply, we present data from four face recognition algorithms (a previous-generation algorithm and three deep convolutional neural networks, DCNNs) for East Asian and Caucasian faces. First, dataset difficulty affected both overall recognition accuracy and race bias, such that race bias increased with item difficulty. Second, for all four algorithms, the degree of bias varied depending on the identification decision threshold. To achieve equal false accept rates (FARs), East Asian faces required higher identification thresholds than Caucasian faces, for all algorithms. Third, demographic constraints on the formulation of the distributions used in the test, impacted estimates of algorithm accuracy. We conclude that race bias needs to be measured for individual applications and we provide a checklist for measuring this bias in face recognition algorithms.
Accurate estimates of face-identification ability are crucial in applied forensic settings. Current face-identification datasets are often large and uncalibrated, making them sub-optimal for pre- and post-training evaluations. To optimize efficient and accurate performance assessments, small sets of well-labelled test items are needed. However, item-wise measures cannot be applied to the common forensic task of identity matching in image pairs, because items are either “matched” or “non-matched” identities. Therefore, in this case, an item response confounds item accuracy and response bias. Here, our goal was to construct flexible, well-calibrated subsets of face-identification items using Item Response Theory (IRT) applied to image triads. These triads were composed of two images of one identity and one image of a different identity; the task was to select the “different” identity. Participants (n=77) were tested on the full item pool of 224 face-image triads. Responses were analyzed using the IRT one-parameter model (Rasch model; Rasch, 1960). This approach provides measurements of subject ability and item difficulty on the same scale. Results of the model demonstrate the probability of endorsing a correct response given an item’s difficulty and a subject’s ability. Using these results, we constructed subsets of items that varied in item difficulty. To test the quality of these item subsets, we used responses to these subsets to estimate participants’ ability and predict accuracy for larger subsets of novel items. Leave-one-out cross validation results showed that we can predict both people’s accuracy on novel items and their individual responses. These calibrated face-identification tests can be used to develop face-identification tests with better flexibility, reliability, and time-efficiency.
Previous generations of face recognition algorithms show differences in accuracy for faces of different races (race bias) (O’Toole et al., 1991; Furl et al., 2002; Givens et al., 2004; Phillips et al., 2011; Klare et al., 2012). Whether newer deep convolutional neural networks (DCNNs) are also race biased is less well studied (El Khiyari et al., 2016; Krishnapriya et al., 2019). Here we present methodological considerations for measuring underlying race bias. We consider two key factors: data-driven and scenario modeling. Data-driven factors are driven by the data itself (e.g., the architecture of the algorithm, image quality, image population statistics). Scenario modeling considers the role of the “user” of the algorithm (e.g., threshold decisions and demographic constraints). To illustrate these issues in practice, we tested four face recognition algorithms: one pre-DCNN (A2011; Phillips et al., 2011) and three DCNNs (A2015; Parkhi et al., 2015), (A2017b; Ranjan et al., 2017), (A2019; Ranjan et al., 2019) on East Asian and Caucasian faces. First, for all four algorithms, the degree of race bias varied as a function of the identification decision threshold. Second, for all algorithms, to achieve equal false accept rates (FARs), Asian faces required higher identification thresholds than Caucasian faces. Third, dataset difficulty affected both overall recognition accuracy and race bias. Fourth, demographic constraints on the formulation of the distributions used in the test, impacted estimates of algorithm accuracy. We conclude with a recommended checklist for measuring race bias in face recognition algorithms.
Traditionally, researchers in automatic face recognition and biometric technologies have focused on developing accurate algorithms. With this technology being integrated into operational systems, engineers and scientists are being asked, do these systems meet societal norms? The origin of this line of inquiry is `trust' of artificial intelligence (AI) systems. In this paper, we concentrate on adapting explainable AI to face recognition and biometrics, and we present four principles of explainable AI to face recognition and biometrics. The principles are illustrated by $\it{four}$ case studies, which show the challenges and issues in developing algorithms that can produce explanations.
We introduce four principles for explainable artificial intelligence (AI) that comprise fundamental properties for explainable AI systems. We propose that explainable AI systems deliver accompanying evidence or reasons for outcomes and processes; provide explanations that are understandable to individual users; provide explanations that correctly reflect the system s process for generating the output; and that a system only operates under conditions for which it was designed and when it reaches sufficient confidence in its output. We have termed these four principles as explanation, meaningful, explanation accuracy, and knowledge limits, respectively. Through significant stakeholder engagement, these four principles were developed to encompass the multidisciplinary nature of explainable AI, including the fields of computer science, engineering, and psychology. Because one-sizefits-all explanations do not exist, different users will require different types of explanations. We present five categories of explanation and summarize theories of explainable AI. We give an overview of the algorithms in the field that cover the major classes of explainable algorithms. As a baseline comparison, we assess how well explanations provided by people follow our four principles. This assessment provides insights to the challenges of designing explainable AI systems.