Reverse engineering (RE) of integrated circuits (ICs) is critical for design verification, security analysis, and intellectual property protection. It involves analyzing IC layouts from high-resolution scanning electron microscopy (SEM) images to identify standard cells from design layouts and reconstruct netlists. However, SEM images are often degraded by noise and random artifacts like edge distortion, occlusion, stitching errors, and layer misalignment, which compromise reliability. This paper presents a novel Vision-Language Model (VLM) approach, built on a gated fusion network, for incorporating robust data quality assurance into this pipeline. The proposed method assesses data integrity by projecting inputs into a shared embedding space, predicting both the standard cell identity and the presence of data corruption. This approach enhances the robustness of the design assurance pipeline, achieving over 99\% data assurance at its operational threshold. Additionally, it is scalable to other designs and technology nodes, providing a practical tool for operators to automatically and interactively localize corruptions, evaluate associated risks, and make informed decisions about allocating resources to improve data quality for trustworthy RE outcomes.
Deepfakes are synthetic media created by deep-generative methods to fake a person’s audio-visual representation. Growing sophistication of deepfake technology poses significant challenges for both machine learning (ML) algorithms and humans. Here we used real and deepfake static face images (Study 1) and dynamic videos (Study 2) to (i) investigate sources of misclassification errors in machines, (ii) identify psychological mechanisms underlying detection performance in humans, and (iii) compare humans and machines in their classification decision accuracy and confidence. Study 1 found that machines achieved excellent performance in classifying real and deepfake images, with good accuracy in feature classification. Humans, in contrast, experienced challenges in distinguishing between real and deepfake static images. Their classification accuracy was at chance level, and this underperformance relative to machines was accompanied by a truth bias and low confidence for the detection of deepfake images. Using dynamic video stimuli, Study 2 found that performance of machines was near chance level, with poor feature classification. Further, machines showed greater lie bias and reduced decision confidence relative to humans who outperformed machines in the detection of video deepfakes. Finally, Study 2 revealed that higher analytical thinking, lower positive affect, and greater internet skills were associated with better video deepfake detection in humans. Combined, the findings across these two studies advance understanding of factors contributing to deepfake detection in both machines and humans; and can inform intervention toward tackling the growing threat from deepfakes by identifying areas of particular benefit from human-AI collaboration to optimize the detection of deepfakes.
This paper presents a lightweight and efficient approach to AI-generated content detection using small autoregressive fine-tuned decoders (AFDs) for secure, on-device deployment. Motivated by resource-efficiency, syntactic awareness, and bias mitigation, our model employs small language models (SLMs) with autoregressive pre-training and loss fusion to accurately distinguish between human and AI-generated content while significantly reducing computational demands. The system achieved highest macro-F1 score of 0.8186, with the submitted model scoring 0.7874-both significantly outperforming the task baseline while reducing model parameters by +/- 60%. Notably, our approach mitigates biases, improving recall for human-authored text by over 60%. Ranking 8th out of 36 participants, these results confirm the feasibility and competitiveness of small AFDs in challenging, adversarial settings, making them ideal for privacy-preserving, on-device deployment suitable for real-world applications.
Editor's notes: Reverse engineering modern complex integrated circuits (ICs) relies on extracting as much helpful information as possible, which requires extensive imaging technology support. One of the key imaging methods, scanning electron microscopy (SEM), offers macroscale resolution values and a wide range of magnification. This tutorial shows how artificial intelligence (AI) techniques facilitate reverse engineering of an IC based on SEM. -Umit Ogras, University of Wisconsin, USA
Access control is already an integral part of the Internet of Things (IoT) to prevent unauthorized use of systems. However, in the post-pandemic world, contactless authentication is desired for multi-user in-person systems. Biometric key-based hardware obfuscation can enable user-specific access control for IoT devices with added protection against piracy, reverse engineering, and hardware tampering. Biometric key-based unlocking of devices relies on various feature generation and classification tasks, for which convolutional neural networks (CNN) have demonstrated state-of-the-art effectiveness. In this regard, CNN-based contactless biometric template submission (e.g., face) can be a potential candidate. However, for resource-constrained devices (i.e., IoT), especially those with limited hardware capabilities, might struggle to efficiently execute CNN-based models due to their intensive memory access patterns, operational delays, communication bandwidth, and power consumption. Exploiting the recent advancements in information flow theory in neural networks and binary features, in this paper, we propose a CNN-based biometric system where binary weights and activations are used to generate compact yet, meaningful binary biometric features to enable computation on resource-constrained edge devices. We implement the proposed framework and a conventional CNN on a face dataset in an FPGA as a proof-of-concept and obtain a classification accuracy of 96
Electronic counterfeiting is a long-lasting problem that continues to cost original manufacturers billions, fund organized crime, and jeopardize national security and mission-critical infrastructures. Manual inspection is a popular and standardized way to detect counterfeit electronic components, but it is time-consuming and requires subject matter experts for classification. State-of-the-art machine learning, deep learning, and computer vision-based physical inspection methods are promising to alleviate these issues. However, the main bottleneck for doing so is a lack of high-quality, publicly available counterfeit image data for training. Producing such datasets is also time-consuming and often requires expensive equipment. In addition, most test labs are not allowed to freely publish images taken from their customer’s chips. One solution to this data shortage bottleneck can be addressed by augmenting synthetic data. In this paper, (i) data multiplication using Progressive GAN, StyleGAN, and classical methods in counterfeit data domain is explored; (ii) a novel framework, named MaGNIFIES, is proposed; and (iii) an efficient Convolutional Neural Network architecture is proposed, which can detect defective parts by training only on the synthetic dataset generated using (i) or (ii). For proof of concept, we have used low-quality images of resistors and capacitors with and without scratch defects as counterfeit and golden components respectively. We have also illustrated how our approach using MaGNIFIES addresses the shortcomings of the existing augmentation methods. Separate data augmentation detection models are trained with each type of augmented data generated using MaGNIFIES, as well as existing techniques, and tested on a test set of real data.
Design-for-test/debug(DfT/D) introduces scan chain testing to increase testability and fault coverage by inserting scan flip-flops. However, these scan chains are also known to be a liability for security primitives. In previous research, dynamically obfuscated scan chains (DOSC) were introduced to protect logic-locking keys from scan-based attacks by obscuring test patterns and responses. In this paper, we present DOSCrack, an oracle-guided attack to de-obfuscate DOSC using symbolic execution and binary clustering, which significantly reduces the candidate seed space to a manageable quantity. Our symbolic execution engine employs scan mode simulation as well as satisfiability modulo theories (SMT) solvers to reduce the possible seed space, while obfuscation key clustering allows us to effectively rule out a group of seeds that share similarities. An integral component of our approach is the use of sequential equivalence checking (SEC), which aids in identifying distinct simulation patterns to differentiate between potential obfuscation keys. We experimentally applied our DOSCrack framework on four different sizes of DOSC benchmarks and compared their run-time and complexity. Our research effectively addresses critical vulnerabilities in scan-chain obfuscation methodologies, offering insights into DfT/D and logic locking for both academic research and industrial applications. Our framework emphasizes the need to craft robust and adaptable defense mechanisms against scan-based attacks.
Text-based automatic personality recognition (APR) operates at the intersection of artificial intelligence (AI) and psychology to determine the personality of an individual from their text sample. This covert form of personality assessment is key for a variety of online applications that contribute to individual convenience and well-being such as that of chatbots and personal assistants. Despite the availability of good quality data utilizing state-of-the-art AI methods, the reported performance of these recognition systems remains below expectations in comparable areas. Consequently, this work investigates and identifies the source of this performance limit and attributes it to the flawed assumptions of text-based APR. These insights are obtained via a large-scale comprehensive benchmark and analysis of text data from five corpora with diverse characteristics and complementary personality models (Big Five and Dark Triad) applied to an assortment of AI methods ranging from hand-crafted linguistic features to data-driven transformers. Finally, the work concludes by identifying the open problems that can help navigate the limitations in text-based automatic personality recognition to a great extent.
Owing to the sophisticated methods in Natural Language Processing and Computational Linguistics, personality assessment has experienced great advancement from being a specialized field in Psychology to the automatic extraction of personality traits from an individual's social media posts. While this advancement could be mainly attributed to the fundamental Lexical Hypothesis, its direct application for the development of Automatic Personality Recognition (APR) methods makes two broad assumptions regarding the perceptibility and ubiquity of personality characteristics in text. This paper presents the first inquiry into the extent to which current text-based personality datasets and modeling approaches align with the aforementioned assumptions. Supported by the first large-scale comparative assessment involving two personality models, the Big Five and Dark Triad, five text datasets, and nine personality modeling approaches, this paper offers insights on accommodating the assumptions of the Lexical Hypothesis for the development of robust APR approaches.
Advances in vision and deep learning have revolutionized feature extraction for face recognition and verification systems, yet, performing kinship verification from such features is still challenging. Ongoing research attempts to imitate a human by identifying features for kinship verification. In this paper, we propose KinfaceNet, a deep learning based kinship feature extractor, capable of extracting kinship features from a single input image independently without requiring its kin pair image. The base model of the method is adopted from face recognition domain which is then transfer learned in the domain of kinship by learning a distance mapping from face images to a compact Euclidean space where distances directly correspond to a measure of kinship similarity. Thus, unlike most of the works in deep learning based kinship domain, the extracted features can be used in many other applications such as image generation and family based clustering, etc. Training is performed by rearranging the data into classes of kin pairs and using a state-of-the-art triplet mining algorithm to address the unbalanced kinship data problem which causes overfitting. Also, one of the major advantages of our framework is that training can be performed on any face feature extractor model pre-trained on large face recognition data, thereby reducing training time by a considerable amount. Comparable verification accuracy is obtained from simple MLP network at only 20th epoch with KinfaceNet features extracted from the Family-In-the-Wild dataset, the largest in the wild kinship dataset available, as well as KinfaceW-I and II datasets.
Artificial intelligence (AI) and machine learning (ML) techniques have been increasingly used in several fields to improve performance and the level of automation. In recent years, this use has exponentially increased due to the advancement of high-performance computing and the ever increasing size of data. One of such fields is that of hardware design—specifically the design of digital and analog integrated circuits, where AI/ ML techniques have been extensively used to address ever-increasing design complexity, aggressive time to market, and the growing number of ubiquitous interconnected devices. However, the security concerns and issues related to integrated circuit design have been highly overlooked. In this article, we summarize the state-of-the-art in AI/ML for circuit design/optimization, security and engineering challenges, research in security-aware computer-aided design/electronic design automation, and future research directions and needs for using AI/ML for security-aware circuit design.
The globalization of printed circuit board (PCB) production has expanded the avenues to introduce vulnerabilities into the electronics supply chain. Malicious parties may modify or alter PCBs to purposely compromise the integrity of a design. In contrast, non-malicious incidents may also change an ideal design; they unintentionally lead to defects. For the reverse engineering process, which requires precise information to replicate a design, these deviations from the intended design present an overwhelming challenge. Although they are primarily unintentional, defects may cause a device to stray from its intended functionality, and present danger to both manufacturers and consumers. Therefore, irrespective of their cause and effect, the PCB inspection industry requires intricate knowledge of their occurrences; defects must be correctly enumerated before satisfactory solutions can be proposed. This article presents an extensive taxonomy of physical defects in PCBs that can be primarily identified using automated optical inspection (AOI), or traditional visual inspection methods where applicable. Our goal is to ultimately improve the state of hardware security by providing a guide to physical defects. To the best of our knowledge, this is the first comprehensive taxonomy of visual PCB defects.
Comprehensive hardware assurance and failure analysis methods require exhaustive validation of the design layout through post-silicon imaging. Overlooked segmentation errors in such images cause significant inflation of resource requirements, both in terms of human resources and process time frame, for fault isolation and debugging in the hardware assurance process. Further, due to their lack of contextual awareness, the existing segmentation measures lack the ability to detect and suppress segmentation errors. In this paper, we introduce the first reference-less context-aware segmentation quality evaluation metric for scanning electron microscopy images, called Structural Equivalency and Connection Uniformity for Reverse Engineering (SECURE), to alleviate this issue by capturing an implicit understanding of design layout from weakly labeled unpaired images. The proposed metric can be seamlessly integrated into the design layout validation workflow for in-line rejection or correction of corrupted segmented images preventing the corrupted data from compromising the reliability of the hardware assurance process while simultaneously supporting process automation. Exhaustive qualitative and quantitative validation is provided to support the efficacy of the metric.
Abstract---Reinforcement learning (RL) has become more popular due to promising results in applications such as chat-bots, healthcare, and autonomous driving. However, one significant challenge in current RL research is the difficulty in understanding which RL algorithms, if any, are practical for a given use case. Few RL algorithms are rigorously tested, and hence understood, for their practical implications. Although there are a number of performance comparisons in literature, many use few environments and do not consider real-world limitations such as run-time and memory usage. Furthermore, many works do not make their code publicly accessible for others to use. This paper addresses this gap by presenting the most comprehensive performance comparison on the practicality of RL algorithms known to date. Specifically, this paper focuses on discrete, model-free deep RL algorithms for their practicality in real-world problems where efficient implementations are necessary. In total, fourteen RL algorithms were trained on twenty-three environments (468 environment instances), which collectively required 224 GB and 766 days CPU time to run all experiments, and 1.7 GB to store all models. Overall, the results indicate several shortcomings in RL algorithms' exploration efficiency, memory/sample efficiency, and space/time complexity. Based on these shortcomings, numerous opportunities for future works were identified to improve the capabilities of modern algorithms. This paper’s findings will help researchers and practitioners improve and employ RL algorithms in time-sensitive and resource-constrained applications such as economics, cybersecurity, and Internet of Things (IoT). Impact Statement---Reinforcement learning (RL) technologies are commonly used in autonomous driving, chat-bot, and business analytic applications. They learn how to adapt to unforeseen situations, reducing the load on human drivers, support teams, and analysts. Although there are a variety of theoretical works in RL literature, very few algorithms are tested and evaluated to facilitate their use in real-life scenarios. The performance comparison introduced in this paper addresses these limitations. The performance analysis framework, re-implemented source code, and findings identified in this study could increase the adoption and speed development of RL technologies in more real-life applications. Moreover, the open challenges, recommendations, and practical implications identified in this paper could facilitate collaboration and development of new technologies among researchers and practitioners in industry and academia.
For successful printed circuit board (PCB) reverse engineering (RE), the resulting device must retain the physical characteristics and functionality of the original. Although the applications of RE are within the discretion of the executing party, establishing a viable, non-destructive framework for analysis is vital for any stakeholder in the PCB industry. A widely regarded approach in PCB RE uses non-destructive x-ray computed tomography (CT) to produce three-dimensional volumes with several slices of data corresponding to multi-layered PCBs. However, the noise sources specific to x-ray CT and variability from designers hampers the thorough acquisition of features necessary for successful RE. This article investigates a deep learning approach as a successor to the current state-of-the-art for detecting vias on PCB x-ray CT images; vias are a key building block of PCB designs. During RE, vias offer an understanding of the PCB’s electrical connections across multiple layers. Our method is an improvement on an earlier iteration which demonstrates significantly faster runtime with quality of results comparable to or better than the current state-of-the-art, unsupervised iterative Hough-based method. Compared with the Hough-based method, the current framework is 4.5 times faster for the discrete image scenario and 24.1 times faster for the volumetric image scenario. The upgrades to the prior deep learning version include faster feature-based detection for real-world usability and adaptive post-processing methods to improve the quality of detections.
Outsourced printed circuit board (PCB) fabrication necessitates increased hardware assurance capabilities. Several assurance techniques based on automated optical inspection (AOI) have been proposed that leverage PCB images acquired using digital cameras. We review state-of-the-art AOI techniques and observe a strong, rapid trend toward machine learning (ML) solutions. These require significant amounts of labeled ground truth data, which is lacking in the publicly available PCB data space. We contribute the FICS PCB Image Collection (FPIC) dataset to address this need. Additionally, we outline new hardware security methodologies enabled by our dataset.