We review and simplify several classical and quantum maps introduced in recent years that have been used in cryptography. For each of these maps, a bitstream is generated and subjected to the NIST test. Leveraging the advantages of FPGA in the loop, these maps are designed in MATLAB Simulink, converted to HDL using HDL Coder, and implemented on FPGA. Vivado software is used for more precise synthesis of the implementation of these maps. The results of a detailed analysis of classical and new quantum maps are compared with each other, as well as with other implementations of chaotic maps in the literature. Implementations related to five classical maps and two quantum maps, with maximum frequencies 125 MHz, and maximum throughputs of 4 Gbps, are confirmed. The suitability of these maps for implementation, leveraging their greater dynamic complexity and larger key space, is evident.
In this paper, a narrowband beamforming technique, widely used in acoustic signal processing, sonar systems, and wireless communications, was implemented on an FPGA (Field Programmable Gate Array) using a model-based design approach. The beamformer was specifically designed to estimate the direction of arrival of acoustic signals and operates in the frequency domain. The model-based design approach facilitated hardware/software (HW/SW) partitioning and system verification within Simulink. Using HDL Coder library, an FPGA Intellectual Property (IP) core was generated that meets all hardware requirements of the reference algorithm and is capable of communicating with both hard processor of a System on Chip (SoC) devices and a soft processor instantiated in the FPGA. This IP was connected with a microblaze processor to enable efficient hw/sw co-design on an Xilinx Artix-7 FPGA. The design was simulated using Xilinx Vivado’s XSIM and verified by comparing simulation results with those from the Simulink model. In addition to accuracy, performance of the design was evaluated in terms of execution speed, latency, resource utilization, power consumption, and compared against its processor-based counterpart. In conclusion, narrowband beamforming algorithm was successfully implemented and tested using a model-based design approach, providing a more streamlined alternative to conventional methods by abstracting low-level FPGA complexities and allowing designers to focus on higher-level system design.
Lattice-based cryptography (LBC) algorithms are considered suitable candidates for post-quantum cryptography (PQC), as they dominate the standardization process put forward by the National Institute of Standards and Technology (NIST). Indeed, three of the four key encapsulation mechanism (KEM) algorithms in the third round of the process are based on computationally hard lattice problems. On the other hand, there is an urgent need for processor designs that can run PQC algorithms efficiently, especially for embedded systems. This study presents an application-specific instruction set processor (ASIP) design for the Kyber, Saber, and NewHope algorithms based on transport triggered architecture (TTA). Custom hardware accelerators are added to the baseline processor architecture for computation-intensive steps without applying any software optimization to the reference code. We compared FPGA and ASIC implementations of our design with the prominent RISC-V cores and instruction set extension studies in the literature. According to the results, the proposed design offers greater efficiency, better performance, and lower resource utilization than its competitors in most cases.
The FPGA-based RISC-V system is a system that utilizes the open-source RISC-V processor architecture and is implemented on FPGA (Field Programmable Gate Arrays). This system provides flexibility and customization options, allowing for tailored solutions for different application domains. Securing the boot process of systems composed of widely used open-source and customizable structures is a critical requirement to maintain data integrity and ensure system security. This study presents a model for implementing secure boot processes using a hardware-based TPM 2.0 module on the FPGA-based RISC-V system, encompassing both software (operating system components) and hardware components. The components required for the RISC-V SoC (System on Chip) and the TPM 2.0 module to be used in secure boot have been implemented on the FPGA development board and functionally verified.
This paper presents a model-based design of AI accelerator following the Vitis TRD flow, implemented on the AMD Kria KV260 Vision AI Starter Kit. The ResNet-18 model, developed in PyTorch, was quantized, compiled with Vitis AI, and deployed to the FPGA via PYNQ. We deeply analyzed different DPU configurations and frequencies, focusing on resource utilization, power, FPS, and energy efficiency. Results show that resource utilization remains constant across frequencies, but lower frequencies increase energy consumption. For optimal performance and energy efficiency, high-MAC DPU configurations should be used at higher frequencies. To our knowledge, no prior research has fully detailed the Vitis TRD flow within Vitis AI. Most rely on Vivado TRD, which requires PetaLinux. This work offers a comprehensive guide for deploying AI models on FPGAs using Ubuntu, eliminating the need for PetaLinux expertise.
As the era of Post-Quantum Cryptography emerges, the demand for efficient and secure cryptographic algorithms has intensified. This conference paper navigates through the details of employing Number Theoretic Transform based modular multiplication algorithm which is K 2 RED tailored for post-quantum cryptographic applications, with a specific emphasis on their Field Programmable Gate Array implementation. NTT is an operation of Lattice Based Cryptography and it converts the numbers to finite field and modular multiplication operation is carried out. So, modular multiplication is the essential operation of NTT and its efficiency is critical for NTT operation. Our findings not only contribute to the growing body of knowledge in PQC but also offer practical guidance for engineers and researchers seeking to implement robust and efficient cryptographic solutions on FPGA platforms. The presented performance metrics serve as a benchmark for evaluating the feasibility and scalability of modular multiplication algorithms in the context of PQC systems. We present a detailed analysis of the area, speed, and latency performance of the implemented modular multiplication algorithms on FPGA platforms. Given the novelty offered by this study, it has demonstrated that the K 2 RED algorithm can be used with different bit lengths. Similarly, Plantard and K 2 RED algorithms have been utilized to create and report different circuit schematics according to various requirements.
Internet of things (IoT) gaining more importance due to its crucial role in pervasive computing and also Industry 4.0. Since the number of IoT devices is scaling up to multiple dozens of billions, the importance of energy efficiency is significantly increased. With the consideration of huge variety of IoT device hardware and software, a comprehensive model and estimation methodology on energy consumption is necessary as an enabler. IoT devices are also frequently updated, upgraded and maintained because of the evolving nature of the requirements and market demands. Each and every such operation has an effect on the power consumption and arose the necessity for a new energy consumption modeling and estimation. This process is applicable for development of IoT devices, as well as the maintenance phase. Since the variety of designs is unlimited, and battery capacity is usually fixed, or a cost factor, a generic, fully simulated, model-based energy consumption estimation of IoT devices is crucial. In this study, we aim to address this problem via proposing fully simulated, model-based, system-level power estimation approaches, as well as their success rate in typical real-life scenarios. It can be seen that the proposed methodology has high accuracy over %97. For the realization of the best-proposed approach, we used Open Virtual Platform (OVP) as an instruction set accurate simulator.
In this paper, we propose a different approach from the traditional design flow by suggesting the use of model-based design tools provided by Matlab Simulink. The first contribution of our paper is the detailed definition of the model-based design approach and the design process. Another contribution of our paper involves the thorough examination and validation of software and hardware integration processes on the Vivado and Vitis platforms for the designed visual cryptography model, following the automatic code generation using HDL Coder and Embedded Coder tools.
In this article, we propose an approach to create a high-quality quantum tent map by utilizing the generalized quantum dot system. Our objective is to determine if its chaos surpasses that of the traditional classic tent. To achieve this, we first introduce a quantum tent map, showcasing its chaotic behavior in relation to control parameters and initial conditions. We validate the presence of chaos by calculating the Lyapunov exponent and analyzing the time series. Next, we design a chaotic S-box based on this new map. Subsequently, we explore the possibility of obtaining strong S-boxes by employing two fractional stochastic models and three components of the proposed quantum tent map. Our primary question is whether the combination of fractional stochastic models and quantum tent maps can result in superior S-boxes. The answer to this question is a resounding “Yes.” The performance of these S-boxes exceeds that of previously proposed models. Finally, we evaluate a mixed S-box as the foundation for achieving highly secure image encryption. Among the models presented in this article, the best-performing S-box is produced through the combination of the fractional gamma distribution (FGD) with parameters α =1.06 and β =2.8, along with the X-quantum tent map dimension. This model achieves an SAC (Strict Avalanche Criterion) value of 0.5, surpassing even AES (Advanced Encryption Standard). Its non-linearity value is 106.625, indicating excellent performance. Additionally, the model has an LP (Linear Property) value of 0.128906 and a DP (Differential Properties) score of 12. Furthermore, we obtain BIC-SAC 0.503209 and BIC-Non-linearity 103.679. Lastly, we present an image encryption algorithm to demonstrate the effectiveness of the mixed S-box, evaluating its performance against various attacks. The results affirm the suitability of this encryption method, with the generated image encryption offering a key space of 2^16960 , ensuring high security. The example image exhibits an entropy value of 7.9311. Moreover, it demonstrates correlations in different orientations: horizontal (0.0076), vertical ( -0.00048659 ), and diagonal (0.00020717). The image also achieves NPCR (Number of Pixel Change Rate) at
This paper describes the design and implementation of a driver drowsiness detection (DDD) system using a modified RiscV processor on a field-programmable gate array (FPGA).To detect drowsiness, Convolutional Neural Network (CNN) is implemented on a RiscV processor.The CNN is trained to classify four primary driver's expressions, including distraction, natural, sleep, and yawn.The trained CNN accuracy is 81.07%on validation data.Furthermore, due to FPGA memory limitations, written C code for the trained CNN is optimized in numerous ways.Optimizations include the usage of dynamic fixed-point data types and dynamic memory allocations.On the other hand, the processor is modified by adding three custom instructions, including custom store, conv2d(2 × 2), and multiply and accumulation (MAC) to enhance the computation rate.As a result, the processor with custom store, conv2d(2 × 2), and MAC as custom instructions achieved the best result in terms of latency, with an improvement factor of 1.7 over the base processor and 1.25 over the processor with only custom store and multiply and accumulation (MAC) in exchange of slight increase in area.
In this study, it is aimed to implement the low-RISC system-on-chip, which is based on the Rocket processor created with the RISC-V instruction set architecture developed by Berkeley University, on FPGA and to run image processing algorithms on this system. While making this implementation, the main target is a system that is very simple, consumes low power, and can be quickly redirected to other purposes. Therefore, it is based on the effective evaluation of the existing system without using any extra customized accelerators. Thus, a free, open source, and powerful enough platform for many embedded system applications is proposed to the designers. For this purpose, a lane detection application designed with standard C libraries such as Gaussian blur filter, Sobel operation filter and other elements, which are widely used in image processing applications, is run with embedded Linux operating system and the results are shared.
The clutter encountered in Ground Penetrating Radar(GPR) systems is an important area of research since it decreases target detection rates. Real-time radar applications and hardware implementation of clutter removal methods in autonomous systems are crucial. In this study, robust non-negative matrix factorization (RNMF) is used, which requires simple mathematical operations and suitable for hardware implementations. FPGA was chosen as the hardware implementation environment due to its re-programmable feature. It has been shown that the hardware implementation results have the same performance as the clutter removal results obtained in the MATLAB environment.
This study shows that a previously published cross correlation based power analysis (CCPA) attack applied to the Montgomery Ladder exponentiation steps of a Rivest Shamir Adleman (RSA) implementation can be improved by working in frequency domain. It is shown that utilizing cross correlation values of discrete Fourier transform (DFT) coefficients instead of time samples, requires lesser power traces to retrieve the key bits of the target implementation. In addition, instead of using DFT coefficients corresponding to the whole measured frequency band, using a few DFT coefficients corresponding to lower bands, even under the first harmonic of the target clock is also an improving factor on the performance of the CCPA. Practical and theoretical results are also compared to both domains. To the best of our knowledge, this is the first study to show the frequency domain applicability and superiorities in terms of horizontal CCPA type attacks.
Channel coding techniques are used to reduce error rates during data transmission in mobile communication systems. LDPC (Low-density Parity Check) coding was first designed by Robert G. Gallager in 1962. It has been accepted and started to be used as a standard for coding data channels in fifth generation mobile communication by member companies of 3GPP (3rd Generation Partnership Project). The QC-LDPC method, which has a more hardware-friendly algorithmic structure, was chosen by 3GPP as the coding method. In this study, the QC-LDPC coding block, which is clearly stated in the 3GPP TS 38–212 standard document, is model-based and implemented on the FPGA.
RPL (IPv6 Routing Protocol for Low-Power and Lossy Networks) is a standardized routing protocol that can organize thousands of resource constraint routers. Although it is an indispensable protocol with its energy-efficient, scalable, and autonomous structure, it is vulnerable to numerous attacks with its sensitive data and mechanisms. In this paper, we intensely analyzed the standard statements of the authenticated security mode of RPL and designed a comprehensive authenticated key establishment scheme extending the BKE (Bilateral Key Exchange). We formally verified our scheme using the Scyther tool. This study will help researchers contribute more to the authenticated security mode of RPL.
Lattice-based structures offer numerous possibilities for post-quantum cryptography. Recently, many post-quantum cryptography algorithms have been built on hard lattice problems. The three of the remaining four algorithms in the final round of the NIST Standardization Process rely on lattice-based methods. However, suitable processor architectures for these algorithms have not been sufficiently investigated. This study examines the potential advantages of transport triggered architecture for these algorithms. We compare popular 64-bit RISC-V processors with our conceptual transport triggered architecture processor over reference software implementations. Our processor provides better results than RISC-V competitors, regardless of the algorithm. It seems to be up to 3x faster, 1.6x-2x smaller, and consumes 1.3x-3.6x less energy than the compared RISC-V cores. Thus, an alternative base architecture is proposed for post-quantum cryptography processor development for embedded systems. The most critical shortcoming of the proposed architecture is the lack of compatible intellectual property core support for system-on-chip designs. We share comparative analyses with test results for different core configurations.
ARINC Specification 664 Part 7 (ARINC-664) defines an Ethernet based deterministic network protocol that provides bounded delay and jitter using redundant communication among the avionics applications. Achieving the end-to-end bounded delay objectives requires that incoming Ethernet frames must be regulated according to the ARINC-664 standard. However, the standard does not specify the details of traffic shaping and scheduling mechanisms. FPGA is one of the most preferred implementation choices for ARINC-664 due to its low power consumption, low latency data transfer, and security advantages. Compared to time consuming FPGA development, a model based hardware design enables faster prototyping and testing environment. In this study, a Hardware Description Language (HDL) convertible simulation environment in Simulink is created for ARINC-664 End System (ES) traffic regulator with several scheduling algorithms, and their performance analysis is reported. In addition, a run-time configurable and hardware convertible dynamic traffic regulator is proposed.
Lightweight cryptography is useful to provide security and privacy in resource constraint embedded devices. Latency and memory consumption are the key elements in performance metrics for lightweight cryptography algorithm implementations. Ascon lightweight cryptography algorithm is one of the finalists in CEASAR competition. In this study, special cryptographic non-standard RISC-V instructions have been developed in order to reduce the required number of clock cycles and instruction memory for the execution of the algorithm on RV32I based processors. A profiling methodology has been developed to choose the best special instruction for achieving the highest benefit in performance. An end-to-end test environment has been formed by extending the GNU Compiler Collection and Spike RISC-V ISA Simulator for the special cryptographic instruction extensions of RISC-V processors. New intrinsic functions and instruction patterns for the new instructions have been integrated into the GCC RISC-V back end. Spike has been modified with the new instructions to run the program. The algorithm has been analysed with the proposed instructions and different optimization flags and improvement results have been shown in this study.
The concepts of “Open-source software” and “Open-source hardware” are thriving in the modern society. A significant part of this effort is directed towards the provision of open-source microprocessor designs. In light of this, researchers at University of California, Berkeley developed a license-free Instruction Set Architecture called “RISC-V”, which essentially defines the vocabulary of the hardware/software interface. Some crucial aspect of an open-source hardware are its extendibility, flexibility, and comprehensibility. Most designs are often extendible and well thought-out, but they are rarely comprehensible, which negatively impacts their extendibility. The common problem is that they either lack documentation or have hastily written ones. With the project that this paper represents, the mentioned problem was tackled by designing an open-source 32-bit Synthesizable RISC-V Core with detailed documentation. The design diagrams and design choices are disclosed, making it easy-to-understand. The RISC-V core is named “Hornet Core”.
The Laplacian filter is one of the fundamental applications in image processing. In our work, the Laplacian filter has been applied to an image, and both hardware and software implementation of the filter has been studied. Our system consists of an OV7670 Camera module, Nexys 4 DDR FPGA board and VGA monitor to display the processed video stream. Mentioned process has forwarding tasks: camera module captures raw RGB data and writes to RAM, Laplacian filter IP processes raw image and the results written back to memory. VGA modules show output images to monitor. The Laplacian filter part considered in hardware and software implementation is compared in terms of time and area.