Zero-knowledge proof (ZKP), one of the most popular privacy-preserving schemes, enables the prover to convince the verifier of a certain statement’s correctness without leaking any private information. Among them, the Pairing-based succinct non-interactive argument of knowledge (zkSNARK), like Groth16 and Plonk, featuring succinct proof size and constant-time verification, is deployed across a promising ecosystem. However, the heavy operations of proof generation that grow rapidly with application sizes and the frequent verification that requires costly Pairing severely hinder broader adoption. Moreover, there are many other cryptographic schemes also constructed based on Pairing, such as Boneh–Lynn–Shacham (BLS) signature, functional encryption (FE), identity-based encryption (IBE), and so on. Unfortunately, existing synthesis-based ZKP accelerators suffer from low area efficiency and incomplete operator coverage, causing poor overall speedup and expensive deployment cost. This work presents the first silicon-proven crypto-processor that supports complete proof generation/verification phases and can adapt to other Pairing-based cryptography (PBC). To achieve high flexibility while maintaining cost efficiency, we employ holistic optimizations across algorithm, architecture, and compiler levels, including the hybrid-grained instruction set, dedicated memory system, reconfigurable datapath, custom Pairing compiler, and utilization-oriented timing scheduling. Fabricated in 28-nm CMOS, it covers all related operators and achieves a $6.4\times $ area reduction compared to prior synthesis-based work LegoZK, which currently makes it the most cost-effective candidate for the practical deployment of ZKP–PBC applications. As a case study, the processor performs $2^{16}$ -gate proof generation/verification with 0.15 J/0.11 mJ, offering two orders of magnitude energy savings over a 14-core 2.4-GHz CPU. For the Pairing operator, it achieves development agility to enable fast design space exploration. It delivers $17\times $ / $1.8\times $ improvements of area-time product (ATP)/energy compared to the prior ASIC work.
Fully homomorphic encryption (FHE) enables privacy-preserving machine learning (PPML) at the cost of intensive computational overhead, which necessitates the use of domain-specific accelerators. To achieve comprehensive support for leveled FHE, this article presents a reconfigurable multi-scheme FHE processor that supports both client-side encryption/decryption and server-side evaluation. First, a reconfigurable processing element (RPE) design for modular arithmetic and a reusable data generator for polynomial sampling are developed to support the various operations in FHE. Second, a configurable RPE array supporting polynomial operations and a decoupled automorphism unit (DAU) necessary for homomorphic rotations are proposed to accelerate the FHE primitives with complex dataflow. Finally, an on-chip data generation strategy and a cache-aware operation scheduling (CAOS) method are introduced to alleviate the memory bottleneck in the end-to-end execution of FHE applications. The chip is fabricated in a 28-nm process and tested with end-to-end execution. Targeting a lightweight parameter set with polynomial degree N=4096 at 128-bit security level, the proposed chip achieves 4.05 mu J per encryption on the client side and provides a throughput of 8.72 kHMul/s on the server side. In terms of the number theory transform (NTT) operation, the chip demonstrates the highest throughput and best area efficiency compared with state-of-the-art solutions.
White-box cryptography (WBC) seeks to protect secret keys (SKs) even under the white-box security model that features adversaries having full control of the execution environment. Due to the ever-growing demand for content protection under security-critical scenarios, the recent progress on WBC has been nothing short of spectacular. However, the security-prioritized strategy also brings massive computation overhead to the current dedicated white-box block cipher (DWBC), seriously hindering its broad application. Compared to traditional software-only solutions, this work proposes a highly efficient domain-specific and silicon-proven processor (named SiWB) to improve the throughput (energy efficiency) by hundreds (several orders) of magnitude. First, not only a standalone crypto-core in prior work, we develop a complete and robust system that integrates multiple hardware cores for accelerating DWBC, user-defined high-speed interface (UPIF) for improving data transmission speed and other synergetic components. Second, to enhance throughput while maintaining energy efficiency, we devise a multi-core architecture based on inter-core load-aware scheduling, ultimately achieving 11.4-Gb/s peak throughput and 14.2-Gb/s/W energy efficiency. Third, a configurable and area-efficient datapath supporting multi-size random linear transformation (LT) and non-LT (NLT) is designed to reduce resource redundancy by means of algorithm-hardware co-optimization. To our best knowledge, this is the first 28-nm silicon-proven hardware accelerator for white-box block ciphers under black-box mode, which occupies 1.8-mm(2) area and achieves 800-MHz peak frequency at 0.9 V. It supports several white-box block ciphers including SPNbox-8/16/24/32, Yoroi-16/32, and white-box addition/rotation/XOR white-box addition/rotation/XOR (WARX), delivering 389x speedup on average versus software programs optimized with advanced encryption standard new instructions (AES-NIs) on a 14-core 3.4-GHz CPU. It enables WBC to be a viable solution for applications with stringent throughput requirements, such as streaming media commonly demanding several Gb/s.
This 0.046mm(2) SHA-3 engine demonstrates overheads of only 17% area, 27% energy, and 0% latency while achieving provable glitch- and transition-robust security (power/EM TVLA MTD>100M). Glitch-free DPL reduces masked states by 3x, resulting in a 72,950 mu m(2) register area reduction and lower logic energy. LMDPL refresh preserves secure asymmetric-share mapping. Masked transitions enable partial replacement of DPL with SRL, saving 50,852 mu m(2) in registers, precharge gates, and multiplexers.
This paper presents the first metastability-based true random number generator achieving Gb/s throughput, supported by a provable stochastic entropy model. The proposed staggered working point ensemble architecture employs four entropy source cores with deliberate offset staggering to ensure PVT-resilient operation without real-time tracking. A stochastic model characterizes the entropy generation dynamics, establishing a conservative min-entropy lower bound compliant with international standards such as NIST SP 800-90B and AIS-31. Fabricated in 28 nm CMOS, the proposed design achieves entropy throughput of 4064.5 Mbps at $997.6 \text{fJ} /$ entropy-bit (1.3 V) and 1462 Mbps at $324.1 \text{fJ} /$ entropy-bit $(0.9 \mathrm{V})$. Validation across 40 test chips confirms full compliance with NIST SP 800-90B and AIS-31 suites. The design maintains a min-entropy exceeding 0.77 after 168 hours aging, demonstrates resilience to 60 mVpp power-injection attacks, and achieves 10.08 mV metastable point drift tolerance, confirming robust high-entropy operation against PVT variations.
This 0.046 mm2 SHA-3 engine demonstrates overheads of only 17% area, 27% energy, and 0% latency while achieving provable glitch- and transition-robust security (power/EM TVLA MTD $>100 \mathrm{M}$). Glitch-free DPL reduces masked states by $3 \times$, resulting in a $72,950 \mu \mathrm{m}^{2}$ register area reduction and lower logic energy. LMDPL refresh preserves secure asymmetric-share mapping. Masked transitions enable partial replacement of DPL with SRL, saving $50,852 \mu \mathrm{m}^{2}$ in registers, precharge gates, and multiplexers.
As the semiconductor industry transitions into an era shaped by ubiquitous artificial intelligence,heterogeneous integration,and the prospective impact of cryptanalytically capable quantum computers,hardware security is increas-ingly extending beyond its traditional role as isolated crypto-graphic engines or standalone blocks[1,2].
Emerging 6G communication systems impose unprecedented requirements on cryptographic primitives, demanding ultra-high throughput, low latency, and strong resistance to implementation-level attacks. LOL2.0 is a recently proposed stream cipher framework that achieves high software efficiency and strong security in post-quantum settings. While several stream ciphers have been proposed to address performance demands, side-channel-protected hardware implementations capable of sustaining 6G-class throughput remain largely unexplored. In this work, we present the first side-channel-protected hardware implementation that meets the 6G-class throughput demand. Focusing on the LOL2.0 stream cipher framework, we leverage Time Sharing Masking to achieve first-order security under the glitch-extended probing model. This design realizes full-phase protection covering initialization, keystream generation, and tag generation. To address diverse deployment requirements, we design two masked architectures: a compact variant optimized for area and randomness efficiency, and a fast variant targeting the maximum achievable throughput. The proposed fast implementations achieve peak throughputs of 183 Gbps and 142 Gbps for the unmasked and masked configurations, respectively. Meanwhile, the compact architecture reduces hardware cost by achieving areas as low as 23.25 kGE and 141.98 kGE in unmasked and masked designs, respectively, while still maintaining competitive throughput. Security is validated through practical side-channel evaluations using Test Vector Leakage Assessment on FPGA platforms. Across up to 100 million measured power traces, no statistically significant first-order leakage is observed for any protected configuration. Overall, this work realizes side-channel-protected stream cipher hardware that sustains ultra-high throughput, providing a concrete path toward secure cryptographic deployment in future 6G communication systems.
HAWK is the only lattice-based candidate in the second round of additional digital signature schemes of the NIST post-quantum cryptography (PQC) standardization process, which is characterized by its low latency, compact storage requirements, and floating-point-free computations. However, research on hardware implementations of HAWK has remained limited, due to its challenging hardware reuse across heterogeneous arithmetic operations and complex computation patterns. This lack of exploration in acceleration and deployment studies currently constrains its broader adoption in practical systems and the future standardization process. To address this issue, HawkPU is proposed, namely HAWK Processing Unit, an efficient and low-latency crypto-processor to accelerate HAWK signature generation and verification on FPGA. Our approach involves a novel vectorized strategy for NTT/FFT, and a high-throughput, reconfigurable arithmetic module that supports various types of integer operations. In addition, several optimized modules are also introduced, including a load-balanced data generation module, a compact GF(2) polynomial multiplier, and an optimized Golomb-Rice encoder/decoder. Furthermore, this paper presents the first full hardware accelerator for HAWK signature generation and verification, achieving a low-latency HAWK hardware implementation. Signature generation and verification performance in security level I are 22.2 μs and 45.2 μs on the Zynq-UltraScale+ platform. Compared to the state-ofthe- art design, our signature generation and verification speed on Zynq-7000 are 88.8x and 33.6x faster. The area-time products of signature generation and verification in terms of LUT/FF/DSP/BRAM are 41.6x/31.4x/51.3x/177.5x and 15.7x/11.9x/19.4x/67.1x lower than the previous work, respectively.
Falcon is a lattice-based quantum-resistant digital signature scheme renowned for its high signature generation/verification speed and compact signature size. The scheme has been selected to be drafted in the third round of the post-quantum cryptography (PQC) standardization process due to its unique attributes and robust security features. Despite its strengths, there has been a lack of research on hardware acceleration, primarily due to its complex calculation flow and floating-point operations, which hinders its widespread adoption. To address this issue, we propose FalconSign, a high-performance, configurable crypto-processor designed to accelerate Falcon signature generation on FPGA/ASIC through algorithmhardware co-design. Our approach involves a new scheduling flow and architecture for Fast-Fourier Sampling to enhance computing unit reuse and reduce processing time. Additionally, we introduce several optimized modules, including configurable randomness generation units, parallel floating-point processing units, and an optimized SamplerZ module, to improve execution efficiency. Furthermore, this paper presents a finely optimized hardware accelerator for the Falcon scheme. Our FPGA implementation results demonstrate a throughput improvement of approximately 5.1 x compared to state-of-the-art designs, with 2.8x/4.5x/4.2x/3.2x fewer in the area (LUTs/FFs/DSPs/BRAMs)-time product, for NIST security level V. The crypto-processor occupies an area of 0.71 mm2 and achieves 5.2k OPS at throughput on the TSMC 28nm process for NIST security level I.
True randomness plays a vital role in modern cryptographic systems, generating keys, and masks. Several existing secure catastrophes can be traced back to inappropriately implemented or utilized TRNG [1, 2]. The design methodology and evaluation of a TRNG had been undergone a revolution toward devising and applying a stochastic model. NIST SP 800-22, the widely used statistical test suit, was rejected by NIST itself to be used for assessing cryptographic random number generators [3]. We followed an entropy evaluation procedure to be compliant with the latest NIST SP800-90B [4]. Among many exploitable random physical phenomena, timing jitter and metastability stand out. Timing-jitter based TRNGs entitle a better compatibility with digital circuit design flow and an extensive formal security analysis, while suffer from limited throughput and cumbersome implementations [5, 6]. On the other hand, metastability-based TRNGs are favored for its unparallel throughput, energy efficiency and compact footprint, at the cost of a higher full-custom design effort and an inevitable mismatch introduced by process variation [7, 8].
A true random number generator (TRNG) is a critical component in ensuring the security of cryptographic systems. Among TRNG implementations, the phase-locked loop-based TRNG (PLL-TRNG) is a widely adopted solution for FPGA platforms due to the availability of a stochastic model. In the previous study, this stochastic model was based on analog noise signals, which potentially led to an oversimplification of the PLL physical process and resulted in an overestimation of entropy. To address this limitation, we extract key platform-specific parameters of the PLL and develop a new stochastic model tailored for multi-output PLL-TRNGs. For the first time, we reveal the effect of the PLL’s bandwidth on the correlation of sampling points and introduce a method for quantitatively controlling sampling point correlations. Finally, we validate the model through on-chip jitter measurements. Experimental results show that the proposed stochastic model accurately describes the behavior of the PLL-TRNG and provides the most conservative entropy lower bound, with a 1.8-fold improvement in jitter resolution.
In cryptographic systems, true random number generation is essential, as a compromised TRNG could lead to a security catastrophe. The raw random numbers are discrete values that are derived at discrete points in time from a noise source of a TRNG. These values often exhibit statistical defects that require post-processing, also called conditioner, to improve uniformity. The two main types of post-processing are algorithmic post-processing and cryptographic post-processing, both of which have pros and cons in theories and applications. However, another type of postprocessing existing between these two types, named entropy extractor, has often been overlooked by the applied cryptographic community. Therefore, we implement two information-theoretically provable entropy extractors: Toeplitz extractor and Trevisan extractor catering to various performance requirements and applications of high-throughput TRNG post-processing. This paper proposes a combination of matrix chunking and FFT acceleration to boost the performance of the Toeplitz extractor, along with a modified Toeplitz matrix design to decrease the hardware consumption. In addition, we introduce a lightweight single-bit extractor to implement an efficient Trevisan extractor. Both algorithms are devised and verified through FPGA hardware simulations. The enhanced Toeplitz extractor achieves a throughput of 42 Gbps, while the Trevisan extractor attains 1.82 Gbps, representing an 84% and 73% improvement in throughput-to-area ratio over the previous best-performing design for each extractor. The standard statistical test suites, such as NIST SP800-22, NIST SP800-90B, and AIS-31, are adopted to evaluate the effectiveness of the proposed post-processing techniques. Naturally, this approach can only serve as a supplementary measure, as modern standards, such as AIS-31, necessitate formal analysis and stochastic models to account for randomness.
Fully homomorphic encryption (FHE) allows arbitrary encrypted computation for privacy enhancement, at the cost of several orders of magnitude slowdown and memory expansion. However, the prohibitive power consumption and fabrication cost associated with the large silicon solutions supporting the computationally expensive bootstrapping operation [1,2,7-11] turn out to be the major barrier to the broader adoption of FHE. Leveled FHE with designated multiplication depth and conjoint encrypted client-server computing [3,4,6,12-17] is considered more desirable and practical for privacy-preserving machine learning (PPML) applications [18]–[21]. As shown in Fig. 17.2.1, three major challenges remain to be addressed for hardware acceleration. First, higher flexibility is required to support prevailing FHE schemes, including BGV [22], BFV [23], and CKKS [24], and diverse primitives on both client side and server side. Second, FHE-specific primitives, including (I)NTT and polynomial automorphism, incur huge computational overhead with complex on-chip data access patterns, which is hard to parallel and tempers the hardware efficiency. Third, FHE-based workloads introduce intensive memory bottlenecks in terms of storage and data movement, making it challenging to execute application-level tasks on-chip. To address these issues, an energy efficient 4.05µJ/Encryption FHE processor with architecture-circuit optimizations to achieve 8.72kHMul/s in the throughput is presented. To the best of our knowledge, this is the first silicon-proven solution to suit well for both client and server computing patterns.
As emerging applications raise ever-boosting and varying computational demand, the reconfigurable accelerator is becoming prevalent due to balanced performance, efficiency, and flexibility. Although the functions of its processing elements (PEs) and interconnections can be defined by a high-level software program, the circuit parameters are mostly managed by hardware or the compiler, leaving opportunities for in-depth optimization and novel features. Considering that the circuit adjustment is of great value in terms of low power and security, this article presents a novel software-defined accelerator named cross-domain software-defined chip (CDSDC), of which most chip properties including circuit parameters, logic, and pipelines can be programmed synergistically by software at runtime. With this design scheme, CDSDC provides new features including: 1) dynamic pipeline-voltage-frequency scaling (DPVFS), which combines pipeline reconfiguration with voltage-frequency scaling to generate an optimal chip configuration for various applications; 2) intrinsic physical unclonable function (PUF), which leverages the uniform PEs as measured circuit delay elements by adjusting the circuit parameters, extracting the entropy of manufacturing process variation, and taking response bits as a PUF; and 3) device-binding confidential configuration, which uses the intrinsic PUF and ASCON algorithm to encrypt/decrypt configuration. The CDSDC has been implemented in silicon as a 28-nm Taiwan Semiconductor Manufacturing Company (TSMC) HPC+ 1P8M 6-mm(2) test chip, operating at peak 400 MHz and 0.9 V, achieving a peak energy efficiency of 51 GOPS/W. It improves performance and efficiency for cross-domain patterns of computation by DPVFS. It improves hardware security against reverse engineering attacks by using intrinsic PUF and device-specific encrypted configuration flow.
Zero-knowledge proof (ZKP) is an attractive cryptographic paradigm that allows a party to prove the correctness of a given statement without revealing any additional information. It offers both computation integrity and privacy, witnessing many celebrated deployments, such as computation outsourcing and cryptocurrencies. Recent general-purpose ZKP schemes, e.g., zero-knowledge succinct non-interactive argument of knowledge (zk-SNARK), suffer from time-consuming proof generation, which is mainly bottlenecked by the large-scale number theoretic transformation (NTT) and multi-scalar point multiplication (MSM). To boost its wide application, great interest has been shown in expediting the proof generation on various platforms like GPU, FPGA and ASIC. So far as we know, current works on the hardware designs for ZKP employ two separated data-paths for NTT and MSM, overlooking the potential of resource reusage. In this work, we particularly explore the feasibility and profit of implementing both NTT and MSM with a unified and high-performance hardware architecture. For the crucial operator design, we propose a dual-precision, load-balanced and fully-pipelined Montgomery multiplier (LBFP MM) by introducing the new mixed-radix technique and improving the prior quotient-decoupled strategy. Collectively, we also integrate orthogonal ideas to further enhance the performance of LBFP MM, including the customized constant multiplication, truncated LSB/MSB multiplication/addition and Karatsuba technique. On top of that, we present the unified, scalable and highperformance hardware architecture that conducts both NTT and MSM in a versatile pipelined execution mechanism, intensively sharing the common computation and memory resource. The proposed accelerator manages to overlap the on-chip memory computation with off-chip memory access, considerably reducing the overall cycle counts for NTT and MSM. We showcase the implementation of modular multiplier and overall architecture on the BLS12-381 elliptic curve for zk-SNARK. Extensive experiments are carried out under TSMC 28nm synthesis and similar simulation set, which demonstrate impressive improvements: (1) the proposed LBFP MM obtains 1.8x speed-up and 1.3x less area cost versus the state-of-the-art design; (2) the unified accelerator achieves 12.1x and 5.8x acceleration for NTT and MSM while also consumes 4.3x lower overall on-chip area overhead, when compared to the most related and advanced work PipeZK.
CRYSTALS-Dilithium has been declared as the first recommended digital signature algorithm in NIST Post-Quantum Cryptography Standardization. The advancement of high-speed hardware research for Dilithium is propelled by the need for real-time processing of extensive data in numerous digital signature applications. To address the slow signature generation speed issue, a two-stage pipeline structure was developed to accelerate the underlying rejection loop, at a cost of substantial resource consumption. In this paper, we present the first analysis on the possibility of leveraging sparse multiplication in the second stage, which can reduce the bit complexity of corresponding multiplications by over 85% and lower the storage requirements for the secret key by over 68%. Building on this, we propose a sparse computing core and a high-speed hybrid architecture for Dilithium, with an efficient scheduling mechanism and optimized modules. Compared to state-of-the-art high-speed implementations on similar platforms, the signature generation speed is at least 2x faster. Meanwhile, the area-time-products of signature generation achieve 3.6x/4.3x/2.0x/2.1x improvement in terms of LUT/FF/DSP/BRAM, respectively.
SHA-3, the latest hash standard from NIST, is utilized by numerous cryptographic algorithms to handle sensitive information. Consequently, SHA-3 has become a prime target for side-channel attacks, with numerous studies demonstrating successful breaches in unprotected implementations. Masking, a countermeasure capable of providing theoretical security, has been explored in various studies to protect SHA-3. However, masking for hardware implementations may significantly increase area costs and introduce additional delays, substantially impacting the speed and area of higher-level algorithms. In particular, current low-latency first-order masked SHA-3 hardware implementations require more than four times the area of unprotected implementations. To date, the specific structure of SHA-3 has not been thoroughly analyzed for exploitation in the context of masking design, leading to difficulties in minimizing the associated area costs using existing methods. We bridge this gap by conducting detailed leakage path and data dependency analyses on two-share masked SHA-3 implementations. Based on these analyses, we propose a compact and low-latency first-order SHA-3 masked hardware implementation, requiring only three times the area of unprotected implementations and almost no fresh random number demand. We also present a complete theoretical security proof for the proposed implementation in the glitch+register-transition-robust probing model. Additionally, we conduct leakage detection experiments using PROLEAD, TVLA and VerMI to complement the theoretical evidence. Compared to state-of-theart designs, our implementation achieves a 28% reduction in area consumption. Our design can be integrated into first-order implementations of higher-level cryptographic algorithms, contributing to a reduction in overall area costs.
White-box cryptography (WBC) seeks to protect secret keys even if the attacker has full control over the execution environment. One of the techniques to hide the key is space hardness approach, which conceals the key into a large lookup table generated from a reliable small block cipher. Despite its provable security, space-hard WBC also suffers from heavy performance overhead when executed on general purpose hardware platform, hundreds of magnitude slower than conventional block ciphers. Specifically, recent studies adopt nested substitution permutation network (NSPN) to construct dedicated white-box block cipher [BIT16], whose performance is limited by a massive number of rounds, nested loop dependency and high-dimension dynamic maximal distance separable (MDS) matrices. To address these limitations, we put forward UpWB, an uncoupled and efficient accelerator for NSPN-structure WBC. We propose holistic optimization techniques across timing schedule, algorithms and operators. For the high-level timing schedule, we propose a fine-grained task partition (FTP) mechanism to decouple the parameteroriented nested loop with different trip counts. The FTP mechanism narrows down the idle time for synchronization and avoids the extra usage of FIFO, which efficiently increases the computation throughput. For the optimization of arithmetic operators, we devise a flexible and vectorized modular multiplier (VMM) based on the complexity-reduced Montgomery algorithm, which can process multi-precision variable data, multi-size matrix-vector multiplication and different irreducible polynomials. Then, a configurable matrix-vector multiplication (MVM) architecture with diagonal-major dataflow is presented to handle the dynamic MDS matrix. The multi-scale (Inv)Mixcolumns are also unified in a compact manner by intensively sharing the common sub-operations and customizing the constant multiplier. To verify the proposed methodology, we showcase the unified design implementation for three recent families of WBCs, including SPNbox-8/16/24/32, Yoroi-16/32 and WARX-16. Evaluated on FPGA platform, UpWB outperforms the optimized software counterpart (executed on 3.2 GHz Intel CPU with AES-NI and AVX2 instructions) by 7x to 30x in terms of computation throughput. Synthesized under TSMC 28nm technology, 36x to 164x improvement of computation throughput is achieved when UpWB operates at the maximum frequency of 1.3 GHz and consumes a modest area 0.14 mm2. Besides, the proposed VMM also offers about 30% improvement of area efficiency without pulling flexibility down when compared to state-of-the-art work.
G. Di Natale合作论文数LIRMM3