
A relaxation oscillator(RO) composed of a voltage and current reference (VCR) and a digital current comparator is presented in this paper. By using two different types of resistors with opposite temperature coefficient (TC) in the VCR, this design successfully stabilizes the operation frequency from -20 to 100° C with a very simple structure. Designed in a standard 180nm CMOS process, the proposed RO consumes 120nW under a 0.8V supply and operates at 121kHz, with a TC as low as 36.2 ppm/°C.
A relaxation oscillator(RO) composed of a voltage and current reference (VCR) and a digital current comparator is presented in this paper. By using two different types of resistors with opposite temperature coefficient (TC) in the VCR, this design successfully stabilizes the operation frequency from -20 to 100° C with a very simple structure. Designed in a standard 180nm CMOS process, the proposed RO consumes 120nW under a 0.8V supply and operates at 121kHz, with a TC as low as 36.2 ppm/°C.
Cu/SiO2 hybrid bonding is a potent tool to effectively mitigate data-movement issues within von Neumann architecture due to the shortening of the distance between the processor and the memory unit. To protect stacked chip performance, the realization of hybrid bonding at low temperatures (<260°C) is paramount. The essence of low-temperature hybrid bonding lies in the construction of desirable chemical structures on Cu and SiO2 surfaces. Therefore, this paper presents two types of feasible surface-activation strategies to achieve selective/non-selective hydrophilization of the Cu/SiO2 surface. Regardless of activation strategy, the Cu-Cu interface with sufficient grain growth and seamless amorphous SiO2-SiO2 interface structure were obtained at 200 °C. Moreover, the non-selective hydrophilization of Cu/SiO2 surface based on Ar/O2→NH4OH activation realized interfacial layer-free SiO2-SiO2 interface, which can provide more reliable mechanical support for next-generation data-centric applications.
Neural Processing Unit (NPU) has become the state-of-the-art solution for accelerating artificial neural networks and is increasingly integrated on the System-on-Chip (SoC) of edge devices such as smartphones and cameras. However, adopting NPU in mission-critical systems, such as aerospace aircraft and autonomous driving demands high reliability, which is currently less explored on industrial NPUs. In this work, we target one of the critical reliability issues - permanent fault - for modern NPUs and provide an instruction-driven fault detection method named LIPFD-NPU. The approach executes dedicated network instructions in a self-testing fashion and generates fine-grained information on the potential fault's location, type and level of impact. An FPGA-based fault emulation framework is used to verify LIPFD-NPU. The results indicate that LIPFD-NPU effectively detects faults with tiny overheads of 0.2% in silicon area and 0.5% in power consumption.
With the rapid growth of data centers, the power supply has shifted from 48/12/1V two-stage architecture to 48/1V single-stage. In this paper, a new two-phase sawtooth voltage mode PWM control is proposed for the double step-down (DSD) converter. In order to solve the problem of inherent cycle delay of PWM control, a fast-transient response scheme is proposed. The converter also has a precharge and soft start scheme, which is designed with a 0.18 µm BCD process. It achieves peak efficiencies of 92.7 % , 90%, 87.8%, and 86% at 250 kHz, 500 kHz, 750 kHz, and 1 MHz, respectively. During a 1A/10 ns load jump, the undershoot is reduced from 200 mV to 102 mV and the setting time is reduced from 5.3 µs to 1.6 µs.
Instruction set randomization has been proposed for many years as a strategy against code injection. However, most of the methods are based entirely on software, which is vulnerable to possible threats like key leakage or bypassing attack. The translation of instructions also brings the loss of performance. Some designs randomize the instruction set based on hardware, but using weak approaches which can be easily bypassed. In this paper, we propose a hybrid instruction set randomization with both compiler support and hardware extension on a RISC-V processor. We adopt AES-128 to randomize RISC-V instruction set with little performance loss. The design has been implemented on Xilinx AV7K325 FPGA board, the results shows that RISC-V instruction set is randomized with no changes in clock frequency, 1377 LUTs increase in resources and 0.38% performance overhead.
An adaptive on-time (AOT) buck converter with constant switching frequency and fast transient response is presented. A frequency-locked loop (FLL) is used to achieve constant switching frequency. The on-time (TON) is adjusted by a TON extender to achieve fast transient response. The proposed AOT buck converter is implemented in 0.18µm CMOS process. The simulation results show that the switching frequency is fixed at IMHz under various load condition and the output voltage undershoot and settling time are only 50m V and 2.5µs, respectively during 4A load transient.
A Ku-band high power amplifier (HPA) is designed based on the 0.15µm GaN HEMT process. To improve the power added efficiency and gain, a high gate-width drive ratio of 1:6:38.4 is selected for a three-stage topology. Multi-order Chebyshev impedance transformers are used for realizing this high impedance transformation ratio match networks. Meanwhile, a compact 8-way power combining network with low insertion loss is adopted to improve the output power and power added efficiency. The measured results under continuous wave (CW) show that the small signal gain exceeds 30 dB over 13–17 GHz, and the input return loss (IRL) is better than -11dB. The output power is between 42–44 dBm and the power-added efficiency (PAE) is more than 30%. The chip size is 2.6 mm×4.3mm.
The S parameter amplitude, latency, resistance, and inductance of TSV-RDL structures with the presence of five kinds of defects are simulated as feature vectors for defect detection and classification. Three nondestructive defect classification schemes for the TSV-RDL structure in advanced packaging are evaluated. Feedforward neural network with rectified linear unit activation function for the backpropagation algorithm is superior for defect classification and may play an important role in design for test and build-in self-repair circuit design.
As one of the most promising energy-efficient paradigms in deploying Neural Network (NN) on hardware, approximate computing ( $A$ xC) has recently gained great traction to replace exact computing. This paper proposes an efficient approximate multiplier design method, which combines the Cartesian Genetic Programming (CGP)-based automatic design method and manual design method. Besides, an error compensation scheme based on the traversal search of truth table is proposed for higher-order multiplier construction. Experiments show that compared to exact multiplier, the proposed approximate multiplier can reduce the area, power consumption, and delay by 54.9%, 55.7%, and 36.86%, respectively. It also shows superiority to the state-of-the-art approximate multiplier. In addition, when deployed in LeNet-5 for MINIST datasets, the proposed multipliers show higher efficiency than exact multiplier with comparable recognition accuracy.
This paper proposes a low input impedance, high output swing column-level analog front-end (AFE) circuit for the frequency modulation continuous wave (FMCW) LiDAR. The AFE circuit adopts a shunt-feedback transimpedance amplifier (TIA) and an output swing compensation limiting amplifier (LA) to amplify the weak echo signal and realize the high output swing. The input direct current cancellation (IDCC) circuit is used to stabilize the output direct operating point, where the lag network is used to extend the frequency range of high gain, and the noise of the direct current is reduced by means of noise transfer. The proposed AFE circuit is implemented in a 55nm CMOS process, and the post-layout simulation results show that the circuit achieves a gain of more than 118.7 dBΩand an output swing of higher than 840 mV in the frequency range of 26 MHz∼400 MHz. The input-referred noise current is 12.45 pA/Hz 0.5 and the power consumption is 7 mW with a 1.2 V power supply.
The optical responses and memory effects of photoelectric synaptic devices based on CdSe quantum dots (QDs) and poly(3-hexylthiophene) (P3HT) are studied in this work. Compared with devices only incorporating CdSe QDs, the devices based on CdSe QDs and P3HT exhibit higher photocurrents because the heterojunction formed by CdSe QDs and P3HT enhances the separation of photogenerated excitons, and the loss of excitons in the QDs reduces. In addition, due to the effect of the surface defect trapping charge of CdSe QDs, the photocurrent of the device can still be maintained for more than 100 seconds under the condition of zero gate voltage. Finally, the device can perform each synaptic activity with a low power consumption of 12.9 pJ by adjusting the concentration of QDs.
This paper is focused on a microwave sensor for the evaluation of the dielectric properties of binary liquid mixtures at RF/microwave frequencies. The sensor consists of a split ring resonator (SRR), built using the microstrip technology. Interdigitated electrodes are integrated into the ring as a sensing element for liquid detection. A proper extraction procedure has been proposed for the accurate evaluation of the resonant frequency of the developed prototype. The resonant extraction procedure is based on the analysis of the frequency-dependent behavior of the complex forward transmission coefficient (S 21 ) that is accurately modeled locally around the resonance by using a fitting function. According to the tests carried out with water-isopropanol liquid mixtures at various volume fractions, the studied device is more sensitive than the more conventional SRR sensor.
A 2.4GHz LC-DCO with low frequency drift is presented to support frequency synthesizer under narrow band system like BLE. In the LC-DCO, a comprehensive temperature compensation scheme, which includes a Proportional To Absolute Temperature (PTAT) current bias and the varactor arrays varying linearly with voltage, is proposed to reduces the frequency drift of LC-DCO as a result of temperature fluctuations. By applying the circuit, frequency drift is reduced from 31MHz (without compensation) to 6MHz within the temperature from -40°C to 120 °C. And the results show that no extra in-band noise is added to LC tank. It consumes 860uW from a 0.9V supply in 40nm CMOS process technology.
This paper presents a statistics-based background capacitor mismatch calibration algorithm for successive approximation register (SAR) analog-to-digital converter (ADC). The calibration algorithm is capable of detecting capacitor mismatch errors based on statistical principles and signal correlation is eliminated by introducing additional dummy capacitors, leading to fast convergence. This calibration increases the signal-to-noise-and-distortion ratio (SNDR) from 64.57dB to 82.03dB and achieves 29dB spurious-free dynamic range (SFDR) improvement. The simulated differential nonlinearity (DNL) and integral nonlinearity (INL) are +0.17/-0.13LSB and +0.36/-0.38LSB respectively.
Exponential calculation is widely used in different algorithms, such as the activation functions of artificial neural networks. However, it is hard to implement on FPGA, consuming much time and resources. In this work, a novel exponential calculation module for fixed-point number is proposed based on the theory of Fast InvSqrt. The proposed exponential unit achieves at most 3.7x throughput while the resource utilization is largely reduced compared with previous works. The efficiency and accuracy are suitable for different applications.
A 0.7-2.5GHz NB-IoT/GNSS/BLE hybrid PLL with a single D/VCO is implemented in 28nm CMOS. With careful frequency planning, the PA pulling effect is mitigated by using a multi-mode divider chain. A divider-by-2.5 relaxes the tuning range requirement of the D/VCO and mitigates the PA pulling for NB-IoT HB band, while a divider-by-6 is designed for NB-IoT LB. With an 8-tap FIR filtering method, a wideband fractional-N PLL is designed without increasing the out-of-band phase noise. The proposed PLL consumes the maximum 4.7mW with 0.9V supply. Experimental results show that the PLL meets the phase noise and spur requirements of the NB-IoT/GNSS/BLE standards.
Large-scale matrix inversion is widely used in massive Multiple Input Multiple Output (MIMO) beamforming systems, but matrix inversion is very complicated in hardware implementation. In this paper, Hermitian matrix decomposition method based on partitioned systolic array is proposed, and the computing structure of the algorithm is improved flexibly by utilizing the partitioned characteristics of large-scale matrix. We compare our method with existing FPGA-based technologies on Xilinx ZCU102 FPGA. The results of the experiment show that our method has better performance than existing techniques in resource utilization, device delay and maximum working frequency when the size of Hermitian matrix is 32 × 32, which is a typical size for MIMO applications.
For edge intelligent applications, this work proposes a tiny neuromorphic hardware core embedding high-speed on-chip synaptic plasticity, by adopting the proposed Temporal-Integrate neuron model and a simplified supervised spike-driven synaptic plasticity rule for on-chip learning. The proposed hardware core was prototyped on a very-low-cost Zybo Zynq-7010 FPGA device, and attained comparably high classification accuracies on many datasets (e.g. 90.4% on MNIST), with a learning and inference speed as high as 11,268 and 11,749 f $r$ ame/s, respectively, while dissipating only 39 mW power under a 250 MHz clock frequency.
GPGPUs utilize multi-dimensional memory subsystems to provide the bandwidth needed by their multi-dimensional parallelism. However, an unfavorable address mapping leads to imbalanced memory request distribution across the memory resources, causing degraded performance and poor power efficiency. The optimal mapping is both application- and hardware-dependent. This paper provides a software-hardware co-design to dynamically reconfigure the address mapping according to the trace of the targeted application. First, a circuit to sample the entropy of address bits is proposed to capture the optimal address mapping. Second, a dynamic reconfiguration mechanism is designed to apply the optimal address mapping. Simulation results show up to 45% performance improvement over fixed address mappings.