
SiC strained nMOSFETs with sub-35 nm gate length can realize superior f T above 400GHz attributed to more than 20% enhancement of the mobility and transconductance. This super-400 GHz f T makes SiC strained nMOS a premium device for mm-wave CMOS circuits design. However, the SiC nMOSFETs reveal a dramatic increase of flicker noise and random telegraph noise (RTN), which may cause worse phase noise and detrimental impact on CMOS oscillator stability. The complex RTN features abnormally long capture and emission time constants (τ c and τ e ) and suggests new mechanism responsible for the anomalously slow trapping and detrapping, due to a significant increase of the relaxation energy from SiC strain. This critical trade-off between high frequency performance and low frequency noise should be considered seriously for an appropriate deployment of high mobility devices and optimization adapted to RF and analog circuits.
Designing lightweight convolutional neural network (CNN) models is an active research area in edge AI. Compute-in-memory (CIM) provides a new computing paradigm to alleviate time and energy consumption caused by data transfer in von Neumann architecture. Among competing alternatives, resistive random-access memory (RRAM) is a promising CIM device owing to its reliability and multi-bit programmability. However, classical lightweight designs such as depthwise convolution incurs under-utilization of RRAM crossbars restricted by their inherently dense weight-to-RRAM cell mapping. To build an RRAM-friendly yet efficient CNN, we evaluate the hardware cost of DenseNet which maintains a high accuracy vs other CNNs at a small parameter count. Observing the linearly increasing channels in DenseNet leads to a low crossbar utilization and causes large latency and energy consumption, we propose a scheme that concatenates feature maps of front layers to form the input of the last layer in each stage. Experiments show that our proposed model consumes less time and energy than conventional ResNet and DenseNet, while producing competitive accuracy on CIFAR and ImageNet datasets.
As a new coding tool of Versatile Video Coding, Luma Mapping with Chroma Scaling (LMCS) maps luma samples and scales chroma residuals based on the mapping function to improve video quality and compression ratio. To achieve real-time decoding and improve the throughput, a hardware decoder needs to be developed. In this paper, a high throughput LMCS hardware decoder design was proposed. Three separate modules were designed to realize the LMCS processes with two implementation schemes of the Luma Mapping module discussed and the better one used. The proposed design reached 535 MHz at 28nm technology, capable of processing videos at 4K@120fps.
The angular displacement sensor is a sensing tool used to detect angle, angular velocity, and angular velocity by converting the amount of angular displacement into an electrical output. Photoelectric encoders, which are currently utilized extensively in engineering practice, have many benefits over other types of angular displacement sensors, including high precision, high durability, small size, wide measuring range, and ease of installation. In this study, a photoelectric chip based on a dual code channel hybrid readout circuit for a high precision Photoelectric encoder is constructed. The circuit uses a single pseudo-random code element that corresponds to a single light and dark stripe of the incremental code mechanism. It can read absolute position information directly within the tolerance of a code plate engraving error of no more than a quarter cycle, and it can also perform sampling at a rate of 1MS/s while maintaining a resolution of 22 bits. The pseudo-random code channel generates two inverse pseudo-random codes using two separate photodetector arrays, which are then passed through two amplifiers and connected to a comparator for comparison. The result is a 1-bit digital signal as the code channel's final output. A 1-bit digital signal is the output's ultimate form. The device is 4.6mm by 3.5mm and uses a 0.35m CMOS technology.
As a highly repetitive logic operation unit in programmable logic modules, the operating speed of a carry adder circuit is crucial to the overall performance of the programmable logic circuit. In order to optimize circuit propagation delay and area overhead, this paper proposes an improved design scheme for carry adder circuits. Firstly, several basic adder structure types and their performance characteristics are compared to determine the most suitable architecture for carry adder circuits in programmable modules. Then, based on the carry-propagate principle of the carry lookahead adder, a novel single MOS transmission gate structure is proposed to optimize the underlying carry logic circuit. Finally, the carry adder circuit is implemented using a full-custom design flow in TSMC 28nm technology. Experimental results show that compared to existing logic designs, the proposed carry adder circuit in this paper reduces propagation delay and area overhead by 14% and 26.5% respectively.
This paper presents a datasheet-driven SPICE model and corresponding automated extraction process for Silicon Carbide (SiC) Metal-Oxide-Semiconductor Field-Effect Transistors (MOSFETs). The proposed SPICE model introduces an equivalent circuit that capture the intrinsic physical mechanisms of the devices and each circuit element is described by non-segmented empirical formulas. A differential evolution (DE) algorithm is adopted for automated model parameters extraction. We have validated the model by comparing the simulation results with datasheets from two mainstream SiC MOSFET devices vendors. Results show that the RMS errors are less that 7% under different bias and temperatures. The model's convergence was further confirmed through simulation of a three-phase full-bridge inverter circuit.
As CMOS technology is approaching atomic feature size, the complexity of SoC (System on Chip) is increasing. As a result, the SoC verification is facing increased challenges, which calls for the development of a new system-level verification methodology. En route to system-level verifications for SoCs, an efficient method for speeding up the SoC verification is the main desideratum. To this end, this paper proposes an efficient method for speeding up subsystem-level functional verification, and takes DDR subsystem in a SoC design as an example to validate the proposed method. A bypass skill is utilized in the speeding up method. Through replacing the DDR configuration interface (a DDR AHB bus connecting to CPU) with another faster designed AHB bus, the proposed speeding up technique takes advantage of a series of tasks within the designed initialization master, which are faster than the same operations which are issued from CPU. The experimental results show that our proposed speeding up approach is efficient, and the running time speed up ratios reach 7.0×, 6.8×, 11.6× for RTL-level simulations, gate-level simulations without back-annotation, and gate-level simulations with back-annotation, respectively.
In this pager, we implement a novel programmable resistance and capacitance network. It can adjust the resistance and capacitance value on the fly when chip is on. The chain architecture makes it very easy for scaling, integration and layout implementation. The area and power costs are low. The IP can be integrated into a SOC chip. By cooperate with other host MCU, it can be widely used in products that need high precision analog circuits in their lifecycle.
Providing end-to-end stochastic computing (SC) neural network acceleration for state-of-the-art (SOTA) models has become an increasingly challenging task, requiring the pursuit of accuracy while maintaining efficiency. It also necessitates flexible support for different types and sizes of operations in models by end-to-end SC circuits. In this paper, we summarize our recent research on end-to-end SC neural network acceleration. We introduce an accurate end-to-end SC accelerator based on a deterministic coding and sorting network. In addition, we propose an SC-friendly model that combines low-precision data paths with high-precision residuals. We introduce approximate computing techniques to optimize SC nonlinear adders and provide some new SC designs for arithmetic operations required by SOTA models. Overall, our approach allows for further significant improvements in circuit efficiency, flexibility, and compatibility through circuit design and model co-optimization. The results demonstrate that the proposed end-to-end SC architecture achieves accurate and efficient neural network acceleration while flexibly accommodating model requirements, showcasing the potential of SC in neural network acceleration.
Edge intelligence hardware requires very high computing energy efficiency because the energy supply is usually limited in edge devices [1] . Memristor crossbars can be the best candidate to perform processing-in-memory applications, where computing energy efficiency can be much better than traditional digital systems based on Von Neumann architecture [2] . The neural network's vector matrix multiplication can be calculated according to the Physical Ohm's law in the memristor crossbars without sending the data back and forth between the computing and memory units. The united architecture of computing and memory units can be very helpful in reducing computing energy consumption in the memristor crossbars.
Current convolutional neural networks (CNN) have achieved a high inference accuracy by deeper architecture, generating a large amount of interlayer data. Limited to on-chip memory, massive feature maps significantly impact the performance of hardware platforms. In this paper, we propose a dynamic codec with adaptive quantization for compressing inlayer feature maps which efficiently reduces the memory storage. Discrete cosine transform (DCT) is utilized to concentrate essential information in the frequency domain while high-frequency components are compressed by adaptive quantization. The quantization tables are dynamically adjusted by the storage of the previous map and the number of the current layer to maintain a relatively high compression ratio and quality for each layer. In addition, Huffman coding and run length encoding (RLE) are used to code compressed data streams from 2-D to 1-D for storage. Furthermore, the codec is implemented on FPGA and synthesized in TSMC 28nm technology. It yields a reduction of on-chip memory by 44.17% and achieves an average compression ratio of 69.18% for feature maps at diverse depths in ResNet-50.
This paper designs and implements a dedicated operator called FUSE, which implements the function of image fusion in any channel and is applied to the noise reduction and super-resolution neural network designed by Witmem. The FUSE operator we designed uses 32 parallel pipeline calculations, which are internally divided into two clock domains based on the proportion of input and output data volume. We design an algorithm to search for weight address index, making the FUSE operator applicable to images of any size. Under the 28 nm standard CMOS technology, We propose the architecture of the operator and implement the hardware design. FUSE’s area is 236460.2 μm 2 , interface frequency is 1GHz, internal working frequency is 500MHz, and the throughput is 13.7GOPS when image channel is 3.
Full-duplex, which aims to simultaneously transmitting and receiving over the same band, play fundamental roles in next-generation wireless communication systems. Effective cancellation of co-band self-interference between transmitters and receivers is a major challenge of full-duplex systems. Here in this work, we present a custom self-interference cancellation chip with the ability of cancelling the analog interference from two transmitters to one receiver. It embeds additional ADC and DAC to realize channel sounding ability. Based on low-delay LVDS interface and parallel FIR filters, the total time delay of the signal reconstruction channel is below 25 ns. Experimental results of the test bench demonstrate a minimum 30 dB self-interference cancellation within 300 MHz bandwidth.
Though hardware auto generators can efficiently generate architectures of different design metrics based on generation formulas, top-level design for digital signal processing remains challenging. In this paper, we propose an automatic timing-driven top-level hardware generation scheme which integrates top-level timing arrangement, code generation and fast evaluation to further alleviate the heavy workload of hardware design for digital systems. To demonstrate the effectiveness of our proposed scheme, a channel impulse response estimator is implemented. It is shown that our scheme can explore design space and optimize hardware architecture automatically with different design constraints.
This paper proposes a high precision capacitive isolation current sensing amplifier. In order to eliminate gain drift with temperature and improve susceptibility to common mode transient interference, second-order sigma-delta modulator and digital isolator interface are used to ensure high-precision and high-stability signal transmission. The proposed capacitive isolation amplifier is designed for signal bandwidth of 100kHz with input voltage range ±250mV. The Simulation results indicate that the proposed amplifier achieves a maximum gain error of 0.3%, a gain error temperature drift of 10.7ppm/°C and an SNR of 63.9dB.
Wireline receiver front-ends increasingly utilize high-speed time-interleaved ADCs for improved equalization through subsequent digital processing and support for higher-order modulation schemes. A 32GS/s 7bit TI-SAR ADC in 28nm with auto error calibration and trimmed clock is proposed for 32Gb/s ADC-Based SerDes Receiver, achieving stable and high performance with relatively low power consumption. Post-layout simulation results show that the proposed TI-SAR ADC achieves 38.65dB SNDR and 61.41mW power consumption with -1dBFS input signal at Nyquist input frequency.
A 6-Gb/s half rate current mode logic (CML) transmitter has been designed in TSMC 28nm CMOS technology, which employs a 3-tap 3-bit feed forward equalizer (FFE), an analog duty cycle correction module (DCC) for half rate clock and a 5-bit output impedance calibration module for 100-ohm differential load. Post-layout simulation indicates that the FFE could compensate for up to 4.7dB channel loss at 3GHz while maintaining 410mVpp eye height. The DCC module could corrects±30% duty cycle mismatch and max deviation of the calibrated impedance is 2.7%. Implemented in 28nm CMOS technology, the transmitter delivers 6-Gb/s data (2 15 -1 PRBS) with 4.7dB channel attenuation at 3GHz. The transmitter (excluding clock generating PLL) consumes 27.67mA from 0.9V supply and occupies a die area of 214μm * 108μm.
Closed-loop neural detection systems pose significant design challenges to neural recording interface circuits. The interface needs to satisfy the required resolution and power budget while meeting the requirements (e.g. DC offset removal, high DC impedance, stimulation artifacts tolerance, etc.) posed by implantable biomedical scenarios. The state-of-the-art analog-to-digital converter (ADC) direct interface employs a delta-sigma modulator (DSM) structure, where a digital integrator (accumulator) is intentionally relocated behind the quantizer to remove the electrode DC offset. However, the power and area consumptions of the multi-bit digital-to-analog converter (DAC) is increased by the accumulator. Based on the DSM architecture with DC offset removal, this paper proposes a DAC scaling technique to decrease the signal amplitude appearing at the DAC input. It results in the significant reduction in the number of required DAC unit cells and therefore the power and area consumptions of DAC. Meanwhile, to obtain high DC input impedance and high linear input-range, we present circuitry design considerations for the first integrator as well as the feedback loop. Finally, behavioral-level simulation results confirm the efficacy of the proposed structure.
CRYSTALS-Kyber is a powerful Post-Quantum Cryptography(PQC) with high resistance to quantum computer attacks. In order to optimize the core sampling structure of Kyber so as to improve the hardware execution efficiency, this paper proposes a high-performance rejection sampling hardware circuit design scheme for Kyber. The scheme first divides the input random numbers by width converter, and then combines four parallel sampling cores to shorten the sampling time period; then on the basis of the high parallelization of sampling cores, this work use the reorganization of random numbers and the bit-weighted value characteristics of binary numbers to design high-efficiency comparator to reduce the rejection rate, and reduce the sampling time and overhead; and finally, splice the valid inputs and set up a buffer to store the generated samples. The scheme is verified for its performance by testing under Taiwan Semiconductor Manufacturing Company(TSMC) 65nm process and FPGA implementation under Vivado platform. The experimental results demonstrate that the optimized sampling circuit reduces the rejection rate from 18.73% to 11.33% and shortens the sampling clock period by 80.84% compared to the basic implementation; comparing with the related work, it reduces the area by 67.77% while controlling the sampling time, and the consumption of hardware resources is greatly reduced, which improves the efficiency of the hardware execution of the Kyber rejection sampling.