Approximate computing has garnered significant attention due to its potential to reduce power consumption, enhance performance, and simplify circuit design. However, the security implications of applying approximate computing techniques remain largely unexplored. Identifying all possible vulnerabilities in approximate designs is very challenging. One of the challenges stems from the lack of insightful methodologies and metrics to perform a precise security evaluation. This paper presents ApproPower, a security-driven framework for pre-silicon evaluation of power side-channel leakage in approximate multipliers. ApproPower enables an analysis of how approximation techniques influence power side-channel leakage and employs symbolic path analysis to estimate delay-dependent power leakage behaviors at design time. Using a set of open-source approximate multipliers, we examine the relationship between data precision, approximate strategies, and measured power leakage. Our results show that symbolic path analysis provides useful guidance for identifying potential power side-channel risks. We also observe that some approximate designs can offer improved resource efficiency while exhibiting reduced leakage; for instance, in the mul7x7u set, a majority of benchmark circuits demonstrate lower leakage after approximation, with mul7x7u_03M achieving a 35% reduction in resource usage compared to mul8x7u_3C6.
This paper presents an inductor-first quadratic step-down (IFQSD) converter that achieves a voltage conversion ratio of D2 and a five-fold extension in the duty cycle compared with a conventional half-bridge converter. The continuous input current effectively alleviates electromagnetic interference (EMI) while maintaining an input current substantially lower than the load current, thereby reducing direct current resistance (DCR) losses. Additionally, the proposed converter achieves soft charging of the flying capacitor to reduce the losses caused by hard charging. Furthermore, an internal ripple compensation circuit is implemented to reduce output voltage ripple. To ensure proper operation of the power transistors under negative voltage conditions, a negative-voltage bootstrap circuit is proposed. The proposed converter is fabricated in a 0.18-μm Bipolar-CMOS-DMOS (BCD) process with a die area of 6.615 mm2. The experimental results validate that the proposed converter achieves an output voltage of 1 V with duty cycles of approximately 0.289 and 0.204 at input voltages of 12 V and 24 V, respectively. The implemented design achieves a peak power conversion efficiency of 90.4%.
Distributed Graph Neural Networks (GNNs) require efficient handling of both fine-grained memory accesses and cross memory-device communication, particularly when scaling to large graphs. However, existing acceleration solutions fail to adequately address low bandwidth utilization and scalability across different models and graph sizes. In this paper, we present OptGNN, a scalable heterogeneous distributed architecture tailored for GNNs. OptGNN addresses the challenges of fine-grained memory access by incorporating a near-memory processing mechanism, which improves internal bandwidth utilization. To optimize external communication, we introduce data packing and scheduling strategies that enhance cross memory-device data transfer efficiency. OptGNN achieves 5.7x performance improvement over baseline distributed GNN acceleration methods and 1.29x performance improvement over SOTA distributed GNN acceleration architecture CLAY. Additionally, the system is designed to support various GNN models and large-scale graphs while ensuring load balancing and high hardware utilization.
Distributed Graph Neural Networks (GNNs) require efficient handling of both fine-grained memory accesses and cross memory-device communication, particularly when scaling to large graphs in multimedia tasks. However, existing acceleration solutions fail to adequately address low bandwidth utilization and scalability across different models and graph sizes. In this paper, we present OptGNN, a scalable heterogeneous distributed architecture tailored for GNNs. OptGNN addresses the challenges of fine-grained memory access by incorporating a near-memory processing mechanism, which improves internal bandwidth utilization. To optimize external communication, we introduce data packing and scheduling strategies that enhance cross memory-device data transfer efficiency. OptGNN achieves 5.7x performance improvement over baseline distributed GNN acceleration methods and 1.29x performance improvement over SOTA distributed GNN acceleration architecture CLAY. Additionally, the system is designed to support various GNN models and large-scale graphs while ensuring load balancing and high hardware utilization.
This paper presents a double step-up-down (DSUD) hybrid converter with single-mode operation, designed for lithium-ion battery applications in mobile devices. By employing solitary $D$ -control, the converter achieves seamless buck and boost operations without mode transitions, allowing output voltage to be seamlessly adjusted to the target voltage. With the inductors permanently connected to the output, the inductor current is always equal to half of the load current, thereby reducing inductor DCR losses and power transistors conduction losses. Furthermore, it achieves continuous output current delivery, which minimizes output voltage ripple and eliminates the right-half-plane zero. Additionally, the introduction of adaptive ramp circuit with fast path improves the transient response, while the proposed dual-stage start-up circuit eliminates inrush current and allows the converter to gradually reach the steady-state. The proposed DSUD converter was fabricated using a 0.18- $\mu $ m BCD process. The chip is designed to deliver a regulated 3.3 V output with a maximum load current of 1 A from an input voltage ranging from 2.7 V to 4.2 V. Measured results show an output voltage ripple of less than 12 mV and a peak efficiency of 96.9%, achieved with $4.7~\mu $ H inductors, $2.2~\mu $ F output capacitor, and switching frequency of 1 MHz.
A global phase-shift control strategy for wireless power transfer system is proposed in this paper. Both the transmitter and the receiver can regulate the output power within a single resonant frequency by employing phase-shift control. In the receiver, the normalized phase shift (DRX) is adjusted to an appropriate value by an error amplifier. Meanwhile, a global power regulation is implemented by adjusting the phase shift of the transmitter (DTX) to constrain DRX within 0.5~0.8, ensuring that the receiver can maintain a high efficiency under a wide load range. To indicate whether the phase shift of the receiver is excessively high or low, a delay-load-shift-keying (DLSK) is presented. By modulating the turn-off time of the receiver, the transmitter can detect these two states to regulate transmitted power. The proposed control system is designed with 0.18μm BCD process. The output voltage is regulated to 5V with only 1.2mV voltage ripple. Simulation results show that the RX efficiency can be maintained above 90% under a load range of 15mA~150mA. Compared to the system without global phase-shift control, the light-load efficiency can be improved by about 25%.
This paper presents a dual-inductor, floating-ground topology for edge computing nodes. The proposed topology minimizes the number of off-chip components while achieving a 12-fold extension in duty cycle for 48-to-1 V conversion compared to the conventional half-bridge topology. To mitigate noise interference between two loads in edge computing nodes, floating ground is implemented without using a transformer. A comprehensive transfer function analysis is presented, demonstrating the robustness of the proposed topology. Furthermore, an eight-stage charge pump is integrated, yielding a 47.38% efficiency improvement for the on-chip voltage regulator. An accelerated assist circuit for the level shifter is proposed to reduce system delay. The proposed converter is fabricated in a 0.18-μm bipolar CMOS DMOS (BCD) process with a die area of 6 mm$^{2}$. Measurement results indicate a minimum duty cycle of 0.25 across a frequency range of 0.3-2 MHz and peak efficiencies of 87.91% and 91.58% at $V_{OUT}$=1 V and $V_{OUT}$=1.8 V, respectively.
This article proposes a symmetrical five-switch hybrid (SFSH) boost converter, which has the same voltage conversion ratio as the conventional boost converter and eliminates the right-half-plane zero. By making the output current path always include an inductor current path and a capacitor current path, the continuity of the output current is enhanced and the ripple of output voltage is reduced. In addition, the SFSH converter solves the flying capacitor inrush current problem in extreme duty cycle, further reducing the conduction loss in capacitor path. An experimental prototype with switching frequency of 200 kHz is designed to verify the performance of the proposed converter. The measurement results demonstrate that the SFSH converter achieves a peak efficiency of 94.5% under the condition of V-IN = 4 V, V-OUT = 6 V, and I-Load = 200 mA. Compared with conventional boost converter, the output voltage ripple is reduced by up to 40 mV. The recovery time is less than 120 mu s under the load transient between 100 and 500 mA.
This paper presents a power efficient twin-track charge balancing method for implantable biphasic neural stimulators. Using a low-power digital optimal-pulse-width searching (OPWS) technique, the required balancing charge can be calculated within only one stimulus cycle. Therefore, the power consumption of the balancer is reduced significantly. In addition, a dual-hysteresis-based window comparator is adopted in order to avoid frequently reactivating balancer in the steady state, which is caused by the residual charge accumulation. The proposed charge balancer has been fabricated in a $0.18 \mu \mathrm{m}$ BCD process. For a stimulus current switching between 1.35 mA and 1.65 mA, the residual charge is compensated fast, and the anodic pulse-width is calibrated within one stimulus cycle. The measured power consumption of the proposed CB is $3.4 \mu \mathrm{W}$.
Design space exploration is essential for optimizing deep neural network accelerators, which face increasing computational and energy demands as model complexity grows. Previous approaches rely heavily on local insights, often neglecting the need for extensive exploration to improve the global perspective. This leads to challenges such as blind exploration and a higher likelihood of getting trapped in local optima. Without dynamic adjustments or adaptive strategies, these methods struggle to navigate large, complex design spaces effectively. In this paper, we propose PRDSE, a design space exploration framework based on reinforcement learning that integrates both intrinsic and extrinsic metrics, guided by prior knowledge. The proposed method incorporates an adaptive adjustment mechanism that dynamically balances intrinsic and extrinsic rewards based on the progress of exploration, improving both search efficiency and optimization performance. Compared to state-of-the-art methods, PRDSE achieves substantial improvements, with latency speedups of up to 3.49x in the cloud environment and 3.43x in the edge environment, respectively. This work demonstrates that PRDSE effectively balances exploration and optimization objectives, providing a more efficient and scalable approach to design space exploration in accelerator design.
This letter presents a linear dynamic voltage scaling (DVS) technique using a dual-loop multistage charge-pump maintaining the minimum voltage headroom for implantable neurostimulation. By adopting an analog DVS with adaptive feedback divider, the stimulus current source could be always kept operating at the boundary of the saturation region and the linear region under different stimulus currents. Furthermore, a simple mode-switched control is introduced to improve the loop response of charge-pump. The design has been fabricated in a 0.18-mu m BCD process. The DVS technique increases the measured stimulus efficiency up to 52.8% higher than a fixed supply voltage with a peak efficiency of 89.6% in the range of the stimulus current from 0.5 mA to 0.9 mA.
This paper presents an inductor-first tri-path (IFTP) buck converter suitable for USB power delivery adapter to charge 1-2 cell battery. The proposed topology adopts the inductor-first strategy and the tri-path strategy of one inductor path and two capacitor paths to extend the output voltage conversion range, realize the continuous input current, eliminate the input EMI noise, and reduce the inductor current. In addition, a phase-interleaved symmetric inductor-first tri-path (PIS-IFTP) buck converter is proposed to alleviate the inrush current in the flying capacitor under extreme duty cycle of IFTP converter, while further reducing inductor current ripple. Two experimental prototypes for 9 V input to 3-8.4 V ouput have been developed, demonstrating excellent ability of IFTP and PIS-IFTP topologies to reduce inductor current and achieve continuous input current over the whole duty cycle and load range. The experimental results validate that the prototypes provide a wide voltage conversion range of 1/3-1 and a maximum output current of 1.8 A. The peak efficiency of IFTP is 93% at VOUT = 6.6 V, while the peak efficiency of PIS-IFTP is 94.5% at VOUT = 3.3 V.
Graph neural networks (GNNs) have emerged as a powerful paradigm for learning on graph-structured data. However, their computational patterns-characterized by irregular aggregation and node updates-pose significant challenges to modern hardware systems, especially in distributed settings where cross-device communication becomes a major bottleneck. This paper presents a formal modeling framework for mapping GNN algorithms onto distributed near-memory computing platforms, with a focus on resource allocation, redundancy elimination, and communication-computation overlap. We propose a holistic approach that includes analytical modeling of computational and communication costs, optimization strategies for load balancing and bandwidth utilization, and a performance evaluation model. Our work provides insights for hardware design space exploration and parallelization strategies in distributed GNN systems.
The design and optimization of deep neural network accelerators necessitates thoughtful consideration of numerous design parameters and various resource/physical constraints that render their design spaces massive in scale and complex in distribution. When faced with these massive and complex design spaces, previous works on design space exploration confront the exploration-exploitation dilemma, struggling to concurrently ensure optimization efficiency and stability. To address the exploration-exploitation dilemma, we present a novel design space exploration method entitled CSDSE. CSDSE implements heterogeneous agents separately accountable for exploration or exploitation to cooperatively search the design space. In order to enable CSDSE to adapt design spaces with various space distributions and expanding scales, we extend CSDSE with mechanism of adaptive agent organization and multi-scale search. Furthermore, we introduce a weighted compact buffer that encourages agents to search in diverse directions and bolsters their global exploration ability. CSDSE is implemented to optimize accelerator design. Compared to former DSE methods, it achieves latency speedups of up to 15.68x and energy-delay-product reductions of up to 16.22x under different constraint scenarios.
In this paper, a fully integrated low -dropout regulator (LDO) using voltage -to-time conversion (VTC) technique is presented for under-1 V supply voltage application. A synchronous VTC technique is proposed using constant-current (CC) charging and discharging to achieve high loop gain. A high gain charge pump (CP) is proposed to improve power-supply rejection (PSR). Furthermore, an asynchronous step detection recovery technique is proposed to achieve fast transient response. A frequency-adaptive oscillator is proposed to remove the noise of the clock signal. The proposed LDO is designed in 28-nm process to achieve a droop voltage of 104 mV at load current transient of 90 mA. The proposed LDO achieves PSR of -77 dB at I-LOAD=100 mA and PSR of -65 dB at I-LOAD)=10 mA for 1 -kHz supply ripple frequency. The quiescent current is 32 mu A and the peak current efficiency is 99.98%.
Wireless power transfer systems have been widely used in implantable medical devices (IMDs). In these systems, multiple supplies are required to power different modules. Generally, a high voltage is used for electrical stimulation, and a low voltage is used for power-hungry circuits including digital signal processor (DSP), telemetry, etc. In prior works, dual-output rectifiers have been well-established to meet the requirement for IMDs, and there are several typical structures. For example, a dual-output rectifier achieves two outputs by adding DC-DC converters in series with the rectifier. However, this structure occupies a large area and the cascade structure greatly reduces the receiver efficiency. Therefore, single-stage dual-output rectifiers (SSDORs) have been widely studied to reduce size and increase efficiency. In [1], a parallel-resonance full-bridge rectifier is proposed to generate dual outputs, but power transistors up to 6 are required. In [2], a series-resonance half-bridge rectifier with 4 transistors is proposed to achieve dual outputs. However, the existing methods are not capable of obtaining two voltages with a large voltage difference. Series-resonance is more appropriate to transfer high power and parallel-resonance is easier to generate a high output voltage [2]. Therefore, a hybrid-resonance rectifier is very suitable for IMDs.
Graph neural networks (GNNs) require a large number of fine-grained memory accesses, which results in inefficient use of bandwidth resources. In this article, we introduce a near-data processing architecture tailored for GNN acceleration, named NDPGNN. NDPGNN provides different operating modes to meet the acceleration needs of various GNN frameworks while ensuring the configurability and scalability of the system. NDPGNN takes advantage of data locality characteristics to repeatedly distribute and utilize data, thereby reducing memory access requirements, and further improving memory access efficiency by combining a subgraph sparse node scheduling strategy with intermediate result reuse. We use data packaging to provide a higher effective data ratio for long-distance data transmission, thereby improving the utilization of the system’s limited bandwidth resources. Compared with the previous method, NDPGNN brings 5.68 times improvement in system performance while reducing energy consumption overhead by 8.49 times.
In this article, vernier effect (VE), which is an effective method to enhance the measurement sensitivity of strain and temperature, was experimentally studied through cascading a Mach-Zehnder interferometer (MZI) with a Sagnac interferometer (SI). The MZI consisted of a tapered single-mode fiber (SMF), while the SI was formed using a panda-shaped polarization-maintained fiber (PMF). Free spectral ranges (FSRs) of the two interferometers were slightly different from each other, and as a result, the VE was successfully excited. The experimental results demonstrated that the measurement sensitivity of strain was improved from 26.7 to -110.92 pm/mu epsilon at the temperature of 24 degrees C, while the measurement sensitivity of temperature was improved from -2978 to 7934 pm/degrees C at the strain of 0 epsilon. The sensitivity was amplified about four times. The cross-sensitivity of temperature and strain measurements is 71.53 mu epsilon/degrees C. The proposed sensor exhibits in advantages of simple in construction, high sensitivity, and multiparameter measurements making it a superior contender for strain and temperature monitoring in engineering
Efficient design space exploration method is crucial for optimizing deep Convolutional Neural Network accelerators, as the large design space comprising hardware architecture and dataflow mapping parameters leads to prohibitive optimization time costs. Prior works employ space compression to improve exploration efficiency but focus only on single layers or partial spaces, failing to prevent expansion as network model depth increases. Therefore, in this brief, we propose the LCDSE framework to address the dilemma of space expansion. LCDSE clusters similar layers into blocks as coarse-grained optimization units to constrain parameters and prune redundant design space. Compared to former frameworks, LCDSE achieves up to 4xdelay speedup, 7.9xenergy-delay-product decrease and 2.5xarea-delay-product reduction along with time cost saving by up to 11.3x, demonstrating its superior optimization efficiency.