To maximize power efficiency and system simplicity, blind equalization is widely utilized in high-speed intensity modulation/direct detection (IM/DD) systems, enabling the cost-efficient deployment in data center interconnects (DCI). Meanwhile, recognized as the theoretically optimal equalizer, maximum likelihood sequence estimation (MLSE) is recommended for broad implementation, particularly in high-speed IM/DD systems. However, conventional MLSE faces challenges due to its inability to blindly estimate both linear and nonlinear channel distortions, coupled with its high implementation complexity. This Letter proposes and experimentally demonstrates a joint blind equalization scheme featuring zero-forcing nonlinear reduced-state sequence estimation (ZF-NL-RSSE) in a real-time 80-Gbps PAM-4 transmission system based on an FPGA chip. Experimental results show that over 5-km single-mode fiber (SMF) transmission, the ZF-NL-RSSE scheme achieves rapid blind convergence, attaining a BER performance superior to traditional FFE and VDFE and comparable to the full-state ZF-NL-MLSE, while significantly lowering computational complexity.
The floating-point multiplier is an important coprocessor component of modern microprocessors and is the core of real-time image processing and deep learning computing. Compared to fixed-point multipliers, the dynamic range is wider, but the complexity is higher. With the rapid development of autonomous driving, artificial intelligence and the Internet of Things, the massive data and complex calculations required for these applications put forward high requirements for the design of the operation unit circuit, in which the multiplier consumes more resources and becomes the bottleneck of the operation unit design. This makes high-efficiency and low-power multipliers the focus of circuit design in recent years. High performance and low power consumption have become the trend in combined cell circuit design for edge computing applications. In this paper, an effective half-precision approximate floating point multiplier design for edge computing (MP4-App-Mul) is proposed. The architecture is based on shift-add approximation (SSA) and sparse processing to optimize power consumption and performance. Compared with the exact multiplier, the proposed approximate multiplier (MP4-App-Mul) has significant advantages in hardware performance, with 86.5
Compute Express Link (CXL) has emerged as a key enabler of memory disaggregation for future heterogeneous computing systems to expand memory on demand and improve resource utilization. However, CXL is still in its infancy stage and lacks commodity products on the market, thus necessitating a reliable system-level simulation tool for research and development. In this article, we propose CXL-DMSim (open-sourced at https://github.com/ferry-hhh/CXL-DMSim), an open-source full-system (FS) simulator to simulate CXL disaggregated memory systems with high fidelity at a gem5-comparable simulation speed. CXL-DMSim incorporates a flexible CXL memory expander model along with its associated device driver, and CXL protocol support with CXL.io and CXL.mem. It can operate in both the app-managed (AM) mode and the kernel-managed (KM) mode, with the latter using a dedicated NUMA-compatible mechanism. The simulator has been rigorously verified against a real hardware testbed with both FPGA- and ASIC-based CXL memory devices, which demonstrates the qualification of CXL-DMSim in simulating the characteristics of various CXL memory devices at an average simulation error of 3.4%. The experimental results using LMbench and STREAM benchmarks suggest that the CXL-FPGA memory exhibits a $\sim 2.88\times $ higher latency than local DDR, while the CXL-ASIC latency is $\sim 2.18\times $ ; CXL-FPGA achieves 45%-69% of local DDR memory bandwidth, whereas the number for CXL-ASIC is 82%-83%. The study also reveals that CXL memory can significantly enhance the performance of memory-intensive applications, improving by $23\times $ at most with limited local memory for Viper key-value database and approximately 60% in memory-bandwidth-sensitive scenarios such as MERCI. Moreover, the simulator's observability and expandability are showcased with detailed case studies, highlighting its great potential for research on future CXL-interconnected hybrid memory pools.
This paper presents an energy-efficient adaptive voltage and frequency scaling (AVFS) system for high-performance processors. The system is designed based on a fast transient response digital low-dropout regulator (DLDO) and a self-calibrating elastic frequency regulator. A fully synthesizable voltage comparator is designed to provide fast voltage resolution via delay differences in DEL cells under different voltage domains. A current-hungry adaptive sampling clock (CKs) generator is designed to provide fast and slow frequencies for voltage comparison and regulation, which effectively reduces the system power. The processor frequency is regulated by a tunable ring oscillator, which is continuously self-calibrated by in-situ critical path timing detection. The proposed system is implemented and simulated in 28-nm CMOS technology, with an area of 0.025 mm2. The system exhibits a voltage droop of 88 mV and a settle time of 15 ns when the load current is increased by 90 mA with a 10 ns edge time. Compared to the baseline design, the droop magnitude and response time are reduced by 30% and 53% separately. The maximum load current and peak current efficiency are 120 mA and 99.8%, respectively. The total energy consumption of the AVFS system is 0.56 pJ/cycle at 0.8 V.
With the growing demand for data center interconnect (DCI) capacity, partial response equalization (PRE), employing either a 1/(1+D) decoder or an maximum likelihood sequence estimation (MLSE) decoder, has emerged as an effective solution to address the critical bandwidth limitations restricting system transmission rates. However, the 1/(1+D) decoder suffers from its inherent feedback loop and error propagation, while the MLSE decoder is limited by its high hardware complexity. Meanwhile, blind adaptive algorithms also encounter bottlenecks in PRE. Severe bandwidth limitations significantly impact the blind convergence of the decision-directed least mean squares (DD-LMS) algorithm. Additionally, PRE adversely affect the traditional zero forcing (ZF) algorithm, leading to extended convergence times. In this paper, we propose and experimentally demonstrate a hardware implementable PRE with reduced-state sequence estimation (RSSE) decoder and delay-ZF algorithm in a real-time 80 Gbps four-level pulse amplitude modulation (PAM- 4) intensity modulation/direct detection (IM/DD) transmission system based on an FPGA chip. Statistical results from various transmission scenarios indicate that the RSSE decoder achieves BER performance comparable to that of the full-state MLSE decoder, while reducing the implementation complexity by up to 75%. Moreover, experimental results demonstrate that under blind equalization, employing the delay-ZF algorithm to update the PAM-7 targeted FFE reduces the blind convergence time by up to 55.8% and 74.9% compared to the traditional ZF and DD-LMS algorithms, respectively.
Merge-based sorting architectures are widely used in FPGA sorting accelerators due to their scalability and high throughput. However, their execution is fundamentally constrained by run-drain stalls, where merge-tree leaves must fully drain their current sorted runs within the same level before admitting new data. This run-level synchronization breaks steady-state pipelined execution and limits sustained throughput under streaming workloads. This paper presents TL-Sort (Triple-Layer Sort), a fully pipelined merge-based sorting architecture that eliminates run-drain stalls. To achieve this, TL-Sort introduces a pipelined virtual tree (PVT) that first explicitly encodes inter-run ordering, thereby eliminating run-level synchronization. Compared with a virtual merge sorter tree, PVT reduces total sorting cycles by up to 2.57× and improves throughput by up to 2.52× with negligible hardware overhead. When sorting 32,768 64-bit records in a multi-pass setting, TL-Sort improves throughput from 0.697 GB/s to 1.474 GB/s compared with a state-of-the-art open-source FPGA sorter.
This paper presents a 7 × 25Gb/s wireline transceiver (TRX) for high-density transmission. A codebook-based correlated none-return-to-zero (CNRZ-7) modulation scheme is employed to achieve high output bandwidth while maintaining strong immunity to crosstalk (XT), simultaneous switching noise (SSN), and common-mode noise (CMN). In the receiver (RX), a single-ended Q-shaped continuous-time linear equalizer (CTLE) replaces the conventional multi-input equalizer to suppress additional noise introduced by finite common-mode gain (ADM-CM). The proposed TRX is fabricated in a 28-nm CMOS process. Measurement results show that the TRX achieves a vertical eye opening of 90 mV and a horizontal eye opening of 27 ps at 25 Gb/s per lane, corresponding to a bit-error rate (BER) o 1 x 1012.
The explosive growth of key-value (KV) cache size in large language model (LLM) inference poses a key challenge to the limited HBM of GPU. Offloading KV cache to host memory has become a prevalent mitigation method. However, the limited host DDR bandwidth, especially in multi-GPU inference scenarios, often leads to offloading bottlenecks, thereby restricting inference speed. Compute express link (CXL) offers a promising alternative to expand host memory capacity and bandwidth on demand. In this article, we present a bandwidth-oriented memory allocation mechanism, named AdaptiveKV, which is self-adaptive to CXL-enabled memory pools and KV cache offloading scales for LLM inference acceleration. Our systematic profiling of CXL-HBM memory bandwidth under GPU workloads reveals that conventional memory strategies neglect dynamic memory bandwidth fluctuations and various CXL memory characteristics, leading to suboptimal memory utilization. Motivated by these insights, AdaptiveKV implements three core designs: (1) a GPU memory conch model to guide memory allocation strategies, (2) a runtime predictor to predict optimal memory allocation ratios, and (3) a dynamic interleaving strategy to allocate memory pages across available NUMA nodes. Experimental results suggest that AdaptiveKV achieves a maximum speedup of 1.90× in LLM inference throughput compared with the state-of-the-art strategies. To further explore AdaptiveKV’s applicability boundary, we also present an FPGA-based CXL memory emulator with configurable performance, revealing that a CXL-to-DDR bandwidth ratio exceeding 8% yields at least a 5% speedup in LLM inference.
Protein spatial structure comparison plays a pivotal role in understanding biological functions, evolutionary relationships, and disease mechanisms. The recent advancements in AI-based prediction tools, such as AlphaFold, have led to the generation of an unprecedented number of protein structures, dramatically expanding structural databases. However, this exponential growth has exposed substantial computational bottlenecks in conventional comparison tools (e.g., CE, Dali), which suffer from poor scalability and efficiency. Even state-of-the-art approaches like ADAMS struggle to process large-scale, continuous structural data effectively. To this end, We propose PSCA, the FPGA-based accelerator specifically designed for protein structure comparison. By employing a non-stagnant sliding window, symmetry simplification and a series of dedicated hardware-optimized designs, we enhance the fluidity and parallelism of the ADAMS algorithm, achieving significant speedups in both descriptor database construction and matching phases. Experimental results conducted on a database comprising 12,414 human proteins and 15,059 C. elegans proteins demonstrated that PSCA achieves 21.4× - 29.9× speedup in database construction and 18× - 265× in matching process compared to the original Golden ADAMS software.
With the large-scale deployment of supercomputing and smart computing centers, high-speed optical devices have become key nodes governing high-performance communications in these systems. The packaging quality of micro-optical components in optical devices directly determines their communication performance. In this article, a new microlens packaging strategy for high-speed optical devices is proposed. First, a multilevel aggregation network (MANet) is constructed for efficient microlens detection, achieving a posture recognition accuracy of +/- 0.15 degrees and an average processing time of 1.2 s. Then, an optical-bus-controlled microlens grasping system and a ceramic gripper are developed, achieving a microlens grasping success rate of more than 95%. Finally, a packaging system is designed and implemented, and the packaged device is experimentally evaluated. Experimental results demonstrate a microlens packaging success rate of 92.64%, a power fluctuation of less than 1.8%, and an extinction ratio fluctuation of less than 0.25 dB.
To address the severe intersymbol interference (ISI) in high-speed optical transmission systems, this paper proposes a memory-augmented neural network equalizer (MANNE). Building upon the architecture of a traditional neural network equalizer (NNE), the scheme introduces an external memory module based on a key-value structure, which dynamically stores and updates the central representations of various 4-level pulse amplitude modulation (PAM-4) symbols in the feature space and their associated uncertainties during network inference. Furthermore, a memory feedback mechanism is incorporated, enabling the MANNE to reference its memory during training, thereby enhancing its feature representation and generalization capabilities. To validate the effectiveness and practicality of the proposed scheme, C-band transmission experiments of a 128 Gbaud PAM-4 signal are carried out over a 2 km standard single-mode fiber (SSMF) transmission link. Experimental results show that at the KP4 forward error correction (KP4-FEC) threshold of 2.4E-4, the receiver sensitivity of MANNE is improved by approximately 1.0 dB compared to the traditional NNE. The Volterra nonlinear equalizer (VNLE) failed to reach this threshold, whereas MANNE not only meets the requirement but also reduces computational complexity by 72.2% compared with VNLE, achieving a good balance between performance and computational efficiency.
This paper presents a ring-amplifier-based pipeline-SAR ADC that incorporates both optimized detect-and-skip (DAS) and weighted-averaging correlated level shifting (WACLS) to improve its power efficiency and accuracy. In the first stage, a tiny level-shifting capacitor array with simple control logic is introduced to replace the conventional large gain-mismatch compensation capacitor. It compensates for the gain mismatch between the coarse and fine quantization DACs by properly shifting the DAC output without any area, power, or speed penalty. The WACLS and second-stage DAS are co-designed by rearranging the second-stage capacitor array into the coarse and fine DACs required for WACLS, while reusing the coarse DAC for the DAS operation. This combined approach simultaneously increases the ring amplifier's equivalent gain, enhances the second-stage operation speed, and reduces the second-stage DAC switching power. Fabricated in 28-nm CMOS, the ADC consumes 2.58 mW from a 1-V supply and occupies 0.01-mm2 core area while achieving 63.16-dB signal-to-noise and distortion ratio and 80.1-dB spurious-free dynamic range at Nyquist, resulting in Schreier and Walden figure of merit of 175.0 dB and 5.49 fJ/conv-step, respectively.
This paper reports a correlated PAM4 (CPAM4) wireline transceiver featuring superior crosstalk cancellation (XTC) and insertion loss (IL) compensation for high-density interconnect. Specifically, thanks to the CPAM4 data transfer matrix, the transmitter (TX) driver reduces the simultaneous switching noise (SSN) and the receiver (RX) eliminates the common-mode noise (CMN) using a reference-less analog front-end (AFE). Prototyped in 28-nm CMOS, the proposed CPAM4 transceiver achieves a data rate density (DRD) of 157 Gb/s/mm and energy efficiency of 1.55 pJ/bit.
Large Language Models (LLMs) built on Transformer architectures have achieved remarkable success in Natural Language Processing (NLP) applications. However, as the scale of model parameters continues to expand, deployment becomes increasingly challenging, particularly for resource-constrained edge devices. Emerging dense architectures such as Monarch Mixer, offer a promising alternative by achieving superior accuracy with fewer parameters and reduced computational demands compared to conventional transformers. Nevertheless, Monarch Mixer’s memory-intensive operators result in inefficient hardware utilization and increased latency, especially in small-scale text inference scenarios. To address this limitation, we propose FPAMM, a fine-grained pipeline accelerator that exploits Monarch Mixer’s operator dependencies through a fine-grained pipeline architecture. Our approach incorporates key optimizations such as operator fusion and dual buffering to enhance data flow efficiency. Concurrently, we optimized resource allocation and ensured load balancing through design space exploration. Deployed on Xilinx Alveo U280 platform, FPAMM achieves 2.21–40.42 × speedup for a 12-layer M2-BERT compared to CPU and GPU baselines. When evaluated against existing FPGA-based Transformer accelerators, FPAMM delivers superior performance, with 7.5 × , 8.35 × and 2.28 × improvements in throughput compared to UPSA, TRAC and FQ-BERT respectively, while maintaining the lowest latency. Through innovative fine-grained pipelining and operator fusion techniques, FPAMM mitigates the memory overhead inherent in Monarch Mixer. Experimental results confirm FPAMM’s potential as a promising solution for low-latency, high-throughput edge deployment.
Conventional heterogeneous computing systems built on PCIe interconnects suffer from inefficient fine-grained host-device interactions and complex programming models. In recent years, many proprietary and open cache-coherent interconnect standards have emerged, among which compute express link (CXL) prevails in the open-standard domain after acquiring several competing solutions. Although CXL-based coherent heterogeneous computing holds the potential to fundamentally transform the collaborative computing mode of CPUs and XPUs, research in this direction remains hampered by the scarcity of available CXL-supported platforms, immature software/hardware ecosystems, and unclear application prospects. This paper presents Cohet, the first CXL-driven coherent heterogeneous computing framework. Cohet decouples the compute and memory resources to form unbiased CPU and XPU pools which share a single unified and coherent memory pool. It exposes a standard malloc/mmap interface to both CPU and XPU compute threads, which share a single per-process page table for user applications, leaving the OS dealing with smart memory allocation, page auto-migration, and management of heterogeneous resources. This design significantly simplifies heterogeneous parallel programming to a level comparable to homogeneous programming. To facilitate Cohet research, we also present a full-system cycle-level simulator named SimCXL, which is capable of modeling all CXL sub-protocols and device types. SimCXL has been rigorously calibrated against a real CXL testbed with various CXL memory and accelerators, showing an average simulation error of 3%. Our evaluation reveals that CXL.cache reduces latency by 68% and increases bandwidth by 14.4x compared to DMA transfers at cacheline granularity. Building upon these insights, we demonstrate the benefits of Cohet with two killer apps, which are remote atomic operation (RAO) and remote procedure call (RPC). Compared to PCIe-NIC design, CXL-NIC achieves a 5.5 to 40.2x speedup for RAO offloading and an average speedup of 1.86x for RPC (de)serialization offloading.
This work presents a codebook-based correlated non-return-to-zero-7 (CNRZ-7) transceiver for high-density printed circuit board (PCB) and substrate links. The proposed signaling scheme maps seven input bits onto eight wires, achieving 87.5% pin efficiency while enforcing a zero-sum constraint across the transmitted wire levels. This coding property provides differential-like suppression of simultaneous switching noise (SSN) and common-mode noise (CMN), and exhaustive transition analysis shows a 33% reduction in worst-case crosstalk-induced jitter (CIJ) compared with single-ended non-return-to-zero (SE-NRZ) signaling. To support high-speed implementation, the transmitter performs codebook mapping at a low-speed stage and drives the channel using eight matched four-level output drivers, reducing high-speed routing complexity and output interconnect parasitics. At the receiver, the multi-input comparator (MIC) is decoupled from the continuous-time linear equalizer (CTLE) to mitigate output degradation caused by common-mode fluctuation in CNRZ-7 pseudo-differential decoding groups. An auxiliary-zero-pair Q-shaped CTLE (AZP-QCTLE) is further introduced to enhance frequency shaping and increase peaking gain. Building on our prior symmetric correlated coding (SCC) transceiver, this work adopts low-speed TX codebook encoding and an RX AFE with reduced CM-to-DM noise conversion, raising the per-lane data rate to 25 Gb/s. The prototype transceiver was fabricated in 28-nm CMOS and characterized over a compact high-density PCB channel. At an aggregate throughput of $7 \times 25$ Gb/s, the proposed CNRZ-7 mode achieves BER down to $10^{-12}$ with a worst-case sampling margin of 0.13 UI, whereas an eight-lane SE-NRZ baseline operated at the same aggregate throughput cannot reach the same BER target. The measured data-rate density is 82 Gb/s/mm, and the energy efficiency is 1.77 pJ/bit.
The traditional 10G/G combo passive optical network (PON) is no longer able to meet the growing traffic demand of the access network, and 50G PON communication has become one of the reliable solutions. In this paper, we study the optical path of a compact wavelength division multiplexing (WDM) optoelectronic device for 50G/10G/G combo PON, which solves the smooth evolution challenge of the coexistence of three rates in an existing network. Firstly, a 50G combo PON beam transmission model with high optical coupling efficiency is established and simulated for optimization. Secondly, the optical coupling law of each laser diode (LD) of the 50G combo PON is studied and analyzed to provide the basis for process guidance for subsequent device packaging. Finally, the performance of the packaged devices is tested and analyzed to study the high-fidelity performance of signal transmission. Experimental results show that the 50G combo PON has a maximum coupling efficiency of 36.7% at total power, a 5% power loss coupling tolerance of +/- 1.5 mu m, and the ability to realize complete signal transmission at three transmit wavelengths 1342 nm, 1490 nm, and 1577 nm through power equalization.
Quantum error correction (QEC) is indispensable for building scalable fault-tolerant quantum computers. Effective QEC demands stringent real-time decoding: the decoder must process syndrome measurements and determine corrections within a time scale–typically on the order of microseconds, to avoid data backlog. Scaling to large number of logical qubits further necessitates significant computational resources. In this work, we propose an architecture, called THQLink, for real-time decoding of quantum error correction codes using high-performance computing (HPC) resources. The network connecting the HPC and the control system of quantum processing unit (QPU) is built on TH-Express and can be adapted to different quantum technologies and their associated control stacks. We report a round-trip latency of 2.944 μs on average, with an incremental overhead of 130 ns per additional hop. Using a parallel window strategy, we demonstrate real-time decoding (1 μs per QEC round) of the surface code up to distance 19 using a matching-based decoder on CPUs. Our work presents a scalable framework for real-time decoding in fault-tolerant quantum computing. It can be readily applied to quantum-centric supercomputers that feature tight integration between QPU and HPC resources, thereby enabling efficient support for hybrid quantum-classical algorithms and computation-intensive workloads offloaded from the QPU.
This paper presents a low power consumption, high current efficiency and fast response low-dropout voltage regulator (LDO) without external capacitors (capless) working at cryogenic temperature (4K). The flipped-voltage-follower (FVF) structure achieves large bandwidth with low power consumption. The voltage detecting unit provides not only extra current but also low voltage overshoot and undershoot when the load steps. The gain bandwidth (GB) enhancing structure with active inductors is adopted to improve loop bandwidth to realize better transient response. The proposed LDO is fabricated in a 28-nm CMOS process, occupying only 0.0026mm2. Measured results show that the LDO provides a voltage of 0.9V and a quiescent current of 15.1μA achieving current efficiency of 99.94 % and loading 25mA at 4K. It achieves 48mV undershoot and 45mV overshoot with a step of 24.95mA/10ns. The results show that the LDO can power circuits at cryogenic temperature and can be applied to large-scale quantum computers.
This paper presents an output capacitor-less low dropout regulator (LDO) with mismatch cancellation for distributed power management. Compared with the traditional LDO, the proposed LDO has an embedded reference and uses current feedback to alleviate the limited vertical voltage margin. Dynamic element matching (DEM) and chopping techniques are adopted to reduce the mismatch between the current replication and thereby reduce the mismatch in output voltage. The LDO is implemented in a 28-nm CMOS process. Simulation results demonstrate that the quiescent power consumption is 9.6μA when power supply is 0.9V and output voltage is 0.6V. It also has 84.9mV undershoot and 70.5ns recovery time when the load steps from 100μA to 10mA in 50ns. Monte-Carlo simulations of output voltage performed on 1500 samples show that the standard deviation decreased by nearly 25 times with DEM and chopping.