This paper presents a critical-timing-relaxed pre-decision feedforward equalizer (pDFFE) designed for high-speed digital signal processor (DSP)-based wireline receivers. Since conventional decision feedback equalizer (DFE) suffers from an increasing timing burden due to recursive feedback loops, which hinder parallelization and scalability, the proposed pDFFE architecture eliminates the feedback path by replacing post-cursor inter-symbol interference (ISI) cancellation with feed-forward equalizer (FFE)-based decisions. The overall system employs a hardware-in-the-loop (HIL) verification platform based on the Versal ACAP, integrating a DSP engine with sign-sign LMS (SS-LMS) adaptation. The power and area of the proposed DSP in 28nm CMOS layout are estimated to be 176.5mW and 0.201mm2, respectively. Both simulation and measurement results under −24dB channel loss at Nyquist frequency demonstrate reliable BER performance comparable to the conventional DFE, while achieving single-clock latency and improving area efficiency, from a minimum of 10 % to a maximum of 47 %, compared to conventional architectures.
While non-return-to-zero (NRZ) signaling and 4-level pulse amplitude modulation (PAM-4) with analog equalizers have served earlier generations of serializer-deserializer (SerDes), they become inadequate for communicating over lossy electrical channels at beyond 100 Gb/s datarate. Consequently, analog-to-digital converter (ADC) and digital signal processor (DSP)-based receivers have emerged as the dominant architecture for long-reach applications with > 100 Gb/s, offering powerful equalization and support for advanced algorithms such as sliding-block DFE (SB-DFE) and maximum-likelihood sequence estimation (MLSE). This paper briefs recent advances in ADC-DSP-based wireline receivers, highlighting overall system and DSP equalizers. In addition, discrete multitone (DMT) modulation is discussed as a promising candidate for next-generation interconnects, with multiple proof-of-concept implementations in advanced CMOS nodes.
This paper presents a SystemVerilog-based modeling and simulation methodology for a 4-level pulse-amplitude-modulation (PAM-4) transceiver. The framework is developed to address verification challenges in complex analog-to-digital converter (ADC)-based mixed-signal systems, where conventional circuit-level simulations suffer from prohibitively long execution time. Two key modeling techniques to significantly improve simulation efficiency are introduced in this work. For the time-interleaved ADC (TI-ADC), a counter-based equivalent circuit model is employed to capture dominant rank-1 mismatch errors while substantially reducing modeling complexity. In addition, for digital clock and data recovery (CDR), a statistical bit-error-rate (BER) evaluation approach based on XMODEL primitives is adopted as an alternative to conventional time-domain analog–mixed-signal (AMS) simulations. This modeling framework enables accurate verification of critical transceiver performance metrics, such as jitter tolerance (JTOL), within a 30-minute simulation window. Compared to conventional timedomain simulations, which require more than 12 hours to perform, the proposed approach achieves approximately a 24× improvement in simulation speed.
This paper proposes a decision-error correction scheme that improves the pre-forward error correction (pre-FEC) bit-error rate (BER) by identifying and correcting symbol errors after the decision feedback equalizer (DFE) without introducing any coding overhead. The proposed method reconstructs the received waveform to localize error-prone symbols without requiring additional redundancy and performs symbol-level correction by exploiting intrinsic properties of pulse-amplitude modulation (PAM) signaling. The error detection and correction scheme is implemented using simple digital arithmetic and logic and completes the correction process within three digital signal processor (DSP) clock cycles, thereby avoiding excessive latency. A 4-level pulse-amplitude modulation (PAM-4) DSP employing a 64-way parallel datapath was implemented in a 14-nm FinFET process, occupying 0.269 mm2 and targeting an 875 MHz clock frequency. The design was validated on a Xilinx ZCU111 RFSoC platform, demonstrating a BER improvement of more than five orders of magnitude at 4.8 Gb/s over a channel with 22.9 dB insertion loss (IL) at Nyquist.
This article presents a 76-Gb/s digital-to-analog converter (DAC)-based discrete multitone (DMT) wireline transmitter (TX) fabricated in 5-nm FinFET. The TX employs a 16-way parallel multi-path delay feedback (MDF) inverse fast Fourier transform (IFFT) processor for area-efficient implementation. The TX digital signal processor (DSP) includes on-chip bit/power-loading and cyclic-prefix (CP) insertion logic and supports 4- to 256-QAM modulation formats across 31 orthogonal subchannels. The time-domain DMT samples generated by the TX DSP are converted into analog waveform using a source-series termination (SST)-based DAC. The prototype is demonstrated over a channel with 9.7-dB insertion loss (IL) at Nyquist, achieving a bit error rate (BER) of 2.1E-4 with a total power consumption of 144 mW from 0.735-V digital and 0.725-V analog supplies, resulting in an energy efficiency of 1.89 pJ/b. The proposed DMT TX improves bandwidth efficiency and signal-to-noise ratio (SNR) over conventional pulse-amplitude modulation (PAM)-based transmitters by employing frequency-domain modulation and equalization. This work is the first demonstration of a DAC-based DMT TX operating at >76 Gb/s fabricated in advanced CMOS technology.
This paper analyzes the fixed-point (FXP) implementation of the maximum-likelihood sequence estimation (MLSE) engine for high-speed PAM-4 wireline transceivers (TRXs). Quantization effects in the ADC/FFE output, branchmetric (BM), and path-metric (PM) computations are systematically evaluated through end-to-end BER simulations at 118Gb/s over a 28.5dB-loss channel. Comprehensive precision sweeps show that insufficient quantization at early DSP stages, particularly in the FFE output and BM computation, causes irreversible BER degradation even when later stages maintain high precision. The analysis identifies the dominant precision bottlenecks and provides quantitative design guidelines for balancing BER performance and hardware efficiency in MLSE receiver DSP implementations for future high-speed wireline transceivers.
This paper presents a $76 ~\text{Gb} / \mathrm{s}$ digital-to-analog converter (DAC)-based discrete multitone (DMT) wireline transmitter (TX) fabricated in 5 nm FinFET. Bit and power loading with 32/64/128-QAM across 31 orthogonal subchannels is demon-strated over a channel with 9.7 dB insertion loss (IL), achieving a bit error rate (BER) of 2.1 E-4. The prototype consumes 144 mW from 0.675 V digital and 0.725 V analog supplies, resulting in an energy efficiency of $1.89 \text{pJ} / \mathrm{b}$. An on-chip DSP performs subchannel-wise bit/power allocation and spectral shaping using a 64-tap inverse fast Fourier transform (IFFT) and cyclic prefix (CP) insertion. Compared to conventional PAM-based TXs, the proposed architecture provides improved bandwidth efficiency and signal-to-noise ratio (SNR) through frequency-domain modulation and equalization. This work is the first demonstration of a DAC-based DMT TX at $76 ~\text{Gb} / \mathrm{s}$ data rate fabricated in advanced CMOS technology.
This article presents an 86.71875-GHz RF transceiver IC featuring a fully integrated clock and data recovery (CDR)-assisted carrier synchronization loop (CSL) for waveguide links. The carrier frequency of 86.71875 GHz is chosen to be the third harmonic of the baseband null frequency of 28.90625 GHz, and the carrier synchronization is achieved using a baseband CDR instead of power-and area-intensive RF circuits. The IC, fabricated in 28-nm CMOS, demonstrates 57.8125-Gb/s pulse-amplitude modulation-4 (PAM-4) data transmission over a 1.5-m waveguide channel while improving the timing margin by 38% compared to conventional methods. The Tx and Rx ICs occupying an area of 1.98 x 0.95 mm(2) consume 190.1 and 117 mW at 57.8125 Gb/s, respectively. The test chip achieves the figure of merit (FoM) of 3.5 pJ/b/m in terms of throughput-distance and energy efficiency.
Multi-agent (MA) simultaneous localization and mapping (SLAM) has been rigorously explored to enhance map accuracy in swarm robotics. Although centralized MA SLAM systems, which depend on a server for complex computations in map optimization, have been extensively studied, the circuit-domain approaches to decentralized MA SLAM systems are still limited due to challenges such as limited memory capacity and security vulnerabilities in wireless inter-agent data transmission. Thus, we propose a BEE-SLAM accelerator, a location-sharing MA neuromorphic SLAM accelerator inspired by bee communication for decentralized MA SLAM systems. The location-sharing-based MA error correction (MAEC) is employed to attain accurate map results without loop closure with a 94.81% reduced number of operations compared to the global map-based MA SLAM. In addition, a 7 x 7 pulsewidth modulation (PWM)-based hybrid mixed-signal/digital pose-cell (HY-PC) array with pseudo pose cells (PPCs) achieves 2.04 x energy efficiency compared to the oscillatory pose-cell array. The test chip fabricated in a 65-nm CMOS technology achieves a peak energy efficiency of 17.96 TOPS/W under 350 x 450 m outdoor exploration.
This paper presents a reconfigurable and energy-efficient digital spectrum shaping signaling (DSSS) for multidrop interfaces, where the output spectrum of the transmitted data is shaped using the 2-times repetitive block transmission to avoid frequency notches in the multidrop channel, thereby achieving a data rate up to 4x the first channel notch frequency. In conventional wireline transceivers (TRX), compensating for frequency notches requires a large number of decision feedback equalizer (DFE) taps at the receiver, resulting in significant area and power overhead. In contrast, the proposed DSSS architecture supports spectrum-efficient reconfigurable dual-mode NRZ/PAM4, reducing required equalization efforts and improving energy efficiency. The proposed scheme and its transmitter (TX) were first validated through event-driven behavioral simulations using XMODEL and verified with equipment-based measurements. Post-layout simulation results in 28nm CMOS process demonstrated 4 Gb/s data rate communicating over a channel having its first > 30 dB notch at 1 GHz, with a 235mV vertical eye opening and a TX energy efficiency of 0.39 pJ/b at 0.8V supply.
This article presents a 2-lane $2 \times 2$ multiple-input, multiple-output (MIMO) 4-level pulse amplitude modulation (PAM-4) minimum mean-squared-error (MMSE)-decision-feedback equalizer (DFE) with far-end crosstalk (FEXT) cancellation for digital-to-analog converter (DAC)-/analog-to-digital converter (ADC)-based high-speed serial links. The receiver (RX) datapath is designed with a 15-tap MIMO feedforward equalizer (FFE) and a one-tap MIMO DFE with the least mean square (LMS), enabling adaptation to channel variation while maintaining the MMSE setting. The RX digital signal processor (DSP) place and route (PnR) in a 28-nm CMOS is estimated to consume 201 mW/lane at a 56-Gb/s/lane data rate while occupying a 0.5-mm2/lane silicon area. We further implement a real-time evaluation platform to verify the functionality of the MIMO PAM-4 MMSE-DFE with rapid bit-error-rate (BER) testing on RFSoC. The measurement result demonstrates that the MIMO MMSE-DFE significantly improves BER performance from 2.75e−3 to 1.31e−7 compared with equalization without FEXT cancellation when communicating over a channel exhibiting 12.4-dB insertion loss (IL) and 13.2-dB IL-to-crosstalk ratio (ICR) at Nyquist.
The growing demand for higher communication bandwidth between processors through wired interconnects in large-scale servers has been driving the need to increase the per-lane data rate beyond the current 112Gb/s. Recently demonstrated analog-to-digital converter (ADC)-based receiver (RX) prototypes with>100Gb/s data rate typically employ a parallel feed-forward equalizer (FFE) with a large number of taps, 1-tap decision feedback equalizer (DFE) [1]–[5], and maximum likelihood sequence estimator (MLSE) as option [6]–[8]. As the data rate grows exponentially, the pulse response length and the number of corresponding inter-symbol interference (lSl) cursors increase accordingly [5], [8]. As the length of the pulse response gets doubled, the FFE tap count also needs to be increased accordingly, which results in substantial area and power overhead. The DFE feedback loop timing closure also gets more stringent as Baudrate increases [9]. With an increased pulse amplitude modulation (PAM) order, the DFE and MLSE design complexity increases exponentially [6]–[8]. While $\mathrm{a} > 100\text{Gb}/\mathrm{s}$ PAM-4 transceiver (TRX) can effectively equalize smooth channels [2]–[5], ripples and notches in the frequency response of the channel can significantly degrade the equalization performance of the current PAM-4 TRX.
We propose a fast algorithm to optimize on-chip equalizing link design utilizing a particle swarm optimization (PSO) method. Finding the optimal design parameters of an equalizing link requires too much computation time, because the dependency between design parameters and performances is too complex, while design space is too large. The proposed algorithm greatly reduces the optimization time by utilizing the superior efficiency of PSO in heuristic search. In experiment, on average, the proposed algorithm optimized a link design 168 $\times$ faster than the previous state-of-the-art result, requiring only 1/256 evaluation counts, and reduced computation time from about 2 h to 45 s.
This paper presents an energy-efficient implementation of frame-level synchronization for discrete multitone (DMT) wireline transceivers (TRX). Simplifying arithmetic precision for computing cross-correlation (Xcorr) can increase the required length of the synchronization sequence (SS) to achieve the target false positive (FP) error probability. From the software simulation, we found that Lss = 128 with 1-bit quantization is sufficient to reach a 1×10−30 error probability while optimizing energy efficiency and area, when communicating over a channel with a 16 dB loss at Nyquist.
This article presents a 2-lane 2 x 2 multiple-input, multiple-output (MIMO) 4-level pulse amplitude modulation (PAM-4) minimum mean-squared-error (MMSE)-decision-feedback equalizer (DFE) with far-end crosstalk (FEXT) cancellation for digital-to-analog converter (DAC)-/analog-to-digital converter (ADC)-based high-speed serial links. The receiver (RX) datapath is designed with a 15-tap MIMO feedforward equalizer (FFE) and a one-tap MIMO DFE with the least mean square (LMS), enabling adaptation to channel variation while maintaining the MMSE setting. The RX digital signal processor (DSP) place and route (PnR) in a 28-nm CMOS is estimated to consume 201 mW/lane at a 56-Gb/s/lane data rate while occupying a 0.5-mm(2)/lane silicon area. We further implement a real-time evaluation platform to verify the functionality of the MIMO PAM-4 MMSE-DFE with rapid bit-error-rate (BER) testing on RFSoC. The measurement result demonstrates that the MIMO MMSE-DFE significantly improves BER performance from 2.75e(-3) to 1.31e(-7) compared with equalization without FEXT cancellation when communicating over a channel exhibiting 12.4-dB insertion loss (IL) and 13.2-dB IL-to-crosstalk ratio (ICR) at Nyquist.
In high-speed links, limited channel bandwidth and frequency-dependent loss increase inter-symbol interference (ISI). Moreover, imperfect synchronization between the transmitter (TX) and receiver (RX) clocks introduces timing offset and jitter, which further degrade sampling accuracy. This paper presents a model of a Mueller–Muller phase detector (MMPD)based clock and data recovery (CDR) system and investigates its performance. For efficient analysis, XMODEL, a high-speed link simulation tool, is used to achieve runtime reduction compared to conventional simulators, enabling efficient evaluation of very low bit error rates (BER). A complete transmitter–receiver (TRX) architecture is constructed under channel loss conditions, with an analog-to-digital converter (ADC)-based receiver architecture employing a phase interpolator (PI)-based Mueller–Muller clock and data recovery (MMCDR) system. Two phase detector structures are modeled for comparative analysis: sign-sign MMPD (SS-MMPD) and linear MMPD.
This paper presents a versatile and fast time-domain architectural modeling framework for high-speed serial data transceivers (TRX) that can employ various analog modulation schemes. We highlight a modeling of TRXs employing an analog multi-tone signaling, which is not straightforward to model and hard to optimize with conventional serial link modeling tools. A method to limit the computing system's memory usage when simulating a data transmission of a long bit-stream, e.g., > 10 Mbits, is also described. The reliability of the modeling framework is proven by some comparisons with a highly-trusted commercial tool for a conventional TRX architecture.
This article presents a 112-Gb/s discrete multitone (DMT) wireline receiver (RX) datapath with a 50-GS/s, 8-bit, 64-way ( 8x 8 ) time-interleaved time-based analog-to-digital converter (TI-TBADC) in a 5-nm FinFET. The TBADC converts the voltage input into a time-domain quantity using a ring oscillator (ROSC). Eight-slice TBADCs, driven from the same first-rank interleaver, share the identical injection-locked ROSC (IROSC) for voltage-to-time conversion (VTC). The DMT digital signal processor (DSP) achieves optimal bit and power loading with 63 orthogonal subchannels by employing a 64-way single-stage multi-path delay feedback (MDF) fast Fourier transform (FFT) core. An on-chip sign-sign least mean square (SS-LMS) engine adapts equalizer coefficients to combat channel fluctuation. The RX prototype demonstrates 4E-4 BER when communicating over the channel, exhibiting 18-dB insertion loss (IL) at Nyquist, while consuming 347-mW power and 0.242-mm(2) silicon area.
This paper presents a novel digital decision feedback equalizer (DFE) design that can relax the feedback timing constraints for analog-to-digital converter (ADC)-based high-speed wireline receivers. The proposed technique breaks the loop-unrolled DFE (LU-DFE) chain by computing multiple LU-DFE chains in parallel with all possible seed symbols, and selecting the appropriate output by the post-processing selection logic. The proposed loop-break DFE (LB-DFE) is functionally equivalent to the conventional DFE with any other implementation techniques such as LU-DFE, look-ahead DFE (LA-DFE), or direct DFE. With topographical synthesis in 28nm CMOS process, the proposed LB-DFE achieved up to 54 % of DFE area saving as compared to LA-DFE with look-ahead factor (LF) of 16 for 112 Gb/s PAM-4 with 875 MHz DSP clock speed. The implementation feasibility and functionality are verified using ZCU111 RFSoC platform at 6 Gb/s (3 GS/s ADC conversion rate) with a channel exhibiting 25 dB loss at 1.5 GHz, demonstrating the same bit error rate (BER) performance between the LB-DFE and the LA-DFE. Equipment-based measurements using arbitrary waveform generator (AWG) and real-time oscilloscope transmitting/receiving 40 GBaud PAM-4 (80 Gb/s) to/from the differential cables with software 21-tap feed-forward equalizer (FFE) and LB-DFE on PC was also conducted.
This article proposes a high-speed framed-pulsewidth modulation (FPWM) transceiver that applies a time-domain modulation scheme for increased spectrum efficiency. The achieved coding gain is 75%, indicating that the minimum pulsewidth is increased by 1.75 times compared to an NRZ scheme with an identical data rate. Such bandwidth reduction renders dispersion tolerance both in copper and optical channels. The encoder and decoder employ successive approximation (SA) and weighted sum (WS) algorithms for power-and-area efficiency. The FPWM demonstrates an 8-dB SNR gain over NRZ at 15-km single-mode fiber (SMF) transmission while maintaining identical back-to-back performance. The FPWM scheme shows 6-dB higher receiver sensitivity at a bit error rate (BER) of 2e-5 than the PAM-4 signaling in 26-Gb/s back-to-back transmission and achieves 4-dB SNR gain at 20-km transmission. The test chip is fabricated in a 28-nm CMOS process and packaged in a flip-chip chip scale package (FCCSP). The test chip occupies 2.2 $\times$ 2.0 mm, including bidirectional two lanes and two phase-locked loops (PLLs), while it consumes 262 mW per lane from a 0.9-V supply. The measured Tx random jitter is 265 fs $_{\mathbf{rms}}$ .