FMCW radars are the key components for contactless range and motion sensing in industrial and healthcare applications. The radar-sensor performance, such as chirp bandwidth $(\text{BW}_{\mathrm{c}\text{hi}\text{rp}})$ , chirp slope, and frequency-modulation (FM) linearity, are determined by the FMCW chirp generator. When battery powered, the radar should be able to operate in a duty-cycled mode with minimal overhead, i.e., fast startup, fast lock at the start of every chirp burst, and minimal reset time in-between chirps, without degrading the radar range and Doppler performance. This work presents a robust fast-lock-acquisition charge-pump (CP)-PLL with a PFO for duty-cycled chirp generation. A fractional-N CP-PLL in a two-point-modulation (TPM) architecture breaks the trade-off between the PLL bandwidth and fast-chirp synthesis [1], [2]. A time-domain sign-extraction by using a 1 b TOC [3] enables the background calibration. A phase-offset-compensating digital-to-time converter (POC-OTC) assists the sign-extraction by compensating the positive/negative phase offsets generated within the type-Il PLL loop.
Beyond 5G systems are expected to approach 1 Tb/s throughput. This poses a significant challenge to the channel decoder. In this paper, we propose a multi-core architecture based on full row parallel layered LDPC decoder with frame interleaving. Compared with conventional partially parallel layered architectures, the proposed architecture increases the throughput by applying frame interleaving into the pipeline architecture and by using multi-core architectures. Two high rate medium size QC LDPC codes are designed with fast decoding convergence speed for this architecture. Both codes are implemented with single core and multi-core architectures to explore different trade-offs between code design, communication performance and implementation. The four decoders are implemented in 16 nm CMOS FinFET technology with a clock rate of 1 GHz. The placement and routing implementation results show that the single core decoder for the LDPC (1027, 856) code is able to provide 114 Gb/s throughput at maximum 3 iterations with an area of 0.173 mm(2) and energy efficiency of 1.56 pJ/bit; the multi-core decoder for the (1032, 860) code is able to provide 860 Gb/s throughput at maximum 2 iterations with an area of 1.48 mm(2) and energy efficiency of 3.24 pJ/bit. The multi-core decoder achieves the highest throughput in the literature for medium size (1-2k) LDPC codes. When compared with other state-of-the-art fully parallel high speed architectures, the proposed architectures bring a significant gain both in area efficiency and energy efficiency while keeping the ability to offer flexibility in code rate, number of iterations and early stop.
This work presents an efficient ASIC implementation of successive cancellation (SC) decoder for polar codes. SC is a low-complexity depth-first search decoding algorithm, favorable for beyond-5G applications that require extremely high throughput and low power. The ASIC implementation of SC in this work exploits many techniques including pipelining and unrolling to achieve Tb/s data throughput without compromising power and area metrics. To reduce the complexity of the implementation, an adaptive log-likelihood ratio (LLR) quantization scheme is used. This scheme optimizes bit precision of the internal LLRs within the range of 1-5 bits by considering irregular polarization and entropy of LLR distribution in SC decoder. The performance cost of this scheme is less than 0.2 dB when the code block length is 1024 bits and the payload is 854 bits. Furthermore, some computations in SC take large space with high degree of parallelization while others take longer time steps. To optimize these computations and reduce both memory and latency, register reduction/balancing (R-RB) method is used. The final decoder architecture is called optimized polar SC (OPSC). The post-placement-routing results at 16nm FinFet ASIC technology show that OPSC decoder achieves 1.2 Tb/s coded throughput on 0.79 mm$^2$ area with 0.95 pJ/bit energy efficiency.
IEEE 802.11ay is the amendment to the 802.11 standard that enables Wi-Fi devices to achieve 100 Gbps using the unlicensed mm-Wave (60 GHz) band at comparable ranges to today’s commercial 60 GHz devices based on the 802.11ad standard. In this paper, we propose a full row-based layered LDPC decoder supporting all the coding rates for 802.11ay. Taking the property of the parity check matrix of 802.11ay, combining multiple layers into single layer improves the hardware utilization hence increases the throughput. The throughput is further increases by interleaving multiple frames to improve the utilization of each pipeline stage. The decoder is synthesized at both 28 nm and 16 nm CMOS technology and power estimated with stimuli at 7db and 3.5db. The 28 nm implementation running at 600 MHz and achieves a throughput of 67 Gbps for coding rate 13/16 at 4 iterations with area efficiency of 160 Gbps/sqmm and consumes an average power consumption of 408 mW and 141 mW, yielding energy efficiency of 6.05 pJ/bit and 2.1 pJ/bit at 3.5db and 7db. The 16 nm implementation running at 1 GHz and achieves a throughput of 112 Gbps at 4 iterations with area efficiency of 589 Gbps/sqmm and consumes an average power of 408mW and 163mW, yielding energy efficiency of 3.64 pJ/bit and 1.45 pJ/bit at 3.5db and 7db.
Project number: 760150 Project website: www.epic-h2020.eu Project start: 1st September, 2017 Duration: 36 months Total cost: EUR 2,966,268.75 EC contribution: EUR 2,966,268.75 Message from the Coordinator EPIC Contributions to ThZ Standards Past/Upcoming Events Next-Generation Channel Coding Project Progress Deliverables Publications In this second edition of the EPIC newsletter, we would like to focus on standardisation efforts and contributions to standards, within the EPIC project.
This paper describes a 28-nm fully integrated 79 GHz Phase Modulated radar SoC including 2TX, 2RX and the mm-wave frequency generation. A custom digital core generates the pseudo random sequence and performs correlation and accumulation of the digitized received data. With 1 W power consumption, 7.5 cm range resolution is achieved while antennas arrangement allows for 5° resolution over ±60° elevation and azimuth scan in 2×2 code domain MIMO operation.
The design of multi-Gbps LDPC decoder has become a hot topic in recent years as the demand of transformation towards 5G. An energy efficient 18Gbps LDPC decoder based on LDPC ASIP with half layer paralleled architecture is proposed. The feasibility of the design is proven by its demonstrator silicon in 28nm CMOS technology, with a record energy efficiency of 18.4 pJ/decoded bit and area efficiency of 23.8 Gbps/mm2 for the ½ coding rate working at 18.4Gbps. With frequency, voltage scaling and multi-core management, the decoder supports a wide range of throughput, from 1.8Gbps to 18.4Gbps. The measurement results show the ASIP based design not only provides an energy efficient high speed solution but also be competitive with published ASIC solution at low and medium throughput scenarios.
We demonstrate a reconfigurable engine for multi-purpose spectrum sensing within the cost and power constraints of mobile devices. The analog part builds up on the Scaldio reconfigurable analog front-end [1]. The digital part is an innovative Digital Front-end for Sensing capable of performing a range of sensing algorithms [3], which has now been fully implemented as a chip. The goal of this demo is the first demonstration of the digital chip, integrated with an analog front-end, enabling real-time validation of the sensing engine. The setup is validated for DVB-T and LTE, two important candidates for future DySPAN networks, as well as for very fast spectrum sweeping. This is the first integrated low power solution that can achieve such a very fast spectrum sweeping, thanks to the integration of two innovative components.
This paper describes the implementation of a flexible Turbo and LDPC outer modem engine which is capable of supporting the WiFi(802.11n), WiMax(802.16e) and 3GPP-LTE standard on the same hardware resources. The chip is implemented in a 65nm CMOS technology and occupies 10.37 mm(2). The decoder flexibility is offered by means of an application-specific instruction-set processor (ASIP), with full datapath reuse between Turbo and LDPC decoding. The encoders are dedicated ASIC datapaths. The maximum clock speed can be set to 320 MHz allowing a decoder output rate for a single iteration in excess of 140 Mbps for Turbo and 640 Mbps for LDPC with a maximum power consumption of 675 mW. The architecture template has been extended to support other standards like the DVB-S2/T2 LDPC decoding as well.
This paper describes the implementation of an energy-efficient digital SDR baseband platform. The multi processor system-on-chip (MPSOC) is implemented in 90nm CMOS technology and occupies 32mm2. It incorporates all digital signal processing required by the physical layer of the WiFi(802.11n), WiMax(802.16e), mobile TV and 3GPP-LTE standards. The heterogeneous architecture with hierarchical wake-up achieves 5mW idle time power, is capable of delivering a net data rate in excess of 200Mbps and consumes 231mW during 108Mbps WLAN 2×2 MIMO Rx, achieving 2.14nJ/b energy efficiency.
Multi-antenna transmission over multi-input, multi-output (MIMO) channels are considered in almost all recent broadband wireless communication standards. Besides, the fast-pacing diversity and evolution of those standards, next to the deep submicron integration cost explosion, urges multi-mode reconfigurable solutions. Software Defined Radio (SDR) is envisioned to enable low-cost, high-volume multi-mode baseband modems both for base-station and user terminals. Yet, supporting high-throughput MIMO standard with limited energy budget as in user terminals is a challenge for SDR architectures. With Space Division Multiplexing (SDM) for instance, N being the number of antennas, the computation load is multiplied by >N2. Capitalizing on a low complexity SDM-OFDM functional architecture, a heterogeneous multi-processor SoC platform with DSP cores delivering 50 to 250MOPS/mW and an integrated software development flow, we demonstrate the SDR implementation of 100Mbps+ SDM-OFDM with 3.6 nJ/bit energy efficiency (383mW average power).
Software defined radios (SDR) requires an application-specific programmable architecture with instruction set targeted toward wireless baseband processing. Enabling this in handhelds terminal asks for both energy-awareness and cost-effectiveness. Micro-architecture efficiency and software mapping productivity must be carefully balanced. Especially in exploiting data-level parallelism, one has to trade off explicit, user-defined subword parallelism and automated, compiler-driven instruction level parallelism. In this paper, we describe an extensive exploration of a scalable subword-parallelism-enabled very long instruction word architecture targeting 100 Mbps SDRs. Coware LISATek tools is used to model a VLIW processor with scalable number of SIMD units. A compilation flow is set up supporting advanced ILP scheduling and subword parallelism encapsulation through intrinsic functions. Based on a set of representative SDR benchmarks, application specific optimization is carried out introducing powerful instruction set extensions. Cost and benefits are evaluated in terms of benchmark execution time and energy. Therefore, RTL is generated from LISATek and synthesized using a 90 nm CMOS library. By varying the architecture and technology parameters the optimal energy-performance tradeoff is derived. The achieved performance is more than sufficient for SISO WLAN such as 802.11 a/g
To assess the performance of forthcoming 4th generation wireless local area networks, the algorithmic functionality is usually modelled using a high-level mathematical software package, for instance, Matlab. In order to validate the modelling assumptions against the real physical world, the high-level functional model needs to be translated into a prototype. A systematic system design methodology proves very valuable, since it avoids, or, at least reduces, numerous design iterations. In this paper, we propose a novel Matlab-to-hardware design flow, which allows to map the algorithmic functionality onto the target prototyping platform in a systematic and reproducible way. The proposed design flow is partly manual and partly tool assisted. It is shown that the proposed design flow allows to use the same testbench throughout the whole design flow and avoids time-consuming and error-prone intermediate translation steps.
The increasing demand for multimodal wireless communication is driving designers towards software defined radio (SDR). Therefore, new high performance reconfigurable platforms for baseband digital signal processing are required. Due to their flexibility, with low reconfiguration overhead, performance and energy efficiency, coarse grain reconfigurable arrays (CGRAs) are good candidates to fulfil this need. ADRES is a CGRA that combines a VLIW processor with a reconfigurable coarse-grain array. In this paper, we analyze the mapping on ADRES of one of the most demanding wireless OFDM DSP algorithms: the space division multiplexing (SDM) receiver. The latter will probably be mandatory in the next WLAN generation (802.11n). We also compare the obtained results with a mapping onto a VLIW processor, showing a gain of 5 in performance and a factor 1.75 in power efficiency.
With multiple-input, multiple-output (MIMO) transmission, impressive capacity or diversity gains can be achieved compared to single antenna systems. We have conceived and implemented a generic platform that enables real-time wireless MIMO transmission. The platform supports the integration and the evaluation of transmit and receive spatial multiplexing and spatial diversity processing techniques for MIMO-OFDM systems in a wireless environment. The platform is based on modular multi-FPGA and multi-board designs with support for high-speed serial connection links between the boards. This paper describes the platform integration of a 2-antenna base station using MIMO transmit processing including special front-end techniques for the exploitation of channel reciprocity in time division duplex (TDD) schemes.
David Novo合作论文数Nomadic Embedded System Division|Department of EE2