There has been a resurgence of interest in vector architectures evident by recent adoption of vector extensions in mainstream instruction set architectures. Traditionally, vector engines leverage this abstraction by exploiting its inherent regularity to increase performance and efficiency. Recent work on SRAM-based compute-in-memory has shown promise in reducing the area overhead of these engines. In this work, we propose ephemeral vector engines (EVE) where we leverage SRAM-based compute-in-memory techniquesas well as bit-peripheral computations to facilitate efficient vector execution. EVE uses a novel approach of bit-hybrid execution, striking a balance between throughput and latency. Evaluated on the Rodinia and RiVEC benchmark suites, EVE achieves almost 8× speed-up compared to an out-of-order processor and 4.59× compared to an integrated vector unit. EVE achieves speed-ups comparable to an aggressive decoupled vector unit and increases the area-normalized performance by over 2 ×. By repurposing SRAM arrays in the L2 cache to create ephemeral vector execution units, EVE is able to efficiently achieve high performance while incurring as little as 11.7% area overhead.
Single-ISA heterogeneous multi-core architectures offer a compelling high-performance and high-efficiency solution to executing task-parallel workloads in mobile systems on chip (SoCs). In addition to task-parallel workloads, many data-parallel applications, such as machine learning, computer vision, and data analytics, increasingly run on mobile SoCs to provide real-time user interactions. Next-generation scalable vector architectures, such as the RISC-V Vector Extension and Arm SVE, have recently emerged as unified vector abstractions for both large- and small-scale systems. In this paper, we propose novel area-efficient high-performance architectures called big.VLITTLE that support next-generation vector architectures to efficiently accelerate data-parallel workloads in conventional big.LITTLE systems. big.VLITTLE architectures reconFigure multiple little cores on demand to work as a decoupled vector engine when executing data-parallel workloads. Our results show that a big.VLITTLE system can achieve $1.6\times$ performance speedup over an area-comparable big.LITTLE system equipped with an integrated vector unit across multiple data-parallel applications and $1.7\times$ speedup compared to an aggressive decoupled vector engine for task-parallel workloads.
Command, Control, Communications, Computers, Intelligence, Surveillance and Reconnaissance (C4ISR) systems currently rely heavily on transponded satellites, both military and commercial at X-, Ku- and Ka-bands, for missions requiring beyond line of sight connectivity. Ku-band commercial satellites are the work-horse for both manned and unmanned Airborne ISR (AISR) operations today. In the future, C4ISR systems will require higher throughput, greater protection and improved affordability to align with future mission needs, mitigate threats, and fit within budgetary constraints. Intelsat EpicNG satellites provide cost-effective, next-generation, Ku-band capabilities including spot beams, higher throughput, improved efficiencies, and protection using deployed infrastructure. In this paper, the EpicNG architecture is described with comparisons to the legacy Ku-band systems and its application to the unique performance and mobility management needs of AISR systems. A detailed performance analysis is presented using representative AISR systems against a set of manned and unmanned mission/platform scenarios, with comparisons to legacy Ku-band and WGS Ka-band performance. It is shown that the EpicNG Ku-band constellation offers unique performance and affordability opportunities for AISR missions and enables the next generation military Ku/Ka-band C4ISR infrastructure.
LDPC codes have attracted much attention recently for their near-capacity performance and high throughput owing to parallel decoding architectures. While simulations are normally done with floating point computation, any practical implementation (ASIC or FPGA) will be built with fixed-point computation. Obviously, decoder speed will increase, and resource requirements will drop, with low-precision implementations, say 3,4, or 5-bit architectures. In this paper we study the effects of quantization in this regime, using density evolution and decoder simulation. Detailed sum-product decoder implementations are given, and performance losses relative to floating point decoding are given. In particular, 4-bit decoder architectures sustain only 0.1 dB penalty.
A novel high-throughput (6 Gb/s), fully-parallel FPGA-based 1200-bit rate-1/2 Low Density Parity Check (LDPC) decoder design is presented. The decoder features a PEG- based regular (6,3) code and a modified min-sum algorithm that improves performance without any additional hardware overhead.