This article provides an efficient hardware generator for high-performance sets-of-real-numbers (SORNs) arithmetic. Complex datapaths are automatically built in VHDL, comprising the generation of arithmetic operations and functions as well as all necessary interconnects. Various SORN datatypes can be easily set up enabling high adaptivity to different applications. For further performance improvements additional optimization and implementation techniques are considered as well. For evaluation, a SORN-based symbol detector for a multiantenna wireless communication scenario is generated that solves a system of linear equations. As SORN-based signal processing allows a quick but rough calculation of the system equations outcome, it can be exploited to reject possible input vectors that do not satisfy the general constraints of the signal detection task. Hence, this generally leads to a heavily reduced set of remaining solutions. Multiple hardware architectures with different SORN datatypes are generated and compared. Further, an analysis is performed considering the signal-to-noise ratio (SNR). Finally, logic synthesis is applied to selected designs and compared to references from the literature highlighting this article to be highly hardware-efficient and suitable for application-specific signal processing.
While centralized linear equalization algorithms can achieve a near optimal uplink detection performance for Massive MIMO systems at the base station, the required interconnect and off-chip input/output data rates surpass the bandwidth of existing interconnect standards. The fully decentralized feedforward equalization architecture can alleviate these bottlenecks by performing the signal equalization in decentralized clusters using only local channel state information. As the effectiveness of the decentralized equalization heavily depends on the fusion of the local signal estimates, we present a novel combiner algorithm based on the multistep method that is capable to reduce the performance loss in contrast to the centralized equalization approach.
In massive multi-user multiple-input-multiple-output (MU-MIMO) communication systems, high interconnect data rates are a major problem when the number of antennas at the base station increases. For some years, decentralized approaches have been proposed which address this issue by distributing the receive samples of clusters of antennas to different dies. In this paper, we present a hardware implementation for a decentralized matrix-splitting signal estimator unit as Gauss-Seidel equalizer which lowers the requirements regarding interconnect data rates. We synthesize our full custom RTL-design for an Artix-7 FPGA-prototype and ASIC-implementation and present the results regarding performance parameters hardware complexity, power, latency and throughput.
We propose an efficient approach for massive MIMO uplink detection using a tree of QR decomposition modules. The receive data is split onto the input modules and reduced during processing. The data rate does not increase while baseband data is combined and forwarded towards the root node of the tree. For the uplink baseband processing of a 128x8 MIMO system a 1.46 speedup is achieved. Our approach improves the applicability and scalability of massive MIMO for wireless industrial communication while requiring very low-latency and jitter.
While linear equalization schemes like zero forcing or minimum mean-square error achieve a near optimal uplink signal estimation performance in large-scale multi-user multiple-input multiple-output systems, the corresponding algorithms lean on centralized processing. To avoid disproportionate interconnect data rates due to the centralized signal estimation, performing a decentralized equalization can mitigate these effects. In this paper, we present a decentralized signal estimation architecture, which combines the ideas of existing decentralized architectures to (i) reduce the overall latency of the signal estimation and (ii) maintain a high data detection performance.
In this paper a new approach for a fast and low-precision number format is explored. The SORN representation uses a fixed set of exact values and open intervals in between to represent real numbers on the digital layer. With SORN arithmetic operations can be carried out using simple lookup tables which results in a very fast and low-complex way of computing. The new arithmetic is applied to an exhaustive search algorithm for maximum likelihood symbol detection in a MIMO transmission environment. The algorithm is implemented in hardware and executed on a Zynq-7000 FPGA. It is shown that the amount of possible solutions for the symbol detection can be reduced by up to 80% using the fast SORN architecture as a preprocessing stage for classic solvers.
In this paper, a block extended coordinate descent algorithm is introduced for MMSE based soft-output massive MIMO signal detection, which exploits the simple inversion of small sub-Gram matrices to allow a low-complexity implementation. We show that the resulting two-coordinates descent approach has a computational complexity comparable to the original coordinate descent signal detector, whereas the latency bottleneck can be relaxed and further the data detection performance can be improved as the simulation results show. Also we show the possibility to approximate the Gram matrix with fewer multiplications while maintaining a near-optimal detection performance.
In this paper an adaptive hardware architecture for high-performance bivariate numerical function approximation is presented. Orthogonal Chebyshev-Polynomials are exploited that cover incremental accuracy refinements. Additionally, switching between two numeric functions is easily deployable by changing the set of polynomial coefficients. For evaluation, different configurations of the proposed hardware function generator are implemented and analyzed considering three bivariate numeric functions. The resulting performance highlights this approach to be a powerful extension for bivariate function approximation.
In this paper, we propose a decentralized feedforward initialization approach for iterative equalization methods, that decomposes a massive MIMO system into multiple smaller subsystems, where successively combining pairs of these systems and their solutions using the Jacobi method results in a high quality initial estimate for a following iterative detection method. In contrast to existing methods, this approach allows a computation of an initial estimate while the Gram matrix and matched filter output vector is not fully available. The BER performances show the effectiveness of this initialization method compared to approaches like maximum-ratio combining at high basestationto- client antenna ratios.
Massive MIMO systems have become more popular in wireless communications due to their improved spectral efficiency compared to existing small-scale MIMO systems. However, current estimation methodes take too long for larger numbers of antennas. In this paper, a near-optimal iterative linear signal detection for massive MIMO is introduced exploiting the random projection method to approximate the channel matrix in a significantly lower dimensional space. This is then used as a preconditioner in the conjugate gradient least squares algorithm to enhance the convergence rate. For evaluation, different scenarios of spatial correlation in a massive MIMO system are considered. In contrast to other low-complexity signal detectors, our approach achieves excellent results in terms of robustness and determined latency.
In this work we present a novel inpainting algorithm to gain reduced acquisition time and high quality data reconstruction for MRI applications. We analyzed MRI recordings of two synthetically generated and of one real measured data set. On the basis of the proposed mask-based sampling trajectories, patients only have to spend a fraction of the recording time in the MRI. Especially, in the range of high k-space coefficient reduction, the reconstruction quality of our Permuted Cubes Wavelet Thresholding (PCWT) approach can compete with standard data compression-focused methods like JPEG2000 or MPEG-4. In all simulations, the proposed algorithm also outperforms state-of-the-art techniques such as BM3D-MRI with respect to accuracy after data reconstruction. Regarding the quality of the approximated MRI data, we mainly focus on the clear recovery of sharp edges without undesirable artifacts and the identification of tumors within MRI frames.
Inpainting-based compression and reconstruction methodology can be applied to systems with limited resources to enable continuously monitor neurological activity. In this work, an approach based on sparse representations and K-SVD is augmented to a video processing in order to improve the recovery quality. That was mainly achieved by using another direction of spatial correlation and the extraction of cuboids across frames. The implementation of overlapping frames between the recorded data blocks avoids rising errors at the boundaries during the inpainting-based recovery. Controlling the electrode states per frame plays a key role for high data compression and precise recovery. The proposed 3D inpainting approach can compete with common methods like JPEG, JPEG2000 or MPEG-4 in terms of the degree of compression and reconstruction accuracy, which was applied on real measured local field potentials of a human patient.
This work presents fast and efficient patch matching and ordering techniques for a novel inpainting-based compression and reconstruction methodology to continuously monitor neural activity. The mask-based compression is especially relevant for the technical realization of fully implantable neural measurement systems (NMS), because of restrictions regarding area and energy consumption. Novel approaches for decompression significantly reduce the number of computations for the procedure of smooth ordering patches (SOP) by a restricted neighboring search along consistent electrode patterns and by a patch group matching technique. Both combined yields a speed-up of 49.2x compared to an unrestricted patch search. With regard to recovered signal quality and compression of up to 95%, the proposed bridge mask achieves accurate results. The fast inpainting-based processing, including the proposed patch matching and ordering approaches, outperforms compression-focused standard techniques like JPEG and JPEG2000 regarding reconstruction quality of real measured neurological signals at high degrees of data reduction.