Finite impulse-response filters (FIR filters) are very commonly used in digital signal processing applications and traditionally implemented using ASICs or DSP-processors. For FPGA implementation, due to the high throughput rate and large computational power required under real-time constraints, they are a challenging subject. Indeed, the limitation of resources on an FPGA, i. e., logic blocks and flip flops, and furthermore, the high routing delays, require compact implementations of the circuits. Hence, in lookup table-based FPGAs, e. g. Xilinx FPGAs, FIR-filters were implemented usually using distributed arithmetic. However, such filters can only be used where the filter coefficients are constant. In this paper, we present approaches for a more flexible FPGA implementation of FIR filters. Using pipelined multipliers which are carefully adapted to the underlying FPGA structure, our FIR filters do not require a predefinition of the filter coefficients. Combining pipelined multipliers and parallely distributed arithmetic results in different trade-offs between hardware cost and flexibility of the filters. We show that clock frequencies of up to 50 MHz are achievable using Xilinx XCAOxx — 5 FPGAs.
One important algorithm for data compression is the variable length coding that often utilizes large code tables.Despite the progress modern FPGAs made, concerning the available logic resources, an ef.cient mapping of those tables is still a challenging task.In this paper, we describe an ef.cient mapping methodology for code trees onto LUT-based FPGAs.Due to an adaptation to the LUT’s number of inputs, for large code tables a reduction of up to 40% of logic blocks is achievable compared with a conventional gate-based implementation.
This paper presents an investigation of LUT-based FPGAs regarding their suitability for a particular application area using circuits which can be described in a typical HDL like Verilog or VHDL, not only synthetic benchmarks. The H.263 video codec has been chosen as benchmark example. In order to compare different FPGA architectures, a generic FPGA model and architecture independent modelling software are used. It is shown that coarse grain FPGAs give better area-speed trade-offs for large circuits whereas more fine grain devices are better suited for smaller designs.
Für die Implementierung von Videosignalverarbeitungsverfahren, ζ. B. Bildtelefonie nach CCITT (H.263) oder MPEG, ist die Auswahl von geeigneten Algorithmen und Architekturalternativen notwendig. Um den Entwickler in einem frühen Design-Stadium in seiner Entscheidungsfindung zu unterstützen, ist es wünschenswert, den Einfluss von Algorithmenund Schaltungsparametern auf die Bildqualität festzustellen. Für lange Videosequenzen ist die üblicherweise verwendete Software-Simulation zur Verifikation solcher Algorithmen nicht geeignet. Daher ist ein flexibles und echtzeitfähiges Hardware-Prototyping von kompletten Videosignalverarbeitungsverfahren notwendig, welches einen hohen Datendurchsatz ermöglicht. Daher wurde das „Video and Image Processing Emulation System" (VIPES) entwickelt, ein Rapid Prototyping System (RPS) zur Untersuchung von kompletten Videosignalverarbeitungsverfahren. VIPES basiert auf einer kommerziellen FPGA-basierten Emulator-Plattform. Der Emulator wurde um Video-Interfaces und spezielle Software erweitert. Durch Einsatz effizienter FPGA-Makros ist eine schnelle und flexible Umsetzung von Videosignalverarbeitungsverfahren möglich. Als Beispiel der Implementierung eines kompletten Verfahrens werden Ergebnisse für einen H.263-Codec vorgestellt. Die Eignung der Methodik wird exemplarisch anhand einer 2-dimensionalen DCT gezeigt. Im Vergleich zum herkömmlichen Designfluss werden die benötigten FPGA-Ressourcen um 4 8 % sowie die Compilierungszeit um 8 0 % reduziert.
A real-time prototyping environment for complete video processing schemes is presented. To realize a real-time processing, a commercial FPGA-based prototyping system is extended by a special video interface, efficient pipelined FPGA macros, and a modified design flow. Reductions of 48% in terms of FPGA resources, and 80% of compilation time are achievable. The feasibility of the prototyping environment is shown for a complete H.263 video codec.