We discuss parallel algorithms for wavelet-based image and video coding. After reviewing fundamentals of the parallel discrete wavelet transform we cover the parallelization of two state-of-the-art compression schemes: a C++ 3-D SPIHT video codec and a JAVA JPEG-2000 implementation.
In this work, we discuss different granularities for parallel wavelet packet video coding using block-based motion compensation and the performance of the corresponding MPI implementations on the HLRS Cray T3-E. Two inter-frame based parallelization methods (group-of-picture parallelization and frame-by-frame parallelization) are compared against intra-frame parallelization. We highlight the advantages and drawbacks of all three approaches.
In this work we describe and analyze algorithms for advanced video coding on distributed memory MIMD architectures. In particular, we consider a wavelet packet based codec using the concept of zerotree encoding. The main contribution of this work is the design of a parallel motion-compensated video coder composed of a wavelet packet decomposition in conjunction with the best basis algorithm followed by zerotree coding. Whereas two sensible parallelization techniques can be employed for the wavelet packet decomposition (subband based partitioning and stripe partitioning), the zerotree coding and motion compensation stages only allow one reasonable parallelization method (stripe partitioning). We investigate the advantages and drawbacks of the resulting different overall data distribution strategies and show experimental results obtained on a Siemens hpcLine cluster and a Cray T3E.
In this work, we describe and analyze algorithms for 2D wavelet packet (WP) decomposition for multicomputers and multiprocessors. In the case of multicomputers, the main goal is the generalization of former parallel WP algorithms which are constrained to a number of processor elements equal to a power of 4. For multiprocessors, we discuss several optimizations of shared-memory algorithms and finally we compare the results obtained on multicomputers and multi-processors employing the message passing (MPI) and shared-memory programming (OpenMP) paradigm, respectively.
In this work we describe and analyze algorithms for advanced image coding on distributed memory MIMD architectures. In particular, we consider a wavelet packet based image codec using the concept of zerotree encoding. The main contribution of this work is the design of a parallel image coder composed of a wavelet packet decomposition in conjunction with the best basis algorithm followed by zerotree coding. Whereas two sensible parallelization techniques can be employed for the wavelet packet decomposition (subband based partitioning and stripe partitioning), the zerotree coding step only allows one reasonable parallelization method (stripe partitioning). We investigate the advantages and drawbacks of all combinations and show experimental results obtained on a Siemens hpcLine cluster and a Cray T3E.
In this work we describe and analyze sequential and parallel algorithms for video coding using the wavelet packet decomposition.Thereby we discuss "efficiency" in a two-fold manner: the first part of this paper analyzes WP algorithms according to their rate-distortion efficiency. We address in this context different subband structures generated by the classical pyramidal decomposition and the best-basis algorithm (BBA), and different block motion estimation algorithms (classical and overlapped) for the generation of displaced frame differences (DFDs).The second part of this paper presents a fine-grained parallelization of a wavelet packet video coder and analyzes efficiency of parallel execution of different data partitioning approaches, comparing the results obtained on a Cray T3E and a Siemens hpcLine cluster.
In this work, we discuss parallel algorithms for three distinct approaches for wavelet-based video coding and the performance of their corresponding MPI implementations on the HLRS Cray T3-E.
This paper deals with different aspects of wavelet packet (WP) based video coding. In introducing experiments we show that WP decomposition and specifically WP decomposition in conjunction with the best basis algorithm are superior in terms of quality as compared to the standard discrete wavelet transform but show prohibitive computational demands (especially for real-time applications). The main contribution of our work is therefore the examination of three parallelization methods for WP based video coding. Two inter-frame based parallelization methods (group-of-picture parallelization and frame-by-frame parallelization) exploit the properties of a videostream (full independence between GOPs and rather high independence between single frames) better than inter-frame parallelization, but show a higher demand in terms of memory and don't respect the frame order defined by the input video stream. We highlight the advantages and drawbacks of all three methods and show experimental results obtained on a Siemens hpcLine cluster and a Cray T3E.
The à trous algorithm represents a discrete approach to the classical continuous wavelet transform. Similar to the fast or pyramidal wavelet transform, the input signal is analysed by using the coefficients of a properly chosen low-pass filter, but in contrast to the latter, all frequency sub bands are retained with full resolution. Therefore, this algorithm is much more demanding in terms of computational complexity compared to the fast wavelet transform and requires some sort of acceleration in order to satisfy real-time constraints. In this paper we develop parallel algorithms for different MIMD architectures for the two-dimensional à trous decomposition. In particular, classical border treatment strategies are discussed and compared in the context of data partitioning. It turns out that in contrast to the fast wavelet transform, the proper choice of a border treatment strategy does not depend on the underlying hardware. Additionally, only low scalability is achieved on multi-computers and multi-processors when employing the message passing paradigm, whereas much better behaviour is observed on multi-processors using the shared memory programming model.
We discuss parallel algorithms for wavelet-based image and video coding. After reviewing fundamentals of the parallel discrete wavelet transform, we cover the parallelization of two state-of-the-art compression schemes: a C++ 3D SPIHT video codec and a Java JPEG-2000 implementation.
In this work we describe and analyze algorithms for 2-D wavelet packet decomposition for MIMD distributed memory architectures. The main goal is the generalization of former parallel WP algorithms which are constrained to a number of processor elements equal to a power of 4. We discuss several optimizations and generalizations of data parallel message passing algorithms and finally compare the results obtained on a Cray T3D.
Strategies for computing the continuous wavelet transform on massively parallel SIMD arrays are introduced and discussed. The different approaches are theoretically assessed and the results of implementations on a MasPar MP-2 are compared.
The 'a trous' algorithm represents a discrete approach to the classical continuous wavelet transform. Similar to the fast wavelet transform the input signal is analyzed by using the coefficients of a properly chosen low-pass filter, but in contradistinction to the latter there follows no concluding decimation step. Examples of practical applications can be found in the field of cosmology for studying the formation of large scale structures of the Universe. In this paper we develop parallel algorithms on different MIMD architectures for the 2D 'a trous' decomposition. We implement the algorithm on several distributed memory architectures using the PVM paradigm and on a SGI POWERChallenge using a parallel version of the C programming language. Finally we investigate experimental results obtained on both of them.
We discuss several issues relevant for parallel wavelet transforms and their possible implications on the choice of a proper programming paradigm for corresponding multiprocessor implementations.
In this work we describe and analyze algorithms for 2-D wavelet packet decomposition for multicomputers and multiprocessors. In the case of multicomputers we especially focus on the question of handling of the border data among the processing elements. For multiprocessors we discuss several optimizations of data parallel algorithms and finally we compare the results obtained on a multiprocessor employing the message passing and data parallelism paradigm, respectively.
In this work we discuss algorithms for 2D wavelet packetdecomposition and best basis selection on massivelyparallel 2D-mesh SIMD arrays. In contrast to the sequentialcase a complete wavelet packet decompositionshows the same computational complexity as a pyramidalwavelet decomposition. Experimental results ona MasPar MP-2 confirm this property.1. INTRODUCTIONWavelet packets [12] represent a generalization of themethod of multiresolution decomposition and comprisethe entire family of ...
Abstract Strategies for computing the continuous wavelet transform on massively parallel SIMD computers are introduced. The different approaches are evaluated from the theoretical as well as from the experimental point of view.
Peter Meerwald合作论文数Dept . of Scientific Computing|University of Salzburg2