The advent of the state-of-the-art video coding standard, High Efficiency Video Coding (HEVC) is expected to bring great changes to relevant fields of broadcasting, storage and communications. HEVC achieves higher coding gains compared to its previous video coding standards in terms of rate-distortion (R-D) performance with various improved coding tools. This leads to heavy computational complexity and costs to HEVC encoders and these comes as strong restrictions specially to develop H/W types of encoders that are more preferred for real-time based applications and services. In particular, the quad-tree based coding unit (CU) structures with various sizes are known to contribute to achieving high coding gains of HEVC. However, RD cost calculation for mode decision with all CU sizes cannot normally be considered in the H/W HEVC encoders for real-time operation. To overcome this, a CU size pre-determination method based on a probabilistic decision model fit to implement H/W HEVC encoders is proposed in this paper. All available CU sizes are checked before inter prediction and unnecessary CU sizes are excluded from inter prediction according to the decision model. Then inter prediction with the reduced number of CU sizes can be performed in parallel with pipeline structures. The experimental results show that the proposed method effectively determines the necessary CU sizes with negligible coding loss of 1.57% for LD (Low-delay) coding structure and 1.08% for RA (Random access) coding structure, respectively in BD-BR.
High efficiency video coding (HEVC) appears due to the demand on high compression video coding beyond H.264/AVC in ultra-high definition (UHD) videos, and it brings high computational complexity with a variety of state of the art coding tools. As for intra prediction, HEVC has 35 prediction modes while H.264/AVC has 9 intra modes. To exploit the spatial correlation, we adopt an edge detection method, establish an edge map, and adaptively select the candidate modes using the edge map for a block. The number of the candidate modes is determined through trade-off between computational complexity and coding efficiency. Besides, the range of coding unit sizes is determined using the uniqueness of the edge directions for the given image block. The proposed scheme reduced the encoding time by 56.8% at the cost of 2.5% BD-BR increase on average compared to Full modes at the HEVC reference software (HM 10.0 [1]).
Bi-predictive motion compensation is an important coding tool in low-delay coding of HEVC. The proposed method reduces computational complexity about 30% by removing exhaustive bi-predictive motion estimation. The coding gain is about 6% in low-delay P coding.
With the increase of image resolution in video application, the memory bandwidth is a critical problem in video coding. An embedded compression algorithm is a technique that can compress the frame data when stored in memory. It is possible to reduce memory requirements. In this paper, we propose a lossless embedded compression algorithm with context-based error compensation to reduce the memory bandwidth requirement. Experimental results have shown at least 50% memory bandwidth reduction on average and the data reduction ratio of the proposed algorithm is up to 5% higher than previously proposed lossless embedded compression algorithm [2].
An embedded compression (EC) is a technique that can compress on-the-fly the frame data when stored in video memory; it has been used to effectively reduce the system memory requirements. In this paper, we propose a lossy EC algorithm that uses hybrid coding scheme of feedforward DPCM and modified four-level BTC, which aims at the target specification of the 4 × 2 block-wise random access at 50 % compression ratio. The algorithmic characteristics of the proposed algorithm are twofold: the first is the error-aware quantization scheme to definitely reduce the total error in a block and the last is the hybrid DPCM/BTC coding scheme to effectively handle various texture patterns that can occur in a block. The architectural advantage stems from the feedforward DPCM that breaks off the inherent dependency between the prediction and quantization present in the conventional DPCM, which helps reduce the implementation burden as well as speed up the overall processing. Comparative evaluation shows that the proposed algorithm outperforms the existing ones by 0.77–6.39 dB.
Proposed is an adaptive scalar quantization scheme. The main advantage is to improve coding performance by suppressing the sum of absolute quantization errors (SAQE) under half the maximal SAQE, assuming that the errors have uniform distribution. This fact was proven and demonstrated, integrated into H.264.
디지털 영상 내의 평탄한 영역에 대한 양자화 과정은 종종 의도하지 않은 의사 윤곽 오차 (false contour artifact)를 발생한다. 본 레터논문에서는 통상적인 블록 기반 비디오 부호화 방식의 양자화 과정에서 발생되는 이러한 오차의 효율적 제거 알고리즘을 보인다. 먼저, 입력 블록에 대해 의사 윤곽의 발생 특성에 기반하여 추출된 특징값들을 이용하여 후보 블록을 선정 한다. 그리고, 해당 블록에 대해 미리 준비된 pseudo-random noise mask를 적용함으로써 의사 윤곽을 제거한다. 이러한 후보 블록 선정을 통한 선택적인 필터링 과정은 불필요한 처리를 최소화함으로써, 화질 열화 억제와 연산 복잡도 감소를 동시에 달성한다. 또한 블록 기반의 처리는 비디오 부호화 루프와의 통합을 용이하게 함으로써, 종래의 코덱 루프 밖에서의 후처리 필터링 방식과 비교하여 구현에 필요한 자원 절감과 프레임 지연 방지라는 측면에서 큰 이점을 갖는다. 제안 알고리즘은 H.264/AVC표준 소프트웨어에 구현되어 율-왜곡 최적화(rate-distortion optimization) 관점에서 압축 성능의 저하없이 낮은 복잡도에서 의사 윤곽을 효율적으로 제거함을 확인하였다. Contour artifact is known as the unintentional result of quantizing a flat area that has smooth gradients. In this letter, a decontouring algorithm is proposed to efficiently remove false contours that occur in typical block-based video coding applications. First, the algorithm goes through a refinement stage to determine candidate blocks probably having noticeable false contours with different kinds of features in a block. Then, pseudo-random noise masking is applied to those blocks to mitigate the contour artifacts. This block-based selective decontouring can efficiently remove the unnecessary processing of those blocks that have no false contour, which incidentally ensures a minor penalty in visual quality and computational complexity. The proposed algorithm was demonstrated, integrated into H.264/AVC, that visual quality can be significantly enhanced with an ignorable rate-distortion (RD) loss and an minor increase in computational complexity.
In this paper, we propose a novel algorithm for block artifact reduction, for which pseudo-random noise masking is selectively used. The experimental results demonstrate that the visual quality of the proposed scheme is significantly improved compared to that of H.264.
In H.264/AVC, motion data can be basically derived by the following two schemes: one is a typical spatial prediction scheme based on the DPCM and the other is a sophisticated spatiotemporal prediction scheme for the skipped motion data, formally referred to as a direct mode. We verified through instruction level profiling that when these schemes are combined with various H.264/AVC coding techniques, the computational burden to derive the motion data could be considerably aggravated. Specifically, its computational complexity amounts to maximally 55% of that of the overall syntax parsing process. In this paper, we aim at an efficient hardware design of the motion data decoding process for H.264/AVC, for which all the key design considerations are addressed in detail and respective rational answers are presented. As comparing the resulting hardware design with the processor-based solution, its effectiveness was clearly demonstrated. The proposed design was implemented with 43.2 K logic gates and three on-chip memories of 3584 bits using Samsung Semiconductor's Standard Cell Library in 65 nm L6LP process technology (SS65LP), and was capable of operating the H.264/AVC high-profile video bitstream of 1080p@60fps at 100 MHz consuming 843 mu W. Crown Copyright (C) 2010 Published by Elsevier B.V. All rights reserved.
The H.264/AVC standard yields higher coding efficiency rates than other video coding standards. This is because it uses the rate-distortion optimization (RDO) technique, which selects the optimal coding mode and a reference frame for each macroblock (MB). In order to achieve this, the encoder has to encode a given block by exhaustively using all kinds of combinations (including different intra and inter-prediction modes). As a result, the computational complexity of video coding in H.264/AVC is extremely high. In this paper, two fast intra-/inter-mode-decision algorithms are proposed to reduce the complexity of the encoder. Both of these algorithms are based on the inter-frame correlation among adjacent pictures. For the fast intra-mode-decision, we used the intra-mode of the most-correlated MB at the reference frame to encode the current MB and the stationary property of the current MB was used for the fast inter-mode-decision. The simulation results show that the proposed algorithms significantly reduced the computational complexity with a negligible loss of PSNR and a slight increase in bitrate.
H.264/AVC는 각 매크로블록에 대해서 최적의 부호화 모드와 참조 프레임을 결정해 주는 RDO (Rate-Distortion Optimization) 기법을 사용하여 기존의 비디오 압축 표준보다 더 좋은 부호화 효율을 얻고 있다. 하지만, RDO 기법은 하나의 매크로블록 모드를 결정할 때마다, 다양한 블록 타입의 화면내 (Intra) 예측을 수행하고 화면간 (Inter) 예측에 대해서도 1/4 화소까지 고려하는 움직임 추정(Motion Estimation)을 수행한 후 발생되는 비트까지 고려하여 최적의 모드를 결정하기 때문에 부호화기의 복잡도가 매우 큰 문제점이 있다. 따라서 영상의 객관적 화질은 유지하면서 부호화기의 복잡도를 낮추기 위한 많은 고속 알고리즘들이 제안되었고 연구 중에 있다. 본 논문에서는, 역 트리 구조의 경계 방향 예측 알고리즘을 이용한 고속 화면내 모드 결정 기법을 제안한다. 제안된 방법은 $4{\times}4$ 블록의 지역 경계 정보를 이용하여 해당 블록의 DE (Dominant Edge)를 찾아내고 DE에 상응하는 화면내 모드를 이용하여 RDO를 수행한다 $8{\times}8$ 블록 (또는 $16{\times}16$ 블록)의 DE는 이전 단계 4개의 $4{\times}4$ 블록 (또는 $8{\times}8$ 블록) DE들로부터 계산되고, 이 단계에서의 RDO 또한 DE에 상응하는 화면내 모드를 이용한다. 실험결과 제안 방법은 화면내 부호화에 사용되는 후보 모드의 수를 줄임으로써 JM12.2와 비교하여 화면내 부호화 시간을 평균 64% 단축시킬 수 있었다. The H.264/AVC standard achieves higher coding efficiency than previous video coding standards with the rate-distortion optimization (RDO) technique which selects the best coding mode and reference frame for each macroblock. As a result, the complexity of the encoder have been significantly increased. In this paper, a fast intra-mode decision algorithm is proposed to reduce the computational load of intra-mode search, which is based on the inverse tree-structure edge prediction algorithm. First, we obtained the dominant edge for each $4{\times}4$ block from local edge information, then the RDO process is only performed by the mode which corresponds to dominant edge direction. Then, for the $8{\times}8$ (or $16{\times}16$) block stage, the dominant edge is calculated from its four $4{\times}4$ (or $16{\times}16$) blocks' dominant edges without additional calculation and the RDO process is also performed by the mode which is related to dominant edge direction. Experimental results show that proposed scheme can significantly improve the speed of the intra prediction with a negligible loss in the peak signal to noise ratio (PSNR) and a little increase of bits.