
In this paper we discuss the application of 3D scene reconstruction techniques in the area of automatic semantic annotation, search and retrieval of unedited video footage. Rather than working with static key-frames we exploit the time-depended dynamic properties of a moving camera. Based on state of the art camera self calibration techniques we develop a powerful analysis chain. We demonstrate, that the reconstructed 3D scene information can be used to generate both, accurate low level scene descriptors as well as meaningful medium and high level semantic information. We show, that the proposed algorithms work even in case of sparse data sets. The proposed algorithms provide a powerful working base for further investigations in the area of low, medium and high level extraction of semantic information for unedited video.
We propose a practical standard-compliant multiple description (MD) image coding technique. Multiple descriptions of an image are generated in the spatial domain by an adaptive prefiltering and uniform down sampling process. The resulting side descriptions are conventional square sample grids that are interleaved with one the other. As such each side description can be coded by any of the existing image compression standards. A side decoder reconstructs the input image by first decompressing the down-sampled image and then solving a least-squares inverse problem, guided by a two-dimensional windowed piecewise autoregressive model. The central decoder is algorithmically similar to the side decoder, but it improves the reconstruction quality by using received side descriptions as additional constraints when solving the underlying inverse problem. Compared with its predecessors the proposed image MD technique offers the lowest encoder complexity, complete standard compliance, competitive rate-distortion performance, and superior subjective quality.
Automatic cell segmentation and tracking in optical microscope images plays a very important role in the study the behaviour of lymphocytes. The variable image contrasts, and especially variable cell densities are major factors to affect the successful cell detection rates. In this paper, two inner and outer cell contours edge detection based cell segmentation algorithms are proposed and used in parallel. Then a detection fusion algorithm is proposed to combine the two detection results and increase the probability of cell detection. Experimental results are used to demonstrate that these algorithms are robust to variations in both image contrast and cell densities. We show that the proposed fusion algorithm can increase cell detection rate significantly to above 90% with the false detection rate about 5%.
License plate recognition usually contains three steps, namely license plate detection/localization, character segmentation and character recognition. When reading characters on a license plate one by one after license plate detection step, it is crucial to accurately segment the characters. The segmentation step may be affected by many factors such as license plate boundaries (frames). The recognition accuracy will be significantly reduced if the characters are not properly segmented. This paper presents an efficient algorithm for character segmentation on a license plate. The algorithm follows the step that detects the license plates using an AdaBoost algorithm. It is based on an efficient and accurate skew and slant correction of license plates, and works together with boundary (frame) removal of license plates. The algorithm is efficient and can be applied in real-time applications. The experiments are performed to show the accuracy of segmentation.
This paper proposes a new approach for extracting highlight scenes from sports broadcasts by using sports news programs. In order to extract the highlight scenes from sports broadcasts without fail, we use sports news programs and identify identical or similar sections between sports broadcasts and sports news programs that cover the sports broadcasts. To extract identical or similar sections between two video data sets efficiently, we developed a two-step method that combines Relay-CDP [1] and Active-search [2]. We evaluated this method from the standpoint of the extraction accuracy of the highlight scenes, and computation time, through experiments using actual broadcast data sets.
A new chroma-based dynamic feature vector is proposed inspired by psychophysical observations that the human auditory system detects reltative pitch changes rather than absolute pitch values. The proposed chroma-based dynamic feature vector describes the relative pitch change intervals. The utility of the proposed feature vector incorporated with a music fingerprint extraction algorithm is experimentally explored within a music cover song identification framework. The results with a classical music database suggest that the proposed biologically plausible dynamic chroma feature vector can be successfully added to the conventional chroma feature vector as a complementary feature; it provides a 5.8% relative performance improvement.
With the increasing popularity of repositories of personal images, the problem of effective encoding and retrieval of similar image collections has become very important. In this paper we propose an efficient method for the joint scalable encoding of image-data and visual-descriptors, applied to collections of similar images. From the generated compressed bit stream, it is possible to extract and decode the visual information at different granularity levels, enabling the so called ldquomidstream content accessrdquo. The proposed approach is based on the appropriate combination of vector quantization (VQ) and JPEG2000 image coding. Specifically, the images are encoded at a first draft level using an optimal visual-codebook, while the residual errors are encoded using a JPEG2000 approach. In this way, the codebook of the VQ is freely available as an efficient visual descriptor of the considered image collection. This scalable representation supports fast browsing and retrieval of image collections providing also a coding efficiency comparable with those of standard image coding methods.
We propose two error resilient techniques for transmission of compressed video over error prone scenarios. The first technique is an improvement on the generalized source channel prediction (GSCP) scheme proposed by Yang and Rose in 2006. The second technique is a modification of weighted prediction (WP) present in the H.264 standard. Both techniques provide greater emphasis on Intra MBs during the formation of a reference frame thereby achieving better error resilience than GSCP with a very modest increase in the bit-rate.
Recent research in speech localization and dereverberation introduced processing of the multichannel linear prediction (LP) residual of speech recorded with multiple microphones. This paper investigates the novel use of intra- and inter-channel speech prediction by proposing the use of a multichannel LP model derived from multivariate autoregression (MVAR), where current LP approaches are based on univariate autoregression (AR). Experiments were conducted on simulated anechoic and reverberant synthetic speech vowels and real speech sentences; results show that, especially at low reverberation times, the MVAR model exhibits greater prediction gains from the residual signal, compared to residuals obtained from univariate AR models for individually or jointly modelled speech channels. In addition, the MVAR model more accurately models the speech signal when compared to univariate LP of a similar prediction order and when a smaller number of microphones are deployed.
We propose a cross-domain correlation in compress images, and introduce a novel spectral prediction algorithm to restore lossy spectral information caused by compression. This cross-frequency spectral predication algorithm is inspired from the spatial correlation and the connection between discrete cosine transform and Hadamard transform. The relationship among cross-frequency coefficients is adopted to predict spectral coefficients. We apply the spectral prediction algorithm in compressed image restoration under the total variation (TV) based regularization. Experimental results of restoration with or without cross-frequency spectral predication are compared, remarkable improvement is observed from the results with cross-frequency spectral prediction.
Multiple-description coding (MDC) provides an effective way to mitigate the effects of packet errors/loses by making use of multiple channels. Perhaps, the most attractive application of MDC is in the peer-to-peer (P2P) scenario to support simultaneous video streaming to a large population of clients. To this end, a number of multiple-description video coding (MDVC) schemes (both non-scalable and scalable) were proposed in the past few years. However, almost all non-scalable schemes would suffer from the prediction mismatch between the references used at the encoder and decoder sides; whereas all scalable schemes (involving a base-layer and some enhancement layers) would suffer from the inter-dependency within the enhancement-layer information. In this paper, we propose a transform-domain MDVC method that can solve these problems and at the same time offer some other interesting features.
In this paper, we addressed the problem of redundancy allocation for protecting packet loss for better quality of service (QoS) in real-time H.264 video streaming. A novel error-resilient approach is proposed for the transmission of pre-encoded H.264 video stream under bandwidth constrained networks. A novel frame importance model is derived for estimating relative importance index for different H.264 video frames. Combining with the characteristics of the network, the optimal resource allocation strategy for different video frames can be determined for achieving improved error resilience. The model uses frame error propagation index (FEPI) to characterize video quality degradation caused by error propagation in different frames in a GOP when suffer from packet loss. This model can be calculated in DCT domain with the parameters extracted directly from the bitstream. Therefore, the complexity of the proposed scheme is very low and much better for real-time video transmission. Simulation results show that the proposed scheme can improve the receiver side reconstructed video quality remarkably under different channel loss patterns.
A novel scheme of scalable video coding (SVC) using super-resolution techniques is proposed in this paper. Utilizing the spatial/temporal scalability of H.264/AVC, we encode half of input high-resolution(HR) frames and their low-resolution (LR) counterparts in a video sequence and employ super-resolution (SR) method to reconstruct skipped frames during decoding. The payload saved from skipped frames can be used to improve the quality of encoded frames. This scheme provides a choice of SVC to improve the quality of HR frames while maintaining low bit rate transmission and reducing encoding complexity. Experiments show that our scheme performs well in both peak signal to noise ratio (PSNR) and subjective visual quality at low bit rate.
Embedded speaker recognition in mobile devices could involve several ergonomic constraints and a limited amount of computing resources. Even if they have proved their efficiency in more classical contexts, GMM/UBM based systems show their limits in such situations, with good accuracy demanding a relatively large quantity of speech data, but with negligible harnessing of linguistic content. The proposed approach addresses these limitations and takes advantage of the linguistic nature of the speech material into the GMM/UBM framework by using clientcustomised utterances. Furthermore, the acoustic structure is then reinforced with video information. Experiments on the MyIdea database are performed when impostors know the client utterance and also when they do not, highlighting the potential of this new approach. A relative gain up to 47% in terms of EER is achieved when impostors do not know the client utterance and performance is equivalent to the GMM/UBM baseline system in other configurations.
Summary form only given: This talk will present some technical challenges in video coding and processing to meet the paradigm shifting trends for future digital entertainment for consumers. Traditional consumer video services have been in the broadcasting mode, from terrestrial TV, to satellite and cable services, in which a single encoder is able to serve millions of decoders. The design principle has been the simple decoder of volume sets at the expense of very complicated encoder. The proliferation of mobile devices with video capture capabilities in the recent years has resulted in a paradigm shift trends that require simple encoder for the mobile devices. The burden of the performance has now shifted to decoder that resides at consumerpsilas home to manage volumetric video captures with desktop computers. This paradigm shift thus created an opportunity for new generations of video coding and processing algorithms and architectures to meet the challenges in more complicated video decoding. In this talk, we will present several examples of new video coding and processing schemes based on distributed source coding. Some detailed analysis and simulation results will be shown to demonstrate that distributed source coding based approach is indeed promising for video decoding and processing for future digital entertainment.
This paper presents a new steganographic method called steganogallery, which means steganographic gallery. It enables us to imperceptibly embed secret information into a sequence of images. The images provided for embedding are just used to form the sequence. In other words, the images are not modified through embedding, while almost all of the image based steganographic methods do modify images for embedding data. The sequence is saved into a file in a plausible form, from which we can restore the embedded information. A prototype system has been developed and tested. It has been shown that the proposed framework of steganogallery does work in principle.
More and more streaming protocols are developed for the multimedia applications. However, many streaming protocols only consider the network stability, but not the characteristics of streaming applications. In order to cooperate with H.264/MPEG-4 AVC scalable extension which can achieve fine granularity of scalability at bit level to the time-vary heterogeneous networks, we design a TCP-friendly congestion control algorithm based on the bandwidth estimation to smoothly change sending rate to avoid unnecessary oscillations so that the subscription decision of SVC layers can be made to better utilize the network resource. In case of the unavoidable network congestion, we unsubscribe scalable video layers according to the packet lost rate and the recently received throughput instead of only dropping one layer at a time to rapidly accommodate the streaming service to the channels and avoid persecuting the other flows at the same bottleneck. In addition, the probing packets for estimating the available bandwidth are encapsulated with RTP/RTCP. The simulations show that the proposed congestion control algorithm for real-time applications efficiently utilizes network bandwidth without hampering the performance of the existing TCP applications.
We propose a new approach for locating forged regions in a video using correlation of noise residue. In our method, block-level correlation values of noise residual are extracted as a feature for classification. We model the distribution of correlation of temporal noise residue in a forged video as a Gaussian mixture model (GMM). We propose a two-step scheme to estimate the model parameters. Consequently, a Bayesian classifier is used to find the optimal threshold value based on the estimated parameters. Two video inpainting schemes are used to simulate two different types of forgery processes for performance evaluation. Simulation results show that our method achieves promising accuracy in video forgery detection.
A new application of exploiting digital watermarking techniques for annotating traffic surveillance videos is presented in this paper. The information of each vehicle collected from other sensors/sources will be embedded into the corresponding pixels in the recorded video to facilitate data management. The traffic scene captured by a video camera will be analyzed first and the individual vehicles are extracted and tracked by using Kalman filtering for effective watermarking. The scheme is integrated with H.264/AVC, which is assumed to be adopted by the visual surveillance system, to achieve an efficient implementation. The issues of payload, effective embedding/detection and rate/distortion are taken into account to fulfill the requirements of such an application. The experimental results demonstrate the feasibility of this potential system.
This paper investigates the optimal PET protection for streaming scalably compressed streams over networks where the delivery time constraints allow limited retransmissions (LR) and the communication channels exhibit both random losses and delays. A key property must be considered in this scenario is the possibility that a packet successfully arrives at the receiver in time, even if its acknowledgment is not received by the sender at certain deadlines. This paper proposes an extended LRPET scheme, namely random-delay LR-PET, in which additional streams may be sent to provide supplemental protection for the packets whose acknowledgments are still missing at a specified time after the transmission. To determine the optimal protection in each transmission opportunity, hypotheses concerning the number of acknowledged packets and the effect of future retransmission are considered. As the key contribution of this paper, we develop a method to derive the effective overall recovery probability versus redundancy characteristic, which significantly simplifies the actual protection assignment procedure. This paper also demonstrates the benefits of the optimization strategy proposed for this random-delay LR-PET scheme and the cruciality of time selection for scheduling retransmission.