
Social images sharing websites, such as Flickr and Picasa, are becoming very popular nowadays. Users are generally recommended to annotate images with free tags, yet these tags are orderless, and thus quite limited for applications like image search, retrieval and management. In this paper, we present a novel semi-supervised learning framework to rank image tags, which learns a ranking projection with theoretic guarantee from visual words distribution to the relevant tags distribution, and then uses it for ranking new image tags. Also as the manual ranking is laborious especially for large scale data collections, we propose an active learning scheme to guide the user ranking process and efficiently obtain the informative tag ranking information. This scheme improves the overall ranking result significantly with few user feedbacks. Experiments on both image benchmark and real Flickr photo collection show the practicability and efficiency of our proposed framework, which also further improves the performance of ranked tag recommendation application.
In this paper, we propose an approach to automatically estimate relationship among people in a family image collection based on results from face analyses technologies including automated face recognition and clustering, demographic assessment, and face similarity measurement, as well as contextual information such as people co-appearance, people's relative positions in photos and image timestamps. As the result, a relation tree can be estimated which provides important semantic information regarding people involved in a photo collection and has numerous applications in photo sharing and browsing, social networking, etc. The methods for deriving and integrating information from photos and the process for estimating a relation tree are described. Experimental results on two typical consumer photo collections and examples of using these results in consumer image retrieval are presented.
Music video is a popular type of entertainment by viewers. Currently, the novel indexing and retrieval approach based on the affective cues contained in music videos becomes more and more attractive to users. Music video affective analysis and understanding is one of the most popular topics in current multimedia community. In this paper, we propose a novel feature importance analysis approach to select most representative arousal and valence features for arousal and valence modeling. Compared with state-of-the-art work by Zhang on music video affective analysis, our main contributions are in the following aspects: (1) Another 3 affect-related features are extracted to enrich the feature set and exploit their correlation with arousal and valence. (2) All extracted features are ordered via feature importance analysis, and then optimal feature subset is selected after ordering. (3) Different regression methods are compared for arousal and valence modeling in order to find the fittest estimation function. Our method achieves 33.39% and 42.17% deduction in terms of mean absolute error compared with Zhang's method. Experimental results demonstrate our proposed method has a considerable improvement on music video affective understanding.
In multi-label learning, an image containing multiple objects can be assigned to multiple labels, which makes it more challenging than traditional multi-class classification task where an image is assigned to only one label. In this paper, we propose a multi-label learning framework based on Image-to-Class (I2C) distance, which is recently shown useful for image classification. We adjust this I2C distance to cater for the multi-label problem by learning a weight attached to each local feature patch and formulating it into a large margin optimization problem. For each image, we constrain its weighted I2C distance to the relevant class to be much less than its distance to other irrelevant class, by the use of a margin in the optimization problem. Label ranks are generated under this learned I2C distance framework for a query image. Thereafter, we employ the label correlation information to split the label rank for predicting the label(s) of this query image. The proposed method is evaluated in the applications of scene classification and automatic image annotation using both the natural scene dataset and Microsoft Research Cambridge (MSRC) dataset. Experiment results show better performance of our method compared to previous multi-label learning algorithms.
In this paper, we formulate the Relative Margin Support Tensor Machines (RMSTMs) problem as an extension of the Relative Margin Machines (RMMs). While the typical Support Tensor Machines (STMs) find a solution that is greatly influenced by the data spread, the proposed RMSTMs maximize the margin in a way relative to the spread of the data. The difference in the obtained solutions can be significant in the cases of badly scaled data, especially in the case of various spreads across different data dimensions. The efficiency of the proposed method is illustrated on the problems of gait and action recognition, where the results acquired verify the superiority of the method in terms of classification performance.
Aurora is the typical ionosphere track generated by the interaction of solar wind and magnetosphere, whose modality and variation are significant to the study of space weather activity. This paper proposes a novel aurora pattern recognition method based on static image classification of day-side aurora. In the feature extraction phase, X-gray level aura matrices (X-GLAMs) are designed to extract the feature of the original aurora images. For classification, models of texture classes are learned using support vector machine (SVM), then a given texture of aurora image can be classified into one of the pre-learned classes. It compares two sets of features: X-GLAMs and basic gray level aura matrices (BGLAMs), both of which are based on different windows on the real aurora image database from Chinese Arctic Yellow River Station. The experimental results illustrate the effectiveness of the proposed dayside aurora classification algorithm.
General image retrieval systems exploit text and link structure to "understand" the content of the web images and lack the discriminative power to deliver visually diverse search results. The result list often contains hundreds of pages, most of which may not be visited, costing a lot of time and energy of users. Unfortunately, many high quality images, containing more visual and semantic information, may appear at these back pages. To tackle this problem, we introduce a re-ranking method called Dual-Rank to improve web image retrieval by clustering and reordering the images retrieved from an image search engine. We first utilize multipartite graph model to represent images and features, then formulate clustering as a constrained multi-objective optimization problem, which can be efficiently solved by semi-definite programming (SDP). The framework of Dual-Rank is composed of Inter-cluster Rank and Intra-cluster Rank, and could rank clusters and images respectively. Our method is evaluated against a standard search engine and significant improvements are reported in terms of MAP, D@n and user experience.
Content-based video retrieval is maturing to the point where it can be used in real-world retrieval practices. One such practice is the audiovisual archive, whose users increasingly require fine-grained access to broadcast television content. We investigate to what extent content-based video retrieval methods can improve search in the audiovisual archive. In particular, we propose an evaluation methodology tailored to the specific needs and circumstances of the audiovisual archive, which are typically missed by existing evaluation initiatives. We utilize logged searches and content purchases from an existing audiovisual archive to create realistic query sets and relevance judgments. To reflect the retrieval practice of both the archive and the video retrieval community as closely as possible, our experiments with three video search engines incorporate archive-created catalog entries as well as state-of-the-art multimedia content analysis results. We find that incorporating content-based video retrieval into the archive's practice results in significant performance increases for shot retrieval and for retrieving entire television programs. Our experiments also indicate that individual content-based retrieval methods yield approximately equal performance gains. We conclude that the time has come for audiovisual archives to start accommodating content-based video retrieval methods into their daily practice.
The performance of human motion classification and recognition systems is highly dependent on the distinctiveness and robustness of the feature descriptor. In this paper, a new descriptor containing motion, shape and spatial layout information is proposed, therefore it is more effective for action modeling and is suitable for detecting and recognizing a variety of actions. Experiments show that the proposed descriptor outperforms other existing methods, such as Moment Invariants and Histogram of Oriented Gradients, on recognizing human motions in an indoor environment with a stationary camera.
In this paper, we explore different ways of formulating new evaluation measures for multi-label image classification when the vocabulary of the collection adopts the hierarchical structure of an ontology. We apply several semantic relatedness measures based on web-search engines, WordNet, Wikipedia and Flickr to the ontology-based score (OS) proposed in [22]. The final objective is to assess the benefit of integrating semantic distances to the OS measure. Hence, we have evaluated them in a real case scenario: the results (73 runs) provided by 19 research teams during their participation in the ImageCLEF 2009 Photo Annotation Task. Two experiments were conducted with a view to understand what aspect of the annotation behaviour is more effectively captured by each measure. First, we establish a comparison of system rankings brought about by different evaluation measures. This is done by computing the Kendall τ and Kolmogorov-Smirnov correlation between the ranking of pairs of them. Second, we investigate how stable the different measures react to artificially introduced noise in the ground truth. We conclude that the distributional measures based on image information sources show a promising behaviour in terms of ranking and stability.
The automatic detection of near duplicate video segments, such as multiple takes of a scene or different news video clips showing the same event, has received growing research interest in recent years. However, there is no agreed way of evaluating near duplicate detection algorithms. This makes it very hard to compare the performance of different algorithms, even if they are applied to the same data set. In this paper we have implemented several evaluation measures found in literature and we apply them to real algorithm outputs and a simulated result data set. We then calculate the correlation between the results obtained with the different measures in order to investigate whether they can be compared or not. The results show that the correlation between the measures is some cases quite low, and some measures are especially sensitive to certain types of deviations from the ground truth. However, a group of precision/recall type measures and two others are clearly correlated, though with moderate correlation coefficients. We also analyze the correlation between these measures and the subjective human judgment of the number of repeated segments in summary videos.
In this paper, we evaluate and compare different feature detection and feature description methods for part-based approaches in human action recognition. Different methods have been proposed in the literature for both feature detection of space-time interest points and description of local video patches. It is however unclear which method performs better in the field of human action recognition. We compare, in the feature detection section, Dollar's method [18], Laptev's method [22], a bank of 3D-Gabor filters [6] and a method based on Space-Time Differences of Gaussians. We also compare and evaluate different descriptors such as Gradient [18], HOG-HOF [22], 3D SIFT [24] and an enhanced version of LBP-TOP [15]. We show the combination of Dollar's detection method and the improved LBP-TOP descriptor to be computationally efficient and to reach the best recognition accuracy on the KTH database.
With increasing the importance of affective computing, it becomes necessary to retrieve and process images according to human affects or preference. However, judging such affective qualities of images is a highly subjective task. In spite of the lack of firm rules, certain features in images are believed to be more related than certain others. In this paper, we suggest predicting certain affective features include in an image using color composition that constitutes the scene. Using such a feature is inspired from Kobayashi's color scale that studies the relation between colors/color compositions and human's affects. Thus, we propose a Probabilistic Affective Model (PAM) to estimate the probabilities that an image is related to certain affective features. For this, we segment an image using mean-shift clustering algorithm, and extract more important regions, which are called seed regions, based on their properties. Thereafter, we find the dominant color compositions among those seed regions and their neighboring regions. Finally, from such color compositions, we infer the numerical ratings for some affective features. To assess the effectiveness of our PAM, we compared its results with 52 users' affective judgments. It was tested with online photo images, then the results show our PAM produced the recall of 85.22% and the precision of 78.16% on average. Potential applications include content-based image retrieval and design of web page interfaces.
In the past few years, multiobjective clustering has been one of the most successful techniques in the field of computer vision and data clustering. This paper proposes a novel unsupervised approach for synthetic aperture radar (SAR) image segmentation, namely, multiobjective immune clustering ensemble technique (MICET). The new technique first divides the image into several regions, and a certain number of pixels are picked out from these regions to form the clustering dataset. Second, artificial immune system (AIS) and multiobjective optimization (MOO) are introduced to generate multiple clustering results, which are then combined together for the following ensemble process. Multiple runs of the multiobjective clustering method with different randomly selected image features are performed to ensure high quality components as well as necessary diversity for an efficient ensemble. Finally, each datum is assigned to one cluster according to the relationship with the clustering dataset. Experimental results show that interesting segmentation performances on SAR images can be achieved by the proposed technique despite its completely unsupervised nature.
Currently Web image search is mostly implemented as text retrieval based on the textual information extracted from the Web page associated with the image. Since the text in the Web page may not match with the image content, image search re-ranking is preferable to refine the text-based search results. In this paper, we propose a novel scheme of latent visual context analysis (LVCA) for image re-ranking. The latent visual context is explored in both latent semantic context and visual link graphs. We argue that the image significance is determined by its contained visual word context, which is analyzed through Latent Semantic Analysis (LSA) and visual word link graph. With the visual word context information, the image context is explored by analysis of image link graph and the significance value for each image can be inferred by VisualRank. In both visual word link graph and image link graph, latent-layer will be incorporated to effectively discover the visual context. We validate our approach on text-query based search results returned by Google Image. Experimental results show improvement of both accuracy and efficiency of our method over the state-of-the-art VisualRank algorithm.
In this paper, an effective shape-based leaf image retrieval system is presented. A new contour descriptor is defined which reduces the number of points for the shape representation considerably. This shape representation is based on the curvature of the leaf contour and it deals with the scale factor in a novel and compact way. A two-step algorithm for retrieval is used. In a first step, the database is reduced using some geometrical features. Then a similarity measure between the contour representations is used to rank conveniently leaf images on the database. We implemented a prototype system based on these features and performed several experiments to show its effectiveness for plant species identification.
The Signature Quadratic Form Distance is an adaptive similarity measure for flexible content-based feature representations of multimedia data. In this paper, we present a deep survey of the mathematical foundation of this similarity measure which encompasses the classic Quadratic Form Distance defined only for the comparison between two feature histograms of the same length and structure. Moreover, we give the benefits of the Signature Quadratic Form Distance and experimental evaluation on numerous real-world databases.
Community structure as an interesting property of networks has attracted wide attention from many research fields. In this paper, we exploit the visual community structure in visual-temporal correlation network and use it to facilitate interactive video retrieval. We propose a hierarchical community-based feedback algorithm (HieCommunityRank) to make full use of the limited user feedback by integrating the most informative context according to visual community semantics. Since it re-ranks video shots respectively through diffusion process in inter-community and intra-community level, HieCommunityRank can guarantee both the global diverse distribution and the local consistency of video shots. Meanwhile it can get fast responsiveness after user feedback, which is rather important facing large amount of video collections. Experiments on TRECVID 09 Search dataset demonstrate the effectiveness and efficiency of the proposed algorithm.
This paper proposes an automatic visual feature weighting method to enhance content-based image retrieval (CBIR). In particular, the proposed method is able to capture user's search intention by identifying the important visual features located at region of interest. Given a query image, the importances of visual features are automatically weighted by a random walk algorithm from a feature association graph, whose association strength is estimated by a localized visual word co-occurrence count among a set of pseudo relevance feedbacks. The visual word here is defined with a bag-of-features model whose visual feature vocabulary is generated by a k-means clustering algorithm. For quantitative evaluation, we implement a prototype CBIR system with weighted visual features (WVF). Extensive experiments on CalTech-101 dataset demonstrate the efficiency and effectiveness of WVF for CBIR.
Social image retrieval has become an emerging research challenge in web rich media search. In this paper, we address the research problem of text-based social image retrieval, which aims to identify and return a set of relevant social images that are related to a text-based query from a corpus of social images. Regular approaches for social image retrieval simply adopt typical text-based image retrieval techniques to search for the relevant social images based on the associated tags, which may suffer from noisy tags. In this paper, we present a novel framework for social image re-ranking based on a non-parametric kernel learning technique, which explores both textual and visual contents of social images for improving the ranking performance in social image retrieval tasks. Unlike existing methods that often adopt some fixed parametric kernel function, our framework learns a non-parametric kernel matrix that can effectively encode the information from both visual and textual domains. Although the proposed learning scheme is transductive, we suggest some solution to handle unseen data by warping the non-parametric kernel space to some input kernel function. Encouraging experimental results on a real-world social image testbed exhibit the effectiveness of the proposed method.