We propose an image-based class retrieval system for ancient Roman Republican coins that can be instrumental in various archaeological applications such as museums, Numismatics study, and even online auctions websites. For such applications, the aim is not only classification of a given coin, but also the retrieval of its information from standard reference book. Such classification and information retrieval is performed by our proposed system via a user friendly graphical user interface (GUI). The query coin image gets matched with exemplar images of each coin class stored in the database. The retrieved coin classes are then displayed in the GUI along with their descriptions from a reference book. However, it is highly impractical to match a query image with each of the class exemplar images as there are 10 exemplar images for each of the 60 coin classes. Similarly, displaying all the retrieved coin classes and their respective information in the GUI will cause user inconvenience. Consequently, to avoid such brute-force matching, we incrementally vary the number of matches per class to find the least matches attaining the maximum classification accuracy. In a similar manner, we also extend the search space for coin class to find the minimal number of retrieved classes that achieve maximum classification accuracy. On the current dataset, our system successfully attains a classification accuracy of 99% for five matches per class such that the top ten retrieved classes are considered. As a result, the computational complexity is reduced by matching the query image with only half of the exemplar images per class. In addition, displaying the top 10 retrieved classes is far more convenient than displaying all 60 classes.
Nowadays, there is a proliferation of available information sources from different modalities-text, images, audio, video and more. Information objects are not isolated anymore. They are frequently connected via metadata, semantic links, etc. This leads to various challenges in graph-based information retrieval. This paper is concerned with the reachability analysis of multimodal graph modelled collections. We use our framework to leverage the combination of features of different modalities through our formulation of faceted search. This study highlights the effect of different facets and link types in improving reachability of relevant information objects. The experiments are performed on the Image CLEF 2011 Wikipedia collection with about 400,000 documents and images. The results demonstrate that the combination of different facets is conductive to obtain higher reachability. We obtain 373% recall gain for very hard topics by using our graph model of the collection. Further, by adding semantic links to the collection, we gain a 10% increase in the overall recall.
Nowadays, there is a proliferation of information objects from different modalities—Text, Image, Audio, Video. Different types of relations between information objects (e.g. similarity or semantic) has motivated graph-based search in multimodal Information Retrieval. In this paper, we formulate a Random Walks problem along our model for multimodal IR, that is robust over different distributions of modalities. We investigate query-dependent and query-independent Random Walks on our model. The results show that the query-dependent Random Walks provides higher precision value than query-independent Random Walks. We additionally investigate the contribution of the graph structure (quantified by the number and weights of incoming and outgoing links) to the final ranking in both types of Random Walks. We observed that query-dependent Random Walks is less dependent on the graph structure. The experiments are applied on a multimodal collection with about 400,000 documents and images.
Finding useful information from large multimodal document collections such as the WWW is one of the major challenges of Information Retrieval (IR). Nowadays the proliferation of available information sources— text, images, audio, video and more—increase the need for multimodal search. Multimodal information retrieval is about the search for information of any modality on the web, with unimodal or multimodal queries. For instance, a unimodal query may contain only keywords, whereas multimodal queries may be a combination of keywords, images, video clips or music files. Users have learnt to explain their information need through keywords and expect the result as a combination of different modalities. Search engines like Google and Yahoo often show related videos or images in addition to the text result to the user. Usually, in a keyword based search, only the metadata information of a video or an image (e.g. tag, caption or description) is used to find relevant results. This approach is limited to textual information only and does not include information from other modalities. There is few options such as Google image search, which considers the image features to perform the image search task based on. In case the user query is an image, or a combination of a video file and keywords, the question arises how can a search engine benefit from different modalities in the query to retrieve multimodal results. Usually, search engines build upon text search by using non-visual information associated with visual content. This approach in multimodal search does not always result in satisfying results, as it completely ignores the information from other modalities in ranking. To address the problem of visual search approaches, multimodal search reranking has received increasing attention in recent years. In addition to the observation that data consumption today is highly multimodal, it is also clear that data is now heavily semantically interlinked. This can be through social networks (text, images, videos of users on LinkedIn, Facebook, or the like), or through the nature of the data itself (e.g. patent documents connected by their metadata inventors, companies, semantic connections via linked data). Structured data is naturally represented by a graph, where nodes denote entities and directed/indirected edges represent the relations between them. Such graphs are heterogeneous, describing different types of objects and links. Connected data poses a challenge is traditional IR method which is based on independent documents. The question arises whether structured IR can be an option for retrieving more relevant data objects.
Currently there is a proliferation of data in different modalities-text, audio, video, and image. Various platforms such as social networks or different online repositories facilitate this trend. In addition, users search more for information from different modalities. This trend has created multimodal Information Retrieval (IR) as a challenge. However, information objects are not independent and there are different types of relations (e.g. semantic or similarity links) between them. In this paper, we evaluate the role of adding semantic links in performance of our graph-based model for multimodal IR. The result analysis shows that we gain 13% recall increase by using inter-lingual semantic links.
Video-to-video linking systems allow users to explore and exploit the content of a large-scale multimedia collection interactively and without the need to formulate specific queries. We present a short introduction to video-to-video linking (also called 'video hyperlinking'), and describe the latest edition of the Video Hyperlinking (LNK) task at TRECVid 2016. The emphasis of the LNK task in 2016 is on multi-modality as used by videomakers to communicate their intended message. Crowdsourcing makes three critical contributions to the LNK task. First, it allows us to verify the multimodal nature of the anchors (queries) used in the task. Second, it enables us to evaluate the performance of video-to-video linking systems at large scale. Third, it gives us insights into how people understand the relevance relationship between two linked video segments. These insights are valuable since the relationship between video segments can manifest itself at different levels of abstraction.
In this paper, we address the shortage of evaluation benchmarks on Persian (Farsi) language by creating and making available a new benchmark for English to Persian Cross Lingual Word Sense Disambiguation (CL-WSD). In creating the benchmark, we follow the format of the SemEval 2013 CL-WSD task, such that the introduced tools of the task can also be applied on the benchmark. In fact, the new benchmark extends the SemEval-2013 CL-WSD task to Persian language.
This paper is concerned with potential recall in multimodal information retrieval in graph-based models. We provide a framework to leverage individuality and combination of features of different modalities through our formulation of faceted search. We employ a potential recall analysis on a test collection to gain insight on the corpus and further highlight the role of multiple facets, relations between the objects, and semantic links in recall improvement. We conduct the experiments on a multimodal dataset containing approximately 400,000 documents and images. We demonstrate that leveraging multiple facets increases most notably the recall for very hard topics by up to 316%.
This paper describes the contributions of Vienna University of Technology (TUW) to the MediaEval 2015 Retrieving Diverse Social Images challenge. Our approach consists of 3 phases: (1) Precision-oriented-phase: in which we focus only on the relevance of the documents; (2) Recall-orientedphase: in which we focus only on the diversity aspect; (3) Merging phase: in which we explore ways to nd a balance between the relevance and diversity factors. We use two fusion methods for this last part. Our best run reached a F1@20 of 0.582.
[1] A. Awan, R. A. Ferreira, S. Jagannathan, and A. Grama. Distributed uniform sampling in unstructured peer-to-peer networks. In HICSS, 2006. [2] T. Berber, A. H. Vahid, O. Ozturkmenoglu, R. G. Hamed, and A. Alpkocak. Demir at imageclefwiki 2011: Evaluating dierent weighting schemes in information retrieval. In CLEF, 2011. [3] R. A. Ferreira, M. Krishna Ramanathan, A. Awan, A. Grama, and S. Jagannathan. Search with probabilistic guarantees in unstructured peerto-peer networks. In P2P, 2005. [4] M. Hlynka and M. Cylwa. Observations on the Metropolis-Hastings Algorithm. University of Windsor, Department of Mathematics and Statistics, 2009. [5] W. H. Hsu, L. S. Kennedy, and S.-F. Chang. Video search reranking through random walk over document-level context graph. MULTIMEDIA, 2007 [6] Y. Jing and S. Baluja. Visualrank: Applying pagerank to large-scale image search. IEEE Trans. Pattern Anal. Mach. Intell., 2008. [7] J. Martinet and S. Satoh. An information theoretic approach for automatic document annotation from intermodal analysis. In Workshop on Multimodal Information Retrieval, 2007. [8] T. Mei, Y. Rui, S. Li, and Q. Tian. Multimedia search reranking: A literature survey. ACM Computing Surveys (CSUR), 2014. References The velocity of multimodal information shared on web has increased significantly.
We present a model for multimodal information retrieval, leveraging different information sources to improve the effectiveness of a retrieval system. This method takes into account multifaceted IR in addition to the semantic relations present in data objects, which can be used to answer complex queries, combining similarity and semantic search. By providing a graph data structure and utilizing hybrid search in addition to structured search techniques, we take advantage of relations in data to improve retrieval. We tested the model with ImageCLEF 2011 Wikipedia collection, as a multimodal benchmark data collection, for an image retrieval task.
Modeling data as a graph of objects is increasingly popular, as we move away from the relational DB model and try to introduce explicit semantics in IR. Conceptually, one of the main challenges in this context is how to “intelligently” traverse the graph and exploit the associations between the data objects. Two highly used methods in retrieving information on structured data are: Markov chain random walks, as is the basic method for page rank, and spreading activation, which originates from the artificial intelligence area. In this paper, we compare these two methods from a mathematical point of view. Random walks have been preferred in information retrieval, while spreading activation has been proposed before, but not really adopted. In this study we find that they are very similar fundamentally under certain conditions. However, spreading activation has much more flexibility and customization options, while random walks holds concise mathematics foundation.
Astera is a model for information retrieval on multimodal collections. We leverage different information sources to improve the effectiveness of a retrieval system. Taking into account both explicit and latent semantics present in the data can fertilize the potentiality to answer complex queries, not currently answerable neither by document retrieval systems, nor by semantic web systems. We provide a hybrid approach combining IR and structured search techniques. This model is under test with a multimodal collection.
We present a generic model for multimodal information retrieval, leveraging different information sources to improve the effectiveness of a retrieval system. The proposed method is able to take into account both explicit and latent semantics present in the data and can be used to answer complex queries, not currently answerable neither by document retrieval systems, nor by semantic web systems. By providing a hybrid approach combining IR and structured search techniques, we prepare a framework applicable to multimodal data collections. To test its effectiveness, we instantiate the model for an image retrieval task.
Finding useful information from large multimodal document collections such as the WWW is one of the major challenges of Information Retrieval (IR). The many sources of information now available - text, images, audio, video and more - increases the need for multimodal search. Particularly important is also the recognition, that each information item is inherently multimodal (i.e. has aspects in its information character that stem from dierent modalities) and forms part of a networked set of related information items. In this paper we propose a graph-based model for multimodal information retrieval based on a faceted view of information objects. For retrieval purposes, we consider both relatedness and similarity relations between objects.
Over-the-air (OTA) software download is a key enabling technology for terminal reconfigurations. With increasing the number of network protocols, variety of software products with different software providers and new software threats, it becomes important to download and install assured software on the mobile terminals. Most of the related research works on downloading assured software, focus on software validation after download, and mainly on the terminal. In this paper, a model is proposed in which an Assurance Agency is exploited. This agency is responsible for performing software examination and validations. Unlike related work, examination and validation operations are centrally performed in the network before downloading the software. Based on a cost analysis, it is shown that the overhead of using the assurance agency in downloading scenario is acceptable.