The large amount and the ubiquitous availability of multimedia information (e.g., video, audio, image, and also text documents) require efficient, effective, and automatic annotation and retrieval methods. As videos start to play an even more important role in multimedia, content-based retrieval of videos becomes an issue, especially as there should be an integrated methodology for all types of multimedia documents.Our approach for the integrated retrieval of videos, images, and text comprises three necessary steps: First, the detection and extraction of shots from a video, second. the construction of a still image from the frames in a shot. This is achieved by an extraction of ky frames or a mosaicing technique. The result is a single image visualization of a shot, which in turn can be analyzed by the ImageMiner double dagger(TM) system.The ImageMiner system was developed in cooperation with IBM at the University of Bremen in the Image Processing Department of the Center for Computing Technologies. It realizes the content-based retrieval of single images through a novel combination of techniques and methods from computer vision and artificial intelligence. Its output is a textual description of an image, and thus in our case, of the static elements of a video shot. In this way, the annotations of a video can be indexed with standard text retrieval systems, along with text documents or annotations of other multimedia documents, thus ensuring an integrated interface for all kinds of multimedia documents. (C) 1999 Elsevier Science Ltd. All rights reserved.
In this paper videos are analyzed to get a content-based decription of the video. The structure of a given video is useful to index long videos efficiently and automatically. A comparison between shots gives an overview about cut frequency, cut pattern, and scene bounds.After a shot detection the shots are grouped into clusters based on their visual similarity. A time-constraint clustering procedure is used to compare only those shots that are positioned inside a time range. Shots from different areas of the video (e.g., begin/end) are not compared. With this cluster information that contains a list about shots and their clusters it is possible to calculate scene bounds. A labeling of all clusters gives a declaration about the cut pattern. It is easy now to distinguish a dialogue from an action scene.The final content analysis is done by the ImageMiner* system. The ImageMiner system developed at the University of Bremen of the Image Processing Department of the Center for Computing Technology realizes content-based image retrieval for still images through a novel combination of methods and techniques of computer vision and artifical intelligence.The ImageMiner system consists of three analysis modules for computer vision, namely for color, texture, and contour analysis. Additionally exists a module for object recognition. The output of the object recognition module can be indexed by a text retrieval system. Thus, concepts like forestscene may; be searched for.We combine the still image analysis with the results of the video analysis in order to retrieve shots or scenes.
The large amount of available multimedia information (e.g. videos, audio, images) requires efficient and effective annotation and retrieval methods. As videos start playing a more important role in the frame of multimedia, we want to make these available for content-based retrieval. The ImageMiner-System, which was developed at the University of Bremen in the AI group, is designed for content-based retrieval of single images by a new combination of techniques and methods from computer vision and artificial intelligence. In our approach to make videos available for retrieval in a large database of videos and images there are two necessary steps: First, the detection and extraction of shots from a video, which is done by a histogram based method and second, the construction of the separate frames in a shot to one still single images. This is performed by a mosaicing-technique. The resulting mosaiced image gives a one image visualization of the shot and can be analyzed by the ImageMiner-System. ImageMiner has been tested on several domains, (e.g. landscape images, technical drawings), which cover a wide range of applications.
The fourth sub-module is responsible for the object recognition. It N based on the previous generated annotations. First the neighbourhood relations of the extracted segments are computed and are represented by Graphs. Second, the object recognition is realized by graph operations triggered by a graph grammar. The graph grammar presents the transformation of the taxonomy, which reflects the domain knowledge.
The abundance of available multimedia information (e.g.videos, audio, images) requires efficient and effective annotation and retrieval methods. The IRIS system is designed for content-based retrieval of single images. Techniques and methods from computer vision and AI are combined in a new way within IRIS. The system has been tested with single images on several domains (e.g. landscape images, technical drawings) covering a wide range of applications As videos become a more important role in the frame of multimedia, we want to include them into the IRIS system. This paper describes the methods to divide the video into shots (scenes with a common content) and to create still images for each scene with the mosaicing technique. These images can be analyzed with the IRIS system. The first results of this approach are also shown in this paper.
The large amount of available multimedia information (e.g. videos, audio, images) requires eecient and eeective annotation and retrieval methods. The System IRIS (Image Retrieval for Information Systems), which was developed at the University of Bremen in the AI group, is designed for content-based retrieval of single images. As videos become a more important role in the frame of multimedia, we want to make videos available for IRIS.The rst step is the detection and extraction of shots from a video using a histogram based method. The second step is to combine the images in a shot to a single image. This image describes the shot (Mosaicing-technique) and can be analyzed with IRIS. 1 The IRIS-System In order to retrieve images it is more sophisticated and usual for human beings to use natural language concepts, e.g. mountainlake, than syntactical features, e.g. red region left up. This leads to a content-based image retrieval. Furthermore, it is unreasonable for any human being to make the content description for thousands of images manually. IRIS combines methods and techniques from computer vision and Artiicial Intelligence to generate content descriptions of images in a textual form automatically. The text retrieval can be done by an ordinary text-retrieval system. We use the IBM 1 SearchManager for AIX 2. The system is implemented on IBM RS6000 3 operating under AIX. The two dominating goals of this system are an automatic processing of the images and a comfortable user interface with a query vocabulary to formulate higher order queries by using concepts. Therefore the IRIS-System is divided into two modules; one for the image analysis to build up the database and a second one for the retrieval.
Christoph Klauck合作论文数University of Applied Sciences Hamburg1