In 2010, the Institute of Computer Science of the Foundation for Research and Technology-Hellas (ICS-FORTH) and the Archaeological Museum of Thessaloniki (AMTh) collaborated towards the creation of a special exhibition of prototypical interactive systems with subjects drawn from ancient Macedonia, named “Macedonia from fragments to pixels”. The exhibition comprises seven interactive systems based on the research outcomes of ICS-FORTH's Ambient Intelligence Programme. Up to the summer of 2012, more than 165.000 people have visited it. The paper initially provides some background information, including related previous research work, and then illustrates and discusses the development process that was followed for creating the exhibition. Subsequently, the technological and interactive characteristics of the project's outcomes (i.e., the interactive systems) are analysed and the complementary evaluation approaches followed are briefly described. Finally, some conclusions stemming from the project are highlighted.
This paper presents a computer vision system that supports non-instrumented, location-based interaction of multiple users with digital representations of large-scale artifacts. The proposed system is based on a camera network that observes multiple humans in front of a very large display. The acquired views are used to volumetrically reconstruct and track the humans robustly and in real time, even in crowded scenes and challenging human configurations. Given the frequent and accurate monitoring of humans in space and time, a dynamic and personalized textual/graphical annotation of the display can be achieved based on the location and the walk-through trajectory of each visitor. The proposed system has been successfully deployed in an archaeological museum, offering its visitors the capability to interact with and explore a digital representation of an ancient wall painting. This installation permits an extensive evaluation of the proposed system in terms of tracking robustness, computational performance and usability. Furthermore, it proves that computer vision technology can be effectively used to support non-instrumented interaction of humans with their environments in realistic settings.
The theme of this paper is an exhibition of prototypical interactive systems with subjects drawn from ancient Macedonia, named "Macedonia from fragments to pixels". Since 2010, the exhibition is hosted by the Archaeological Museum of Thessaloniki and is open daily to the general public. Up to now, more than 165.000 people have visited it. The exhibition comprises 7 interactive systems which are based on some research outcomes of the Ambient Intelligence Programme of the Institute of Computer Science, Foundation for Research and Technology - Hellas. The digital content of these systems includes objects from the Museum’s permanent collection and from Macedonia.
We present work on exploiting modern graphics hardware towards the real-time production of a textured 3D mesh representation of a scene observed by a multicamera system. The employed computational infrastructure consists of a network of four PC workstations each of which is connected to a pair of cameras. One of the PCs is equipped with a GPU that is used for parallel computations. The result of the processing is a list of texture mapped triangles representing the reconstructed surfaces. In contrast to previous works, the entire processing pipeline (foreground segmentation, 3D reconstruction, 3D mesh computation, 3D mesh smoothing and texture mapping) has been implemented on the GPU. Experimental results demonstrate that an accurate, high resolution, texture-mapped 3D reconstruction of a scene observed by eight cameras is achievable in real time.
This paper describes the outcomes stemming from the work of a multidisciplinary R&D project of ICS-FORTH, aiming to explore and experiment with novel interactive museum exhibits, and to assess their utility, usability and potential impact. More specifically, four interactive systems are presented in this paper which have been integrated, tested and evaluated in a dedicated, appropriately designed, laboratory space. The paper also discusses key issues stemming from experience and observations in the course of qualitative evaluation sessions with a large number of participants.
This paper presents a system that supports the exploration of digital representations of large-scale museum artifacts in through non-instrumented, location-based interaction. The system employs a state-of-the-art computer vision system, which localizes and tracks multiple visitors. The artifact is presented in a wall-sized projection screen and it is visually annotated with text and images according to the location as well as walkthrough trajectories of the tracked visitors. The system is evaluated in terms of computational performance, localization accuracy, tracking robustness and usability.
In this paper, the design and implementation of a hardware/software platform for parallel and distributed multiview vision processing is presented. The platform is focused at supporting the monitoring of human presence in indoor environments. Its architecture is focused at increased throughput through process pipelining as well as at reducing communication costs and hardware requirements. Using this platform, we present efficient implementations of basic visual processes such as person tracking, textured visual hull computation and head pose estimation. Using the proposed platform multiview visual operations can be combined and third-party ones integrated, to ultimately facilitate the development of interactive applications that employ visual input. Computational performance is benchmarked comparatively to state of the art and the efficacy of the approach is qualitatively assessed in the context of already developed applications related to interactive environments.
We present the development of a multi-touch display based on computer vision techniques. The developed system is built upon low cost, off-the-shelf hardware components and a careful selection of computer vision techniques. The resulting system is capable of detecting and tracking several objects that may move freely on the surface of a wide projection screen. It also provides additional information regarding the detected and tracked objects, such as their orientation, their full contour, etc. All of the above are achieved robustly, in real time and regardless of the visual appearance of what may be independently projected on the projection screen. We also present indicative results from the exploitation of the developed system in three application scenarios and discuss directions for further research.
In this paper, the application of computer vision techniques to the localization of multiple persons in a relatively wide gaming terrain is presented. Multiple views are employed both for terrain coverage, but most importantly, for treatment of occlusions. Through the appropriate selection of lightweight operations and acceleration strategies, an adequate frame rate is achieved despite the large volume of input data. The resulting system is employed in the development of multiplayer entertainment applications, which are demonstrated and evaluated.
Camera networks are increasingly employed in a wide range of Computer Vision applications, from modelling and interpretation of individual human behaviour to the surveillance of wide areas. In most cases, the evidence gathered by individual cameras is fused together, making the synchronization of acquired images a crucial task. Cameras are typically hosted on multiple computers in order to accommodate the large number of acquired images and provide the computational resources required for their processing. In the application layer, vision processing is thus supported by multiple processing nodes (CPUs, GPUs or DSPs). The proposed platform is able to handle the considerable technical complexity involved in the synchronous acquisition of images and the allocation of processes to nodes. Figure 1 illustrates an overview of the proposed and implemented architecture.
This paper presents the process and tangible outcomes of a rapid prototyping activity towards the creation of a demonstrator, showcasing the potential use and effect of Ambient Intelligence technologies in a typical office environment. In this context, the hardware and software components used are described, as well as the interactive behavior of the demonstrator. Additionally, some conclusions stemming from the experience gained are presented, along with pointers for future research and development work.
3D head pose estimation constitutes a special problem of human motion modeling. An accurate and robust solution to this problem is of particular interest, because the 3D head pose of a human conveys important information on his/her behavior. Significant advances have been achieved in human head pose estimation for relatively close-range images, but the related available methods are not directly applicable in wider-range imaging conditions. The proposed method is overviewed in Fig. 1. The visual hull of a person is obtained from images acquired synchronously from multiple viewpoints. While moving, the person’s head is tracked in 3D employing a variant of the Mean-Shift algorithm and a spherical kernel. The texture on the surface of the hull is collected from multiple views and projected on a hypothetical sphere S that is concentric to the person’s head. This gives rise to spherical image Is within which face detection is simplified, because exactly one frontal face is guaranteed to appear in it at a known spatial scale.
Computational Vision and Robotics LaboratoryInstitute of Computer ScienceFoundation for Research and Technology — Hellas (FORTH)N. Plastira 100, Vassilika VoutonHeraklion, Crete, 700 13 GreeceWeb: http://www.ics.forth.gr/cvrlE-mail: {sarmis | zabulis | argyros}@ics.forth.grTel: +30 2810 391600, Fax: +30 2810 391601
We present a new approach for the detection of events in image sequences. Our method relies on a number of logical sensors that can be defined over specific regions of interest in the viewed scene. These sensors measure time varying image properties that can be attributed to primitive events of interest. Thus, the logical sensors can be viewed as a means to transform image data to a set of symbols that can assist event detection and activities interpretation. On top of these elementary sensors, temporal and logical aggregation mechanisms are used to define hierarchies of progressively more complex sensors, able to detect events having more complex semantics. Finally, scenario verification mechanisms are employed to achieve process monitoring, by checking whether events occur according to a predetermined order. The proposed framework has been tested and validated in an application involving monitoring of automated processes. The obtained results demonstrate that the proposed approach, despite its simplicity, provides a promising framework for vision based event detection in the context of such applications.