Emerging cost-efficient depth sensor technologies reveal new possibilities to cope with difficulties in action recognition. Depth information improves the quality of skeleton detec- tion process, hence, pose estimation can be done more efficiently. Recently many studies fo- cus on temporal analyses over estimated skeleton poses to recognize actions. In this paper we have an inclusive study of the spatiotemporal kinematic features and propose an action recognition framework with feature selection capability to deal with the multitudinous of features by leveraging data mining capabilities of random decision forests. We describe human motion via a rich collection of kinematic feature time-series computed from the skel- etal representation of the body in motion. We discriminatively optimize a random decision forest model over this collection to identify the most effective subset of features, localized both in time and space. Later, we train a support vector machine classifier on the selected features. This approach improves upon the baseline performance obtained using the whole feature set with a significantly less number of features (one tenth of the original). To justify our method we test the framework on various datasets and compared it with state-of-the- art. On MSRC-12 dataset (25) (12 classes), our method achieves 94% accuracy. On the WorkoutSU-10 dataset (28), collected by our group (10 physical exercise classes), the ac- curacy is 98%. On MSR Action3D dataset (9) (20 classes) we obtain 87% average accura- cy and for UTKinect-Action dataset (10) (10 classes) the accuracy is 92%. Other than regular activities, we also tried our approach to detect a falling person using the dataset which we recorded as an extension to our original dataset. We test how our method adjusts on different types of actions and we obtained promising results for this type of action. We discuss that our approach provides insights on the spatiotemporal dynamics of human ac- tions and can be used to as part of different applications especially for rehabilitation of pa- tients.
In this paper, we present an action recognition framework leveraging data mining capabilities of random decision forests trained on kinematic features. We describe human motion via a rich collection of kinematic feature time-series computed from the skeletal representation of the body in motion. We discriminatively optimize a random decision forest model over this collection to identify the most effective subset of features, localized both in time and space. Later, we train a support vector machine classifier on the selected features. This approach improves upon the baseline performance obtained using the whole feature set with a significantly less number of features (one tenth of the original). On MSRC-12 dataset (12 classes), our method achieves 94% accuracy. On the WorkoutSU-10 dataset, collected by our group, the accuracy is 98%. The approach can also be used to provide insights on the spatiotemporal dynamics of human actions.
Real-time rendering of large animated crowds consisting of thousands of virtual humans is important for several applications including simulations, games, and interactive walkthroughs but cannot be performed using complex polygonal models at interactive frame rates. For that reason, methods using large numbers of precomputed image-based representations, called impostors, have been proposed. These methods take advantage of existing programmable graphics hardware to compensate for computational expense while maintaining visual fidelity. Thanks to these methods, the number of different virtual humans rendered in real time is no longer restricted by computational power but by texture memory consumed for the variety and discretization of their animations. This work proposes a resource-efficient impostor rendering methodology that employs image morphing techniques to reduce memory consumption while preserving perceptual quality, thus allowing higher diversity or resolution of the rendered crowds. Results of the experiments indicated that the proposed method, in comparison with conventional impostor rendering techniques, can obtain 38 % smoother animations or 87 % better appearance quality by reducing the number of key-frames required for preserving the animation quality via resynthesizing them with up to 92 % similarity on real time.
For the past two decades, the need for three-dimensional (3-D) scanning of industrial objects has increased significantly and many experimental techniques and commercial solutions have been proposed. However, difficulties remain for the acquisition of optically non-cooperative surfaces, such as transparent or specular surfaces. To address highly reflective metallic surfaces, we propose the extension of a technique that was originally dedicated to glass objects. In contrast to conventional active triangulation techniques that measure the reflection of visible radiation, we measure the thermal emission of a surface, which is locally heated by a laser source. Considering the thermophysical properties of metals, we present a simulation model of heat exchanges that are induced by the process, helping to demonstrate its feasibility on specular metallic surfaces and predicting the settings of the system. With our experimental device, we have validated the theoretical modeling and computed some 3-D point clouds from specular surfaces of various geometries. Furthermore, a comparison of our results with those of a conventional system on specular and diffuse parts will highlight that the accuracy of the measurement no longer depends on the roughness of the surface.
Real-time rendering of large animated crowds consisting thousands of virtual humans is important for several applications including simulations, games and interactive walkthroughs, but cannot be performed using complex polygonal models at interactive frame rates. For that reason, several methods using large numbers of pre-computed image-based representations, which are called as impostors, have been proposed. These methods take the advantage of existing programmable graphics hardware to compensate the computational expense while maintaining the visual fidelity. Making the number of different virtual humans, which can be rendered in real-time, not restricted anymore by the required computational power but by the texture memory consumed for the variety and discretization of their animations. In this work, we proposed an alternative method that reduces the memory consumption by generating compelling intermediate textures using image-morphing techniques. In order to demonstrate the preserved perceptual quality of animations, where half of the key-frames were rendered using the proposed methodology, we have implemented the system using the graphical processing unit and obtained promising results at interactive frame rates.
Digital music performance requires a high degree of interaction with input controllers that can provide fast feedback on the user's action. One of the primary considerations of professional artists is a powerful and creative tool that minimizes the number of steps required for the speed-demanding processes. Nowadays, mobile devices have become popular digital instruments for musical performance. Most of the applications designed for mobile devices use touch screen, keypad, or accelerometer as interaction modalities. In this paper, we present a novel interface for musical performance that is based on a magnetic interaction between a user and a device. The proposed method constitutes a touchless interaction modality that is based on the mutual effect between the magnetic field surrounding a device and that of a properly shaped magnet. Extending the interaction space beyond the physical boundary of a device provides the user with higher degree of flexibility for musical performance which, in turn, can open doors to a wide spectrum of new functionalities in digital music performance and production.
Future’s environments will be sensitive and responsive to the presence of people to support them carrying out their everyday life activities, tasks and rituals, in an easy and natural way. Such interactive spaces will use the information and communication technologies to bring the computation into the physical world, in order to enhance ordinary activities of their users. This paper describes a speech-based spoken multimedia retrieval system that can be used to present relevant video-podcast (vodcast) footage, in response to spontaneous speech and conversations during daily life activities. The proposed system allows users to search the spoken content of multimedia files rather than their associated meta-information and let them navigate to the right portion where queried words are spoken by facilitating within-medium searches of multimedia content through a bag-of-words approach. Finally, we have studied the proposed system on different scenarios by using vodcasts in English from various categories, as the targeted multimedia, and discussed how it would enhance people’s everyday life activities by different scenarios including education, entertainment, marketing, news and workplace.
In this paper, we introduce a revolutionary interaction framework that is based on the idea of around device interaction. The proposed method constitutes a touch-less data entry system that is based on the interaction between the magnetic fields around a device and a magnet. The magnetic field that surrounds the device is generated by a magnetic sensor (compass) that is embedded in the new generation of mobile phones. The movements of a permanent magnet in front of the device deforms the sensor's original magnetic field pattern whereby we can constitute a new means of communication between the user and the device. Thus, the magnetic field encompassing the device plays the role of a communication channel and encodes the hand-movement patterns of the user into temporal changes of the sensor's magnetic field. In the back-end of the communication, an engine samples the momentary status of the field during a trial and recognizes the user's pattern by matching it against some pre-recorded templates. The proposed method has been tested in a variety of applications (such as micro-interaction, handwriting recognition, user authentication, etc) and concluded in very promising results.
In this work, we present a hands clapping rhythm analysis module of a video analytics framework, which monitors elderly patients and automatically collect statistical data about patient activities. Hands clapping activity is analyzed in terms of frequency of clapping, extent of clapping, and direction change. A severe level Alzheimer patient was chosen from an elderly house. The main idea makes use of optical flow vectors which represent the motion change of image features in consecutive frames. The algorithm steps are composed of detecting optical flow vectors in skin regions, clustering based on the direction, calculating the average flow vector in each cluster and observing these vectors over time. The magnitude of the average flow represents the speed of motion. In the supplementary figure, handsclapping.png, the experimental results are presented. Hands motion of the patient on the right has been observed for 100 frames (4 secs). Input hands region, detected optical flows are demonstrated, followed by the two resultant motion flow groups depicted by black and white regions. The patient is active during 100 frames and claps hands eight times, two of which are long extent clapping, when the patient is very happy. In the graphs, blue lines represent the motion of right hand, while red lines represent left hand. The occurrence of clapping hands is detected by finding the instant, when right hand moves in (+) direction and changes direction to (-); and left hand moves in (-) direction and changes to (+); and the speed of each hand is greater than 2 units. It happens at frames: 4, 11,17,28,53,63,68,87. The graph in the bottom shows the distance traveled by each hand per frame. The symmetry in motion waves of right and left hand depicts the clapping motion characteristics and validates effectiveness of the proposed method.
Advances in the medical imaging technology has lead to an exponential growth in the number of digital images that needs to be acquired, analyzed, classified, stored and retrieved in medical centers. As a result, medical image classification and retrieval has recently gained high interest in the scientific community. Despite several attempts, such as the yearly-held ImageCLEF Medical Image Annotation Challenge, the proposed solutions are still far from being sufficiently accurate for real-life implementations. In this paper we summarize the technical details of our experiments for the ImageCLEF 2009 medical image annotation challenge. We use a direct and two ensemble classification schemes that employ local binary patterns as image descriptors. The direct scheme employs a single SVM to automatically annotate X-ray images. The two proposed ensemble schemes divide the classification task into sub-problems. The first ensemble scheme exploits ensemble SVMs trained on IRMA sub-codes. The second learns from subgroups of data defined by frequency of classes. Our experiments show that ensemble annotation by training individual SVMs over each IRMA sub-code dominates its rivals in annotation accuracy with increased process time relative to the direct scheme.
It is important for drowsiness detection systems to identify different levels of drowsiness and respond appropriately at each level. This study explores how to discriminate moderate from acute drowsiness by applying computer vision techniques to the human face. In our previous study, spontaneous facial expressions measured through computer vision techniques were used as an indicator to discriminate alert from acutely drowsy episodes. In this study we are exploring which facial muscle movements are predictive of moderate and acute drowsiness. The effect of temporal dynamics of action units on prediction performances is explored by capturing temporal dynamics using an over complete representation of temporal Gabor Filters. In the final system we perform feature selection to build a classifier that can discriminate moderate drowsy from acute drowsy episodes. The system achieves a classification rate of .96 A' in discriminating moderately drowsy versus acutely drowsy episodes. Moreover the study reveals new information in facial behavior occurring during different stages of drowsiness.
This paper presents a new active contour-based, statistical method for simultaneous volumetric segmentation of multiple subcortical structures in the brain. In biological tissues, such as the human brain, neighboring structures exhibit co-dependencies which can aid in segmentation, if properly analyzed and modeled. Motivated by this observation, we formulate the segmentation problem as a maximum a posteriori estimation problem, in which we incorporate statistical prior models on the shapes and intershape (relative) poses of the structures of interest. This provides a principled mechanism to bring high level information about the shapes and the relationships of anatomical structures into the segmentation problem. For learning the prior densities we use a nonparametric multivariate kernel density estimation framework. We combine these priors with data in a variational framework and develop an active contour-based iterative segmentation algorithm. We test our method on the problem of volumetric segmentation of basal ganglia structures in magnetic resonance images. We present a set of 2-D and 3-D experiments as well as a quantitative performance analysis. In addition, we perform a comparison to several existent segmentation methods and demonstrate the improvements provided by our approach in terms of segmentation accuracy.
The puzzle-assembly problem has many application areas such as restoration and reconstruction of archeological findings, repairing of broken objects, solving jigsaw type puzzles, molecular docking problem, etc. The puzzle pieces usually include not only geometrical shape information but also visual information such as texture, color, and continuity of lines. This paper presents a new approach to the puzzle-assembly problem that is based on using textural features and geometrical constraints. The texture of a band outside the border of pieces is predicted by inpainting and texture synthesis methods. Feature values are derived from these original and predicted images of pieces. An affinity measure of corresponding pieces is defined and alignment of the puzzle pieces is formulated as an optimization problem where the optimum assembly of the pieces is achieved by maximizing the total affinity measure. A Fast Fourier Transform based image registration technique is used to speed up the alignment of the pieces. Experimental results are presented on real and artificial data sets.
This paper develops a recursive method for computing moments of 2D objects described by elliptic Fourier descriptors (EFD). To this end, Green's theorem is utilized to transform 2D surface integrals into 1D line integrals and EFD description is employed to derive recursions for moments computations. A complexity analysis is provided to demonstrate space and time efficiency of our proposed technique. Accuracy and speed of the recursive computations are analyzed experimentally and comparisons with some existing techniques are also provided.
Advances in the medical imaging technology has lead to an exponential growth in the number of digital images that need to be acquired, analyzed, classified, stored and retrieved in medical centers. As a result, medical image classification and retrieval has recently gained high interest in the scientific community. Despite several attempts, the proposed solutions are still far from being sufficiently accurate for real-life implementations. In a previous work, performance of different feature types were investigated in a SVM-based learning framework for classification of X-Ray images into classes corresponding to body parts and local binary patterns were observed to outperform others. In this paper, we extend that work by exploring the effect of attribute selection on the classification performance. Our experiments show that principal component analysis based attribute selection manifests prediction values that are comparable to the baseline (all-features case) with considerably smaller subsets of original features, inducing lower processing times and reduced storage space.
Gwen Littlewort合作论文数Machine Perception Laboratory;;Institute for Neural Computation10
Selim Balcisoy合作论文数Sabanci University3
David Fofi合作论文数Laboratoire Le2i UMR CNRS 6306
IUT Le Creusot (dpt MP)
Universite de Bourgogne3