In this paper we present a system that detects and tracks objects and agents, computes spatial relations, and communicates those relations to the user using speech. Our system is able to detect multiple objects and agents at 30 frames per second using a RGBD camera. It is able to extract the spatial relations in, on, next to, near, and belongs to, and communicate these relations using natural language. The notion of belonging is particularly important for Human-Robot Interaction since it allows the robot ground the language and reason about the right objects. Although our system is currently static and targeted to a fixed location in a room, we are planning to port it to a mobile robot thus allowing it explore the environment and create a spatial knowledge base.
Home robots may come with many sophisticated built-in abilities, however there will always be a degree of customization needed for each user and environment. Ideally this should be accomplished through one-shot learning, as collecting the large number of examples needed for statistical inference is tedious. A particularly appealing approach is to simply explain to the robot, via speech, what it should be doing. In this paper we describe the ALIA cognitive architecture that is able to effectively incorporate user-supplied advice and prohibitions in this manner. The functioning of the implemented system on a small robot is illustrated by an associated video.
We describe the effective use of online learning to enhance the conversational capabilities of a concierge robot that we have been developing over the last two years. The robot was designed to interact naturally with visitors and uses a speech recognition system in conjunction with a natural language classifier. The online learning component monitors interactions and collects explicit and implicit user feedback from a conversation and feeds it back to the classifier in the form of new class instances and adjusted threshold values for triggering the classes. In addition, it enables a trusted master to teach it new question-answer pairs via question-answer paraphrasing, and solicits help with maintaining question-answer-class relationships when needed, obviating the need for explicit programming. The system has been completely implemented and demonstrated using the SoftBank Robotics [34] humanoid robots Pepper and NAO, and the telepresence robot known as Double from Double Robotics [4].
Big data refers to huge information whose quantity is beyond the capacity of the system to manage, process and capture. On the basis of the biological and behavioral characteristics of the human, information gathered from which person can be recognized and referred as biometrics. Example includes finger print, face, voice and behavioral analysis etc. Among these, face recognition will not make use of any physical contact with the biometrics system; it is more secure and effective. Hence face recognition becomes one of the best technologies in the biometrics field. Although plenty of efforts have been carried out in the field of face recognition, this work has mark with many challenges in general settings. The criteria behind the successful face recognition systems are developed only under constrained situations with small sized data bases. Hence there is necessary to work under general setting that fit to big data for face recognition. This paper, aims to increase performance and accuracy of face recognition by considering face spoofing attack (veracity). Result analysis of our proposed work ensures increase in accuracy compared to other methods.
Protecting the privacy of the fingerprint in authentication systems is become a major issue now-a-days because of the widespread use of fingerprint recognition systems. Traditional encryption and transformation techniques are shown to be more vulnerable to attacks. Therefore, fingerprint combination at the image and feature level has been proposed. This paper introduces two approaches for protecting fingerprint privacy by combining two different fingerprints into a new identity. This paper compares two systems that were introduced to protect the privacy of fingerprint. First is a novel system for fingerprint privacy protection by mixing features of two different fingerprints and thus generate a new identity. During enrolment, the system captures left and right thumb impression from a user. The new identity contains minutiae points of right thumb and has an orientation of left thumb impression. Second is a technique that combines minutiae features of two different fingerprints of a user. The minutiae points of each fingerprints is protected in the new identity. In addition minutiae filtering is done in order to remove spurious minutiae for improving the performance of both the systems. Finally the performance of each technique in terms of FRR, ERR and FAR is compared. For evaluating the performance of two techniques, this work uses same algorithms for the pre-processing and post-processing of fingerprint image.
In this paper we present the modeling strategies that were applied by the IBM Research team to the medical modality classification, retrieval and compound figure separation tasks of ImageCLEF 2013. We present our methods for each task and discuss our submitted textual, visual, and mixed runs, as well as their results, the use of external resources and human supervision. The key components of our modality classification submissions were: 1)fusion of multiple low level image descriptors extracted at different spatial granularities 2) use of pre-existing medical categories classifiers trained from web sources 3) pseudo-probabilistic analysis of modalityspecific derived text patterns and 4) multiple fusion strategies to combine visual and textual information. For the case based retrieval task, we applied topic modeling on top of text extracted from the full set of Pubmed articles, using three different corpuses to guide an expansion: one from UMLS keywords, one from WordNet keywords, and one from the union of the previous two. Retrieval was performed using Lucene indexing on top of such representations. For the image based retrieval task, we adopted visual image similarity based on CHI square distance between low level visual descriptors. In the compound figure separation task we tried a combination of two approaches, one based on a connected components analysis in a binarized image, the other adopting the common notation of subfigures using text.
Management and monitoring of data centers is a growing field of interest, with much current research, and the emergence of a variety of commercial products aiming to improve performance, resource utilization and energy efficiency of the computing infrastructure. Despite the large body of work on optimizing data center operations, few studies actually focus on discovering and tracking the physical layout of assets in these centers. Such asset tracking is a prerequisite to faithfully performing administration and any form of optimization that relies on physical layout characteristics. In this work, we describe an approach to completely automated asset tracking in data centers, employing a vision-based mobile robot in conjunction with an ability to manipulate the indicator LEDs in blade centers and storage arrays. Unlike previous large-scale asset-tracking methods, our approach does not require the tagging of assets (e.g., with RFID tags or barcodes), thus saving considerable expense and human labor. The approach is validated through a series of experiments in a production industrial data center.
In this study, we propose a novel biometric signature for human identification based on anatomically unique structures of the left ventricle of the heart. An algorithm is developed that analyzes the 3 primary anatomical structures of the left ventricle: the endocardium, myocardium, and papillary muscles. Comparisons of these analyses between probe and gallery images produces a similarity score that is used as the basis of the biometric. The performance of the algorithm is tested on a cohort of 10 de-identified subjects imaged by Cardiac MRI. Perfect matching between individuals is obtained with good separation between the genuine and impostor classes. In summary, this study demonstrates using anatomy of the left ventricle of the human heart for the purposes of a biometric signature.
This paper describes our Extensible Language Interface (ELI) for robots. The system is intended to interpret far-field speech commands in order to perform fetch-and-carry tasks, potentially for use in an eldercare context. By "extensible" we mean that the robot is able to learn new nouns and verbs by simple interaction with its user. An associated video [1] illustrates the range of phenomena handled by our implemented real-time system.
We will demonstrate a robot for data center energy management, in action, on a simulated data center floor. We shall highlight the robot's navigation, tile and obstacle classification, event scheduling and preemption capabilities, along with its ability to discover charging docks, and successfully dock with extreme precision. We shall also show simulations on real data center layouts evincing navigational efficiency gains obtained by our latest heuristic enhancements.
We describe an inexpensive robot that serves as a physical autonomic element, capable of navigating, mapping and monitoring data centers with little or no human involvement, even ones that it has never seen before. Through a series of real experiments and simulations, we establish that the robot is sufficiently accurate, efficient and robust to be of practical benefit in real data center environments. We demonstrate how the robot's integration with Maximo for Energy Optimization, a commercial data center energy management product, supports autonomic management at the level of the data center as a whole, particularly self-diagnosis of emerging thermal problems.
We describe an inexpensive autonomous robot capable of navigating previously unseen data centers and monitoring key metrics such as air temperature(1). The robot provides real-time navigation and sensor data to commercial IBM software, thereby enabling real-time generation of the data center layout, a thermal map and other visualizations of energy dynamics. Once it has mapped a data center, the robot can efficiently monitor it for hot spots and other anomalies using intelligent sampling. We demonstrate the robot's effectiveness via experimental studies from two production data centers.
The texture in a human iris has been shown to have good individual distinctiveness and thus is suitable for use in reliable identification. A conventional iris recognition system unwraps the iris image and generates a binary feature vector by quantizing the response of selected filters applied to the rows of this image. Typically there are 360 angular sectors, 64 radial rings, and 2 filter responses. This produces a full-length iris code (FLIC) of about 5760 bytes. In contrast, this paper seeks to shrink the representation by finding those regions of the iris that contain the most descriptive potential. We show through experiments that the regions close to the pupil and sclera contribute least to discrimination, and that there is a high correlation between adjacent radial rings. Using these observations we produce a short-length iris code (SLIC) of only 450 bytes. The SLIC is an order of magnitude smaller the FLIC and yet has comparable performance as shown by results on the MMU2 database. The smaller sized representation has the advantage of being easier to store as a barcode, and also reduces the matching time per pair.
There are a wide variety of approaches to Artificial Intelligence. Yet interestingly we find that these can all be grouped into four broad categories: Silver Bullets, Core Values, Emergence, and Emulation. We will explain the methodological underpinnings of these categories and give examples of the type of work being pursued in each. Understanding this spectrum of approaches can help defuse arguments between practitioners as well as elucidate common themes.
One of the main challenges in building an efficient and scalable automatic fingerprint identification system is to identify features which are highly discriminative and are reproducible across different prints of the same finger. Most existing fingerprint matching approaches rely on minutiae geometry. Relatively, little effort has gone into analyzing ridge flow patterns present in the fingerprint, partly due to difficulty in extracting robust discriminative features from the fingerprint images. In this paper, we analyze the usefulness of ridge curvature information for fingerprint matching and classification applications. Specifically, for an indexing framework, we explore whether the curvature information can be utilized along with the existing minutiae geometry-based features for further reducing the number of potential candidates for fingerprint identification. Experimental results indicate the robustness of the proposed curvature-based characterization and its usefulness in improving the efficiency of existing fingerprint-based identification systems.
AI has many techniques and tools at its disposal, yet seems to be lacking some special “juice” needed to create a true being. We propose that the missing ingredients are a general theory of motivation and an operational understanding of natural language. The motivation part comes largely from our animal heritage: a real-world agent must continually respond to external events rather than depend on perfect modeling and planning. The language part, on the other hand, is what makes us human: competent participation in a social group requires one-shot learning and the ability to reason about objects and activities that are not present or on-going. In this paper we propose an architecture for self-motivation, and suggest how a language interpreter can be built on top of such a substrate. With the addition of a method for recording and internalizing dialog, we sketch how this can then be used to impart essential cultural knowledge and behaviors.