This paper presents SPIKE, an automated algorithm to generate best practices by analyzing Storage Area Network (SAN) configuration errors. Best practices are a useful tool in problem diagnosis as most configuration problems are caused by the violation of best practices in the storage network domain. However, the manual generation of best practices is tedious, error-prone and costly in terms of time and manpower. SPIKE uses a combination of information-retrieval principles, entity ranking and decision-tree classification to statistically infer the best practices for the prevention of SAN configuration problems. Preliminary results from an initial implementation of SPIKE indicate speed and accuracy improvements over manually generating best practices.
Multivariate time series (MTS) datasets are common in various multimedia, medical and financial applications. In order to efficiently perform k nearest neighbor searches for MTS datasets, we present a similarity measure, Eros (extended Frobenius norm), an index structure, Muse (multilevel distance-based index structure for Eros), and a feature subset selection technique, Ropes (recursive feature elimination on common principal components for Eros). Eros is based on principal component analysis, and computes the similarity between two MTS items by measuring how close the corresponding principal components are using the eigenvalues as weights. Muse constructs each level as a distance-based index structure without using the weights, up to z levels, which are combined at the query time with the weights. Ropes utilizes both the common principal components and the weights recursively in order to select a subset of features for Eros. The experimental results show the superiority of our techniques as compared to earlier approaches.
Time series is a series of observations over time. When there is one observation at each time instance, it is called a univariate time series (UTS), and when there are more than one observations, it is called a multivariate time series (MTS). While UTS datasets have been extensively explored, MTS datasets have not been broadly investigated. The techniques for UTS datasets, however, cannot be simply extended for MTS datasets, since multivariate time series is different from multiple univariate time series. That is, an MTS item may not be broken into multiple univariate time series and be separately analyzed, because this will result in the loss of the correlation information within the multivariate time series. In this dissertation, we introduce a set of techniques for multivariate time series analysis based on principal component analysis (PCA). As a similarity measure for MTS datasets, we present Eros (Extended Frobenius norm). Eros computes the similarity between two MTS items by comparing the corresponding principal components and using the variances that the principal components represent as weights. For efficient retrieval of MTS items using Eros, we introduce an index structure for Eros, termed Muse (Multilevel distance-based index structure for Eros). Given a query item, Muse first utilizes the lower bound of Eros to filter out the MTS items that are not to be in the set of k Nearest Neighbors. Subsequently, Muse refines the MTS items that are not filtered out by employing Eros in order to exactly identify the k Nearest Neighbors of the given query item. Inherently, an MTS item is very high dimensional. Hence, it is, in general, beneficial to reduce the dimension of the dataset before applying data mining techniques, e.g., classification and clustering, which results in the elimination of irrelevant and/or redundant data. For Eros, we present a feature subset selection technique, termed Ropes ( Recursive Feature Elimination on Common Principal Components for Eros). Ropes utilizes the common principal components and the weights recursively in order to select a subset of features for Eros. In addition, utilizing the correlation information and Eros, we introduce a set of feature subset selection and feature extraction techniques for multivariate time series datasets, such as Corona (Correlation as Features), CLeVer (descriptive Common principal component Loading based Variable subset selection) and KEros. Corona is a supervised feature subset selection technique, which first represents an MTS item using the correlation coefficients, and recursively eliminates at each time one of the features based on the contribution to the classification decision boundary. CLeVer is an unsupervised feature subset selection technique, which performs the feature subset selection based on the contribution to the common principal components. KEros performs the feature extraction based on the Kernel PCA technique using Eros as the similarity measure between two MTS items. With the advent of various sensing techniques, there are cases where each data is represented in an n-way array, where n is greater than 2. One of the examples would be the functionalMagnetic Resonance Imaging (fMRI) data, where each data is represented in a 3-way array, and an fMRI stream is represented in a 4-way array. An n-way array may be flattened into a matrix, where, for example, Eros can be applied. However, this flattening may result in the loss of the spatial correlation. In order to address this problem, we extended Eros to these n-way array datasets, termed nEros (n-way Eros). Intuitively, for an n-way array, there are n ways of unfolding it into a matrix. For each fold, we perform Eros, and sum up the n results into one similarity value. Our experimental evaluation employing various real-world and synthetic datasets shows that the presented techniques based on the correlation information within the MTS items perform better than traditional approaches that do not utilize the correlation information, e.g., Euclidean distance.
Four real-world case studies demonstrate the effectiveness of using immersidata - the normally untapped data from interchanges between users and immersive 3D environments such as computer games and virtual reality - to help understand user behavior and experiences.
We have used two aminoglycosides, G418 and paromomycin, to develop a reliable selection system fornptll transgenic sweet-potato (Ipomoea batatas (L.) Lam.). Embryogenic calli derived from shoot apical meristems were bombarded with gold particles coated with pCAMBIA2301, which contained thenptll andgusA genes. When compared on a kill curve that was based on calli proliferation and cell viability, G413-selection proved to be more efficient and had fewer escapes than kanamycin. These bombarded expiants were then selected on G418-containing media. The total time required from bombardment to plant establishment in soil was seven to nine months. Multiple copies of the transgene were integrated into the sweetpotato genome. Northern analysis confirmed transgene expression in the regenerated plants, and a paromomycin assay demonstrated that thenptll gene was functionally expressed in transformed sweetpotato. These molecular analyses and assays all showed that selection with G418 and paromomycin is reliable. So far, we have produced 69 transgenic events with this system, at a transformation frequency of approx. 1.1%. That efficiency is based on the number of transgenic plants obtained and the amount of calli bombarded. Thus, this selection method that combines G418 with paromomycin is now available for selectingnptll transgenic sweetpotato.
and addresses the challenges involved in its acquisition, storage, query, and analysis. To illustrate immersidata’s importance, we report four real-world case studies in the medical and educational domains. We show that immersidata can be as revealing, if not more so, as the results obtained with other experimental monitoring and observation, data collection, and analysis techniques. Our purpose, however, isn’t to replace these techniques or the human experimenters, but to provide a complementary approach that allows for data that would normally be lost to be captured and analyzed. These four case studies represent our preliminary work to showcase immersidata’s usefulness. This collection of case studies demonstrates that a general data architecture for acquisition, query, and analysis of immersidata can be beneficial to many fields, regardless of their focused applications.
We describe two approaches to aid in game design, evaluation and development for user-players staying there , continuing to engage in game activities. The first is the hierarchical activity-based scenario (HABS) approach providing a theoretical framework to support the design of game narrative and scenario, model and reason about user-players' behavior and experience from acting in the scenario, and help identify problematic aspects of game design. The second is a continuous and unobtrusive approach that supports evaluation. Central to our approach is a tool called ISIS (Immersidata analySIS) to query and identify data of interest and to index events within virtual or video recordings, or graphical visualizations of game sessions. Used in conjunction with HABS, analysis of the associated data and indexed events helps us to understand user-players' behaviour and experience and aids in the detection of design problems to inform game development. To demonstrate our approaches we describe how they have been utilized in the development of an educational serious game.
Integrated Media Systems Center University of Southern California Los Angeles, CA 90089-0781, USA marsht@usc.edu Shamus P. Smith Department of Computer Science Durham University Durham DH1 3LE, United Kingdom shamus.smith@durham.ac.uk Kiyoung Yang Computer Science Department University of Southern California Los Angeles, CA 90089-0781, USA kiyoungy@usc.edu Cyrus Shahabi Integrated Media Systems Center & Computer Science Department University of Southern California Los Angeles, CA 90089-0781, USA shahabi@usc.edu We describe a continuous and unobtrusive approach to capture data amassed from user-player interactions with virtual or game environments. Central to this is a tool called ISIS (Immersidata analySIS) to query and identify data of interest and to index events within video recordings of game sessions. Analysis of the associated data and video clips help us to understand user-players’ behaviour and experience to assess and inform the design and development of games. ISIS supports six queries to identify: actions and activities, breaks in interaction caused by reflection or ineffective and problematic design, navigation problems caused by user disorientation, and events or tasks that are the most difficult to perform in a game. In the development of an educational serious game, we illustrate how our approach can help inform redesign.
New and emerging interactive digital media present many new challenges to human-computer interaction (HCI). With digital media designed to captivate user’s attention through stimulating experience and so encourage them in staying there engaged in pursuing activities, one of the main challenges is to devise effective and appropriate evaluation and development approaches that don’t disrupt the user. While standard evaluation methods can provide insightful information that can be used to inform development, a range of limitations have been identified. In this paper we describe the seamless capture and management of data that represents user’s interchanges with interactive digital media. Next we describe two complementary tools that we have developed to help analyze and interpret these interchanges and help us to understand user’s experience and behavior to inform development throughout phases of the life cycle. We demonstrate the effectiveness of our approaches through their application to an educational serious game. Categories & Subject Descriptors H5.2 [Information Interfaces and Presentation]: User Interfaces Evaluation/Methodology; H3.3 [Information Storage and Retrieval]: Information Search and Retrieval
Feature subset selection (FSS) is one of the data pre-processing techniques to identify a subset of the original features from a given dataset before performing any data mining tasks. We propose a novel FSS method for Multivariate Time Series (MTS) based on Common Principal Components, termed CL e V er. It utilizes the properties of the principal components to retain the correlation information among original features while traditional FSS techniques, such as Recursive Feature Elimination (RFE), may lose it. In order to evaluate the effectiveness of our selected subset of features, classification is employed as the target data mining task. Our experiments show that CL e V er outperforms RFE and Fisher Criterion by up to a factor of two in terms of classification accuracy, while requiring up to 2 orders of magnitude less processing time.
We describe our experiences of the 2020Classroom, an on-going project to develop a three-dimensional immersive learning environment through a game called Metalloman to teach bioscience concepts to engineering undergraduate students. Specifically, this paper focuses on work towards the development of a methodology through the refinement of techniques from HCI, activity theory and digital game design to inform design and evaluation for enjoyable, engaging and motivating, as well as usable immersive learning environments. Using these methods, studies to evaluate usability and user experience to inform redesign are described and preliminary studies to assess learning outcomes are outlined.
This paper describes an approach towards automating the identification of design problems with three-dimensional mediated or gaming environments through the capture and query of user-player behavior represented as a data schema that we have termed "immersidata". Analysis of data from a study of an educational computer game that we are developing shows that this approach is an effective way to pinpoint potential usability or design problems occurring in unfolding situational and episodic events that can interrupt or break user experience. As well as informing redesign, a key advantage of this cost-effective approach is that it considerably reduces the time evaluators spend analyzing hours of videoed study material.
It is argued that the greater a user perceives him/herself to be vicariously in character or is able to empathize with other characters/humans, the more they have a sense of being connected to a mediated environment. The term coined to describe this sense of user engagement is “vicariously there”. In this article we provide a framework of vicarious and empathic experience in mediated environments and review previous work and their measures. Focusing on three-dimensional interactive mediated environments (IME: digital games, virtual reality, virtual environments, etc.), we describe on-going research towards the development of ways to reason about the extent to which users feel a sense of engagement with, or connection to, characters or users. Limitations of this work are identified and future research directions towards an unobtrusive and continuous method are discussed.
Feature subset selection (FSS) is a known technique to pre-process the data before performing any data mining tasks, e.g., classification and clustering. FSS provides both cost-effective predictors and a better understanding of the underlying process that generated data. We propose Corona, a simple yet effective supervised feature subset selection technique for Multivariate Time Series (MTS). Traditional FSS techniques, such as Recursive Feature Elimination (RFE) and Fisher Criterion (FC), have been applied to MTS datasets, e.g., Brain Computer Interface (BCI) datasets. However, these techniques may lose the correlation information among MTS variables, since each variable is considered separately when an MTS item is vectorized before applying RFE and FC. Corona maintains the correlation information by utilizing the correlation coefficient matrix of each MTS item as features to be employed for SVM. Our exhaustive sets of experiments show that Corona consistently outperforms RFE and FC by up to 100% in terms of classification accuracy, and takes more than one order of magnitude less time than RFE and FC in terms of the overall processing time.
Multivariate time series (MTS) data sets are common in various multimedia, medical and financial application domains. These applications perform several data-analysis operations on large number of MTS data sets such as similarity searches, feature-subset-selection, cluster ing and classification. Inherently, an MTS item has a large number of dimensions. Hence, before applying data mining techniques, some form of dimension reduction, e.g., feature extraction, should be performed. Principal Component Analysis (PCA) is one of the techniques that have been frequently utilized for dimension reduction. However, traditional PCA does not scale well in terms of dimensionality, and therefore may not be applied to MTS data sets. The Kernel PCA technique addresses this problem of scalability by utilizing the kernel trick. In this paper, we propose a PCA based kernel to be employed for the Kernel PCA technique on the MTS data sets, termed KEros, which is based on Eros, a PCA based similarity measure for MTS data sets. We evaluate the performance of KEros using Support Vector Machine (SVM), and compare the performance with Kernel PCA using linear kernel and Generalized Principal Component (GPCA). The experimental results show that KEros outperforms these other techniques in terms of classificati on
We present a continuous and unobtrusive approach to analyze and reason about users' personal experiences of interacting with virtual and game environments. Focusing on an immersive educational game environment that we are developing, this is achieved through the capture and storage of user's movements and events that occur as a result of interactions with and within immersive environments. Termed immersidata , we then query and analyze immersidata to make sense of user behavior.Two example approaches are described. The first describes an application ISIS (Immersidata analySIS) that provides a tool for analysis of user behavior/experience through the indexing of immersidata with video clips of students' gaming sessions. This approach is described by way of an example to identify the causes of interruptions or breaks in interactions/focus of attention to facilitate the identification of problematic design. In our second example we describe our work towards classifying students' performance through immersidata. To this aim, we describe one example of transforming immersidata into multivariate time series and then by applying feature subset selection techniques we identify the features that differentiate students. We describe the application of this approach to identify novice and expert players with 90\% accuracy. One proposal is to use this to customize the game environment appropriate to the students' ability. Finally, we present future directions for the continuation of the work presented herein and also, the application of the immersidata system to capture, store and analyze personal behavior/experiences and provide appropriate feedback in our work and home environments.
Multivariate time series (MTS) datasets are common in various multimedia, medical and financial applications. We propose a similarity measure for MTS datasets, Eros Extended Frobenius norm), which is based on Principal Component Analysis (PCA). Eros applies PCA to MTS datasets represented as matrices to generate principal components and associated eigenvalues. These principal components and eigenvalues are then used to compare the similarity between MTS matrices. Though Eros in itself does not satisfy the triangle inequality, without which existing multidimensional indexing structures may not be utilized, the lower and upper bounds to satisfy the triangle inequality are obtained. In order to show the validity of Eros for similarity search on MTS datasets, we performed several experiments on three datasets (2 real-world and 1 synthetic). The results show the superiority of our approaches as compared to the traditional similarity measures for MTS datasets, such as Euclidean Distance (ED), Dynamic Time Warping (DTW), Weighted Sum SVD (WSSVD) and PCA similarity factor (SPCA) in precision/recall.
Cyrus Shahabi合作论文数Department of Computer Science, Viterbi School of Engineering, University of Southern California15