Data-based process monitoring in injection molding plays an important role in compensating disturbances in the process and the associated impairment of part quality. Selecting appropriate features for a successful online quality prediction based on machine learning methods is crucial. Time series such as the injection pressure and injection flow curve are particularly suitable for this purpose. Predicting quality as early as possible during a cycle has many advantages. In this paper it is shown how the recording length of the time series affects the prediction performance when using machine learning algorithms. For this purpose, two successful molding quality prediction algorithms (k Nearest Neighbors and Ridge Regression) are trained with time series of different lengths on extensive data sets. Their prediction performances for part weight and a geometric dimension are evaluated. The evaluations show that recording time series until the end of a cycle is not necessary to obtain good prediction results. These findings indicate that early reliable quality prediction is possible within a cycle, which speeds up prediction, allows timely part handling at the end of the cycle and provides the basis for automated corrective interventions within the same cycle.
Process-data-supported process monitoring in injection molding plays an important role in compensating for disturbances in the process. Until now, scalar process data from machine controls have been used to predict part quality. In this paper, we investigated the feasibility of incorporating time series of sensor measurements directly as features for machine learning models, as a suitable method of improving the online prediction of part quality. We present a comparison of several state-of-the-art algorithms, using extensive and realistic data sets. Our comparison demonstrates that time series data allow significantly better predictions of part quality than scalar data alone. In future studies, and in production-use cases, such time series should be taken into account in online quality prediction for injection molding.
Finding structural similarities in graph data, like social networks, is a far-ranging task in data mining and knowledge discovery. A (conceptually) simple reduction would be to compute the automorphism group of a graph. However, this approach is ineffective in data mining since real world data does not exhibit enough structural regularity. Here we step in with a novel approach based on mappings that preserve the maximal cliques. For this we exploit the well known correspondence between bipartite graphs and the data structure formal context (G, M, I) from Formal Concept Analysis. From there we utilize the notion of clone items. The investigation of these is still an open problem to which we add new insights with this work. Furthermore, we produce a substantial experimental investigation of real world data. We conclude with demonstrating the generalization of clone items to permutations.
It is well known that any bipartite (social) network can be regarded as a formal context $(G,M,I)$. Therefore, such networks give raise to formal concept lattices which can be investigated utilizing the toolset of Formal Concept Analysis (FCA). In particular, the notion of clones in closure systems on $M$, i.e., pairwise interchangeable attributes that leave the closure system unchanged, suggests itself naturally as a candidate to be analyzed in the realm of FCA based social network analysis. In this study, we investigate the notion of clones in social networks. After building up some theoretical background for the clone relation in formal contexts we try to find clones in real word data sets. To this end, we provide an experimental evaluation on nine mostly well known social networks and provide some first insights on the impact of clones. We conclude our work by nourishing the understanding of clones by generalizing those to permutations of higher order.
Researchers must face the exponential growth of the body of available scholarly literature, which makes it ever harder to keep track with one’s own community, especially for newcomers. In this thesis, we explore different means of supporting researchers with that task. For this purpose, we follow two approaches: We provide analyses of research communities and of researchers’ interactions through data that can be obtained from the phases in the life cycle of scholarly publications (creation, dissemination, usage, and citation in other publications). The resulting statistics and visualizations allow researchers to better understand their own communities, to identify the most important players and publications, and to find valuable conversational partners at conferences. For the analysis of publication usage and connections to citations, we turn to social bookmarking systems and investigate the actions of users in BibSonomy. The provided insights can help operators of such systems improve them. Our second approach is more proactive, focusing on supporting researchers by pointing them directly to important publications – through automatically computed personalized recommendations and through social peer review. The analysis of research and researchers often relied on studying scholarly publications and their metadata. Such studies can reveal insights into how scientific work is conducted, they can shed light on communities and research topics, and they allow the measurement of certain forms of impact, a publication, an individual researcher, or a venue had. The exploited data – publication metadata– is generated when publications are created. The life cycle of a scholarly publication, however, just begins with a publication’s creation: Publications are disseminated (e.g., presented at conferences), they are used (e.g., acquired, stored, collected, marked as to-read, and, of course, read), and they are cited. With the advent of the Web 2.0, traces of the activities in these phases have become observable. In this thesis, we collect and analyze datasets from all four stages of the publication life cycle. We thus go beyond traditional means of scientometrics, touching such fields as altmetrics, web log analysis, and role discovery. We not only present new insights into communities that have not been investigated before, but we also demonstrate new means of analysis that are generalizable to other communities as well. Among them are formal concept analysis to visualize influences between groups of authors and social network analyses of interaction networks. Our datasets comprise – next to a traditional publication corpus containing metadata and references – a face-to-face contact network, gathered from real-live interactions of researchers during a conference, and datasets from the scholarly social bookmarking system BibSonomy. Social bookmarking services allow their users to publicly store and annotate resources, like web links, photos, videos, or publications. As representatives of the Web 2.0, social
For the popular task of tag recommendation, various (complex) approaches have been proposed. Recently however, research has focused on heuristics with low computational e ort and particularly, a time-aware heuristic, called BLL, has been shown to compare well to various state-of-the-art methods. Here, we follow up on these results by presenting another time-aware approach leveraging userinteraction data in an easily interpretable, on-they computable approach that can successfully be combined with BLL. We investigate the in uence of time as a parameter in that approach, and we demonstrate the e ectiveness of the proposed method using two datasets from the popular public social tagging system BibSonomy.
Social tagging systems have established themselves as an important part in today’s Web and have attracted the interest of our research community in a variety of investigations. Henceforth, several aspects of social tagging systems have been discussed and assumptions have emerged on which our community builds their work. Yet, testing such assumptions has been difficult due to the absence of suitable usage data in the past. In this work, we thoroughly investigate and evaluate four aspects about tagging systems, covering social interaction, retrieval of posted resources, the importance of the three different types of entities, users, resources, and tags, as well as connections between these entities’ popularity in posted and in requested content. For that purpose, we examine live server log data gathered from the real-world, public social tagging system BibSonomy. Our empirical results paint a mixed picture about the four aspects. Although typical assumptions hold to a certain extent for some, other aspects need to be reflected in a very critical light. Our observations have implications for the understanding of social tagging systems and the way they are used on the Web. We make the dataset used in this work available to other researchers.
Social tagging systems have established themselves as an important part in today’s Web and have attracted the interest of our research community in a variety of investigations. Henceforth, several aspects of social tagging systems have been discussed and assumptions have emerged on which our community builds their work. Yet, testing such assumptions has been difficult due to the absence of suitable usage data in the past. In this work, we thoroughly investigate and evaluate four aspects about tagging systems, covering social interaction, retrieval of posted resources, the importance of the three different types of entities, users, resources, and tags, as well as connections between these entities’ popularity in posted and in requested content. For that purpose, we examine live server log data gathered from the real-world, public social tagging system BibSonomy. Our empirical results paint a mixed picture about the four aspects. Although typical assumptions hold to a certain extent for some, other aspects need to be reflected in a very critical light. Our observations have implications for the understanding of social tagging systems and the way they are used on the Web. We make the dataset used in this work available to other researchers.
Communities can intuitively be defined as subsets of nodes of a graph with a dense structure in the corresponding subgraph. However, for mining such communities usually only structural aspects are taken into account. Typically, no concise nor easily interpretable community description is provided.For tackling this issue, this paper focuses on description-oriented community detection using subgroup discovery. In order to provide both structurally valid and interpretable communities we utilize the graph structure as well as additional descriptive features of the graph's nodes. A descriptive community pattern built upon these features then describes and identifies a community, i.e., a set of nodes, and vice versa. Essentially, we mine patterns in the "description space" characterizing interesting sets of nodes (i.e., subgroups) in the "graph space"; the interestingness of a community is evaluated by a selectable quality measure.We aim at identifying communities according to standard community quality measures, while providing characteristic descriptions of these communities at the same time. For this task, we propose several optimistic estimates of standard community quality functions to be used for efficient pruning of the search Space in an exhaustive branch-and-bound algorithm. We demonstrate our approach in an evaluation using five real-world data sets, obtained from three different social media applications. (C) 2015 Elsevier Inc. All rights reserved.
In social tagging systems, like Mendeley, CiteULike, and BibSonomy, users can post, tag, visit, or export scholarly publications. In this paper, we compare citations with metrics derived from users' activities (altmetrics) in the popular social bookmarking system BibSonomy. Our analysis, using a corpus of more than 250,000 publications published before 2010, reveals that overall, citations and altmetrics in BibSonomy are mildly correlated. Furthermore, grouping publications by user-generated tags results in topic-homogeneous subsets that exhibit higher correlations with citations than the full corpus. We find that posts, exports, and visits of publications are correlated with citations and even bear predictive power over future impact. Machine learning classifiers predict whether the number of citations that a publication receives in a year exceeds the median number of citations in that year, based on the usage counts of the preceding year. In that setup, a Random Forest predictor outperforms the baseline on average by seven percentage points. (C) 2016 Elsevier Ltd. All rights reserved.
Social bookmarking systems have established themselves as an important part in today’s Web. In such systems, tag recommender systems support users during the posting of a resource by suggesting suitable tags. Tag recommender algorithms have often been evaluated in offline benchmarking experiments. Yet, the particular setup of such experiments has rarely been analyzed. In particular, since the recommendation quality usually suffers from difficulties such as the sparsity of the data or the cold-start problem for new resources or users, datasets have often been pruned to so-called cores (specific subsets of the original datasets), without much consideration of the implications on the benchmarking results. In this article, we generalize the notion of a core by introducing the new notion of a set-core , which is independent of any graph structure, to overcome a structural drawback in the previous constructions of cores on tagging data. We show that problems caused by some types of cores can be eliminated using set-cores. Further, we present a thorough analysis of tag recommender benchmarking setups using cores. To that end, we conduct a large-scale experiment on four real-world datasets, in which we analyze the influence of different cores on the evaluation of recommendation algorithms. We can show that the results of the comparison of different recommendation approaches depends on the selection of core type and level. For the benchmarking of tag recommender algorithms, our results suggest that the evaluation must be set up more carefully and should not be based on one arbitrarily chosen core type and level.
Social tagging systems have established themselves as a quick and easy way to organize information by annotating resources with tags. In recent work, user behavior in social tagging systems was studied, that is, how users assign tags, and consume content. However, it is still unclear how users make use of the navigation options they are given. Understanding their behavior and differences in behavior of different user groups is an important step towards assessing the effectiveness of a navigational concept and improving it to better suit the users' needs. In this work, we investigate navigation trails in the popular scholarly social tagging system BibSonomy from six years of log data. We discuss dynamic browsing behavior of the general user population and show that different navigational subgroups exhibit different navigational traits. Furthermore, we provide strong evidence that the semantic nature of the underlying folksonomy is an essential factor for explaining navigation.
Communities can intuitively be defined as subsets of nodes of a graph with a dense structure. However, for mining such communities usually only structural aspects are taken into account. Typically, no concise and easily interpretable community description is provided. For tackling this issue, we focus on fast description-oriented community detection using subgroup discovery, cf. [1, 2]. In order to provide both structurally valid and interpretable communities we utilize the graph structure as well as additional descriptive features of the contained nodes. A descriptive community pattern built upon these features then describes and identifies a community given by a set of nodes, and vice versa. Essentially, we mine for patterns in the “description space” characterizing interesting sets of nodes in the “graph/community space”; the interestingness of a community is then evaluated by a selectable quality measure. We aim at identifying communities according to standard community quality measures, while providing characteristic descriptions of the respective communities at the same time. In order to implement an efficient approach, we propose several optimistic estimates of standard community quality functions. Together with the proposed exhaustive branch-and-bound algorithm, these estimates enable fast description-oriented community detection. This is demonstrated in an evaluation using five real-world data sets, obtained from three different social media applications.
Semantic relations that closely resemble the human intuition of semantic relatedness, have been extracted automatically from sources, like text corpora, Wikipedia, or folksonomies, e.g., for constructing ontologies or for enhancing website navigation. Thereby, folksonomies are especially interesting since often, rich semantic structures emerge from the annotation of resources through users. In previous work however, proxies, like WordNet-based relatedness measures, have been used for evaluation, rather than relying directly on captured human intuition. Here, we critically examine this form of evaluation and compare it to evaluations relying directly on human intuition. We find that WordNetbased measures hardly correlate with state-of-the-art human intuition datasets. Moreover, for meaningful results, the evaluation datasets must be domain-specific to provide a reasonably high overlap of words with the source for the extracted semantics. We demonstrate our results on two real world folksonomy datasets, using well-known evaluation datasets, as well as new, crowdsourced word similarity estimations. Overall, we argue that although directly evaluating semantic relatedness measures on human intuition may require collecting an adapted set of annotated samples, this form of benchmarking has clear advantages over the currently used WordNet-based measures, presenting a more realistic evaluation.
Scholarly success is traditionally measured in terms of citations to publications. With the advent of publication management and digital libraries on the web, scholarly usage data has become a target of investigation and new impact metrics computed on such usage data have been proposed -- so called altmetrics. In scholarly social bookmarking systems, scientists collect and manage publication meta data and thus reveal their interest in these publications. In this work, we investigate connections between usage metrics and citations, and find posts, exports, and page views of publications to be correlated to citations.
The combination of ubiquitous and social computing is an emerging research area which integrates different but complementary methods, techniques, and tools. In this paper, we focus on the Ubicon platform, its applications, and a large spectrum of analysis results. Ubicon provides an extensible framework for building and hosting applications targeting both ubiquitous and social environments. We summarize the architecture and exemplify its implementation using four real-world applications built on top of Ubicon. In addition, we discuss several scientific experiments in the context of these applications in order to give a better picture of the potential of the framework, and discuss analysis results using several real-world data sets collected utilizing Ubicon.
Social tagging systems have established themselves as an important part in today's web and have attracted the interest from our research community in a variety of investigations. The overall vision of our community is that simply through interactions with the system, i.e., through tagging and sharing of resources, users would contribute to building useful semantic structures as well as resource indexes using uncontrolled vocabulary not only due to the easy-to-use mechanics. Henceforth, a variety of assumptions about social tagging systems have emerged, yet testing them has been difficult due to the absence of suitable data. In this work we thoroughly investigate three available assumptions - e.g., is a tagging system really social? - by examining live log data gathered from the real-world public social tagging system BibSonomy. Our empirical results indicate that while some of these assumptions hold to a certain extent, other assumptions need to be reflected and viewed in a very critical light. Our observations have implications for the design of future search and other algorithms to better reflect the actual user behavior.
Social tagging systems have established themselves as an important part in today’s web and have attracted the interest of our research community in a variety of investigations. Henceforth, several assumptions about social tagging systems have emerged on which our community also builds their work. Yet, testing such assumptions has been difficult due to the absence of suitable usage data in the past. In this work, we investigate and evaluate four assumptions about tagging systems by examining live server log data gathered from the public social tagging system BibSonomy. Our empirical results indicate that while some of these assumptions hold to a certain extent, other assumptions need to be reflected in a very critical light.
Robert Jäschke合作论文数Leibniz Universitat Hannover, Fakultat fur Elektrotechnik und Informatik, Fachgebiet Wissensbasierte Systeme13
Michelangelo Ceci合作论文数University of Bari, Italy1