In this study, different machine learning algorithms were analysed to predict casting defects in a cold chamber magnesium high-pressure die casting process. Based on component-related process and quality data from 7982 casting cycles, models were trained using Support Vector Machine, Random Forest and AutoML algorithms from Auto-Sklearn. The aim was to predict the presence of defect classes ("cold flow", "shrinkage cavity", "blister", "soldering point", "scrap"). Random Forest models achieved the best prediction quality overall, especially for the defect class "soldering points", which is the least frequent detected defect. An analysis of training data amounts showed that the prediction quality only improves slightly beyond 1000 training cycles, except for the "soldering point" defect class, which showed further improvements with more training data. Furthermore, the scope of data in terms of measurement sources affected the prediction quality significantly. Random Forest prediction models that were trained exclusively with casting machine data generally provide a solid basis for predicting casting defects. The highest increase in prediction performance was achieved by adding die sensor data. Overall, the prediction quality of all models was always above the statistically expected values (Balanced Accuracy of 50%), and soldering points in particular were predicted with a Balanced Accuracy of more than 80%. It was found that block temperature sensors in the shot sleeve and force measurements (for measuring the cavity pressure via the ejector pin) had comparatively high correlations with all defect classes and were weighted highly by the Random Forest models for decision-making.
For qualitative data analysis (QDA), researchers assign codes to text segments to arrange the information into topics or concepts. These annotations facilitate information retrieval and the identification of emerging patterns in unstructured data. However, this metadata is typically not published or reused after the research. Subsequent studies with similar research questions require a new definition of codes and do not benefit from other analysts’ experience. Machine learning (ML) based classification seeded with such data remains a challenging task due to the ambiguity of code definitions and the inherent subjectivity of the exercise. Previous attempts to support QDA using ML rely on linear models and only examined individual datasets that were either smaller or coded specifically for this purpose. However, we show that modern approaches effectively capture at least part of the codes’ semantics and may generalize to multiple studies. We analyze the performance of multiple classifiers across three large real-world datasets. Furthermore, we propose an ML-based approach to identify semantic relations of codes in different studies to show thematic faceting, enhance retrieval of related content, or bootstrap the coding process. These are encouraging results that suggest how analysts might benefit from prior interpretation efforts, potentially yielding new insights into qualitative data.
Abstract The Future Lab Production demonstrates the potentials of digitalisation by using the die casting process as an example process. The project shows how manufacturing companies can digitalise their existing machines, analyse their data and exchange information along the supply chain while maintaining data sovereignty. The aim is to support companies with digitalisation from the machine to data platforms. The article describes the methods used, the concepts developed and their benefits.
The war in Ukraine seems to have positively changed the attitude toward the critical societal topic of migration in Europe – at least towards refugees from Ukraine. We investigate whether this impression is substantiated by how the topic is reflected in online news and social media, thus linking the representation of the issue on the Web to its perception in society. For this purpose, we combine and adapt leading-edge automatic text processing for a novel multilingual stance detection approach. Starting from 5.5M Twitter posts published by 565 European news outlets in one year, beginning September 2021, plus replies, we perform a multilingual analysis of migration-related media coverage and associated social media interaction for Europe and selected European countries. The results of our analysis show that there is actually a reframing of the discussion illustrated by the terminology change, e.g., from "migrant" to "refugee", often even accentuated with phrases such as "real refugees". However, concerning a stance shift in public perception, the picture is more diverse than expected. All analyzed cases show a noticeable temporal stance shift around the start of the war in Ukraine. Still, there are apparent national differences in the size and stability of this shift.
For a broader adoption of AI in industrial production, adequate infrastructure capabilities and ecosystems are crucial. This includes easing the integration of AI with industrial devices, support for distributed deployment, monitoring, and consistent system configuration. IIoT platforms can play a major role here by providing a unified layer for the heterogeneous Industry 4.0/IIoT context. However, existing IIoT platforms still lack required capabilities to flexibly integrate reusable AI services and relevant standards such as Asset Administration Shells or OPC UA in an open, ecosystem-based manner. This is exactly what our next level Intelligent Industrial Production Ecosphere (IIP-Ecosphere) platform addresses, employing a highly configurable low-code based approach. In this paper, we introduce the design of this platform and discuss an early evaluation in terms of a demonstrator for AI-enabled visual quality inspection. This is complemented by insights and lessons learned during this early evaluation activity.
The development of intelligent solutions for manufacturing is a challenging task. Industry 4.0 platforms can provide a unifying layer here. However, flexible AI support, openness for evolving service and components from different vendors and adaptability to the diverse and changing requirements is required from such a platform to boost IIoT development. For this purpose, our approach combines - as a "power trio" - (1) wide use of Asset Administration Shells (AAS) for targeting device, component and service heterogeneity, with (2) configuration support for dealing with the diverse and changing requirements and (3) code generation for cost-effective creation of customer specific platform instances, AAS and AI-based Industry 4.0 applications on top of the IIP-Ecosphere platform. The platform has been implemented based on vertically scaled AAS and evaluated with two Industry 4.0 demonstrators. In this context, we discuss the experiences we made with our approach.
Abstract Für die erfolgreiche Digitalisierung in der Produktion ist die IT-Infrastruktur, zum Beispiel zur einfachen Anbindung von Geräten und Steuerung von Datenflüssen, von zentraler Bedeutung. Bisherige Lösungen basieren jedoch meist auf proprietären Protokollen und bieten wenig Konfigurations- und Kontrollmöglichkeiten. Basierend auf industriellen Standards, wie z. B. Verwaltungsschalen, wird im Projekt IIP-Ecosphere daher eine offene Code-Basis für die ganzheitliche Umsetzung von Digitalisierungsprojekten in der Produktion entwickelt.
Maximizing the benefits of AI for Industry 4.0 is about more than just developing effective new AI methods. Of equal importance is the successful integration of AI into production environments. One open challenge is the dynamic deployment of AI on industrial edge devices within close proximity to manufacturing machines. Our IIP-Ecosphere 1 platform was designed to overcome limitations of existing Industry 4.0 platforms. It supports flexible AI deployment through employing a highly configurable low-code based approach, where code for tailored platform components and applications is generated. In this paper, we measure the performance of our platform on an industrial demonstrator and discuss the impact of deploying AI from a central server to the edge. As result, AI inference automatically deployed on an industrial edge is possible, but in our case three times slower than on a desktop computer, requiring still more optimizations.
Understanding the factors related to migration, such as perceptions about routes and target countries, is critical for border agencies and society altogether. A systematic analysis of communication and news channels, such as social media, can improve our understanding of such factors. Videos and images play a critical role in social media as they have significant impact on perception manipulation and misinformation campaigns. However, more research is needed in the identification of semantically relevant visual content for specific queried concepts. Furthermore, an important problem to overcome in this area is the lack of annotated datasets that could be used to create and test accurate models. A recent study proposed a novel video representation and retrieval approach that effectively bridges the gap between a substantiated domain understanding - encapsulated into textual descriptions of Migration Related Semantic Concepts (MRSCs) - and the expression of such concepts in a video. In this work, we build on this approach and propose an improved procedure for the crucial step of the concept labels' textual augmentation, which contributes towards the full automation of the pipeline. We assemble the first, to the best of our knowledge, migration-related videos and images dataset and we experimentally assess our method on it.
Our perception of the situation in a country or a region is strongly influenced by the reflection of this situation in mass and social media channels. This effect is even more pronounced for geographically and culturally distant regions, for which no firsthand experience is available. To avoid information overload, news outlets typically filter the available news from foreign countries based on the expected interest of the target audiences. Such filtering imposes an inherent bias in the reporting and can create a distorted perception of a region among the consumers of news of other regions. This might lead to misunderstandings between countries and unsubstantiated political and individual decisions (e.g., in the context of migration). In this article, we systematically analyze the bias created in news reports. We consider Europe, or more precisely the European Union (EU) as our zone of concern, and examine its image in the media (news outlets) of other regions, Europe(NON-EU), Africa, Asia, Middle-East, America, and Oceania. An analysis of the year 2018 (January–December 2018) of news published in those regions reveals marked differences in the editorial policies and presented narrative when dealing with EU-related news. We observe a significant variation in the sentiment polarity of the reported EU-related stories between the European and other regional news outlets. We further analyze the polarity variation among different subregions of large geographical areas, such as Africa, Asia, and America. We observe a contrasting difference in their editorial policies. This trend also holds for news related to different topics, such as politics, business, economy, health, and international relation.
For their attractiveness, comprehensiveness and dynamic coverage of relevant topics, community-based question answering sites such as Stack Overflow heavily rely on the engagement of their communities: Questions on new technologies, technology features as well as technology versions come up and have to be answered as technology evolves (and as community members gather experience with it). At the same time, other questions cease in importance over time, finally becoming irrelevant to users. Beyond filtering low-quality questions, "forgetting" questions, which have become redundant, is an important step for keeping the Stack Overflow content concise and useful. In this work, we study this managed forgetting task for Stack Overflow. Our work is based on data from more than a decade (2008 - 2019) - covering 18.1M questions, that are made publicly available by the site itself. For establishing a deeper understanding, we first analyze and characterize the set of questions about to be forgotten, i.e., questions that get a considerable number of views in the current period but become unattractive in the near future. Subsequently, we examine the capability of a wide range of features in predicting such forgotten questions in different categories. We find some categories in which those questions are more predictable. We also discover that the text-based features are surprisingly not helpful in this prediction task, while the meta information is much more predictive.
Migration, and especially irregular migration, is a critical issue for border agencies and society in general. Migration-related situations and decisions are influenced by various factors, including the perceptions about migration routes and target countries. An improved understanding of such factors can be achieved by systematic automated analyses of media and social media channels, and the videos and images published in them. However, the multifaceted nature of migration and the variety of ways migration-related aspects are expressed in images and videos make the finding and automated analysis of migration-related multimedia content a challenging task. We propose a novel approach that effectively bridges the gap between a substantiated domain understanding - encapsulated into a set of Migration-related semantic concepts - and the expression of such concepts in a video, by introducing an advanced video analysis and retrieval method for this purpose.
Trends like digital transformation even intensify the already overwhelming mass of information knowledge workers face in their daily life. To counter this, we have been investigating knowledge work and information management support measures inspired by human forgetting. In this paper, we give an overview of solutions we have found during the last 5 years as well as challenges that still need to be tackled. Additionally, we share experiences gained with the prototype of a first forgetful information system used 24/7 in our daily work for the last 3 years. We also address the untapped potential of more explicated user context as well as features inspired by memory inhibition, which is our current focus of research.
Inhibition is one of the core concepts in Cognitive Psychology. The idea of inhibitory mechanisms actively weakening representations in the human mind has inspired a great number of studies in various research domains. In contrast, Computer Science only recently has begun to consider inhibition as a second basic processing quality beside activation. Here, we review psychological research on inhibition in memory and link the gained insights with the current efforts in Computer Science of incorporating inhibitory principles for optimizing information retrieval in Personal Information Management. Four common aspects guide this review in both domains: 1. The purpose of inhibition to increase processing efficiency. 2. Its relation to activation. 3. Its links to contexts. 4. Its temporariness. In summary, the concept of inhibition has been used by Computer Science for enhancing software in various ways already. Yet, we also identify areas for promising future developments of inhibitory mechanisms, particularly context inhibition.
When a major event such as a crisis situation occurs, people post messages on social media sites such as Twitter, in order to exchange information or to share emotions. These posts can provide useful information to raise situation awareness and support decision making, e.g., by aid organizations. In this paper, we propose a novel method for social media crawling, which exploits a Bayesian inference framework to keep track of keyword changes over time and uses a counter-stream to gauge the inclusion of noise and irrelevant information. In addition, we present a framework to evaluate real-time adaptive social search algorithms in a reproducible manner, which relies on a semi-automated approach for ground-truth construction. We show that our method outperforms previous methods for very large scale events.
The idea of the Preserve-or-Forget (PoF) approach introduced in this book is to follow a forgetful, focused approach to digital preservation, which is inspired by human forgetting and remembering. Its goal is to ease the adoption of preservation technology especially in the personal and organizational context and to ensure that important content is kept safe, useful, and understandable in the long run. For this purpose, it stresses the smooth interaction between information management and preservation management. Leveraging the PoF approach, in this chapter we introduce a reference model, which will be referred to in the following as PoF Reference Model. The model pays special attention to the functionality which bridges between Information Management System (Active System) and Digital Preservation System (DPS), such as the selection of content for preservation and the transfer of content between the systems. The model aims to encapsulate the core ideas of the PoF approach, which considers Active System and DPS as a joint ecosystem into a re-usable model, and is inspired by the core principles of this approach: synergetic preservation, managed forgetting, and contextualized remembering. The design of the PoF Reference Model was driven by the identification of five required characteristics for such a reference model: it has to be integrative, value-driven, brain-inspired, forgetful, and evolution-aware. The PoF Reference Model consists of a functional part (Functional Model) and of an associated Information Model. The Functional Model is made up of three layers: Core Layer, Remember and Forget Layer, and Evolution Layer. For each layer, we discuss the main functional entities and the representative workflows, also relating them to existing standards and practices in digital preservation. The functionality required to mediate between the Active System and the DPS has been encapsulated into the PoF Middleware, which has been designed and implemented as part of the ForgetIT project. The Information Model describes the preservation entities and their relationships, also discussing the interoperability with existing digital preservation standards.
This unique monograph advocates a novel forgetful approach to the long-term handling of personal multimedia content, inspired by human psychology.
Current trends, like digital transformation and ubiquitous computing, yield in massive increase in available data and information. In artificial intelligence (AI) systems, capacity of knowledge bases is limited due to computational complexity of many inference algorithms. Consequently, continuously sampling information and unfiltered storing in knowledge bases does not seem to be a promising or even feasible strategy. In human evolution, learning and forgetting have evolved as advantageous strategies for coping with available information by adding new knowledge to and removing irrelevant information from the human memory. Learning has been adopted in AI systems in various algorithms and applications. Forgetting, however, especially intentional forgetting, has not been sufficiently considered, yet. Thus, the objective of this paper is to discuss intentional forgetting in the context of AI systems as a first step. Starting with the new priority research program on 'Intentional Forgetting' (DFG-SPP 1921), definitions and interpretations of intentional forgetting in AI systems from different perspectives (knowledge representation, cognition, ontologies, reasoning, machine learning, self-organization, and distributed AI) are presented and opportunities as well as challenges are derived.
Themis Palpanas合作论文数Department of Computer Science, Universite Paris Cite;French University Institute7
Heiko Maus合作论文数Knowledge Management Group;German Research Center for Artificial Intelligence (DFKI) GmbH5