Out-of-distribution states in robot manipulation often lead to unpredictable robot behavior or task failure, limiting success rates and increasing risk of damage. Anomaly detection (AD) can identify deviations from expected patterns in data, which can be used to trigger failsafe behaviors and recovery strategies. Prior work has applied data-driven AD on time series data for specific robotic tasks, however the transferability of an AD approach between different robot control strategies and task types has not been shown. Leveraging time series data, such as force/torque signals, allows to directly capture robot-environment interactions, crucial for manipulation and online failure detection. As robotic tasks can have widely signal characteristics and requirements, AD methods which can be applied in the same way to a wide range of tasks is needed, ideally with good data efficiency. We examine three industrial robotic tasks, robotic cabling, screwing, and sanding, each with multi-modal time series data and several anomalies. Several autoencoderbased methods are compared, and we evaluate the generalization across different robotic tasks and control methods (diffusion policy-, position-, and impedance-controlled). This allows us to validate the integration of AD in complex tasks involving tighter tolerances and variation from both the robot and its environment. Additionally, we evaluate data efficiency, detection latency, and task characteristics which support robust detection. The results indicate reliable detection with AUROC exceeding 0.96 in failures in the cabling and screwing task, such as incorrect or misaligned parts and obstructed targets. In the polishing task, only severe failures were reliably detected, while more subtle failure types remained undetected.
AI-driven computer vision applications require a profound database to ensure predictable behaviors and performance. Such predictable behaviors are especially important for industrial applications in gaining trust from users. However, such a database is not readily available in industrial applications, and its acquisition is not trivial either. Active learning methods can be applied to ramp up data within a project deployment to iteratively increase the database, and thus the application predictability. Unfortunately, we observe that this often leads to a loss of user trust in the application, which is difficult to regain once lost. This leads to a “chicken-and-egg” dilemma in which neither the database nor the application is developed. In this work, we review state-of-the-art methods and approaches to further boost the database the initial active data ramp-up phase. Here, we focus on recent advancements in GenAI-based data generation and augmentation methods and review their adaptability on an industrial computer vision classification use case. Although we observe a potential for automatic data ramp-up, we also see a domain miss match in between the source (training environment) and target (industrial use-case) – regarding context defined in natural language and object characteristics.
Turbine blades of industrial gas turbines exhibit sub-coating cracks after a certain period of operation. These cracks stem from thermal and mechanical strains. Turbine blades are complex parts with high material and manufacturing costs, making a repair attractive from economical and sustainability perspectives. After removal of the coating, to reach the crack depths, the base material needs to be removed. The conventional repair process is dominated by manual work, including defect detection, geometry parametrization and CAM path planning. The scope of this work is to replace time and labor-intensive work steps by an adaptive and automated subtractive repair process. The proposed digital process chain includes 3D scanning of the turbine blade, highly accurate 3D geometry reconstruction and subsequent automated path planning for the milling operation. Post-processing and milling are executed, and the actual results are compared to the reconstructed and planned geometry, showing that satisfactory accuracy can be achieved. The overall time is drastically reduced compared to the conventional process. Finally, the potential for optimization and industrialization is discussed.
Augmenting large language models (LLMs) with external tools has proven to be effective for producing consistent, deterministic results. While capabilities of agentic AI systems are growing rapidly and the integration into various domain application speeds up, their resource consumption finds little attention in literature. With continuously more capable models requiring more energy, coupled with ever increasing numbers of users, there is a need for alternative solutions. We present an approach of using small-scale open-source models locally powering an agentic AI system in the context of production planning. By defining and iteratively executing experiments, the system can optimize discrete-event simulation models – while running entirely on-device. With significant drawbacks in speed and performance, the local configuration consumes only a fraction of the estimated resources of state-of-the-art models. By offering further benefits in data sovereignty, operational independence and cost control, the concept of locally deployed models presents an attractive alternative for industry-grade manufacturing applications.
Visuomotor imitation policies can achieve high performance under in-distribution visual conditions yet fail when visually similar objects or receptacles are introduced. We study this behavior as a problem of conditional visual grounding: the visual target required for successful control changes with the manipulation phase and, in more complex tasks, with the observed task state. Using Action Chunking with Transformers (ACT), we systematically introduce distractor objects and receptacles with controlled color and shape similarity and localize failures to picking and placement. We find that distractor sensitivity is specific to both the type of visual similarity and the manipulation stage. Guided by this diagnosis, we evaluate distractor augmentation, phase-dependent attention regularization, and appearance-based visual prompting as complementary interventions for improving target selection while preserving spatial information required for control. These interventions substantially improve robustness in simulation and on a physical UR3e. We further examine the same failure pattern in a pretrained vision-language-action policy on a state-conditioned instrument-handling task, where the observed state of a medical instrument determines the correct destination. Together, the results show that visual distractors can cause incorrect object or destination selection even when the underlying manipulation skill remains intact, and that explicitly improving target selection can substantially recover performance across distinct visuomotor policy-learning regimes.