Deformable object manipulation remains a major challenge in robotics due to unstructured shapes, complex surface textures, and highly variable mechanical properties. Fish entrails represent an extreme case of such objects, exhibiting substantial biological variability, high fragility, and slippery, poorly defined contact conditions. This paper explores the development and testing of robotic handling of cod innards, specifically focusing on liver, roe, and milt. The study enhances the knowledge base on handling innards and further explores the possibility of using rudimentary and readily available robotic grippers to validate the feasibility of automation for these delicate tasks. The experiments demonstrated high success rates for roe and substantial improvement under teleoperation for liver, while milt handling revealed clear mechanical limitations. Liver handling presented variable success, while milt handling highlighted the need for specialized grippers due to its slippery nature. Roe handling proved to be robotically successful, and represents an immediately viable automation candidate (572 runs were recorded.) The findings suggest that further development of customized grippers and control algorithms could enable a wider application of robotic systems in the seafood industry, leading to increased productivity, a more consistent product quality, and improved health, safety, and environment (HSE). This paper provides practical insights and foundational data for future advancements in robotic automation for seafood processing/handling of innards. (Video compilation from experiments at: bit.ly/GutOutMov, raw data results: bit.ly/GutOutRawData.)
Advanced robotic manipulation of deformable, volumetric objects remains one of the greatest challenges due to their pliancy, frailness, variability, and uncertainties during interaction. Motivated by these challenges, this article introduces Sashimi-Bot, an autonomous multi-robotic system for advanced manipulation and cutting, specifically the preparation of sashimi. The objects that we manipulate, salmon loins, are natural in origin and vary in size and shape, they are limp and deformable with poorly characterized elastoplastic parameters, while also being slippery and hard to hold. The three robots straighten the loin; grasp and hold the knife; cut with the knife in a slicing motion while cooperatively stabilizing the loin during cutting; and pick up the thin slices from the cutting board or knife blade. Our system combines deep reinforcement learning with in-hand tool shape manipulation, in-hand tool cutting, and feedback of visual and tactile information to achieve robustness to the variabilities inherent in this task. This work represents a milestone in robotic manipulation of deformable, volumetric objects that may inspire and enable a wide range of other real-world applications.
In this letter, we present a novel dual-task, closed-loop, visual servoing-based active vision framework in an eye-in-hand configuration. The proposed active vision framework continuously drives the camera motion by coupling continuous Next-Best-View (NBV) planning and visual servo control within a unified formulation, is NBV-objective-agnostic, and enables real-time, closed-loop exploration of objects. We demonstrate how this approach can be applied to the 3D reconstruction of static volumetric objects. The approach is validated in the real world with a diverse set of relevant objects and we observe that the visual servo scheme produces smooth exploration trajectories that keep the camera focused at the object. We also show that our gradient-based continuous NBV-strategy is highly competitive with baseline strategies that leverage global viewpoint sampling and results in efficient exploration with strong object coverage.
Deep reinforcement learning (RL) often relies on simulators as abstract oracles to model interactions within complex environments. While differentiable simulators have recently emerged for multi-body robotic systems, they remain underutilized, despite their potential to provide richer information. This underutilization, coupled with the high computational cost of exploration-exploitation in high-dimensional state spaces, limits the practical application of RL in the real-world. We propose a method that integrates learning with differentiable simulators to enhance the efficiency of exploration-exploitation. Our approach learns value functions, state trajectories, and control policies from locally optimal runs of a model-based trajectory optimizer. The learned value function acts as a proxy to shorten the preview horizon, while approximated state and control policies guide the trajectory optimization. We benchmark our algorithm on three classical control problems and a torque-controlled 7 degree-of-freedom robot manipulator arm, demonstrating faster convergence and a more efficient symbiotic relationship between learning and simulation for end-to-end training of complex, poly-articulated systems.
We present a novel framework for non-prehensile shape manipulation of deformable objects using Deep Reinforcement Learning. Unlike previous approaches that rely on grasping, our method employs a sequence of gentle pushing actions to deform objects into target shapes. We introduce a continuous parametrization of pushing actions that allows for precise control over pushing trajectories, enabling more flexible and efficient manipulation. The framework is applicable to a wide range of objects by representing them as sampled boundary coordinates, removing the need for predefined object partitions. Trained entirely in simulation, our controller demonstrates zero-shot transfer to real-world scenarios without additional training. Extensive evaluations show that our approach not only matches but substantially exceeds the performance of previous methods, while being more gentle and efficient. We demonstrate successful manipulation across various deformable objects and materials, including food items like salmon and pork loin. This work represents a significant advancement in robotic manipulation of deformable objects, with potential applications in food processing, manufacturing, and beyond.
The use of robotic vision to cut deformable food objects is a challenge in robotics that has the potential to improve autonomy and create new opportunities in industries such as medicine, the food industry, and services. While cutting rigid objects is relatively simple, cutting deformable objects like food items, which change shape during the cutting process, is a significant challenge that requires advances in various aspects of robotics, including vision, modeling, hardware design, and control. This paper discusses recent developments in vision-based approaches for robots cutting food items and highlights the main challenges that must be overcome to succeed in this task and outline some potential future research directions.
We investigate the active manipulation of objects using model-free and long-horizon DRL (Deep Reinforcement Learning) to achieve target shapes. Our proposed approach uses visual observations consisting of segmented images, to mitigate the sim-to-real gap. We address a long-horizon manipulation task requiring a sequence of accurate actions to achieve the target shapes using a robot arm with an RGB-D camera in eye-in-hand configuration, and an elongated, volumetric, elastoplastic object. We find similar objects in food, marine, and manufacturing domains. The aim is to actively manipulate the object into an arbitrary target shape using image observations. We trained a DRL agent using PPO (Proximal Policy Optimization) by running 768 parallel actors in simulation, for a total of 1,2M environment interactions, and tested this on 200 unseen target deformations. In three attempts, 82% of the trials achieved a greater than 90% overlap with the 200 target shapes. By relying on segmentation images as a visual observation space, we successfully transferred the agent to the real world without supplementary training. Our approach does not need any real-world manipulation examples nor fine-tuning in the real world. The robustness of our approach was demonstrated in simulation, and experimentally validated in the real world for specific manipulation tasks, achieving a 94.2% mean zero-shot overlap success rate on previously unseen target shapes.
We present a novel vision-based, 6-DoF grasping framework based on Deep Reinforcement Learning (DRL) that is capable of directly synthesizing continuous 6-DoF actions in cartesian space. Our proposed approach uses visual observations from an eye-in-hand RGB-D camera, and we mitigate the sim-to-real gap with a combination of domain randomization, image augmentation, and segmentation tools. Our method consists of an off-policy, maximum-entropy, Actor-Critic algorithm that learns a policy from a binary reward and a few simulated example grasps. It does not need any real-world grasping examples, is trained completely in simulation, and is deployed directly to the real world without any fine-tuning. The efficacy of our approach is demonstrated in simulation and experimentally validated in the real world on 6-DoF grasping tasks, achieving state-of-the-art results of an 86% mean zero-shot success rate on previously unseen objects, an 85% mean zero-shot success rate on a class of previously unseen adversarial objects, and a 74.3% mean zero-shot success rate on a class of previously unseen, challenging "6-DoF" objects.Raw footage of real-world validation can be found at https://youtu.be/bwPf8Imvook
Application of robotics on production lines often involves handling flexible objects (such as items of natural origin or plastic bags containing liquid/bulk substances), which makes it crucial to consider the shape of an item before and after it has been affected by robotic manipulation. Most of the time deformable items are challenging for the robot in such operations as grasping, cutting, or packaging. The objective of this paper is to track object deformations and perform a task based on this information. The paper addresses issues in tracking object deformation and proposes a solution for deformation tracking to form preliminary knowledge and scene awareness on the robot side. A curve-fitting-based method was implemented to define a region of interest using images from a RealSense D415 camera. The developed approach identifies the maximum number of aligned points and uses it to determine where the deformation occurred. The results of this research show that the deformations are efficiently tracked. Utilising the algorithm proposed in this paper, an efficient method capable of making the robot aware of the deformation present in the scene is demonstrated. This approach is applicable in domains such as food processing, healthcare, and other fields where gentle and precise manipulations are required. The method is useful in industrial applications in which deformation cannot be completely avoided but still needs to be tracked.
An important aspect of robotic grasping is the ability to detect incipient slip based on real-time information through tactile sensors. In this paper, we propose to use Video Vision Transformers to detect the onset of slip in grasping scenarios. The dynamic nature of slip makes Video Vision Transformers well-suited for capturing temporal correlations with relatively small datasets. The training data is acquired through two GelSight tactile sensors attached to the generic finger grippers of a Panda Franka Emika robot arm that grasps, lifts and shakes 30 everyday objects in order to induce slip. We further conducted an ablation study by considering 5, 4, 3, and 2 frames prior to slip onset, revealing consistent prediction accuracy. Our approach demonstrates the capability to predict slips well in advance, even up to the 5th frame before the onset. This underscores the predictive capability of our approach, indicating its effectiveness in slip detection well before of its occurrence. This advance prediction capability may be a valuable tool for undertaking preemptive corrective actions, such as implementing a more secure gripper closure. We evaluate the efficiency of our approach to predict onset of slip on 10 previously-unseen objects and achieve a zero-shot mean prediction accuracy of 99%.
Research on robotic manipulation of fragile, compliant objects, such as food items, is gaining traction due to its game-changing potential within the food production and retailing sectors, currently characterized by manually intensive and highly repetitive tasks. Food products exhibit high levels of frailness, biological variation, and complex 3-D shapes and textures. For these reasons, introducing greater levels of robotic automation in the food and agricultural sectors remains an important challenge. This article addresses this challenge by developing a human-centered, haptic-based, learning from demonstration (LfD) policy that enables pretrained autonomous grasping of food items using an anthropomorphic robotic system. The policy combines data from teleoperation and direct human manipulation of objects, embodying human intent and interaction areas of significance. We evaluated the proposed solution against a recent state-of-the-art LfD policy as well as against two standard impedance controller techniques. Results show that the proposed policy performs significantly better than the other considered techniques, leading to high grasping success rates while guaranteeing the integrity of the food at hand.
If we are to develop robust robot-based automation in primary production and processing in the agriculture and ocean space sectors, we have to develop solid vision-based perception for the robots. Accurate vision-based perception requires fast 3D reconstruction of the object in order to extract the geometrical features necessary for robotic manipulation. To this end, we present an accurate, real-time and high-resolution ICP-based 3D registration algorithm for eye-in-hand configuration using an RBG-D camera. Our 3D reconstruction, via an efficient GPU implementation, is up to 33 times faster than a similar CPU implementation, and up to eight times faster than a similar library implementation, resulting in point clouds of 1 mm resolution. The comparison of our 3D reconstruction with other ICP-based baselines, through trajectories from 3D registration and reference trajectories for an eye-in-hand configuration, shows that the point-to-plane linear least squares optimizer gives the best results, both in terms of precision and performance. Our method is validated for the eye-in-hand robotic scanning and 3D reconstruction of some representative examples of food items and produce of agricultural and marine origin.
In this paper, we propose a novel approach for transferring a deep reinforcement learning (DRL) grasping agent from simulation to a real robot, without fine tuning in the real world. The approach utilises a CycleGAN to close the reality gap between the simulated and real environments, in a reverse real-to-sim manner, effectively "tricking" the agent into believing it is still in the simulator. Furthermore, a visual servoing (VS) grasping task is added to correct for inaccurate agent gripper pose estimations derived from deep learning. The proposed approach is evaluated by means of real grasping experiments, achieving a success rate of 83 % on previously seen objects, and the same success rate for previously unseen, semi-compliant objects. The robustness of the approach is demonstrated by comparing it with two baselines, DRL plus CycleGAN, and VS only. The results clearly show that our approach outperforms both baselines.
Recent developments have shown that Deep Learning approaches are well suited for Human Action Recognition. On the other hand, the application of deep learning for action or behaviour recognition in other domains such as animal or livestock is comparatively limited. Action recognition in fish is a particularly challenging task due to specific research challenges such as the lack of distinct poses in fish behavior and the capture of spatio-temporal changes. Action recognition of salmon is valuable in relation to managing and optimizing many aquaculture operations today such as feeding, as one of the most costly operations in aquaculture. Inspired by these application domains and research challenges we introduce a deep video classification network for action recognition of salmon from underwater videos. We propose a Dual-Stream Recurrent Network (DSRN) to automatically capture the spatio-temporal behavior of salmon during swimming. The DSRN combines the spatial and motion-temporal information through the use of a spatial network, a 3D-convolutional motion network and a LSTM recurrent classification network. The DSRN shows an accuracy that is suitable for industrial use in prediction of salmon behavior with a prediction accuracy of 80%, validated on the task of predicting Feeding and NonFeeding behavior in salmon at a real fish farm during production. Our results show that the DSRN architecture has high potential in feeding action recognition for salmon in aquaculture and for applications domains lacking distinct poses and with dynamic spatio-temporal changes.
The robotic handling of compliant and deformable food raw materials, characterized by high biological variation, complex geometrical 3D shapes, and mechanical structures and texture, is currently in huge demand in the ocean space, agricultural, and food industries. Many tasks in these industries are performed manually by human operators who, due to the laborious and tedious nature of their tasks, exhibit high variability in execution, with variable outcomes. The introduction of robotic automation for most complex processing tasks has been challenging due to current robot learning policies. A more consistent learning policy involving skilled operators is desired. In this paper, we address the problem of robot learning when presented with inconsistent demonstrations. To this end, we propose a robust learning policy based on Learning from Demonstration (LfD) for robotic grasping of food compliant objects. The approach uses a merging of RGB-D images and tactile data in order to estimate the necessary pose of the gripper, gripper finger configuration and forces exerted on the object in order to achieve effective robot handling. During LfD training, the gripper pose, finger configurations and tactile values for the fingers, as well as RGB-D images are saved. We present an LfD learning policy that automatically removes inconsistent demonstrations, and estimates the teacher's intended policy. The performance of our approach is validated and demonstrated for fragile and compliant food objects with complex 3D shapes. The proposed approach has a vast range of potential applications in the aforementioned industry sectors.
Despite advances in computer vision and segmentation techniques, the segmentation of food defects such as blood spots, exhibiting a high degree of randomness and biological variation in size and coloration degree, has proven to be extremely challenging and it is not successfully resolved. Therefore, in this paper, we propose an approach for robust automated pixel-wise classification for segmentation of blood spots, focusing specifically on challenging texture-uniform cod fish fillets. A multimodal vision system, described in this paper, enables perfectly aligned RGB and D-depth images for localization of segmented blood spots in 3D. Classification models based on (1) Convolutional Neural Networks - CNN and (2) Support Vector Machines - SVM for the classification of defective fillets were developed. A colour based, pixel-wise and SVM-based model was developed for accurate segmentation and localisation of blood spots resulting in 96% overall accuracy when tested on whole fillet images. Classification between normal and defective fillets based on GPU (Graphical Processing Unit)- accelerated CNN classification model achieved 100% accuracy, versus the SVM-based model achieving 99%. We present a novel data augmentation approach that desensitizes the CNN towards shape features and makes the CNN to focus more on colour. We show how pixel-wise classification is,used for an accurate localization of blood spots in 3D space and calculation of resulting 3D gripper vectors, as an input to robotic processing. (C) 2017 Elsevier B.V. All rights reserved.
In Norway, the final stage of front half chicken harvesting is still a manual operation due to a lack of automated systems that are suitably flexible with regard to production efficiency and raw material utilisation. This paper presents the ‘GRIBBOT’ – a novel 3D vision-guided robotic concept for front half chicken harvesting. It functions using a compliant multifunctional gripper tool that grasps and holds the fillet, scrapes the carcass, and releases the fillet using a downward pulling motion. The gripper has two main components; a beak and a supporting plate. The beak scrapes the fillet down the rib cage of the carcass following a path determined by the anatomical boundary between the meat and the bone of the rib cage. The supporting plate is actuated pneumatically in order to hold the fillet. A computer vision algorithm was developed to process images from an RGB-D camera (Kinect v2) and locate the grasping point in 3D as the initial contact point of the gripper with the chicken carcass for harvesting operation. Calibration of camera and robot was performed so that the grasping point was defined using 3D coordinates within the robot’s base coordinate frame and tool centre point. A feed-forward Look-and-Move control algorithm was used to control the robot arm and generate the motion trajectories, based on the 3D coordinates of the grasping point as calculated from the computer vision algorithm. The results of an experimental proof-of-concept demonstration showed that GRIBBOT was successful both in scraping the carcass, grasping chicken fillets automatically and in completing the front half fillet harvesting process. It demonstrated a potential for the flexible robotic automation of the chicken fillet harvesting operation. Its commercial application, with further development, can result in automated fillet harvesting, while future research may also lead to optimal raw material utilisation. GRIBBOT shows that there is potential to automate even the most challenging processing operations currently carried out manually by human operators.