
Image segmentation is a fundamental step in computer vision and a cornerstone of robotic perception, serving as the foundation for interpreting data acquired from vision sensors, enabling robots to analyze complex visual environments, identify and localize objects, and support intelligent decision-making and autonomous control. It plays a critical role in applications such as autonomous navigation, robotic manipulation, medical robotics, agricultural robotics, autonomous vehicles, and human–robot interaction. Image segmentation has evolved from classical methods, which relied on handcrafted rules and mathematical models, to deep learning approaches that learn complex visual patterns directly from data. This evolution reflects advances in algorithms, computational power, and the theoretical foundations of mathematics and data science. Modern deep learning methods rely heavily on large, well-annotated datasets to train sophisticated neural networks. Yet, classical techniques remain valuable in certain scenarios, offering faster, reliable results without extensive computational requirements. Understanding the strengths and limitations of both approaches is key to selecting the right method. This paper surveys image segmentation techniques, comparing them in terms of accuracy, computational cost, and processing speed to guide informed method selection.
Parallel Delta robots are widely used in high-speed industrial automation, particularly in food-processing operations that require short cycle times and precise manipulation. Their performance under complex loading conditions, however, remains critical for reliable integration into mechatronic production systems. In many technological processes, the end effector experiences not only gravitational forces but also additional loads caused by product contact, lateral disturbances, and dynamic gripping actions. These conditions influence actuator torque demand, structural behaviour, and workspace feasibility. To address this, the present study develops a complete kinematic formulation and a force analysis model for the ABB IRB 360 3/1130 FlexPicker® Delta robot sourced from ABB Robotics, Västerås, Sweden. The models enable computer-based simulation of the robot’s behaviour under both vertical and horizontally oriented external forces. The results show how combined loading affects torque distribution and may restrict feasible workspace regions. These findings support motion planning, task feasibility assessment, and the reliable integration of high-speed parallel manipulators in industrial mechatronic environments.
This study presents the design, embedded implementation, and screening-level experimental terrain-performance evaluation of an RHex-inspired hexapod robot using a fixed encoder-assisted alternating-tripod state-machine gait. The work aims to provide an experimentally grounded assessment of how terrain properties and practical leg-thickness variation influence the locomotion of a fabricated low-complexity legged platform. A 2 × 2 × 2 full-factorial screening design was adopted to evaluate terrain roughness, terrain compliance, and leg thickness. Four terrain conditions were tested: concrete, rocky terrain, foam mats, and grass, corresponding to smooth–rigid, rough–rigid, smooth–soft, and rough–soft surfaces, respectively. Each treatment combination was evaluated in two replicate runs using final forward displacement, lateral displacement, absolute displacement, and peak current as the response variables. The full-factorial analysis showed that terrain roughness had the clearest significant effect on forward displacement and absolute displacement, while terrain compliance significantly affected absolute displacement and showed observable trends in lateral displacement and peak current. Leg thickness did not produce a statistically significant main effect within the tested 2.5 mm and 5.0 mm configurations, fixed gait, and terrain set. The findings indicate that, for this platform and experimental scope, terrain roughness and compliance affected locomotion more strongly than the tested morphology variation. The study contributes a reproducible baseline workflow for terrain-performance evaluation in low-complexity legged robots and identifies directions for future work involving stronger replication, quantified terrain characterization, improved energy measurement, and closed-loop terrain-adaptive control.
Generalist robots need to perform diverse tasks while operating in dynamic, uncertain, and unstructured environments, often around human beings. Vision-language-action (VLA) models have recently emerged as a promising and flexible framework for integrating perception, reasoning, robotic control, and action execution to develop generalist robotic policies. This systematic literature review (SLR) examines more than 140 VLA-related publications between 2020 and 2025 following the Preferred Reporting Items for Systematic Reviews and Meta-Analyses (PRISMA) guidelines. To the best of our knowledge, it is the first PRISMA-compliant systematic review dedicated to VLA models, offering a structured discussion of robotic policies, VLA architectures, and inference optimization methods. The review also presents descriptive analyses of the included studies and a glossary defining the terminology commonly used in VLA and generalist robotic policy research. The findings reveal substantial diversity among VLA models in terms of their supported modalities, robotic embodiments, training strategies, and architectural designs. Despite the rapid growth of VLA research, several important areas remain underexplored, including the execution of complex, long-horizon tasks, effective integration of speech, and deployment on low-cost hardware, while ensuring robust, safe, and secure operation.
In studies of intelligent agents, the pursuit–evasion problem, commonly related to the predator–prey paradigm, has been used as a reference setting for examining decision-making, coordination, and adaptation in multi-agent systems. In this paper, pursuit and evasion strategies are reviewed from their theoretical foundations to the gradual incorporation of adaptive methods based on evolutionary algorithms and machine learning. Analytical formulations drawn from control theory, game theory, and graph-based models are considered together with learning-oriented methods, including multi-agent reinforcement learning, neuroevolution, evolutionary robotics, and competitive coevolution. The literature is arranged by methodological paradigm so that the scope, limitations, and applicability of each approach can be discussed in relation to dynamic and uncertain environments. Through this organization, classical models, algorithmic developments, and recent research trends are brought into the same discussion, while relevant gaps and possible future directions in pursuit–evasion research are identified. The contribution of this work is a structured synthesis in which analytical and adaptive perspectives are brought together within a unified reference for researchers and practitioners in robotics, artificial intelligence, and multi-agent systems.
Upper-limb motor dysfunction resulting from neurological disorders severely limits patients’ activities of daily living and social participation. Pneumatic upper-limb rehabilitation robots have emerged as a promising intervention owing to their inherent compliance, lightweight design, and high power-to-weight ratio, which facilitate safe, repetitive, and home-based training. Despite these advantages, extensive clinical translation remains hindered by challenges including actuator hysteresis, nonlinear dynamics, limited accuracy in intention recognition, and inconsistent clinical evaluation metrics. This review systematically examines recent advancements in pneumatic upper-limb rehabilitation robots across four critical dimensions: structural design, human–robot interaction, control strategies, and clinical translation. We comparatively analyze rigid exoskeletons, soft wearable devices, and rigid–soft hybrid configurations based on output capability, motion accuracy, comfort, and clinical applicability. The findings suggest that while rigid systems offer high precision and soft systems maximize safety, rigid–soft hybrid architectures represent a critical developmental trend for balancing motion accuracy with interaction compliance. Furthermore, the review evaluates multimodal sensing techniques (e.g., EMG, EEG, and IMUs) for motion intention decoding and training state monitoring, alongside conventional, adaptive, and artificial intelligence-driven control methods aimed at compensating for pneumatic nonlinearity and improving real-time response. Current clinical evidence indicates that these systems effectively enhance upper-limb function and muscle strength, particularly in post-stroke rehabilitation; however, existing trials are frequently constrained by small sample sizes, short interventions, and heterogeneous protocols. Future research must prioritize rigid–soft hybrid architectures, robust multimodal sensor fusion, digital twin-assisted assessment, adaptive intelligent control, and standardized home-based rehabilitation platforms. Ultimately, this comprehensive review provides a concise reference for the design optimization and clinical deployment of next-generation pneumatic rehabilitation systems.
Service robots often receive natural language instructions in changing workspaces where multiple visible objects may match one description. Relying on detector confidence, random selection, or direct vision–language model (VLM) prediction can lead to a wrong action. This paper presents ActivAsk, a zero-shot framework for resolving referential ambiguity before robotic grasping. ActivAsk constructs open-vocabulary candidates from red-green-blue-depth (RGB-D) input, asks candidate-grounded yes/no questions when needed, updates the candidate state from the user’s answer, and grasps after target resolution. It selects among VLM-proposed candidate partitions using an expected free energy (EFE) criterion motivated by active inference; with neutral response preferences, this reduces to information gain over candidate partitions. Offline experiments showed that interactive clarification improved target accuracy from about 53–54% for noninteractive baselines to about 90–92%. ActivAsk matched the best interactive accuracy (92.13%) while asking 15.47–19.71% fewer questions on asked trials and 21.43–23.88% fewer for ambiguous instructions. In online real robot experiments, ActivAsk achieved 92.98% target selection accuracy and 87.72% full correct object grasp success; unresolved or wrong targets were not physically executed after operator-controlled verification and were counted as task failures.
Accurately measuring the six degrees of freedom (6-DoF) pose of a sample is a critical prerequisite for applications in robotic sample handling, optical metrology, and industrial quality control. Our pose estimation problem requires matching a small, partial observation of 4–20 discrete 3D points to a reference point set of up to 50 points. Popular registration algorithms such as Go-ICP, TEASER++, and MAC are designed for large-scale correspondence problems and do not perform reliably in our operating regime. Therefore, we propose an improved geometric hashing method that robustly estimates the rigid transformation in the presence of noise and outliers, and demonstrate its effectiveness and speed using both simulated and real-world datasets.
Logarithmic-spiral soft grippers couple tendon actuation, variable-curvature morphology, distributed contact, and frictional load support. This study evaluates a finite-candidate predictive force-shape (PFS) controller within a control-oriented reduced-order surrogate of a two-tendon gripper. PFS is compared with open-loop, fixed-tension, position-only, and hybrid force-shape controllers across multiple object geometries, simultaneous uncertainty, payload-friction conditions, transient loads, ablations, and parameter variations. In the nominal study, PFS produced a mean force RMSE of 7.26 N and a mean peak local force of 8.07 N, while the comparison implementations produced force RMSE values from 57.9 N to 99.3 N and peak forces near 30.6 N. This behavior involved a geometric tradeoff: PFS position RMSE was 0.088 m, compared with 0.073 m for PO and 0.074 m for HFS. The remaining numerical studies characterize how this tradeoff changes inside the surrogate. A matched higher-resolution verification at (N,Nc)=(60,48) preserved the principal force–position tradeoff: PFS yielded a mean force RMSE of 10.25 N and peak local force of 8.84 N, while PO and HFS retained lower position RMSE. Because the morphology and contact relations are phenomenological, the results are interpreted as reproducible numerical evidence rather than experimental validation or proof of physical superiority.
Remote teleoperation interfaces for precision manipulation should support operator awareness and usability while preserving human control, consistent with the human-centred goals of Industry 5.0. This study conducted a within-subject user evaluation with 18 participants using a UR5 arm. Three interface conditions were compared: an enhanced first-person interface with visual overlays (FPV+), a multi-view interface combining first-person and third-person views (PiP), and a multi-view interface with enhanced visual overlays (PiP+). Participants completed a ball task and a pen insertion task in pseudo-randomised order. Binary task success was analysed using generalised linear mixed-effects models, while System Usability Scale scores were analysed using a linear mixed-effects model. Neither adding a third-person view to the enhanced first-person interface nor adding the minimap and tactile-dot overlays to the multi-view interface was significantly associated with task success. A post hoc sensitivity analysis showed that these binary-success comparisons were sensitive only to large effects, so they are interpreted as inconclusive rather than as evidence of absence. Success was significantly lower for the pen task, and its relationship with prior human–robot interaction experience differed from that observed for the ball task. Adding a third-person view was associated with an approximately 16-point higher System Usability Scale score, whereas the minimap and tactile-dot overlays were not significantly associated with usability. These findings illustrate a human-centred evaluation principle relevant to Industry 5.0: an interface feature can improve the operator’s experience of a teleoperation system even when no corresponding change in objective task success is detectable, revealing a benefit that a performance-only evaluation might overlook.
Efficient robotic teleoperation for industrial inspection requires interfaces that maximize intuitive control and situational awareness. While traditional mixed reality (MR) systems offer alternatives, standard controller-based methods lack environment customization and demand external peripheral hardware. This study introduces a peripheral-hardware-free, customizable 3D spatial virtual cockpit application deployed on a Meta Quest 3 headset for long-distance, non-line-of-sight teleoperation of the Improbability Roller-2, a mobile robotic platform that dynamically adjusts its wheel geometry to traverse varied terrain over a Virtual Private Network (VPN). The system’s novel spatial cockpit architecture allows operators to scale multi-channel parameters dynamically without relying on physical hardware controllers. The framework was evaluated using transmission-quality benchmarks, an active industrial machine shop deployment, and a 20-participant user study assessing usability and cognitive load. Experimental results yielded a mean System Usability Scale (SUS) score of 82.75, corresponding to an excellent usability rating, and a low mean operator workload, with a NASA Task Load Index (NASA-TLX) score of 5.19 out of 21. Participants also rated the workspace customization feature positively, assigning it a rating of 4.6 out of 5. In terms of communication performance, the system successfully completed all inspection tasks even under severely degraded VPN conditions. These findings demonstrate that the proposed customizable 3D spatial interface provides a robust, scalable alternative to traditional controller-based setups for long-distance remote robotic inspection in real-world industrial settings.
Safe human–robot collaboration remains a critical challenge in manufacturing. Traditional safety approaches, such as cages and proximity sensors, are often insufficient for dynamic human interaction. This paper presents a digital twin-based collision avoidance framework for industrial collaborative robot manipulation. The system integrates RGB-D sensing, human pose estimation using Ultralytics YOLO26s-pose, Kalman-filter-based 3D arm tracking, short-term motion prediction, and QP-based reactive motion control. Human arm keypoints detected from RGB-D images are reconstructed in 3D, transformed into the robot base frame, and tracked during temporary occlusion using Kalman filtering with kinematic constraints. Predicted human–robot clearance is evaluated to trigger speed reduction, stopping, or collision avoidance commands. The framework was implemented with a UR10e robot, an Intel RealSense D435 camera, a Unity3D digital twin, and ROS communication. Controlled laboratory experiments demonstrated the proof-of-concept feasibility of the integrated framework for tracking human arm motion, anticipating proximity risk, and triggering protective robot responses. The results do not establish deployment readiness in complex industrial or multi-participant environments.
Most gynecological interventions do not take advantage of the possible access through the natural orifice to the operating zone and/or use rigid tools, which leads to more invasive procedures. The purpose of this research is to reduce invasiveness by creating a natural orifice endoscopic surgical tool. By analyzing the varied anatomies present in patients to extract functional requirements, we propose a conceptual design that allows for better navigation of the environment thanks to a custom design with active control over endoscope shape. We manufactured and tested this new design of a tendon-driven continuum robot in a phantom that is representative of the geometrical properties and variability of a uterus, validating its operation and functionality.
This paper presents a distributed coordination framework for simultaneous multi-target tracking using a mobile wireless sensor network (MWSN) based on discrete-event-system principles. The proposed framework employs a finite-state-machine architecture, where autonomous mobile sensors sequentially process detection and tracking events. Unlike passive tracking approaches that react to target loss after it occurs, the proposed strategy implements predictive handover through Extended-Kalman-Filter-based uncertainty propagation. This enables sensors to anticipate target loss and to reposition auxiliary sensors in advance, acquiring targets along their predicted trajectories. A bidding-based allocation mechanism coordinates sensor assignments by evaluating four competing objectives: network preservation, spatial proximity to handover points, temporal mission feasibility, and estimation uncertainty. The proposed framework integrates four components: EKF-convergence-triggered proactive handover, multi-objective competitive bidding, distributed min–max conflict resolution, and fusion-driven proportional navigation. Unlike existing methods, auxiliary sensors navigate using confidence-weighted EKF estimates shared by neighboring sensors rather than their own measurements. An ablation study over ten Monte Carlo trials confirms that each component contributes independently, with EKF-based predictive triggering identified as the dominant performance driver.
This study presents a decentralized shared actor–critic framework for cooperative multi-robot coverage in continuous two-dimensional simulation. The method combines permutation-invariant local observations, continuous differential-drive control, and reward shaping based on stepwise Hungarian assignment distances, collision penalties, and time efficiency. Homogeneous teams of four, five, and six agents are evaluated in an obstacle-free environment using five independent training seeds. In the final training window, the full reward configuration achieved full-team success rates of 98.2 ± 2.9% for four agents, 85.1 ± 18.0% for five agents, and 96.3 ± 2.0% for six agents, with mean landmark coverage above 96% in all cases. The lower mean in the five-agent setting was associated with higher seed-level variability dominated by one low-success seed. Reward ablations without assignment shaping or collision penalties remained viable, and seed-level tests did not show a statistically significant final-window advantage of the full reward configuration. The full configuration reached the 80% rolling-success threshold earlier in median terms, with the clearest seed-level support in the four-agent setting. Within-environment comparison showed higher full-team success than MADDPG and MAPPO under the matched training horizon and final-window protocol. Deterministic arena-size transfer from 15×15 to 30×30 showed decreasing full-team success as arena size increased, while partial landmark coverage remained higher than strict full-team completion. The results support the method for small homogeneous teams in the tested obstacle-free simulation, while larger teams, external obstacles, aerial-robot dynamics, formal safety guarantees, and hardware deployment remain future work.
In recent years, hand exoskeleton robots have attracted extensive attention from researchers and practitioners due to their potential to rehabilitate, assist, and enhance hand movements, particularly for stroke patients. With an ageing population increasingly affected by strokes, there is a growing demand for patient-centred interventions which place less demand on clinicians, especially wearable devices that can enhance hand function. Advances in artificial intelligence have opened new avenues for developing more reliable and adaptive assistive systems. This study presents a systematic literature review, following the PRISMA protocol on the design elements of hand exoskeleton robots, acknowledging the emerging perspectives on AI integration and ethical considerations. The study provides a comprehensive foundation for future research and development in rehabilitation technologies by systematically synthesising the current mechanical architecture, actuation, sensors, material, weight, and cost aspects of soft hand exoskeleton robots for rehabilitation. The results show important patterns and trade-offs in various design dimensions, providing useful information to direct the development of more accessible and efficient rehabilitation solutions in the future.
Long-term localization in dynamic and changing environments remains a key challenge for autonomous vehicles. Semantic Simultaneous Localization and Mapping (SLAM) enhances traditional SLAM by integrating high-level semantic understanding, enabling robust mapping and localization even under complex scenarios. In this context, multi-modal sensor fusion—particularly the combination of LiDAR and camera data—has proven essential in leveraging complementary strengths: the geometric accuracy of LiDAR and the rich semantic cues from images. A significant advancement in this domain is the adoption of graph-based semantic localization frameworks, where semantic entities and spatial relationships are encoded in graph structures to improve map consistency, loop closure detection, and data association over time. This review presents a comprehensive survey of recent developments in Semantic SLAM, with a focus on long-term localization for autonomous vehicles using multi-modal fusion strategies. We categorize existing methods into traditional SLAM, vision-based, point-cloud-based, and graph-based techniques, emphasizing the role of semantic data association and loop closure in maintaining long-term consistency. Additionally, we discuss the integration of deep learning techniques for semantic segmentation and feature extraction. Finally, we analyze widely used datasets and evaluation metrics, identifying current limitations and proposing directions for future research on robust, scalable, and semantically enriched localization.
This paper highlights the state of the art in Cooperative Dual-Manipulation (CDM) and Cooperative Multi-Manipulation (CMM), comparing advances in modeling, control, planning, sensing, vision, and end-effector technologies. Methods originally established in CDM have been extended or adapted to support higher complexity of CMM. A historical timeline visualizes the steady growth of cooperative manipulation (CM) and the recent acceleration of CMM driven by rising process complexity and the need for more flexible automation strategies. CM is becoming increasingly relevant as industrial processes demand higher payload capacity, larger workspaces, and greater flexibility. In addition, this paper categorizes existing applications by cooperation type and application domain. Here, a clear dominance of simultaneous object manipulation tasks is visible (fixation-fixation). However, fixation-tooling tasks, where one manipulator grasps the product while another performs a tool operation, and tooling-tooling tasks, where multiple manipulators perform tool operations simultaneously, remain significantly underrepresented. A similar imbalance is found for rigid/non-deformable object manipulation and flexible/deformable object manipulation, respectively. Based on this review, several research gaps are identified: (i) reliable flexible object manipulation methods; (ii) CM strategies for disassembly (e.g., battery pack deconstruction); (iii) complexity in control and planning for multi-manipulator systems; (iv) pathways to industrial deployment beyond laboratory demonstrators; and (v) task-specific tooling and end-effector innovation.
Sampling control trajectories from standard distributions—a foundation for Model Predictive Control (MPC) and Model Predictive Path Integral (MPPI) methods—although probabilistically complete, in practice leads to poor exploration of the configuration space, resulting in catastrophic outcomes in the robot’s operation. Recent works have proposed a family of approaches for sampling in control space that lead to uniform coverage of the configuration space, introducing the notion of C-Uniformity. However, those methods suffer from the need for state and action space discretization and level-sets, pre-computation. In this work, we introduce a proof-of-concept training strategy for deep generative models to achieve diverse and near C-Uniform sampling capabilities. Two variational inference approaches—information maximization algorithm and Stein variational gradient descent—are used as a foundation for training normalizing flow and flow matching models to sample control sequences that lead to a wider coverage of the configuration space, compared to the standard distributions, without state/action space discretization. Qualitative and quantitative evaluations support the advantage of the methods in terms of state space coverage and success rate in downstream social navigation tasks.