
IntroductionHuman-drone interaction is an emerging field investigating ways to extend the conventionally passive role of drones with autonomous, socially interactive functionalities. Drones offer unique opportunities for applications, such as exercise support, due to their mobility and flexibility. While previous research has studied drones for, e.g., pacing, gait analysis, and accessibility in running contexts, we propose that an equally crucial part of such applications is the socially interactive role of the drone. MethodsWe present a study in two phases. In the first phase, we conducted a preliminary study that informed the design of a co-design workshop through eight semi-structured interviews exploring drone support for runners. In the second phase, we designed and conducted a co-design workshop with five experienced recreational runners to reflect on the practical and ethical challenges of using drones and to envision ways in which drones could support and interact with runners in future scenarios. The collected materials were analyzed using reflexive thematic analysis and annotated visual analysis.ResultsThe results present five low-fidelity social drone prototypes, including participants’ reflections tied to each prototype. Furthermore, we present analytical themes that demonstrate the necessity of extending and re-framing approaches to designing social drones.DiscussionThe paper makes three key contributions. First, we offer an approach to co-designing social drones that focus on eliciting the experiences of participants and manifesting those in concrete, embodied design concepts. Second, we validate the existing Design Space for Social Drones framework through a concrete application domain. Our validation shows that the dimensions can be further detailed, and that non-user perspectives are a potentially crucial extension. Third, we present three recommendations that encapsulate concerns for further designs of social drones.
Concentrating on various types of grasped objects, this work presents a distributed coordinated control scheme for a dual-arm reconfigurable manipulator (DRM) via adaptive dynamic programming (ADP). Kinematic and force analyses are performed to formulate the dynamics of a single-arm manipulator and the object using joint torque feedback technology and the Newton–Euler formulation. To address uncertainties associated with the grasped object, a gradient model-based adaptive algorithm is developed for online estimation of the grasp matrix relating the object to the DRM. By employing an adaptive observer to identify unknown dynamic terms, an enhanced optimal performance index is constructed that incorporates both object tracking performance and internal force effects within the manipulator. The performance index function is approximated by only a critic neural network, and the optimal control policy is obtained by policy iteration. Lyapunov stability analysis demonstrates that internal force error, position tracking error, and critic network weight approximation error are all uniformly ultimately bounded. Experimental results confirm the effectiveness of the developed coordinated control approach.
Introduction:Tendon-driven exoskeletons are often limited by mechanical bulk, complex cable routing, and large externally mounted actuators. Similarly, conventional linear actuators used in wearable exoskeleton systems are often limited by low stroke-to-body-length ratios that restrict the range of motion. This study presents the mechanical design of a novel Rail-guided Inline DirEct dRive (RIDER) with a high stroke-to-length ratio for wearable exoskeletons. Methods:The proposed system creates a linearized RIDER system using a "motor-driven cart" layout that travels along a toothed twin-rail structure. The prototype system was developed using 3D-printed materials and open-source electronics, and the resultant force-velocity performance was experimentally evaluated under vertical loading. Vertical displacement was measured using an ultrasonic sensor (HC SR-04), and the corresponding velocity was computed via a finite-difference approximation. Results:The actuator demonstrated a maximum free-load velocity of 7.27 cm/s and a maximum force output of 103.7 N prior to slipping. When integrated into a tendon-driven system, the duo configuration produced a 6.22 Nm torque output with a range of motion of 120° or a 10.37 Nm torque output with a range of motion of 72°. The actuator also demonstrates a 190% (i.e., 0.58 vs 0.2) higher stroke-to-length ratio than a representative telescopic actuator of similar geometry. Conclusion:In this study, we show that the RIDER system demonstrates a wearable, rail-based actuation strategy capable of adjustable torque amplification and improved spatial-efficiency integration for human-robot interaction in wearable robotics. The current study focuses on the benchtop mechanical validation of the RIDER system as a proof-of-concept for future integration into wearable exoskeletons. This research is part of a larger study to develop adaptive exoskeletons that enable scalable performance improvements for assistive robotics in the workplace and neurorehabilitation.
An insole-type active assist device has been developed as a robotic system to dynamically correct ankle alignment at heel contact in patients with medial knee osteoarthritis. Although our previous feasibility study demonstrated that the device could be safely used during an on-the-spot stepping task, its effects on loading behavior remain unclear. This study aimed to investigate whether dynamic ankle alignment correction using the device alters horizontal ground reaction force variability and center-of-pressure sway during stepping. Force-plate data obtained from six ambulatory patients with medial knee osteoarthritis were analyzed as a secondary biomechanical analysis. Each participant performed repeated stepping trials under two conditions: a non-control condition, in which the device was worn without motor control, and a control condition, in which heel-eversion assistance was provided. Stance phases were extracted using vertical ground reaction force, and horizontal ground reaction force components, center-of-pressure sway measures, peak vertical ground reaction force, and vertical impulse were calculated. Compared with the non-control condition, the control condition reduced the ranges of anterior–posterior and medial–lateral ground reaction force components in all participants. Center-of-pressure rectangular sway area also decreased consistently, and trajectory length tended to decrease, whereas peak vertical ground reaction force and vertical impulse showed little change. These findings suggest that dynamic ankle alignment correction using a robotic insole-type device may reduce horizontal loading variability and center-of-pressure sway during stepping without substantially altering vertical loading. Because this study was based on a small sample and an exploratory secondary analysis, the findings should be interpreted as preliminary biomechanical evidence rather than evidence of clinical effectiveness.
Introduction:Autistic children are increasingly engaging with social robots in educational and support contexts, but limited research has compared the perceptual and design preferences of autistic and neurotypical children. Methods:This mixed-methods study examined the robot design preferences of 43 children aged 3-15 years, including 21 autistic and 22 neurotypical children, and included semi-structured interviews with 11 adult stakeholders, comprising six experts and five non-experts. Quantitative analyses explored relationships between robot preference and perceptual features, including size, color, movement, voice, and anthropomorphism, while qualitative thematic and co-occurrence analyses examined perceived benefits, ethical issues, accessibility challenges, and ideal design characteristics. Results:Perceptual characteristics, particularly movement intensity and color multiplicity, influenced robot preference. Across groups, children generally preferred humanoid robots with soft voices and eye movement. Gender-related differences appeared more consistent than diagnostic differences in preferences for size and movement speed. Autistic children highlighted playful and social interaction functions and showed greater variability in preferred personality-related traits. Qualitative findings indicated strong interest in robot-supported interactions, while also raising concerns about accessibility and possible emotional dependency. Discussion:These findings suggest that inclusive social robot design should prioritize modificable sensory features, moderated anthropomorphism, predictable interaction patterns, and personalization mechanisms that account for individual, gender-related, and diagnostic differences.
Robot-assisted physiotherapy has attracted increasing attention for its potential to provide repeatable, stable, and controllable physical interaction during rehabilitation-oriented therapy. However, contact-rich physiotherapy tasks remain challenging because the robot must reproduce therapist-demonstrated massage skills while adapting to non-planar and deformable body surfaces, suppressing impact during contact transition, and maintaining stable force regulation. This paper proposes a contact-aware robot-assisted physiotherapy framework that integrates task-space skill generalization, contact state estimation, dynamic velocity adjustment, force-error compensation, and bounded variable impedance control. Therapist-guided demonstrations are encoded using Dynamic Movement Primitives (DMPs) to construct a physiotherapy skill library, enabling typical massage skills, including kneading, patting, pushing, and pressing, to be generalized to new start and goal points in the robot Cartesian task space. During execution, force/torque feedback is used to estimate the contact point and surface normal, update the local task frame, regulate the approach velocity, and compensate for force-tracking errors. Experimental validation was conducted on a silicone abdominal model and in a preliminary healthy-volunteer back physiotherapy test. The results show that the proposed dynamic contact strategy suppresses excessive impact during contact transition, avoiding the 165.86 N peak impact observed under high-speed contact. The force compensation strategy reduces the force-tracking RMSE from 0.835 N to 0.395 N, corresponding to a 52.69% improvement. In addition, the bounded variable impedance strategy improves motion-force coordination compared with fixed-stiffness control. These results demonstrate that the proposed framework improves contact transition safety, force regulation accuracy, and task adaptability at the control-performance level, providing a feasible basis for further development of robotic physiotherapy systems.
Hydraulic booms used in material handling are typically commanded in cylindrical task-space coordinates, while dynamic models used for control are commonly formulated in joint coordinates. This mismatch leads to unnecessarily high-dimensional system representations when predictive control is applied to underactuated suspended loads. This paper addresses this modeling inconsistency by introducing a reduced-order state-space reformulation of grapple sway dynamics directly in cylindrical task-space velocity coordinates. The proposed input-consistent state transformation eliminates the dependence of the dynamics on input accelerations by embedding actuator–sway coupling within modified velocity states, yielding a nonlinear model driven solely by cylindrical velocity commands. Based on this representation, a nonlinear model predictive control (NMPC) framework is developed for simultaneous goal reaching, sway suppression, and obstacle avoidance. The NMPC cost function is designed to capture the practical trade-off between rapid boom motion and suppression of suspended load oscillations. In addition, approximate hydraulic actuator dynamics identified from a high-fidelity AMESim simulator model of a forwarder crane are incorporated to better reflect realistic boom behavior. Simulation studies in representative boom operation scenarios demonstrate that the proposed formulation enables effective sway regulation and accurate task-space motion control using a reduced set of states without requiring full joint-space models or payload sensing. The results indicate that the proposed cylindrical state-space reformulation simplifies predictive control design for underactuated hydraulic booms while preserving the essential dynamics of the suspended load.
Learning from demonstration (LfD) has become a popular approach with the emergence of modern transformer-based algorithms. However, the performance of these policies is limited by the quality of the demonstrations. Combining imitation and exploration promises to train policies that perform better and are more reliable. However, this requires a robotic system that can explore safely without damaging itself or the environment, especially in contact-rich tasks during which the robot must exert force on its environment to solve the task. In this study, we investigate the combination of a state-of-the-art reinforcement learning (RL) algorithm with human demonstrations to learn how to open a door with minimal task-specific engineering on an articulated soft robot arm. We found that learning from both exploration and demonstration data stored in separate buffers makes the algorithm not only more sample-efficient and robust but also allows the policy to reach a higher performance level than the provided expert demonstrations. We also show that using an articulated soft robotic arm allows us to perform RL on a real robotic system without any pretraining and with a simple safety system that does not require any additional sensors, such as force-torque sensors. Additionally, we can implicitly learn the nonlinearities stemming from the soft materials in the actuator. Our findings show that combining LfD with RL results in both better performance and more robust behaviors and indicate that articulated soft robots allow for learning contact-rich tasks safely on a real system.
Heterogeneous multi-robot teams require systems that can interpret natural-language goals, allocate tasks, and adapt to unexpected events. We developed CoMuRoS (Collaborative Multi-Robot System), a generalizable hierarchical architecture combining a centralized task-manager LLM with decentralized robot-level LLMs for executable Python generation from primitive ROS2 skills. The task manager uses static planning rules and dynamic context, including task history, robot/task status, and detected events, while onboard perception using VLM/image processing classifies events as relevant or irrelevant and triggers replanning. Hardware experiments demonstrated recovery from disruptive events, filtering of irrelevant distractions, and coordinated transport with emergent human-robot cooperation, achieving success rates of 9/10 for collaborative object recovery, 8/8 for coordinated transport, and 5/5 for human-assisted recovery. Simulation studies demonstrated intention-aware replanning. A curated benchmark of 22 scenarios, 54 tasks, and around 20 robots evaluated task allocation, classification, IoU, executability, and correctness across multiple LLMs, with correctness up to 0.91 ± 0.053; a 20-scenario replanning benchmark achieved Correctness = 0.948 ± 0.034 using Grok 3. CoMuRoS enables runtime, event-driven replanning on physical robots and supports flexible multi-robot and human-robot collaboration across diverse scenarios.
Industry 5.0 cobotics calls for collaborative robots that adapt to the operator’s physical and cognitive state in real time. Most current industrial deployments remain reactive, responding only after explicit commands and ignoring the operator’s psychophysiological condition. The three-module Human-Centric Digital Twin (HCDT) framework establishes the upstream perception-and-reasoning architecture for such systems using Vision-Language Models, and identifies closed-loop feedback and physical-assistance mechanisms as the natural next layer in the framework’s development. Building on that foundation, this paper presents a complementary, real-time instantiation specialised for the resource-constrained, vision-plus-biosignal setting typical of ergonomics-and-safety HRC, deployed on a standard collaborative robot (Universal Robots UR3 with a Robotiq 2F-85 gripper) and extended with closed-loop operator-state adaptation. The Perception module pairs MediaPipe hand-landmark tracking and a Random Forest gesture classifier with Tobii Pro Glasses 3 pupillometry. The Reasoning module combines a ten-state finite-state machine, a weighted composite fatigue score (blink rate, blink duration, PERCLOS, hand-jerk) banded into FRESH, MILD, MODERATE and SEVERE, a saccade-gated commitment fixation, and a Task-Evoked Pupillary Response (TEPR) detector for acute-stress soft-emergency-stop. The Motion module provides a closed-form analytical inverse-kinematics solver, workspace clamping and a slew-rate-limited joint-velocity controller. Velocity-based linear extrapolation of hand position over a 0.5 s horizon enables anticipatory robot pre-positioning. The system is presented as a single-operator feasibility study rather than a generalisation claim. On a leak-free split-then-augment evaluation (520 raw test samples not seen by the training fold in any orientation), the Random Forest classifier reaches 99.62% accuracy with both errors falling in the Okay-Neutral confusion pair that the geometric override layer is designed to catch; a comparison against MLP and KNN baselines on the same split shows all three classifiers within the real-time control budget. Saccade-gated early-commit is designed to recover the modelled 150–300 ms gaze-to-motion lead, and the TEPR Soft E-Stop operates on an independent 300–500 ms channel. End-to-end gesture-to-motion latency has a conservative worst-case ceiling of approximately 830 ms, formed by a camera-buffer term, the 10-frame debounce and one control tick, inside the 1-s interactive response-time limit adopted as the design target. Owing to a Tobii Pro Glasses 3 hardware fault, the physiological pathways (fatigue banding, TEPR soft-stop and gaze commitment) were verified through their downstream control responses to simulated triggers rather than validated on organic pupil data; organic, multi-participant validation is identified as future work.
Robots that execute language-conditioned tasks in dynamic environments often rely on feedback only after an action has failed, which can be insufficient when failures involve collisions or workspace conflicts. This paper presents a predictive monitoring framework that uses Vision-Language Models (VLMs) to assess near-future execution risk during robot task execution. The framework first generates structured plans with action execution conditions and a plan-level fallback action. During execution, a monitoring module combines visual observations, the current action, and the relevant execution conditions to estimate whether a condition is likely to be violated within a short future time window. When the predicted risk exceeds a task-specific threshold, the robot halts the current action, executes the fallback behavior, and replans from the updated state. We evaluate the approach in Gazebo simulation on mobile navigation with a moving human obstacle and manipulation with two robot arms sharing a workspace. Across controlled collision-risk settings, the proposed method achieves higher task success rates than reactive VLM-based baselines while requiring fewer replanning events than conservative current-state precondition checking. The results indicate that predictive vision-language monitoring can improve task completion in simulated dynamic robot tasks, while remaining subject to limitations such as VLM latency, prompt sensitivity, and evaluation beyond simulation.
This paper presents an explainable AI analysis of brake control in CARLA closed-loop driving. Longitudinal braking is studied through a threshold policy, a risk-aware reference policy, and a distilled policy learned from reference rollouts. The framework combines structured traffic-state features, high-fidelity XGBoost surrogates, and SHAP to analyze deployed brake behavior across Town10HD and Town05. Results show that brake generation is consistently dominated by front-vehicle distance, relative speed, and time to collision, while lateral variables contribute weakly. The distilled controller retains a forward-risk-oriented explanation structure rather than behaving as an arbitrary black box, but the degree of apparent semantic alignment with the reference policy is environment-dependent. Fallback-aware analysis shows that this alignment is substantially reinforced by the deployed safety fallback in Town10HD, while Town05 provides a cleaner view of the learned component itself. These findings show that learned brake control can remain interpretable when grounded in semantically structured state representations.
This work details a method for manipulating the horizontal position of free-drifting ocean-profiling floats using vertical depth control. Profiling floats are often used to study oceanographic processes of varying spatial and temporal scales. Since deploying floats is logistically demanding and expensive, active floats must collect measurements most relevant to the sampling objectives. To that end, we present the FloatCast framework for managing drift in a fleet of floats via vertical depth control. The FloatCast software selects float dive commands that are anticipated to align best with sampling objectives by combining tools from machine learning, optimization, and feedback control. First, an echo state network is used to generate forecasts of ocean surface flow. These forecasts are used to compute float trajectory predictions for each set of candidate dive commands. The performance of each set of dive commands is then ranked using a mapping error scoring metric, and the highest-ranking command set is provided to the floats for their next dive. Reported here are simulation and real-time experimental results from FloatCast control in the Philippine Sea. The simulation results provide insight into the broader FloatCast methodology; experimental results show that this control framework has promise, though it offers no performance guarantees. Although these floats are inherently underactuated, implementing FloatCast can increase the sampling effectiveness of a float array and, in turn, improve scientific data products.
Embodied picking for e-commerce fulfillment remains vulnerable to dense clutter, reflective packaging, and background variation, which can undermine the effectiveness and robustness of visuomotor policies learned from demonstrations. A key limitation is the absence of explicit object grounding, causing policies to exploit spurious contextual cues rather than task-relevant visual evidence. To address this issue, we propose the Foveated Diffusion Policy (FDP), which integrates object-centric visual grounding into diffusion-based action generation. FDP adopts a panorama–fovea dual-stream visual encoder: a low-resolution panorama stream captures global scene context, while a detection-guided fovea stream uses ROI Align to extract high-resolution object tokens. At each denoising step, these tokens are incorporated into the diffusion denoiser through object-guided cross-attention, thereby grounding action generation in target-specific visual evidence instead of undifferentiated global representations. We evaluate FDP on ten benchmark datasets spanning six simulated manipulation tasks with both proficient-human and multi-human demonstrations. FDP achieves average success rates of 81.0% and 76.8%, respectively, outperforming diffusion-based baselines and delivering particularly strong performance on object-centric tasks, including Push-T (86.7 ± 2.5%) and Lift (100.0 ± 0.0%). Ablation studies on a diagnostic subset show that foveated conditioning improves the average success rate by 11.3 percentage points over the non-foveated variant. FDP also maintains an 84.0% success rate under severe localization perturbations (σ = 10 pixels), demonstrating resilience to detection noise. Furthermore, accelerated DDIM sampling reduces policy-inference latency to approximately 55 ms with 10 denoising steps, supporting the feasibility of closed-loop control. Collectively, these results demonstrate that explicit object grounding is a promising and computationally practical approach to improving the robustness of visuomotor policies in warehouse-relevant simulated manipulation scenarios.
Collaborative robots (cobots) are increasingly deployed in industrial as well as non-industrial domains to support human-centered operation. While physical safety and task efficiency have received considerable attention, less is known about how feedback modality influences operator experience and physiological responses under different collaboration demands. This study examines the effects of feedback modalities in two human–cobot collaboration scenarios representing distinct coordination structures: a low-collaboration turn-taking task and a high-collaboration synchronous shared-control task. Across these scenarios, auditory, visual, and combined auditory–visual feedback modalities were evaluated, and an additional projected assistive interface (PAI) condition, designed specifically for the given scenario, was introduced in the high-collaboration task. Outcomes included task completion time, subjective user assessments of interaction quality and comfort, and physiological responses based on heart rate variability (HRV) indicators reflecting autonomic activity. The results indicate that feedback modality significantly influences performance, subjective experience, and physiological activation within each collaboration scenario. In the low-collaboration task, the no-feedback condition yielded the shortest completion times but was associated with lower perceived interaction quality and descriptively higher autonomic activation. In contrast, multimodal feedback improved perceived comfort and task understanding while slightly increasing completion time. In the high-collaboration task, the examined PAI guidance condition was associated with the shortest task completion times and the highest subjective ratings, while also being accompanied by increased sympathetic-dominant physiological activation. These findings indicate a design trade-off in which richer guidance can enhance coordination efficiency but may be associated with increased cognitive and physiological demands. Importantly, the results suggest that physiological activation should be interpreted in relation to task demands and performance outcomes, rather than as a direct indicator of detrimental stress. More broadly, the study suggests that collaboration configuration and feedback modality should be co-designed as interdependent parameters within the specific coordination demands of collaborative human–robot interaction tasks. Although the PAI was designed for this specific scenario, future research may examine whether similarly tailored interfaces can support positive user evaluations in other complex collaborative tasks. By integrating subjective and physiological measures, this work provides evidence-informed guidance for designing feedback strategies that support efficient, safe, and sustainable human–cobot collaboration.