
This paper proposes a novel navigation method for the autonomous weeding robot R4 (Ritsumeikan Road-weed Removal Robot), designed to reduce labor costs in roadside weeding operations. In visual navigation systems, misrecognition of road regions often generates paths that deviate from the actual road geometry. To address this, the proposed method employs road-region detection and vanishing-point detection to identify road geometry and generates trajectories using linear or curved mathematical models depending on the road conditions. This study focuses on the preliminary verification of the proposed system to explore its feasibility for roadside environments. The fundamental characteristics and potential of the navigation method are evaluated through preliminary driving experiments, providing insights for future robust autonomous navigation.
High sensitivity, large area tactile sensing, and reliable contact localization are required for safe and adaptive human-robot interaction. Robots need high sensitivity and precise localization to distinguish harmless contact from potentially dangerous interactions and operate safely near humans. Microphone based tactile sensors are studied for their wide sensing range and for capturing dynamic contact events. Earlier studies used the high sensitivity and wide dynamic range of microphones to fabricate large area tactile skins. Yet, most localization algorithms assume a planar sensor surface, which limits their application to robots based on non-planar and curved geometries. This work proposes a contact localization framework for microphone-embedded tactile sensors on non-planar surfaces. This method uses geometric mapping of the sensor surface to compute the geodesic distance matrix of the shortest surface paths between candidate contact points and embedded microphones. Contact vibrations travel along this contact layer and are measured in intensity patterns across microphones, which are compared with precomputed geodesic matrices to estimate the most probable contact location. The approach is verified with a hemispherical sensor having positive Gaussian curvature. The performance is evaluated through indentation experiments using root mean squared error analysis. By extending microphone-based tactile sensing beyond planar assumptions, this work achieved 87.20% localization accuracy on non-zero Gaussian surfaces and laid the foundation for accurate contact localization on curved robotic structures.
Autonomous ground robots operating in agricultural environments require reliable perception to navigate safely under unstructured and visually challenging conditions. Oil palm plantations present difficulties due to dense vegetation, uneven terrain, and highly variable illumination. This paper presents a deployment-oriented multi-class semantic segmentation pipeline based on YOLOv8, designed for ground-level navigation in plantation-like outdoor environments. Five semantic classes relevant to robot navigation which are oil palm trees, drains, drain edges, uphill terrain, and puddles, are defined and annotated to support terrain understanding and navigation safety. The pipeline integrates multi-class YOLOv8 segmentation with a flexible inference design and illumination-aware preprocessing, where Contrast Limited Adaptive Histogram Equalisation (CLAHE) is conditionally applied during inference under low-light conditions. Qualitative evaluation in a controlled outdoor environment resembling plantation settings demonstrates stable and meaningful segmentation outputs across varying illumination and terrain scenarios, with improved reliability under low-light conditions when conditional preprocessing is enabled. The proposed approach emphasises practical deployment considerations and provides a foundation for integrating semantic perception into autonomous ground robot navigation in plantation environments.
This paper proposes an integrated object-logging system for indoor environments that automatically detects moved objects and acquires detailed information using a mobile robot. The proposed method first detects motion through background subtraction and identifies the person and the manipulated object. A mobile robot then navigates to the target location to capture close-up images. Based on the acquired data, the system estimates the moved object, action location, human behavior (pick or place), and timestamp, and records them as structured logs. Experiments were conducted in an environment under both single-person and multi-person conditions. The results demonstrate stable object detection performance in near-range scenarios. In one-minute evaluation trials, the system achieved over 84% accuracy in recognizing moved objects and human actions. Furthermore, the complete processing pipeline achieved more than 78% accuracy when objects were placed on a desk. These findings indicate that the proposed system is effective for automatic logging of object manipulation events in indoor environments.
Accurate underwater navigation remains challenging because GPS signals are unavailable underwater, leading to unbounded drift over time. Bathymetric simultaneous localization and mapping (SLAM) addresses this challenge by using relative pose constraints from multibeam echosounder (MBES) measurements. However, the reliability of these constraints can change greatly depending on the terrain and sensor settings. Fixed or manually tuned covariance values can make the optimization unstable if a few bad constraints dominate and distort the estimated trajectory. In this paper, we propose a learning-based uncertainty modeling framework that predicts constraint reliability from registration quality indicators and assigns adaptive weights for pose-graph optimization. Our method keeps a physics-based error-state EKF front-end and standard geometric registration, using learning only for modeling uncertainty. We test the approach on real-world ocean data collected in Pohang, South Korea, using GPS as ground truth. In a 2 km mission with 78 submaps, our method achieves a drift of 0.52% distance traveled, avoiding catastrophic failures seen with naive uniform weighting (3.72% drift) and reaching near the performance of expert-tuned heuristics (0.42% drift). These results show a practical path to robust bathymetric SLAM while keeping the estimation structure interpretable.
Research in robot navigation has been accelerating rapidly with the advent of vision–language models (VLMs) that can explicitly reason about the diverse variables present in real-world driving scenarios. Despite the growing success of AI-based robot navigation, existing datasets and policies remain tightly coupled to specific hardware embodiments, which limits cross-platform generalization. As a result, platforms for which data are substantially harder to collect and far less publicly available than for wheeled robots, such as legged robots, remain chronically under-represented, aggravating the dataset imbalance across robot morphologies. To resolve this data discrepancy, we propose XNav-Pipe, a 2-staged synthetic data generation framework that expands specific-robot driving datasets to general-purpose navigation across robot types. In Stage 1, we imitate prior wheeled-robot driving behaviors using lightweight PPO agents to re-generate data on novel platforms. In Stage 2, we leverage commonsense-aware VLMs to synthesize new navigation behaviors from given instructions and trajectories. Our framework, built on NVIDIA Isaac Lab, allows scalable and diverse simulation of robot navigation across morphologies.
We study multi-robot formation reconfiguration under adjacency exclusion constraints, modeled as Independent Set Reconfiguration under token sliding - a problem that is PSPACE-complete on general graphs. We show that on Ptolemaic networks (gem-free chordal graphs), the problem becomes efficiently tractable by exploiting strong perfect elimination orderings. We present a linear-time feasibility planner (O(|V | + |E|)) based on canonical reduction, and an exact optimal planner (O(|V |4)) via dynamic programming over the clique tree. We further interpret these results through the lens of statistical mechanics, showing that Ptolemaic topologies correspond to an exactly solvable hard-core lattice gas whose ergodic structure and phase-space geodesics are fully characterized by our algorithms.
This paper presents a T-spoke wheel design for stair-climbing mobile robots and a reinforcement learning (RL)-based locomotion control framework developed in NVIDIA Isaac Lab. Unlike conventional circular wheels that fail to gain traction on stair edges, the proposed T-spoke wheel features radially protruding spoke structures with a T-shaped cross-section, enabling mechanical engagement with stair steps. A parametric wheel model is constructed within Isaac Lab, allowing systematic variation of wheel radius while maintaining fixed T-profile dimensions. A four-wheel-drive (4WD) robot equipped with independently driven T-spoke wheels is trained using Proximal Policy Optimization (PPO) to autonomously climb stairs. A parametric study across five wheel radii (65–145mm) reveals that the flange gap between adjacent spokes is the critical geometric factor for stair engagement. All T-spoke configurations with sufficient flange gap successfully climbed 7–8 steps on average, while the circular-equivalent baseline (65 mm) achieved only 1% success rate. The results confirm the geometric advantage of the T-spoke design and identify heading deviation as the primary remaining challenge for future improvement.
The diffuse and specular reflection of plant leaves, like many materials, exhibits strong angular dependencies that pose challenges for quantitative spectroscopy in agricultural applications. For quantitative use of spectral data, for example in chemometric models of chlorophyll content, such angular effects need to be understood and controlled. We present a robotic point spectroscopy system that integrates a UV–VIS spectrometer with relay optics on a 6-DoF manipulator, enabling angular sweeps of the observation direction from 30° to 150° while maintaining a 1 mm measurement spot within a 5 mm radius on the leaf surface at a working distance of 100 mm. Using this system, we acquire angle-resolved spectra from basil and coleus leaves and show that, for basil, the normalized difference vegetation index (NDVI) for a fixed leaf location can vary by up to 71% depending on observation angle, whereas coleus exhibits markedly weaker angular dependence. These results indicate that measurement geometry alone can induce significant index variations and highlight multi-angle, single-point robotic spectroscopy as a useful basis for collecting controlled datasets for future quantitative modeling.
Ensuring functional safety in autonomous driving systems necessitates the rapid diagnosis of sensor detachment and malfunctions. This paper proposes a novel sensor monitoring framework integrated with a Common Data Representation (CDR) partial parsing scheme within the ROS 2 middleware environment. The standard subscription mechanism in ROS 2 possesses structural limitations, as it requires the complete de-serialization of messages before executing callback functions—a process that induces unavoidable latency when handling high-bandwidth sensor data such as LiDAR point clouds. To circumvent this bottleneck, this study implements a technique that directly extracts the specific 12-byte segment including CDR encapsulation and timestamp. This approach enables the instantaneous determination of sensor timeouts by bypassing the entire deserialization overhead. To validate the efficacy of the proposed method, experiments were conducted in the Hwaseong Living Lab, a designated autonomous driving testbed in Korea, using a Kia Sportage research vehicle equipped with a multi-sensor suite, including LiDAR and cameras. Experimental results under high CPU-load conditions revealed that while the deserialization overhead of conventional Typed Subscriptions increased approximately 2.8-fold (from 2, 066µs to 5, 696µs at the P99 latency), the proposed CDR parsing method remained largely unaffected by computational load due to its minimal 12-byte read operation. These findings underscore the deterministic performance advantages of the proposed technique, particularly in resource-constrained environments. Furthermore, by integrating a LiDAR-camera fusion-based distance detection algorithm with a Behavior Tree, we established a decentralized architecture capable of providing a dynamic Fault Tolerant Time Interval (FTTI). This framework ensures flexible and rapid safety responses to system malfunctions in strict adherence to the ISO 26262 international standard for functional safety in road vehicles.
Evaluating quadruped locomotion under payload conditions remains challenging due to the lack of systematic and robot-agnostic assessment methodologies. While prior research has focused primarily on developing payload-aware controllers or increasing load-carrying capacity, comparatively little attention has been given to rigorously analyzing how different payload distributions affect locomotion performance. This paper presents a systematic experimental framework for evaluating quadruped robots under varied payload distributions across controlled terrains. The proposed methodology is robot-agnostic and emphasizes standardized test design, repeatable evaluation procedures, and quantitative performance metrics. We validate the framework using the Boston Dynamics Spot platform in both simulation and real-world experiments. Performance is assessed using two primary metrics: Torque Symmetry Ratio (TSR), which captures joint-level load balance, and curvature-based trajectory smoothness, which quantifies deviations from nominal motion. Our results demonstrate that payload distribution significantly influences locomotion stability, joint loading patterns, and path consistency. In both environments, rear-weighted configurations degrade performance, while forward and centrally distributed payloads maintain overall stability and smoothness in robot path following, with nuanced differences between simulation and hardware results. Additionally, we analyze the impact of a payload-aware controller on load handling performance.
Despite a widespread fascination from the public, microrobots remain confined to specialized laboratories due to the need for expensive and bespoke equipment. Here, we present a low-cost platform that makes programmable microrobots accessible to a broad audience. The system combines off-the-shelf electronics, a 3D-printed stage, and a smartphone-based optical system to both visualize and reprogram robots, with a total cost under $70. Using this platform, users can reliably power electronically-integrated microrobots, encode behaviors through light-based signals, and observe their motion across centimeter-scale arenas. We validate the system’s ease of use in a high school classroom, where first-time users successfully programmed robots and explored their behaviors. Because every element of the system, including the robots, can be mass manufactured, our platform opens new doors in education, outreach, and research, helping to move these tiny machines out of labs and into the world at large.
Robotic connector assembly is challenging due to complex geometries and uncertainties in contact interactions, and conventional approaches often rely on vision sensors or predefined assembly sequences. In this study, we propose a data-driven pose estimation model to estimate the positional error between connectors in a connector assembly task using a 6-axis force/torque (F/T) sensor without predefined contact states or assembly sequences. Impedance-based control was used to perform robotic connector assembly experiments and acquire six-axis contact reaction force data during the assembly process. A transformer-based deep learning model was trained on the measured force, torque time-series data to estimate contact pose and positional errors between male and female connectors without predefined contact states. Experimental results show that the proposed model achieves stable convergence and estimates positional errors with an average error of 0.291mm. The proposed model enables efficient reassembly without repetitive search motions and is well-suited for industrial environments where the use of vision sensors is challenging.
This paper presents a robotic gripper designed to enable the simultaneous acquisition of visual and visuotactile information within a shared field of view using a single camera. The proposed gripper employs a passive linkage mechanism that changes the mounting angle of the fingertip camera in coordination with the opening and closing motions of the gripper. Consequently, global visual information is obtained during the approach phase, whereas local visual and visuotactile information can be captured after grasping without additional actuators or sensors. Based on the geometric relationship between the passive linkage structure and camera field of view, design constraints were formulated, and the design parameters were determined numerically. A prototype gripper was fabricated, and experiments were conducted to verify the view transition and the simultaneous observation of an object-side marker and tactile-surface markers within a shared field of view. The experimental results confirmed that the proposed mechanism achieved the intended view transition and that, under a simple vertical contact condition, marker displacements could be measured repeatably.
Vine robots can navigate deeply into confined environments by extending their bodies through eversion. In many applications, a tip-mounted camera provides visual feedback for robot control. However, as the robot grows, the tip-mounted camera can undergo unintended rolling about the growth axis due to its weight, tether drag, and environmental friction. This rolling motion causes misalignment between the camera orientation and robot body, limiting reliable operation. To address this issue, we propose a fully passive tip stabilization mechanism that maintains camera orientation during growth. An internal wing-shaped structure is inserted between the robot’s three retraction channels, forming a rigid co-rotating unit with the camera assembly and mechanically coupling it to the robot body. This geometric engagement passively suppresses undesired rolling motion without additional sensors, active control, or changes to the growth mechanism. We analyze the geometric and frictional trade-offs of the wing design and experimentally evaluate multiple configurations. The results show that the proposed mechanism effectively suppresses camera roll while preserving smooth eversion. Demonstrations in straight, curved, and narrow paths further validate passive geometric coupling as a practical approach for tip orientation stabilization in vine robots.
Traditional propeller-driven underwater robots rely on continuous high-speed rotation, which may cause environmental disturbance, pose risks to aquatic organisms, and suffer from entanglement in cluttered environments. In contrast, bio-inspired undulating fins offer an environmentally friendly propulsion alternative with superior performance in confined spaces. Inspired by cuttlefish locomotion, this paper presents a bio-inspired robotic platform driven by elastic continuum fins, which exploit the tendency of elastic structures to adopt minimum-energy configurations.Thrust and locomotion experiments were conducted in an indoor pool to systematically investigate the effects of key kinematic parameters, including oscillation frequency and amplitude. The results demonstrate that undulatory frequency plays a dominant role in propulsion performance, with both thrust and swimming velocity increasing as frequency rises. In addition, an optimal combination of kinematic parameters is identified within the tested range. These findings provide practical guidelines for the design and control optimization of undulating-fin-based aquatic robots.
Deep reinforcement learning has enabled quadrupedal robots to traverse challenging terrains, yet energy efficiency remains a limiting factor for prolonged autonomous operation. Most existing frameworks rely on fixed or adaptively tuned proportional-derivative (PD) controllers that operate exclusively in the joint space. Such approaches typically lack an explicit mechanism for contact compliance, often applying excessive torque on benign terrains while providing insufficient absorption of reaction forces on irregular surfaces. To address these limitations, we propose HIP, a hybrid impedance and PD control framework that fuses joint-space PD control for trajectory tracking with task-space impedance control for contact compliance. The impedance term, mapped to joint torques via the Jacobian transpose, models compliant foot-tip behavior that absorbs impact energy during ground contact rather than resisting it through rigid control. To coordinate the two control modalities, we further introduce the attention for representation combiner (ARC) network. The ARC network employs a cross-attention mechanism between a gain actor and a joint actor, enabling control gains and desired joint positions to be generated in a coordinated manner. A state estimator augmented with a per-leg stumble estimator provides additional proprioceptive context to both actors for proactive gain adaptation. Simulation experiments across diverse terrains demonstrate that HIP achieves velocity tracking accuracy comparable to existing baselines while delivering improved energy efficiency. Torque decomposition analysis further confirms that the impedance component effectively reduces torque peaks during contact events.
Conventional robotic grippers such as suction cups and parallel-jaw claws often struggle with the diverse shapes, surface textures, and fragility of real-world objects. Inspired by the elephant’s trunk, we present the TRUNK-gripper, a soft multifunctional system that integrates gripping, twisting, wrapping and absorption to enhance grasp stability and surface adaptability. Its soft-skirt structure enables vacuum sealing on irregular surfaces, twisting motions to extract embedded items, and compliant contact for gently enclosing fragile objects. We introduce a 100-object benchmark comprising 49 adversarial shapes, 20 fragile items, and 31 household goods, and use it to evaluate the TRUNK-gripper under real-world, multi-challenge grasping scenarios. Achieving a 93.4% success rate (467/500), the TRUNK-gripper significantly outperformed a parallel-jaw gripper (80.6%, 403/500) and a standard suction gripper (1.2%, 6/500). These results underscore the potential of biologically inspired, contact-based multifunctional grasping for robust performance across a wide range of manipulation tasks.
Tactile roughness perception is commonly formulated as either classification or regression, implicitly assuming that tactile signals can be grounded to a stable absolute scale. In practice, however, tactile measurements are highly sensitive to sensing variability, making this numerical grounding unreliable under limited supervision or domain shift.In this paper, we investigate how the choice of learning objective affects robustness when absolute grounding is unstable. We compare classification, regression, and ranking-based learning on a sandpaper dataset with 5 grit levels and 9,595 samples. Under well-conditioned in-domain settings, all paradigms achieve comparable performance. However, when supervision is scarce (e.g., 5% training data) or when substantial domain shift is introduced through changes in sensor hardware and exploration protocol, their behavior diverges.Under domain shift, regression exhibits a pronounced collapse in ordinal consistency, whereas ranking-based learning preserves relative ordering more effectively, outperforming regression by 0.290 in Spearman rank correlation. These findings suggest that directly optimizing ordinal consistency yields more robust generalization when tactile roughness labels provide reliable ordering but unstable numerical scales.
Autonomous navigation toward target ultrasound views remains challenging in robotic ultrasound imaging due to the absence of global anatomical registration and explicit geometric models. This paper presents a similarity-based local iterative navigation method for coarse localization of target abdominal ultrasound views. The proposed approach is formulated as a local search problem within a bounded, image-defined reference space, where navigation decisions are derived from ultrasound image similarity without recovering absolute anatomical coordinates or relying on global path planning.A robotic ultrasound system is developed to iteratively estimate the relative displacement between the current probe position and a user-specified target image. At each navigation step, relative position estimation is performed using ultrasound image similarity with respect to a preconstructed local reference image set, and the estimated displacement is used to guide discrete planar probe motion under closed-loop control. To reduce the dependence on single-point similarity estimation, a Top-K neighborhood aggregation strategy is adopted for position estimation, in which multiple high-similarity candidates are combined through similarity-weighted averaging.Experiments conducted on an abdominal ultrasound phantom provide an initial proof-of-concept that the proposed method can guide the probe toward the target image from different initial positions within a local region. The results show that similarity-based relative displacement estimation provides directionally consistent guidance for coarse localization, while probe–surface contact conditions are regulated throughout navigation. These findings suggest the feasibility of similarity-driven local navigation for coarse ultrasound view localization within a constrained search space.