Modular robots offer adaptability and reconfigurability, yet their application in aquatic environments and dynamic multi-tasking—particularly for manipulation—remains underexplored. We hypothesize that incorporating soft-bending capabilities into modular designs can significantly enhance versatility in such settings. In this work, we introduce a variable-stiffness soft modular robot that integrates rigid 3D-printed components, soft foam, a cable-driven actuation mechanism, and a propeller for aquatic propulsion. Permanent magnets enable fast, passive inter-module connections. This robot can bend, steer, and connect with others, supporting a variety of functions. It acts as a gripper to retrieve debris from water surfaces, assembles into a floating raft for drone landings, and forms a snake-like chain that transitions seamlessly between land and water. Additionally, multiple robots can collaborate in swarm-like behaviors to transport payloads. Our findings demonstrate that combining soft deformation with modularity enables a multifunctional robotic platform capable of navigating and interacting in complex, aquatic environments.
Vision–Language–Action (VLA) controllers are often built by extending vision–language models (VLMs) with action supervision, relying on multimodal backbones with large data and compute requirements. We demonstrate that a text-only large language model (LLM) can be adapted into a VLA-style controller when visual observations are rendered into a text input using an ASCII representation. This ASCII-as-vision interface enables existing training and deployment stacks for LLMs to efficiently condition on visual state, follow natural-language instructions, and produce constrained, executable actions. We fine-tune and compare multiple LLMs and VLMs across model families and scales, using both expert demonstrations from a planning-based teacher, as well as DAgger for iterative improvement. In a 2D manipulation benchmark, in both simulation and on a physical manipulator, the resulting controllers can identify task-relevant entities and plan feasible action sequences. Our results suggest that ASCII rendering can serve as a lightweight, interpretable modality bridge from images to text, complementing conventional VLA pipelines, and opening directions for VLA research with text-only backbones.
The Manipulider is a buoyancy-actuated underwater robot that enables thrusterless, glide-like locomotion and attitude-based manipulation, while providing a magnetic modular interface for rapid payload swapping (e.g., a gripper or sensors). Four syringe-based buoyancy engines distributed around the body jointly regulate net buoyancy and the center of buoyancy, allowing the vehicle to maintain large tilt angles through static force balance without continuous thrust and to avoid propeller entanglement risks. We present the mechanical and electrical design, calibration procedure, and control architecture. Experiments with a gripper attached (no external payload) show a controllable buoyancy-displacement range of 40 mL per engine (≈160 g total buoyancy authority), maximum statically stable tilts of 64.6° (single-engine) and 61.8° (dual-engine), and representative vertical and tilt-transition dynamics. We further demonstrate tilt regulation, controlled ascent/descent primitives, and a proof-of-concept gripper-based payload-transport sequence without thrusters.
World models in robot learning predict future states from visual observations and actions, enabling agents to reason about the consequences of their controls. However, many action-conditioned models are evaluated in settings where motion is dominated by immediate control, whereas aquatic surface vehicles and other real-world objects continue moving under inertia and are displaced by hidden ambient drift, such as water currents or wind. We propose FlowMo-WM, an end-to-end trainable visual world model that infers object-centric motion state and a predictive long-history context associated with hidden drift from image-action histories without direct supervision of flow fields. FlowMo-WM factorizes image-action history into a short-history latent state, trained to summarize object-centric motion, and a longer-history context, trained to summarize slowly varying exogenous influences. A zero-context residual transition separates action-conditioned base dynamics from context-dependent drift effects during latent rollout. In simulated aquatic surface-vehicle environments with diverse hidden flows, disturbances, and randomized vehicle dynamics, FlowMo-WM improves long-horizon rollout accuracy over representative action-conditioned latent world models. Prediction-time context ablations, in which the inferred context is zeroed or shuffled during rollout, show that the ambient context is important for stable prediction under hidden drift, while frozen linear probes characterize information encoded in the learned factors.
This paper presents the first steps toward a soft dolphin robot using a bio-inspired approach to mimic dolphin flexibility. The current dolphin robot uses a minimalist approach, with only two actuated cable-driven degrees of freedom actuated by a pair of motors. The actuated tail moves up and down in a swimming motion, but this first proof of concept does not permit controlled turns of the robot. While existing robotic dolphins typically use revolute joints to articulate rigid bodies, our design - which will be made opensource - incorporates a flexible tail with tunable silicone skin and actuation flexibility via a cable-driven system, which mimics muscle dynamics and design flexibility with a tunable skeleton structure. The design is also tunable since the backbone can be easily printed in various geometries. The paper provides insights into how a few such variations affect robot motion and efficiency, measured by speed and cost of transport (COT). This approach demonstrates the potential of achieving dolphin-like motion through enhanced flexibility in bio-inspired robotics.
Soft robots offer adaptable, safe interactions in complex environments, with the potential for diverse applications, such as mimicking biological motions. One major challenge is designing and prototyping soft robots with varying deformation modes, which can be a time-consuming process. To address this hurdle, reconfigurable modular robots have emerged as a solution, allowing reusable and rapid prototyping into different soft robots. However, balancing simplicity in design with extensive deformation capabilities remains an open problem. Existing reconfigurable soft robotic modules have demonstrated adaptability, often relying on modular stacking to achieve a wide range of deformations. Typically, achieving complex deformations, such as forming a continuous curve, requires multiple modules connected in a chain, as each individual module can only transition between a limited set of predefined deformation states. We introduce SoftSnap modules: snap-together components that enable the rapid assembly of a class of untethered soft robots. Each SoftSnap module integrates computation, motor-driven string actuation, and a flexible thermoplastic polyurethane (TPU)-printed deformable structure, allowing a vast deformation range through different pre-wired string configurations. These modules connect seamlessly with other SoftSnap units or customizable connectors. Demonstrated configurations include starfish-like, brittle star, snake, 3D gripper, and ring-shaped robots, showcasing ease of assembly, adaptability, and functional diversity. The scalable, reconfigurable design of SoftSnap provides researchers with an efficient and flexible platform for rapidly prototyping untethered soft robotic systems.
Traditional swarm robots rely on specific communication and planning strategies to coordinate particular tasks. Human swarms exhibit distinctive characteristics due to their capacity for language-based communication and active reasoning. This paper presents an exploratory approach to robotic swarm intelligence that leverages Large Language Models (LLMs) to emulate human-like active problem-solving behaviors. We introduce a decentralized multi-robot system where each robot initially only has its local information and does not know of the existence of the other robots. The robots utilize LLMs for reasoning and natural language for inter-robot communication, enabling them to discover peers, share information, and coordinate actions dynamically. In a series of experiments in zero-shot settings, we observed human-like social behaviors, including mutual discovery, identification, information exchange, collaboration, negotiation, and error correction. While the technical approach is straightforward, the main contribution lies in exploring the interactive societies that LLM-driven robots form - a form of robot social dynamics (or robotic social behavior analysis), examining how human-like communication protocols and collaborative structures emerge among robots through language-based interaction. In this context, we use the term "robot social dynamics" to describe the interaction patterns that arise within robot collectives, inspired by, but distinct from traditional human anthropology.
We hypothesize that online movement videos have untapped potential for teaching physical skills, and we developed a platform that automatically generates practice plans from raw TikTok dance videos. The practice plans teach one segment at a time using fading guidance and part-learning principles and are presented using a web-based interface featuring concurrent visual aids. Two user studies (n=54, n=38) were conducted. The first showed significant improvements in learning outcomes compared to standard tutorials, underscoring the importance of well-structured practice plans and offering nuanced insights into the design and effectiveness of visual aids. The second study found that segmentation and emoji-based dual-coding only benefit learning when integrated into a well-designed lesson structure. We provide a set of practical recommendations for enhancing online movement learning, focusing on the need for substantive part-learning activities and careful use of visual aids to prevent cognitive overload.
This paper explores the design and fabrication of a robotic airfoil based on a tensegrity morphing structure. We begin by introducing a family of tensegrity morphing airfoil designs that convert a continuous airfoil shape into a discrete configuration. The airfoil structure is divided into two main components: a rigid section (the D-section head) and a flexible section (the tensegrity morphing tail). To calculate the aerodynamic forces acting on the airfoil, we utilize the panel method. Following the analysis, we developed a CAD model of the airfoil, integrating all necessary electronics-such as the battery, PCB board (as shown in Fig. 5), and motors-within the rigid D-section. The flexible tail is actuated using strings, allowing for adaptive morphing. Our approach combines lightweight tensegrity structural principles with adaptive design, thorough analysis, and precise fabrication techniques. This integration aims to develop highly efficient structural systems suitable for morphing airfoil applications. Furthermore, the design methodology presented can be applied to the creation of morphing wings, robotic fingers, grippers, and various other soft robotic systems.
Modular robots are currently designed to perform a variety of tasks, primarily focusing on locomotion or manipulation through the reconfiguration of rigid modules. However, the potential to integrate multiple functions, such as making each robot deployable and capable of building lattice structures for self-construction and infrastructure creation, remains largely unexplored. To advance the field, we hypothesize that combining tensegrity principles with modular robotics can create lightweight, deformable units capable of integrating three critical functions within a single design: navigating varied terrains, manipulating arbitrary shape objects, and assembling weight-sustainable, active large infrastructures. Here, we designed untethered modular robots that are deformable, lightweight, deployable, outdoor-scale, capable of bearing loads, and capable of 3D attachment and detachment. With these characteristics, the system can form various 3D structures using different assembly methods, such as walking into position or being transported by rotorcraft. The deformability and lightweight nature of each block enable the assembled structures to dynamically change shape, providing capabilities such as added compliance during locomotion and manipulation and the ability to interact with the environment in tasks like tent and bridge assemblies. In summary, we suggest that integrating lightweight and deformable properties into modular robot design offers potential improvements in their adaptability and multi-functionality.
We present a new class of curved block-based line structures whose component chains are flexible when separated, and provably rigid when assembled together into an interlocking double chain. The joints are inspired by traditional zippers, where a binding fabric or mesh connects individual teeth. Unlike traditional zippers, the joint design produces a rigid interlock with programmable curvature. This allows fairly strong curved structures to be built out of easily stored flexible chains. In this paper, we introduce a pipeline for generating these curved structures using a novel block design template based on revolute joints. Mesh embedded in these structures maintains block spacing and assembly order. We evaluate the rigidity of the curved structures through mechanical performance testing and demonstrate several applications.
This paper presents the energy-optimal trajectories for skid-steer rovers on hard ground, without obstacles. We obtain 29 trajectory structures that are sufficient to describe minimum-energy motion, which are enumerated and described geometrically; 28 of these structures are composed of sequences of circular arcs and straight lines; there is also a special structure called whirls consisting of different circular arcs. Our analysis identifies that the turns in the trajectory structures (aside from whirls) are all circular arcs of a particular turning radius, R′, the turning radius at which the inner wheels of a skid-steer rover are not commanded to turn. This work demonstrates its paramount importance in energy-optimal path planning. There has been a lack of analytical energy-optimal trajectory generation for skid-steer rovers, and we address this problem by a novel approach. The equivalency theorem presented in this work shows that all minimum-energy solutions follow the same path irrespective of velocity constraints that may or may not be imposed. This non-intuitive result stems from the fact that with this model of the system the total energy is fully parameterized by the geometry of the path alone. With this equivalency in mind, one can choose velocity constraints to enforce constant power consumption, thus transforming the energy-optimal problem into an equivalent time-optimal problem. Pontryagin’s Minimum Principle can then be used to solve the problem. Accordingly, the extremal paths are obtained and enumerated to find the minimum-energy path. Furthermore, our experimental results by using Husky UGV provide the experimental support for the equivalency theorem.
Recent large language models (LLMs) have demonstrated promising capabilities in modeling real-world knowledge and enhancing knowledge-based generation tasks. In this paper, we further explore the potential of using LLMs to aid in the design of soft modular robots, taking into account both user instructions and physical laws, to reduce the reliance on extensive trial-and-error experiments typically needed to achieve robot designs that meet specific structural or task requirements. Specifically, we formulate the robot design process as a sequence generation task and find that LLMs are able to capture key requirements expressed in natural language and reflect them in the construction sequences of robots. To simplify, rather than conducting real-world experiments to assess design quality, we utilize a simulation tool to provide feedback to the generative model, allowing for iterative improvements without requiring extensive human annotations. Furthermore, we introduce five evaluation metrics to assess the quality of robot designs from multiple angles including task completion and adherence to instructions, supporting an automatic evaluation process. Our model performs well in evaluations for designing soft modular robots with uni- and bi-directional locomotion and stair-descending capabilities, highlighting the potential of using natural language and LLMs for robot design. However, we also observe certain limitations that suggest areas for further improvement.
We present a scalable combined localization infrastructure deployment and task planning algorithm for underwater assembly. Infrastructure is autonomously modified to suit the needs of manipulation tasks based on an uncertainty model on the infrastructure's positional accuracy. Our uncertainty model can be combined with the noise characteristics from multiple devices. For the task planning problem, we propose a layer-based clustering approach that completes the manipulation tasks one cluster at a time. We employ movable visual fiducial markers as infrastructure and an autonomous underwater vehicle (AUV) for manipulation tasks. The proposed task planning algorithm is computationally simple, and we implement it on AUV without any offline computation requirements. Combined hardware experiments and simulations over large datasets show that the proposed technique is scalable to large areas.
Developing physical motion skills is an essential and empowering aspect of the human experience, and dance learning in particular has been shown to be widely useful with benefits ranging from promoting socioemotional learning in children [3] to increasing motivation and curiosity in college students [2] and improving quality of life for patients with Parkinson's disease [5]. Yet human dance instructors are expensive and electronically available guided dance instruction often requires a subscription. Many dances are shared on video-first social media sites such as TikTok and now Instagram, however a limitation to expanding access to dance instruction is the human effort that must be put into creating high quality dance tutorials.
We present the first free-floating autonomous underwater construction system capable of using active ballasting to transport cement building blocks efficiently. It is the first free-floating autonomous construction robot to use a paired set of resources: compressed air for buoyancy and a battery for thrusters. In construction trials, our system built structures of up to 12 components and weighing up to 100Kg (75Kg in water). Our system achieves this performance by combining a novel one-degree-of-freedom manipulator, a novel two-component cement block construction system that corrects errors in placement, and a simple active ballasting system combined with compliant placement and grasp behaviors. The passive error correcting components of the system minimize the required complexity in sensing and control. We also explore the problem of buoyancy allocation for building structures at scale by defining a convex program which allocates buoyancy to minimize the predicted energy cost for transporting blocks.
In this demo we showcase an interactive application to support the learning of “TikTok dance challenge” short dance choreographies. Our system utilizes dance challenge videos as the information source, performing music analysis and pose estimation to segment the dance into learnable chunks and generate a practice plan that implements motor learning techniques such as incremental part-learning and fading guidance. These plans are presented in a web app that implements video demonstration, augmented webcam mirroring, practice recording/review functionality, and both concurrent and terminal feedback. By operating on a ubiquitous information source, generating the lessons automatically, and requiring only a web browser and webcam in the user interface, our system is a step towards significantly expanding the reach of dance choreography learning and a platform for further research into dance HCI.
In this paper, we present a soft modular block inspired by tensegrity structures that can form load-bearing structures through self-assembly. The block comprises a stellated compliant skeleton, shape memory alloy muscles, and permanent magnet connectors. We classify five deformation primitives for individual blocks: bend, compress, stretch, stand, and shrink, which can be combined across modules to reason about full-lattice deformation. Hierarchical function is abundant in nature and in human-designed systems. Using multiple self-assembled lattices, we demonstrate the formation and actuation of 3-dimensional shapes, including a load-bearing pop-up tent, a self-assembled wheel, a quadruped, a block-based robotic arm with gripper, and non-prehensile manipulation. To our knowledge, this is the first example of active deformable modules (blocks) that can reconfigure into different load-bearing structures on-demand.