3D Gaussian Splatting (3DGS) has significantly advanced real-time novel view synthesis by representing scenes as dense collections of anisotropic 3D Gaussian primitives. However, the irregular spatial distribution of Gaussians often leads to poor GPU utilization, as warp divergence and redundant computation degrade rendering performance. To address this, we present Local-GS, a warp-coherent rendering paradigm that, organizes Gaussian primitives with respect to SIMT (Single Instruction, Multiple Threads) execution boundaries rather than scene geometry. Specifically, we propose three warp-coherent stages: a hoisting stage that precomputes shared parameters at tile level, a culling stage that discards warps with no contribution, and a blending stage that replaces per-pixel branching with a uniform instruction stream. Across extensive benchmarks on multiple datasets, Local-GS improves efficiency without compromising quality. As a plug-and-play optimization, it provides additional performance gains to all tested baselines, culminating in a $7.76\times$ speedup on Deep Blending scenes.
Standard on-policy reinforcement learning relies on heuristic clipping to enforce trust regions, but this mechanism imposes a severe cost by indiscriminately truncating high-return yet high-divergence updates. We demonstrate that explicitly constraining the policy ratio variance provides a principled local approximation to trust-region constraints, eliminating the need for binary hard clipping. By acting as a distributional “soft brake”, this approach preserves critical gradient signals from novel discoveries while naturally down-weighting and enabling the reuse of stale, off-policy data. We introduce R^2 VPO (Ratio-Variance Regularized Policy Optimization), which implements this constraint via a primal-dual optimization framework. Extensive evaluations across 7 LLM scales, spanning both fast and slow reasoning paradigms, and 10 robotic control tasks demonstrate the generality of the proposed approach. R^2VPO achieves substantial performance gains on mathematical reasoning benchmarks, with particularly pronounced improvements on smaller models, while significantly improving sample efficiency. Furthermore, it consistently outperforms PPO baselines in continuous control domains, particularly in sparse-reward and dynamic environments. Together, these findings establish ratio-variance regularization as a principled foundation for stable and data-efficient policy optimization.
Reliable geometric perception and contact localization are essential for tactile robots operating in unstructured environments. Conventional point cloud filtering and contact planning are often developed independent of tactile hardware, which leads to resolution mismatch and unstable contact. This article presents a sensor-driven framework that embeds the specifications of capacitive tactile sensors into both point cloud filtering and geometric contact localization. We develop a tactile perception-guided filtering method that uses a two-scale design to suppress depth artifacts while selectively preserving geometry that is resolvable and relevant to contact formation. This design reduces the artifact metric theta(95) by 58% on average compared with representative baselines. For contact relevant structures, the height loss rate is kept below 5%, while 85% of contact irrelevant structures are removed, which reduces subsequent computation without discarding actionable geometry. On the filtered point cloud, we propose a geometry-aware contact localization algorithm that incorporates sensor geometry and identifies regions consistent with the sensor structure by using principal curvature. Repeated robot trials on planar, limited area, and high curvature contact tasks show a contact success rate of 92% and a mean contact coverage of 90.5%, outperforming representative baselines. System ablation further confirms that TPF and GCL provide complementary gains, and the full pipeline achieves the best end-to-end performance. Runtime evaluation reports an average per-frame latency of 0.19 s for filtering and 0.72 s for localization, supporting online contact planning
A foundational humanoid motion tracker is expected to be able to track diverse, highly dynamic, and contact-rich motions. More importantly, it needs to operate stably in real-world scenarios against various dynamics disturbances including terrains, external forces, and physical property changes for general practical usage. To achieve this goal, we propose Any2Track (Track Any motions under Any disturbances), a two-stage RL framework to track various motions under multiple disturbances in the real world. Any2Track reformulates dynamics adaptivity as an additional capability on top of basic action execution and consists of two key components: AnyTracker and AnyAdapter. AnyTracker is a general motion tracker with a series of careful designs to track various motions within a single policy. AnyAdapter is a history-informed adaptation module that endows the tracker with the online dynamics adaptivity to overcome sim2real gap and multiple real-world disturbances. We deploy Any2Track on Unitree G1 hardware and achieve successful sim2real transfer in a zero-shot manner. Any2Track performs remarkably well in tracking various motions under multiple real-world disturbances.
Intention recognition of space non-cooperative targets serves as the key for orbital game with proximity satellites. Different from traditional pose-based methods prevalent in previous works, this paper establishes the first end-to-end framework for implicit orbital motion intention recognition of space non-cooperative targets directly from original image observations captured by spaceborne camera on the spacecraft. Specifically, we first propose an end-to-end neural network for implicit intention recognition consisting of a spatial-domain image feature module and a time-domain sequence feature module, and also compare the performance of various network structures for the latter, including long short-term memory, gated recurrent unit and their bidirectional variants with optional self-attention mechanism. A hierarchical intention set is designed to characterize the mapping from image sequence data to target orbital motion intention, which is delineated as an upper-level mission type along with the corresponding lower-level motion mode. To fill the blank of space-based image datasets for intention recognition, we generate a Space target Intention Recognition Image (SIRI) dataset based on a high-fidelity simulation platform developed in the ROS/Gazebo environment. Experimental results identify the CNN-BiGRU as the optimal network structure for our proposed model reaching the accuracy of 99.92% at a moderate computation cost, and also demonstrate the generalization of our method to orbital motion parameters and the comparable recognition performance to traditional pose-based methods. Our work eliminates the requirement for intermediate pose estimation results and facilitates the design of a more compact intention recognition module in future space situational awareness systems.
Few-shot learning seeks to recognize novel classes from limited examples. Model-agnostic meta-learning (MAML), known for its simplicity and flexibility, learns an effective initialization for fast adaptation in data-scarce settings. However, MAML-based methods face challenges when there is a significant distributional shift between training and testing tasks, leading to inefficient learning and poor generalization across domains. In this work, we identify the core issues: inflexible weight update rules and limited adaptive learning capabilities. Instead of focusing solely on better initialization, we aim to enhance the adaptation process. Consequently, we propose a novel Layer-Adaptive Proportional-Integral-Derivative (LA-PID) optimizer integrated into a meta-learning framework. This design incorporates classical control theory, utilizing PID control to dynamically adjust task-specific gains at each network layer. Additionally, the theoretical conditions for optimal hyperparameter initialization and global model convergence are addressed from both control and optimization perspectives. Experiments on benchmark datasets show that LA-PID achieves state-of-the-art performance in few-shot classification, cross-domain, and regression tasks, while requiring fewer training steps.
Musculoskeletal robots offer intrinsic compliance and flexibility, providing a promising paradigm for versatile locomotion. However, existing research typically relies on models with fixed muscle physiological parameters. This static physical setting fails to accommodate the diverse dynamic demands of complex tasks, inherently limiting the robot's performance upper bound. In this work, we focus on the morphology and control co-design of musculoskeletal systems. Unlike previous studies that optimize single physiological attributes such as stiffness, we introduce a Complete Musculoskeletal Morphological Evolution Space that simultaneously evolves muscle strength, velocity, and stiffness. To overcome the exponential expansion of the exploration space caused by this comprehensive evolution, we propose Spectral Design Evolution (SDE), a high-efficiency co-optimization framework. By integrating a bilateral symmetry prior with Principal Component Analysis (PCA), SDE projects complex muscle parameters onto a low-dimensional spectral manifold, enabling efficient morphological exploration. Evaluated on the MyoSuite framework across four tasks (Walk, Stair, Hilly, and Rough terrains), our method demonstrates superior learning efficiency and locomotion stability compared to fixed-morphology and standard evolutionary baselines.
In intelligent biochemical laboratories, accurate non-invasive identification of liquid properties is essential for autonomous robotic operation. To overcome the limitations of vision-based methods caused by transparent liquids, opaque containers, and physical occlusions, this paper proposes a liquid property measurement and recognition method based on capacitive tactile sensing. To validate the method, a series capacitance model of the sensor–container–liquid system is established, revealing the mapping mechanism between permittivity and capacitance response, which is verified through COMSOL multiphysics simulations. A commercial capacitive tactile sensor is integrated into a robotic gripper to construct a non-contact liquid measurement system, and the correlation between output signals and measurement parameters is systematically analyzed. To handle high-dimensional non-stationary signals generated during measurement, the STNet spatiotemporal hybrid architecture is proposed, enabling high-precision recognition of complex liquid media. Experimental results show that the proposed measurement scheme exhibits excellent repeatability and robustness in real experimental scenarios. This study provides a safe and reliable measurement solution for online monitoring of hazardous or transparent reagents in automated laboratory environments.
Classifiers fusion can be seen as a kind of multisource information fusion (MSIF), and classifiers fusion based on Dempster-Shafer (DS) evidence theory is an effective approach to improve the accuracy of classification tasks. However, different classifiers usually exhibit varying performances, making it challenging to achieve enhanced classification accuracy through direct fusion. Simultaneously, when the frame of discernment (FoD) of the target class expands, the number of focal elements involved in the fusion increases, resulting in a rapid growth in computational complexity. To enhance the classification performance while reducing the time cost of fusion, a novel weighted fusion of classifiers method based on approximate reasoning and reliability evaluation (WFC-AR-RE) is proposed in this article. Specifically, at first, the key focal elements are determined based on the outputs of classifiers, and an approximate basic belief assignment (BBA) is generated. Subsequently, the validation set is utilized to evaluate the performance of each classifier, thus obtaining the self-reliability of each BBA. Afterward, a novel divergence measure is introduced to quantify the discrepancy between BBAs, determining the relative reliability of each BBA. Finally, the fusion weight of each BBA is derived from its self-reliability and relative reliability, and Dempster's rule is applied to combine the weighted BBA. The proposed WFC-AR-RE algorithm is applied to the MSIF system, and its effectiveness is demonstrated on 12 public datasets.
Here, we report a high-frequency, interaction-induced pneumatic oscillator (HIPO) that generates self-sustained oscillations through internal mechanical-fluidic coupling in soft materials. Under constant airflow, reversible contact switching between a holed membrane and a flexible cover enables periodic oscillations up to 97 Hz without rigid valves or electronic control. The oscillator directly converts pneumatic energy into mechanical motion and drives diverse bioinspired soft robots, including crawling, flapping, and swimming systems. By integrating oscillation generation, actuation, and environmental interaction within a single compliant structure, HIPO establishes a compact and scalable approach to embodied actuation in soft robotics.
Operating household appliances by reading and understanding user manuals remains a fundamental and challenging problem in robotics. Recent works leverage large language models (LLMs) and vision-language models (VLMs) to interpret manuals, improving appliance operation success. However, these approaches fail when manuals are unavailable or incomplete. In this paper, we introduce an autonomous assistant for robotic appliance operation, built upon an LLMs/VLMs-powered multi-agent collaborative framework. Our system can read, comprehend, and summarize manuals, autonomously infer operational logic, and execute actions on appliances with a robotic arm. Importantly, for unseen appliances without manuals, it can acquire operational knowledge from generalized manuals and on-demand web search. Extensive evaluations on over one thousand tasks show that our framework substantially outperforms baselines and achieves robust performance in simulation and real-world experiments.
In this paper, we propose a real-time control method for unmanned aerial vehicles to execute missions involving time-sensitive targets (TSTs) in dynamic environments with obstacles. The nonlinear dynamics inherent to unmanned aerial vehicles (UAVs) often limit the effectiveness of traditional control methods that rely on simplified models, leading to infeasibility in practical applications. In TST missions, stringent temporal constraints further restrict maneuverability, while the presence of obstacles narrows the feasible control space and increases the difficulty of achieving timely and safe task completion. To address these challenges, We integrate Signal Temporal Logic and Linear Temporal Logic to formalize complex task requirements, and utilize control barrier functions to enforce real-time safety constraints. These components are unified with a quadratic programming framework that generates continuous control inputs for nonlinear systems. Simulation results validate the effectiveness of the proposed method in satisfying time constraints, ensuring obstacle avoidance, and enabling adaptive and feasible control in dynamic scenarios.
Abstract Embodied agents must continuously adapt to the physical world using interaction data collected across varying timescales, controllers, and environmental conditions. However, standard reinforcement learning assumes stationary dynamics and on-policy data, a premise often violated in reality where physical parameters drift and historical data becomes heterogeneous. The central challenge lies in the compound distribution shift: replayed transitions follow an occupancy distribution that diverges fundamentally from the current physical reality, leading to biased value estimation and catastrophic learning collapse. In this work, we propose Transition Occupancy Matching as a unifying principle to resolve policy and dynamics shifts within a single mathematical framework. We introduce Occupancy-Matching Policy Optimization (OMPO), a novel algorithm that optimizes a surrogate objective explicitly correcting for transition discrepancies. By leveraging a dual reformulation with a sign-free logarithmic link, OMPO transforms the intractable matching problem into a stable min-max optimization, amenable to arbitrary reward structures. Crucially, OMPO integrates a distributional critic and a multimodal encoder with a small-scale local buffer, allowing the agent to anchor massive historical data to the immediate physical context for rapid adaptation. Extensive evaluations across diverse benchmarks—including MuJoCo locomotion, DM-Control, Meta-World, and high-fidelity Panda robot manipulation—demonstrate that OMPO consistently outperforms specialized baselines in stationary, domain-shifting, and non-stationary settings. By unifying distribution correction across policy and dynamics shifts, OMPO addresses a fundamental bottleneck in transfer learning, providing a robust algorithmic framework for continual adaptation in changing physical conditions.
In the automated co-design of soft robots, precisely adapting the material stiffness field to task environments is crucial for unlocking their full physical potential. However, mainstream platforms (e.g., EvoGym) strictly discretize the material dimension, artificially restricting the design space and performance of soft robots. To address this, we propose EvoGymCM (EvoGym with Continuous Materials), a benchmark suite formally establishing continuous material stiffness as a first-class design variable alongside morphology and control. Aligning with real-world material mechanisms, EvoGymCM introduces two settings: (i) EvoGymCM-R (Reactive), motivated by programmable materials with dynamically tunable stiffness; and (ii) EvoGymCM-I (Invariant), motivated by traditional materials with invariant stiffness fields. To tackle the resulting high-dimensional coupling, we formulate two Morphology-Material-Control co-design paradigms: (i) Reactive-Material Co-Design, which learns real-time stiffness tuning policies to guide programmable materials; and (ii) Invariant-Material Co-Design, which jointly optimizes morphology and fixed material fields to guide traditional material fabrication. Systematic experiments across diverse tasks demonstrate that continuous material optimization boosts performance and unlocks synergy across morphology, material, and control.
Musculoskeletal robots provide superior advantages in flexibility and dexterity, positioning them as a promising frontier towards embodied intelligence. However, current research is largely confined to relative simple tasks, restricting the exploration of their full potential in multi-segment coordination. Furthermore, efficient learning remains a challenge, primarily due to the high-dimensional action space and inherent overactuated structures. To address these challenges, we propose Diff-Muscle, a musculoskeletal robot control algorithm that leverages differential flatness to reformulate policy learning from the redundant muscle-activation space into a significantly lower-dimensional joint space. Furthermore, we utilize the highly dynamic robotic table tennis task to evaluate our algorithm. Specifically, we propose a hierarchical reinforcement learning framework that integrates a Kinematics-based Muscle Actuation Controller (K-MAC) with high-level trajectory planning, enabling a musculoskeletal robot to perform dexterous and precise rallies. Experimental results demonstrate that Diff-Muscle significantly outperforms state-of-the-art baselines in success rates while maintaining minimal muscle activation. Notably, the proposed framework successfully enables the musculoskeletal robots to achieve continuous rallies in a challenging dual-robot setting.
Autonomous vehicles rely on continuous environmental perception to assess obstacle distribution to ensure safe driving. However, current perception technologies face substantial challenges, particularly under adverse weather conditions and when encountering long-tail scenarios. With the advent of transformer-based attention mechanisms, large language models (LLMs), exemplified by GPT, have exhibited emergent intelligence, offering new possibilities for achieving high-performance perception. This technological advancement has led to the development of multimodal LLMs (MLLMs), which incorporate multimodal encoders. These models enable a single LLM to process multisource data while performing advanced understanding and reasoning tasks, enhancing complex environmental perception capabilities. Despite significant progress in MLLMs, there remains a notable gap in systematic research on their optimal application to perception tasks. Therefore, this article presents a comprehensive survey of recent advancements in MLLM-based perception. First, we introduce the mainstream vision-language perception tasks, widely adopted evaluation metrics, and existing language-enhanced autonomous driving datasets. Next, we outline the general architectural design principles of MLLMs. Subsequently, we provide a taxonomy and in-depth analysis of current MLLMs, focusing on three dimensions: input modality, alignment technique, and scene representation, elucidating their underlying implementation paradigms. Finally, we summarize the key challenges and emerging research directions in MLLM-driven perception. This survey aims to facilitate further progress in MLLMs by synthesizing these insights.
In this article, we propose a multimodal robotic agent that addresses the limitations of passive perception and fixed-model deployment through a flexible large-model architecture. This architecture enables dynamic selection of multimodal large language models based on computational constraints. Building on this foundation, we introduce a logic-guided active perception strategy that decides which skills (e.g., knock and weigh) to employ based on intermediate reasoning, rather than exhaustively executing all possible actions. Our work focuses on cohesive skill integration within a unified control loop, optimizing both perception and action. This allows the agent to strategically probe objects’ visual, auditory, tactile, and weight attributes for accurate material inference and robust task completion. Extensive evaluations in the simulation environment highlight the efficiency and adaptability of our method, especially in active perception and latent information inference for robotic systems.
In this work, we present CollabVLA, a self-reflective vision-language-action framework that transforms a standard visuomotor policy into a collaborative assistant. CollabVLA tackles key limitations of prior VLAs, including domain overfitting, non-interpretable reasoning, and the high latency of auxiliary world models, by integrating VLM-based reflective reasoning with diffusion-based action generation under a mixture-of-experts design. Through a two-stage training recipe of action grounding and reflection tuning, it supports explicit self-reflection and proactively solicits human guidance when confronted with uncertainty or repeated failure. It cuts normalized Time by ?2× and Dream counts by ?4× vs. explicit-reasoning agents, achieving higher success rates, improved interpretability, and balanced low latency compared with existing methods. This work takes a pioneering step toward shifting VLAs from opaque controllers to genuinely assistive agents capable of reasoning, acting, and collaborating with humans.
This study introduces an innovative framework for anomaly detection and handling in multi-robot collaboration within scientific laboratories, addressing challenges posed by complex procedures and precision equipment. We propose a hierarchical approach utilizing large language models (LLMs) for task planning, action decomposition, and dynamic adjustments. A key innovation is the development of a cross-behavior tree (BT) extension algorithm triggered by anomaly nodes, enabling efficient localization and management of anomalies. Additionally, we present a spatiotemporal coupling strategy to suppress anomaly propagation, minimizing the need for task re-planning. To validate our approach, we constructed a high-fidelity simulation environment and conducted extensive experiments, demonstrating the algorithm’s effectiveness across various scenarios. Experimental results confirm that our method reduces BT node expansion by up to 17.7% and task duration by up to 18.9%, with strong stability and operational efficiency even under high anomaly probabilities. Real-world testing further affirms the system’s flexibility and fault tolerance, underscoring significant advancements in laboratory automation. This work lays the groundwork for pioneering automated anomaly handling mechanisms in future intelligent laboratories.
Zengqi Sun (孙增圻)合作论文数Department of Computer Science and Technology, Tsinghua University25