6D pose estimation is a crucial subfield of visual measurement, extensively applied in robotics, autonomous driving, and augmented reality. Existing instance-level and category-level 6D pose estimation methods typically rely on prior object models or category labels, limiting their generalization and scalability in complex and dynamic environments. To address this limitation, we propose GCM-Pose, a general 6D pose estimation framework that performs inference using object-specific reference data. Specifically, a sparse structure-from-motion (SfM) model is first constructed from a set of RGB images or a short video. Then, GCM-Pose estimates the object's 6D pose by establishing 2-D-3-D correspondences between the query image and the SfM model. This approach can be applied to novel objects without any CAD models. To handle cross-modal discrepancies and multiscale variations in the 2-D-3-D feature matching process, we introduce a cross-modal multiscale deformable attention (CMMDA) network trained with a triplet loss learning strategy. CMMDA addresses these issues by aligning feature descriptors from different modalities within a multiscale latent high-dimensional space, enabling the establishment of precise and robust 2-D-3-D correspondences. We also construct a large-scale dataset consisting of 200 objects under various real-world conditions, including lighting variation, occlusion, and transparent objects. Experimental results on LineMOD, OnePose-LowTexture, and real-world scenes demonstrate that GCM-Pose achieves state-of-the-art (SOTA) performance, highlighting its effectiveness and wide applicability in 6D pose estimation tasks.
Current methods for 3D semantic segmentation propose training models with limited annotations to address the difficulty of annotating large, irregular, and unordered 3D point cloud data. They usually focus on the 3D domain only, without leveraging the complementary nature of 2D and 3D data. Besides, some methods extend original labels or generate pseudo labels to guide the training, but they often fail to fully use these labels or address the noise within them. Meanwhile, the emergence of comprehensive and adaptable foundation models has offered effective solutions for segmenting 2D data. Leveraging this advancement, we present a novel approach that maximizes the utility of sparsely available 3D annotations by incorporating segmentation masks generated by 2D foundation models. We further propagate the 2D segmentation masks into the 3D space by establishing geometric correspondences between 3D scenes and 2D views. We extend the highly sparse annotations to encompass the areas delineated by 3D masks, thereby substantially augmenting the pool of available labels. Furthermore, we apply confidence- and uncertainty-based consistency regularization on augmentations of the 3D point cloud and select the reliable pseudo labels, which are further spread on the 3D masks to generate more labels. This innovative strategy bridges the gap between limited 3D annotations and the powerful capabilities of 2D foundation models, ultimately improving the performance of 3D weakly supervised segmentation.
Training robust vision-based robotic manipulation policies requires large-scale data, yet real-world collection is costly and unsafe. We present a high-fidelity simulation framework for data generation and policy learning, built on Isaac Lab with a closed-loop teleoperation pipeline for a UR5 manipulator and Robotiq 2F-85 gripper. The framework unifies demonstration collection, domain randomization, and policy training in a single pipeline. A central component is a semanticaware parameter identifier that helps overcome a limitation of conventional DR by preventing unconstrained sampling from decoupling physical parameters from visual appearance, which otherwise yields semantically inconsistent scenarios that introduce spurious visual-physical correlations. To resolve this, we leverage a Vision-Language Model to infer material categories from RGB observations and assign physically plausible nominal values with confidence-aware bounds for friction and density. We validate the framework by training a Vision-LanguageAction policy on the generated data. The results show that the parameter identifier attains high hit rates on material property prediction, and that the proposed DR meaningfully improves average task success over a position-only baseline.
Traditional iterative reconstruction methods are accurate but computationally expensive, limiting their use in high-throughput and real-time ptychography. Recent deep learning approaches improve speed, but often predict phase as a Euclidean scalar despite its 2π periodicity, which can introduce wrapping artifacts, discontinuities at , and a mismatch between the loss and the underlying signal geometry. We present a deep learning framework for ptychographic reconstruction that models phase on the unit circle using cosine and sine components. Phase error is optimized with a differentiable geodesic loss, which avoids branch-cut discontinuities and provides bounded gradients. The network further incorporates saturation-aware dual-gain input scaling, parallel encoder branches, and three decoders for amplitude, cosine, and sine prediction, together with a composite loss that promotes circular consistency and structural fidelity. Experiments on synthetic and experimental datasets show consistent improvements in both amplitude and phase reconstruction over existing deep learning methods. Frequency-domain analysis further shows better preservation of mid- and high-frequency phase content. The proposed method also provides substantial speedup over iterative solvers while maintaining physically consistent reconstructions.
Event-based human action recognition has gained increasing attention due to its efficiency in dynamic scenarios. Contemporary methodologies for event-based action recognition predominantly treat the problem as a one-hot classification task, which limits their ability to leverage the semantic relationships among various actions. To address this limitation, we propose a Spiking Event-Text Feature Fusion (SETFF) framework, which enhances recognition performance by integrating event and text modalities through a dual-stream architecture. SETFF leverages generative large language models to produce action descriptions, serving as semantic prompts that guide event feature learning. Specifically, a contrastive loss function is employed to align the features of both modalities, enriching the model's capacity to distinguish intricate and subtle actions. Extensive experiments on neuromorphic datasets, including PAF, DailyAction-DVS, DVS128 Gesture, Bullying10K, and UCF101-DVS, demonstrate that SETFF achieves state-of-the-art accuracy, with top-1 accuracy rates of up to 99.65% on the DailyAction-DVS dataset and 98.39% on the PAF dataset. Experimental results underscore the effectiveness of multimodal fusion in SNNs, advancing event-based action recognition while preserving the energy efficiency characteristic of SNNs.
Neural Radiance Fields (NeRF) have shown remarkable success in image novel view synthesis (NVS), inspiring extensions to LiDAR NVS. However, most methods heavily rely on accurate camera poses for scene reconstruction. The sparsity and textureless nature of LiDAR data also present distinct challenges, leading to geometric holes and discontinuous surfaces. To address these issues, we propose SG-NLF, a pose-free LiDAR NeRF framework that integrates spectral information with geometric consistency. Specifically, we design a hybrid representation based on spectral priors to reconstruct smooth geometry. For pose optimization, we construct a confidence-aware graph based on feature compatibility to achieve global alignment. In addition, an adversarial learning strategy is introduced to enforce cross-frame consistency, thereby enhancing reconstruction quality. Comprehensive experiments demonstrate the effectiveness of our framework, especially in challenging low-frequency scenarios. Compared to previous state-of-the-art methods, SG-NLF improves reconstruction quality and pose accuracy by over 35.8% and 68.8%. Our work can provide a novel perspective for LiDAR view synthesis.
Depth estimation from stereo or multi-view images is of substantial interest due to a wide range of applications. Recently, deep learning based approaches have been shown to be promising for stereo matching. However, existing stereo matching approaches are mostly data-driven, which often converge to local minima biased toward the training data. In this paper, we propose a simple but effective regularization framework to improve the training of the stereo matching networks. More specifically, we propose using low-level structure detection such as edge detection and keypoint detection as constraints for the regularization of the stereo matching network via multi-task learning. By introducing the low-level structure detection as an auxiliary task, we are able to improve the model training of stereo matching. In addition, a disparity aggregation module is also proposed to consider the association between the stereo matching and low-level structures. We apply the proposed structure regularization on four different CNN-based stereo matching algorithms. The experimental results on four public datasets, including Scene Flow, KITTI 2012, KITTI 2015 and Middlebury, verify our assumptions and show the effectiveness and generality of the proposed framework.
Navigation assistance is critical for blind and visually impaired (BVI) individuals. However, effective solutions remain underdeveloped due to limited research attention. This gap originates from two fundamental constraints: (1) GPS-based positioning lacks sufficient precision in indoor environments. (2) Navigation approaches oversimplify BVI users as sensor-driven robots. We argue that existing methods fail to address real-world navigation complexities since BVI individuals have more specific navigation needs indoors. Moreover, data-collecting devices commonly used in robotic systems, such as LiDAR and depth cameras, are not easily available to them. Consequently, there is an absence of effective and user-friendly indoor navigation systems for BVI users. To address this, we propose PhoneGuide-SLAM: a smartphone-based navigation method that leverages visual simultaneous localization and mapping (SLAM), purely adopting a stereo camera as a sensor. The navigation system can provide real-time positioning, dynamic obstacle avoidance, and incremental path planning, guiding users to reach desired destinations along straight and accessible walking paths. Extensive experiments demonstrate the effectiveness of our system, showcasing its ability to enhance the independent mobility of BVI people.
Expressing and controlling fine-grained spatial attributes of objects in large-scale models presents significant challenges, as these spatial attributes are often difficult to describe textually and exhaustive enumeration is impractical. This hinders effective alignment with user preferences regarding spatial attribute-object relationships in fine-grained synthesis tasks. To tackle this problem, we propose AttrObjDiff, a novel framework built on the pre-trained Stable Diffusion model to integrate spatial attribute maps. Firstly, AttrObjDiff constrains the denoising step using trainable cross-attention fusion modules, attribute-enhancing cross-attention and LoRAs. The fusion modules take layout features extracted by a frozen ControlNet and corresponding fine-grained attribute maps as inputs to generate joint constraint features of spatial attribute-object relationships. We leverage attribute-enhancing cross-attention within the U-Net to further refine these spatial attributes. Finally, LoRAs are employed to align with these joint constraint features of finegrained relationships. Secondly, AttrObjDiff enhances the reverse process with lightweight noise reranking models to improve spatial object-attribute alignment. The reranking models select semantic noises related to fine-grained relationships, improving synthesis quality without significantly increasing computational costs. Experimental results demonstrate that our method can generate high-quality images guided by fine-grained spatial object-attribute relationships, improving synthesis controllability and semantic consistency.
Adapting frozen vision-language models (VLMs) to downstream tasks via prompt tuning has emerged as a parameter-efficient alternative to full fine-tuning. However, existing prompt tuning methods still lack explicit control over how strongly prompts should adapt to different samples, despite the significant variation in sample difficulty and prediction reliability-a critical limitation for trustworthy deployment in complex real-world scenarios such as autonomous perception and intelligent control systems. We propose UGPT (Uncertainty-Guided Dynamic Prompt Tuning), a framework that explicitly leverages visual prediction uncertainty as a sample-level control signal for prompt modulation. UGPT computes normalized entropy from a frozen zero-shot CLIP branch, maps it through a lightweight Uncertainty Injection Network (UIN) into prompt-space perturbation vectors, and applies broadcast injection to generate sample-adaptive dynamic prompts. Under the 16-shot few-shot setting, UGPT uses only 41 K trainable parameters in total, achieves 81.58% Top1 accuracy on ImageNet-1K, and obtains consistent improvements across four out-of-distribution benchmarks (Avg. OOD: 62.0%). Ablation studies and interpretability analyses confirm that the perturbation magnitude increases monotonically with uncertainty, validating the principled “harder samples receive stronger modulation“ mechanism.
Interplanetary scintillation (IPS) phenomenon behaves as twinkling of a compact radio source due to scattering in the solar wind. Here we report the first identification of IPS signal by a new parabolic dish antenna with its diameter of 30 m, observed on 2024 March 1. The radio source of 3C446 at solar elongation of 6.4° was targeted during its 5 minute transit across our antenna beam at ArQi (115.13°E, 44.73°N). Our receiver is characterized as simultaneous measurements of dual polarizations from 3C446 emission, integrated over a 40 MHz bandwidth around its central frequency of 1.4 GHz. The horizontally and vertically polarized components, though independently and simultaneously measured, show a strong correlation in terms of both time-series and power spectral analyses. Using a spectra-fitting method on basis of the classic IPS theory, the solar wind speed in our observation case is inferred to be 283 km s ^−1 . Such a solar wind speed inferred at the piercing point of 24 solar radii along our IPS ray-path reaffirms the interplanetary diagnostic capability of ground-based IPS technique. Our first-light IPS observation at ArQi represents a significant advance in the ongoing commissioning phase of our IPS-dedicated telescope facility within the Chinese Meridian Project–Phase II.
Industrial assembly documents—operator Standard Operating Procedures (SOPs) from Manufacturing Execution Systems and material specifications from Product Lifecycle Management / Enterprise Resource Planning (PLM/ERP) systems—contain the key parameters needed for downstream execution, yet no existing method recovers a process-oriented core knowledge graph (hereafter core execution graph) from such heterogeneous sources while systematically abstaining on non-target inputs. We present POKG, a framework that reformulates document-to-KG construction under a dual-constraint mechanism: core execution graph recovery on target press-fit assembly documents and conservative abstention (returning an empty graph) on non-target, cross-domain, or evidence-sparse documents. We construct a 20-bundle evaluation corpus whose target bundles mirror the document packaging conventions of real industrial production sites and whose non-target bundles include cross-domain, near-miss, and corrupted inputs. The pipeline integrates process-family routing, multi-source extraction with type-specific prompts, evidence-gated candidate adjudication, and topology guards. A dual-agent review-and-repair extension addresses the omission-dominant error profile through structure–evidence disentangled auditing with bounded recovery. On the 20-bundle corpus (10 target, 10 non-target), the base pipeline achieves target node F1 = 0.95 and target edge F1 = 0.60; the dual-agent extension raises edge F1 to 0.83 (mean ± 0.03 over 5 runs); 9 of 10 non-target bundles correctly produce empty graphs (degradation accuracy = 0.90).
After an older adult misses a medication dose, a software agent can send another reminder and an embodied agent can bring the medication. Yet neither explains whether the person forgot, is confused, has side effects, or deliberately refused, nor what support is appropriate. This reveals a structural gap in Agentic AI: Digital Agents primarily transform software states, while Embodied Agents transform physical states; neither makes a person's evolving state and agency the primary object of modeling, intervention, and evaluation. We introduce Combodied Agents, a human-centered paradigm that perceives, models, predicts, and supports individual human-state trajectories over time, using software tools, sensors, wearables, robots, and human services as action channels rather than end goals. We unify fragmented capabilities across personal assistants, health agents, AI companions, and adaptive human–AI systems into a closed loop: event-based multimodal perception reconstructs meaningful personal events; longitudinal, correctable memory provides temporal context; Personal World Models estimate future personal states and outcomes under alternative decisions and interventions; and an admissible intervention policy selects proportionate support under consent, uncertainty, safety, reversibility, and user control. Feedback from the person and environment updates the loop. Rather than requiring an exhaustive Human Digital Twin, the framework uses purpose-bounded, uncertainty-aware, user-correctable representations. We organize the design space by human-state targets, relational contexts, and agent roles, and propose scenario-centered evaluation, agency-preservation metrics, benchmark requirements, edge-native personal models, and governance directions. Combodied Agents shift Agentic AI from external task completion toward sustained human benefit.
Recent advances in vision-language-action (VLA) models enable end-to-end robotic manipulation by learning from demonstrations. However, existing approaches have predominantly focused on scaling model architectures or expanding dataset size, often overlooking the interaction between training data distribution and inference strategies. To address this gap, we jointly study data distribution and inference strategy for robotic grasping, emphasizing coordinated distribution design and adaptive inference. Our approach integrates a compositional data collection strategy that decomposes the task state space into orientation and spatial factors, efficiently expanding coverage with a limited number of demonstrations. Additionally, we propose an adaptive inference mechanism that dynamically adjusts the execution horizon during critical task phases, thereby enhancing task performance. Embeddingspace analysis suggests that performance saturates once the relevant state-space dimensions are sufficiently covered. Real-robot experiments validate our approach, demonstrating a 98% grasp success rate with only 81 demonstrations. Furthermore, applying adaptive inference to a moderately covered dataset improves the success rate from 90% to 97%, suggesting that inference-level refinement can complement data coverage, though it cannot fully compensate for insufficient distribution design.
Recent advances in speech-driven facial animation have attracted significant interest across computer graphics, human-computer interaction systems, and immersive virtual reality applications. However, existing methods remain constrained by dependencies on specific reference videos or proprietary face mesh structures, limiting their applicability across diverse production pipelines and reducing compatibility with industry-standard animation workflows. To overcome these fundamental limitations in generalization and deployment flexibility, we propose Speech2Blend-an end-to-end hybrid convolutional-recurrent network that directly learns nonlinear speech-to-blendshape parameter mappings. This novel approach enables markerless speech-driven facial animation generation without restrictive inputs like video references or specialized facial rigs. Trained on the largest available digital human dataset (BEAT) and rigorously evaluated using three benchmark datasets with photorealistic visualization tools, Speech2Blend achieves state-of-the-art performance. It delivers superior audio-visual synchronization through learned temporal dynamics and reduces lip vertex error by 30% compared to existing baseline methods. These advances significantly lower production costs for virtual human speech animation while enabling cross-platform compatibility with common game engines and animation software.
This article investigates the robust nonfragile leaderless consensus control issues of nonlinear multiagent systems (MASs) in the presence of controller gain perturbations, external interferences, and switching directed networks. A novel distributed nonfragile consensus controller is first devised. Subsequently, on the basis of the property that an MAS directed network’s Laplacian matrix can be broken down into the product of two particular matrices, the conversion from the consensus control issue to the asymptotic stability control issue is achieved via two variable substitutions related to the above property. Additionally, a sufficient condition, which can guarantee the MASs’ asymptotic stability, is proposed and proved by Lyapunov stability theory and algebraic graph theory. Finally, the validity of the devised method is demonstrated by a simulation example.
Neural ordinary differential equations (ODEs) provide a continuous-depth framework for neural networks, offering advantages such as structural continuity and memory efficiency. Despite these appealing properties, neural ODEs remain vulnerable to adversarial perturbations, limiting their use in safety-critical applications. Existing robustness strategies typically rely on local sensitivity reduction or flow regularization, which provide only limited guarantees of long-term stability. To address this limitation, we propose an attractor dynamics network (ADN), a lightweight module that leverages attractor mechanisms to enhance classification robustness. ADN assigns each ODE output to a target attractor by combining learnable attractor anchors, allowing models to learn stable latent representations that are resilient to input noise. We design a training objective that enforces finite-time convergence toward attractors while jointly optimizing the original task loss. Extensive experiments on image classification benchmarks demonstrate that incorporating attractor dynamics into neural ODEs via ADN not only reduces sensitivity to input perturbations but also promotes a more structured distribution of features. Moreover, ADN can be readily integrated into neural ODE variants, offering an efficient way to improve classification robustness.
Multispectral object detection aims at detecting targets with multiple spectral modalities, i.e. RGB and infrared images, to improve its reliability and robustness in harsh environments. Existing methods overlook the decoupled modeling of cross-modal fusion, making it difficult to select task-specific information from the multiple modals, resulting in false negatives. To this end, we propose a two-stage modal feature enhancement method that decouples cross-modal fusion into two parts, i.e. intra-modal and inter-modal feature interaction, to enhance modal features for multispectral object detection. First, for intra-modal feature interaction, we design a low-rank feature enhancement module that enhances the difference between target and background by suppressing irrelevant information in low-rank space. Second, for inter-modal feature interaction, we introduce a query-guided cross-modal feature enhancement module, which leverages modal-specific queries to retrieve modality-relevant information from intermediate features. Experimental results demonstrate that our method outperforms existing methods on benchmark datasets.
Solar flares, coronal mass ejections (CMEs) and enegertic particles, etc., are the driving sources that may cause catastrophic space weathers. It is desirable to obtain information of solar eruptions like flares and CMEs, etc., propagating from the Sun to the near-Earth space. The Chinese Meridian Project includes the interplanetary scintillation (IPS) telescopes to investigate the structures and properties of the solar wind throughout the inner heliosphere. From IPS observations one can obtain disturbance information on CME speeds and thus on CME arrival times well off the Sun-Earth line. When combined with modeling techniques and/or in situ data, other parameters such as CME masses can also be obtained, along with CME propagation directions and arrival times. Therefore, a radio telescope array with three 140 m 40 m parabolic cylinder antennas at the main site and two 30 m antennas at two subsites about 200 km away from each other, featuring a multi-site array with the highest sensitivity dedicated to IPS observations in the world, has been supported as a major facility of the Chinese Meridian Project. The detailed description of the final optimized design and implementation of this IPS radio telescope array is introduced. The antennas and array configuration, the analog and digital receiving systems for the main site and subsites, the calibration of the IPS telescope array and data processing are described. Finally the overall performance of the IPS telescope array is provided. The detailed information on the IPS radio telescope array will facilitate the use of its data serving space weather research and applications.
Diffusion-based methods have achieved remarkable success in photorealistic image generation, leveraging iterative denoising steps to improve image quality. However, multi-step denoising often suffers from error accumulation—similar to exposure bias in autoregressive models—due to suboptimal noise estimation, which can lead to degraded semantic alignment and image fidelity. To tackle the challenge of suboptimal inner latent representations in generation and improve the inner latent, this paper introduces a novel method NoisePO, an efficient semantic noise preference optimization framework. NoisePO employs a semantic noise preference optimization generative adversarial network (NPO-GAN) and noise ranking methods to search for semantically relevant noises based on textual conditions, thus eliminating undesired semantic features while emphasizing the necessary semantic ones. Specifically, NoisePO utilizes a light NPO-GAN to generate semantic noises that encourage the latent at the previous step to incorporate more semantic information from the caption. Then, light ranking models are employed to filter out low-quality noises and select the best noise. Experimental results demonstrate that NoisePO consistently outperforms the baselines across widely used frameworks, achieving notable improvements in image quality, semantic consistency, and user-specific alignment as measured by IS, FID, CLIP, and other metrics. These results indicate that NoisePO effectively enhances synthesis quality and strengthens text-image alignment.