In the deformation measurement of high-temperature structures, image degradation caused by thermal radiation and random errors introduced by heat haze restrict the accuracy and effectiveness of deformation measurement. The purpose of this study is to suppress thermal radiation and heat haze using fusion-restoration image processing methods, thereby improving the accuracy and effectiveness of Digital Image Correlation (DIC) in the measurement of high-temperature deformation. For image degradation caused by thermal radiation, based on the image layered representation, the image is decomposed into positive and negative channels for parallel processing, and then optimized for quality by multi-exposure image fusion. To counteract the high-frequency, random errors introduced by heat haze, we adopt the Feature Similarity Index (FSIM) as the objective function to guide the iterative optimization of model parameters, and the grayscale average algorithm is applied to equalize anomalous gray values, thereby reducing measurement error. The proposed multi-exposure image fusion algorithm effectively suppresses image degradation caused by complex illumination conditions, boosting the effective computation area from 26 ε _xx is reduced by 85.3 ε _yy and γ _xy are reduced by 36.0
Accurate measurement of shock wave motion parameters with high spatiotemporal resolution is essential for applications such as power field testing and damage assessment. However, significant challenges are posed by the fast, uneven propagation of shock waves and unstable testing conditions. To address these challenges, a novel framework is proposed that utilizes multiple event cameras to estimate the asymmetry of shock waves, leveraging its high-speed and high-dynamic range capabilities. Initially, a polar coordinate system is established, which encodes events to reveal shock wave propagation patterns, with adaptive region-of-interest (ROI) extraction through event offset calculations. Subsequently, shock wave front events are extracted using iterative slope analysis, exploiting the continuity of velocity changes. Finally, the geometric model of events and shock wave motion parameters is derived according to event-based optical imaging model, along with the 3D reconstruction model. Through the above process, multi-angle shock wave measurement, motion field reconstruction, and explosive equivalence inversion are achieved. The results of the speed measurement are compared with those of the pressure sensors and the empirical formula, revealing a maximum error of 5.20% and a minimum error of 0.06%. The experimental results demonstrate that our method achieves high-precision measurement of the shock wave motion field with both high spatial and temporal resolution, representing significant progress.
High dynamic range (HDR) imaging under extreme illumination remains challenging for conventional cameras due to overexposure. Event cameras provide microsecond temporal resolution and high dynamic range, while spatially varying exposure (SVE) sensors offer single-shot radiometric diversity.We present a hardware--algorithm co-designed HDR imaging system that tightly integrates an SVE micro-attenuation camera with an event sensor in an asymmetric dual-modality configuration. To handle non-coaxial geometry and heterogeneous optics, we develop a two-stage cross-modal alignment framework that combines feature-guided coarse homography estimation with a multi-scale refinement module based on spatial pooling and frequency-domain filtering. On top of aligned representations, we develop a cross-modal HDR reconstruction network with convolutional fusion, mutual-information regularization, and a learnable fusion loss that adaptively balances intensity cues and event-derived structural constraints. Comprehensive experiments on both synthetic benchmarks and real captures demonstrate that the proposed system consistently improves highlight recovery, edge fidelity, and robustness compared with frame-only or event-only HDR pipelines. The results indicate that jointly optimizing optical design, cross-modal alignment, and computational fusion provides an effective foundation for reliable HDR perception in highly dynamic and radiometrically challenging environments.
Real-time monitoring of high-energy propellant combustion is difficult. Extreme high dynamic range (HDR), microsecond-scale particle motion, and heavy smoke often occur together. These conditions drive saturation, motion blur, and unstable particle extraction in conventional imaging. We present a closed-loop Event-SVE measurement system that couples a spatially variant exposure (SVE) camera with a stereo pair of neuromorphic event cameras. The SVE branch produces HDR maps with an explicit smoke-aware fusion strategy. A multi-cue smoke-likelihood map is used to separate particle emission from smoke scattering, yielding calibrated intensity maps for downstream analysis. The resulting HDR maps also provide the absolute-intensity reference missing in event cameras. This reference is used to suppress smoke-driven event artifacts and to improve particle-state discrimination. Based on the cleaned event observations, a stereo event-based 3D pipeline estimates separation height and equivalent particle size through feature extraction and triangulation (maximum calibration error 0.56%). Experiments on boron-based propellants show multimodal equivalent-radius statistics. The system also captures fast separation transients that are difficult to observe with conventional sensors. Overall, the proposed framework provides a practical, calibration-consistent route to microsecond-resolved 3D combustion measurement under smoke-obscured HDR conditions.
During warhead detonation, high-density, high-speed, and mutually occluded fragments are generated. Their mechanical parameters (position, velocity, kinetic energy) directly determine the lethality of the warhead fragment field. However, high-intensity flash and smoke in detonation scenarios severely hinder the accurate acquisition of these mechanical parameters. To address this challenge, this paper integrates experimental mechanics approaches and presents an event-driven method for reconstructing the dynamic trajectories of fragments and measuring their mechanical parameters. As a novel brain-inspired visual sensor, event cameras offer microsecond-level temporal resolution and high dynamic range lighting change perception, overcoming the difficulty of accurately measuring high-speed targets under strong flash interference. The method constructs a multi-event-camera vision system, adopting three geometric constraints: time-correlated epipolar constraint to find potential matching event point pairs, and trifocal tensor line constraint plus local homography constraint to eliminate mismatches. A comprehensive probability model is established, with entropy weight method determining the weight of each constraint's probability to quantitatively filter mismatches. 3D trajectory reconstruction is achieved via spatial line-line intersection and nonlinear optimization. Finally, the velocity and kinetic energy of the fragments are calculated based on the reconstructed trajectory. This method provides reliable technical support for the mechanical damage evaluation of warhead fragment fields and the tactical protection design.
Despite the rapid advancements in event-based motion estimation, current geometric methods primarily focus on velocity estimation. However, absolute pose estimation, which is equally crucial for key applications such as robotic navigation and augmented reality, remains relatively underexplored. Consequently, the simultaneous recovery of absolute pose and velocity from event streams remains an open and challenging problem. To address this gap, we propose a geometric framework for absolute pose and velocity estimation by leveraging 3D lines in the scene and the events they trigger. At the core of the framework lie two key geometric constraints: the orthogonality between a 3D line and the normal vector of its corresponding event plane, and the collinearity of an event with the 2D projection of its associated line. Based on these constraints, we present both linear and polynomial solvers for absolute pose estimation. The former enables efficient computation, while the latter provides a globally optimal solution for rotation. For velocity estimation, we develop an efficient linear solver and a more accurate optimization-based solver to recover both angular and linear velocities. Notably, our methods require a minimum of three event-line correspondences to determine the 6-DoF absolute pose or velocities independently. Extensive experiments in simulation and on real-world datasets demonstrate that our methods achieve state-of-the-art performance, with significant improvements in accuracy and computational efficiency compared to existing methods. The demo code is publicly available at https://github.com/Zibin6/EventPoseVelocity.
Accurate body-to-body relative pose measurement is a fundamental requirement for close-range cooperative moving-platform operation. Monocular vision offers a passive, lightweight, low-cost, and low-power sensing route for such tasks. However, the pose recovered from visual observation is first expressed in the camera frame and cannot directly provide the relative pose between platform body frames. Recovering body-to-body relative pose therefore requires solving the unknown camera-to-body transformations. For this setting, a closed kinematic constraint AiX=YBi=Ci is formulated from mutual monocular observations of known body-fixed 3D feature points, where Ci is the frame-wise body-to-body relative pose. To handle feature mismatches, noisy Perspective-n-Point estimates, and unreliable pose pairs, we first construct a hierarchical robust initialization that combines feature-level PnP--RANSAC with frame-level consensus verification under bidirectional reprojection consistency. The accepted frames and retained feature observations are then used in an image-domain nonlinear optimization, which jointly optimizes the camera-to-body transformations and the relative pose sequence by minimizing bidirectional reprojection errors. Numerical simulations, Blender-rendered image experiments, and laboratory tests with OptiTrack references are conducted to evaluate accuracy, robustness, and computational efficiency. The results show that the frame-level consensus stage supplies a reliable initialization for reprojection optimization and that the proposed mutual monocular framework can provide practical close-range body-to-body relative pose measurements for cooperative camera-body platforms.
Quantitative optical measurement of critical mechanical parameters—such as plume flow fields, shock wave structures, and nozzle oscillations—during rocket launch faces severe challenges due to extreme imaging conditions. Intense combustion creates dense particulate haze and luminance variations exceeding 120 dB, degrading image data and undermining subsequent photogrammetric and velocimetric analyses. To address these issues, we propose a hardware-algorithm co-design framework that combines a custom spatially varying exposure (SVE) sensor with a physics-aware dehazing algorithm. The SVE sensor acquires multi-exposure data in a single shot, enabling robust haze assessment without relying on idealized atmospheric models. Our approach dynamically estimates haze density, performs region-adaptive illumination optimization, and applies multi-scale entropy-constrained fusion to effectively separate haze from scene radiance. Validated on real launch imagery and controlled experiments, the framework demonstrates superior performance in recovering physically accurate visual information of the plume and engine region. This offers a reliable image basis for extracting key mechanical parameters, including particle velocity, flow instability frequency, and structural vibration, thereby supporting precise quantitative analysis in extreme aerospace environments.
Camera calibration is a crucial first step in 3D computer vision. We propose a new method for camera calibration with a designed collimator, which can be applied to cameras with different focal lengths. Unlike the traditional collimator methods that need to know the direction of the collimated rays, we design a new system where the calibration pattern is rigidly attached to the reticle, enabling calibration through multi-view observations. The optical geometric principle of the collimator indicates that the relative motion between the calibration patterns fits the spherical motion model. We propose a calibration algorithm for our collimator system based on generic camera models. The spherical motion constraint reduces the motion parameters to be solved, so our method can achieve more accurate and robust calibration than the generic algorithm. We validate the performance of our method on the synthetic data. Our experiments on actual calibration data demonstrate that our method is feasible and returns accurate calibration parameters.
Objective Monocular pose estimation of non-cooperative targets is a key technique for space missions, including on-orbit servicing, rendezvous and docking, and autonomous target monitoring. With the rapid development of deep neural networks, semantic-keypoint-based pose estimation methods have achieved high pose accuracy by combining robust image feature extraction with Perspective-n-Point (PnP) solvers. However, in practical aerospace scenarios, obtaining only a deterministic pose estimate is insufficient. Non-cooperative targets often suffer from weak texture, partial occlusion, illumination variation, specular reflection, and large viewpoint changes. These factors introduce significant uncertainty into keypoint localization and subsequently affect the reliability of the estimated 6D pose. Therefore, a pose estimation system should not only output a pose value but also provide a confidence region that characterizes the possible range of the pose error. Existing uncertainty-aware pose estimation methods often rely on sampling-based uncertainty propagation or statistical assumptions. Although such methods can provide probabilistic information, they usually suffer from high computational cost, loose confidence regions, or insufficient coverage reliability. In particular, directly propagating sampled keypoint confidence regions through PnP may generate overly conservative pose confidence regions, while conventional covariance propagation may fail to provide reliable coverage under complex non-Gaussian keypoint errors. To address these problems, this paper proposes an uncertainty-driven confidence region estimation method for monocular pose estimation of non-cooperative targets. The goal is to construct compact and reliable 6D pose confidence regions while preserving high pose estimation accuracy and computational efficiency. Methods The proposed framework consists of three main stages: uncertainty information extraction, uncertainty calibration, and uncertainty propagation from image keypoints to 6D pose confidence regions. First, a lightweight neural network is adopted to predict semantic keypoint heatmaps from monocular images. Instead of only extracting deterministic keypoint locations, the method further estimates keypoint uncertainty from the predicted heatmap distribution. The heatmap response is used to compute both the expected keypoint coordinate and its associated uncertainty, thereby providing a probabilistic description of each semantic keypoint. This design enables the subsequent confidence estimation process to start from image-domain uncertainty rather than from manually assumed noise models. Second, the paper introduces inductive conformal prediction (ICP) to calibrate the predicted keypoint uncertainty. The predicted heatmap variance may not naturally correspond to the actual localization error because neural networks are often miscalibrated. To bridge this gap, the method defines a nonconformity score that measures the relationship between the predicted keypoint uncertainty and the real keypoint localization error. A calibration set is then used to estimate the quantile of the nonconformity scores. Based on this quantile, each predicted keypoint is assigned a confidence region with a user-specified coverage condition of the PnP objective, the Jacobian between keypoint perturbations and pose perturbations is obtained. The image-domain keypoint uncertainty can then be propagated to the pose domain through this Jacobian. The final 6D pose confidence region is represented by separate rotational and translational confidence regions, whose covariance matrices are obtained through deterministic propagation. This avoids repeated sampling and improves computational efficiency while maintaining a probabilistic interpretation. The method is evaluated on public non-cooperative target pose estimation datasets, including the SPEED satellite dataset and the LMO dataset. The evaluation covers keypoint localization accuracy, pose estimation accuracy, confidence region coverage rate, confidence region volume, calibration-set sensitivity, visualization results, and computational efficiency. The main comparison is conducted against existing statistical pose confidence estimation methods and their uncertainty propagation variants. Results and Discussions Experimental results verify that our method improves pose accuracy and confidence region performance. On the SPEED dataset, the proposed heatmap-based network reaches an average percentage of correct keypoints (PCK) of 92.75 degrees o, clearly exceeding the baseline of 83.79 degrees o. The lightweight structure ensures accurate keypoint localization, supporting reliable pose solving and uncertainty propagation. Our method maintains competitive translation accuracy while significantly lowering rotation error. Accurate keypoints reduce PnP input noise, and the uncertainty-aware mechanism suppresses unreliable keypoint disturbances, effectively stabilizing 6D pose estimation. On SPEED, our method achieves rotational and translational coverage of 96.4 degrees o and 98.7 degrees o, well constraining real pose errors. Meanwhile, its confidence region volumes are reduced by 63.0 degrees o (translation) and 93.5 degrees o (rotation) compared with the baseline. The IFT-based deterministic propagation avoids the over-conservatism of sampling methods and yields compact, high-quality confidence regions. CDF, boxplot and visualization results further demonstrate our advantages: smaller region volume, fewer outliers, and better stability. Uncertain keypoints correspond to larger local bounds, while precise keypoints generate tighter constraints. The propagated 6D confidence regions closely wrap pose errors without redundant expansion. Evaluated on the LMO dataset, our method exhibits strong generalization to complex object appearances and pose distributions, extending applicability from satellite scenarios to general monocular pose estimation. Calibration experiments show sufficient representative samples are indispensable for stable coverage. Insufficient calibration data reduces region volume but degrades coverage, conforming to conformal prediction theory. Calibration sets should cover diverse viewpoints, illumination and keypoint uncertainty conditions. Benefiting from the lightweight network and analytical IFT propagation instead of repeated sampling, our method lowers computational cost, making it suitable for resource-limited spaceborne and embedded real-time systems. Overall, the proposed method achieves an optimal trade-off among pose accuracy, coverage reliability, region compactness and computational efficiency. It converts image-level uncertainty into calibrated keypoint confidence regions and deterministically propagates them to interpretable 6D pose bounds. Monocular pose estimation of non-cooperative targets is fundamental to on-orbit servicing, rendezvous docking and autonomous monitoring. Deep learning-based semantic keypoint methods combined with PnP achieve high-precision pose estimation. However, deterministic results are insufficient for practical aerospace tasks. Challenges including weak texture, occlusion, illumination variation, specular reflection and viewpoint change induce severe keypoint uncertainty, undermining 6D pose reliability. Hence, practical systems need both pose values and error confidence regions. Current uncertainty-aware methods rely on sampling or statistical assumptions, suffering from high computation, loose regions and poor coverage. Sampling-based propagation yields over-conservative bounds, while traditional covariance methods fail under non-Gaussian errors. This work proposes an uncertainty-driven confidence region estimation method to attain compact, reliable 6D pose regions with high accuracy and efficiency. Conclusions This paper proposes an uncertainty-driven confidence region estimation method for monocular pose estimation of non-cooperative targets. The method builds a complete uncertainty estimation chain from image-domain semantic keypoint prediction to 6D pose confidence region construction. By introducing ICP, the predicted keypoint uncertainty is statistically calibrated to satisfy a preset coverage probability. By applying the implicit function theorem to the PnP optimization process, the calibrated keypoint uncertainty is analytically propagated to the pose domain without expensive sampling. As a result, the method can generate reliable and compact rotational and translational confidence regions. Experiments on public datasets demonstrate that the proposed method improves keypoint localization accuracy, reduces pose estimation error, increases confidence region coverage, and significantly decreases confidence region volume. Compared with existing methods, it provides more practical pose confidence estimation for non-cooperative target perception. The proposed framework is especially valuable for space missions, where autonomous systems need not only accurate pose estimates but also reliable uncertainty information for decision-making and risk control. Nevertheless, the current method mainly focuses on aleatoric uncertainty caused by image observation noise and keypoint localization ambiguity. Future work can further consider epistemic uncertainty from neural network models, non-Gaussian uncertainty modeling, temporal pose confidence estimation, and multi-view confidence fusion. These extensions may further improve the reliability of pose confidence estimation in complex real-world aerospace scenarios. probability. This procedure transforms raw neural-network uncertainty into statistically calibrated keypoint confidence regions. Third, the calibrated keypoint uncertainty is propagated to the pose space. Instead of using repeated sampling to propagate keypoint confidence regions through PnP, the paper derives an analytical uncertainty propagation formulation based on the implicit function theorem (IFT). The PnP pose is regarded as the implicit solution of a nonlinear least-squares optimization problem. By differentiating the optimality
The accuracy of photomechanics measurements critically relies on image quality, particularly under extreme illumination conditions such as welding arc monitoring and polished metallic surface analysis. High dynamic range (HDR) imaging above 120 dB is essential in these contexts. Conventional CCD/CMOS sensors, with dynamic ranges typically below 70 dB, are highly susceptible to saturation under glare, resulting in irreversible loss of detail and significant errors in digital image correlation (DIC). This paper presents an HDR imaging system that leverages the spatial modulation capability of a digital micromirror device (DMD). The system architecture enables autonomous regional segmentation and adaptive exposure control for high-dynamic-range scenes through an integrated framework comprising two synergistic subsystems: a DMD-based optical modulation unit and an adaptive computational imaging pipeline. The system achieves a measurable dynamic range of 127 dB, effectively eliminating saturation artifacts under high glare. Experimental results demonstrate a 78
Multi-camera systems are increasingly adopted in robotics and autonomous navigation for their wide field of view, flexibility, and fault tolerance. Nevertheless, existing PnP solvers fail to handle multiple projection centers. This paper introduces a virtual point formulation that bridges the standard PnP and generalized pose problems, enabling a unified pipeline that transforms existing PnP solvers into generalized pose solvers. Based on this framework, we derive three Virtual-point-based Generalized Pose solvers, namely VGPc, VGPq, and VGPr, leveraging Cayley, quaternion, and rotation-matrix parameterizations, respectively. Extensive experiments demonstrate that the proposed solvers inherit the accuracy and efficiency of original PnP algorithms while significantly outperforming existing generalized solvers. Specifically, VGPc achieves higher estimation accuracy under heteroscedastic noise conditions, VGPq maintains global optimality, whereas VGPr provides superior computational efficiency without accuracy degradation.
The characterization of mechanical properties for high-dynamic, high-velocity target motion is essential in defense testing. It provides crucial data for validating weapon systems and precision manufacturing processes etc. However, existing measurement methods face challenges such as limited dynamic range, discontinuous observations, and high costs. This paper presents a new approach leveraging an event-based multi-view photogrammetric system, which aims to address the aforementioned challenges. First, the monotonicity in the spatiotemporal distribution of events is leveraged to extract the target's leading-edge features, eliminating the tailing effect that complicates motion measurements. Then, reprojection error is used to associate events with the target's trajectory, providing more data than traditional intersection methods. Finally, a target velocity decay model is employed to fit the data, enabling accurate motion measurements via our multi-view data joint computation. In a light gas gun fragment test, the proposed method showed a measurement deviation of 4.47% compared to the electromagnetic speedometer.
Objective The rapid development of the low-altitude economy has expanded the application scenarios of unmanned aerial vehicles (UAVs). In formation flight and cooperative transportation, accurate inter-UAV relative pose estimation is essential for stable coordination. Compared with global navigation satellite systems (GNSS), which are vulnerable to signal blockage in denied environments, and simultaneous localization and mapping (SLAM), which usually requires considerable computation and communication resources, monocular vision offers a lightweight, passive, and low-power solution for close-range relative measurement. However, visual pose estimates are expressed in the onboard camera frame, whereas cooperative control requires poses in the UAV body frame. When two moving UAV platforms observe each other, the unknown static hand-eye transformations between the body frames and mounted cameras, together with the visual pose observation sequences, form a kinematic closed loop. Although this constraint can be formulated as a robot-world/hand-eye calibration problem, robotics-domain calibration solvers generally assume observations directly acquired from high-precision sensors with low noise. In contrast, UAV visual poses are intermediate estimates derived from feature extraction and Perspective-n-Point (PnP)-based pose estimation. Feature mismatches, PnP uncertainty, high-frequency vibration, illumination changes, and motion blur may introduce non-Gaussian noise and outlier-contaminated pose sequences. Directly applying robotics-domain solvers to this dual-moving UAV system can therefore cause severe error propagation and unstable calibration. To address this problem, this study proposes a hierarchical robust hand-eye calibration framework from the feature level to the frame level. The framework first integrates random sample consensus (RANSAC) into single-frame PnP estimation to reject pixel-level mismatches, then introduces a frame-level RANSAC scheme for hand-eye calibration. An uncertainty-aware weighted sampling strategy based on the PnP reprojection root mean square error (RMSE) guides the solver toward high-confidence pose subsets, reducing the required iterations. A reprojection-consistency-based inlier criterion is further designed to identify reliable frames, and the final hand-eye parameters are estimated from the maximum consensus set using an algebraic solver. Methods The proposed framework consists of feature-level robust pose estimation and frame-level joint hand-eye estimation. In the front-end stage, RANSAC is embedded in the PnP solver. Minimal 2D-3D correspondence sets are repeatedly sampled to generate pose hypotheses, which are evaluated using reprojection error. Pixel-level mismatches are thereby removed before pose refinement. The inlier correspondences are then optimized by minimizing reprojection error with the Levenberg-Marquardt algorithm, producing an initial sequence of relative pose observations. In the back-end stage, the method addresses the low efficiency of uniform RANSAC sampling by assigning each pose observation a sampling weight according to its uncertainty. Specifically, the reprojection RMSE from single-frame PnP estimation is used as a reliability indicator. After normalization, the inverse RMSE is converted into a non-uniform sampling probability, with a small positive parameter introduced to prevent weight explosion for near-zero residuals and maintain numerical stability. This design increases the probability that high-quality observations enter the minimal sample set and improves the expected convergence speed. To evaluate candidate hand-eye models, the proposed method avoids using algebraic residuals directly in the three-dimensional pose space SE(3), where rotation and translation have different physical dimensions and are only loosely coupled. Instead, the closed-loop constraint is projected back to the image plane. For each candidate model, the predicted pose is used to reproject known three-dimensional feature points, and point-wise reprojection errors are computed against the observed image features. A frame is accepted as an inlier only when both an adaptive reprojection error threshold and a feature-consistency ratio threshold are satisfied. This criterion preserves the native visual geometric constraint and improves sensitivity to feature-level deviations. Finally, the maximum consensus set obtained by RANSAC is used for hand-eye parameter estimation. Following algebraic hand-eye calibration principles, a Kronecker-product formulation constructs an overdetermined linear system in which rotation and translation are decoupled, and the two components are jointly solved to obtain the final calibration result. Results and Discussions The proposed method is evaluated using Monte Carlo numerical simulations and Blender-based rendering simulations. It is compared with several robotics-domain hand-eye calibration baselines, including Kronecker-product, stochastic global optimization, orthogonal-approximation, iterative, and probabilistic solvers, as well as with a uniform-sampling RANSAC variant. The results show that when PnP-derived visual pose observations contain non-Gaussian noise and outlier frames, robotics-domain solvers suffer from significant accuracy degradation and, in some cases, estimation divergence. By contrast, both uniform and weighted RANSAC variants substantially suppress outlier interference, confirming the necessity of a frame-level robust mechanism in UAV visual calibration. Under representative noise settings, the proposed weighted sampling strategy maintains pose estimation accuracy at the same order of magnitude as uniform sampling while reducing computation time by about 70 %. Across different numbers of pose observations, feature points, and noise levels, the weighted strategy consistently requires fewer iterations, indicating that PnP reprojection RMSE is an effective prior for selecting high-quality samples. Ablation experiments further verify the advantage of the reprojection-consistency-based inlier criterion. Compared with an algebraic-distance criterion based on rotation and translation residuals, the proposed criterion achieves higher average accuracy under different noise levels. Algebraic residuals are measured in SE (3) and do not explicitly preserve the image-domain geometric constraints that generate the visual poses. Under complex non-Gaussian noise, such criteria may be misled by pseudo-consensus sets with similar local noise distributions. In contrast, reprojection consistency evaluates candidate models in the pixel domain, maintains dimensional consistency, and directly reflects the influence of feature deviations on the closed-loop constraint. The data-dependency analysis shows that estimation accuracy gradually stabilizes when the number of pose observations exceeds 20 or when the number of feature points per frame exceeds 25. The weighted strategy improves computation efficiency by approximately 50 %-70 % under these conditions while maintaining comparable accuracy. Spatial-scale experiments indicate that the proposed method effectively suppresses error propagation in short-and medium-range cooperative scenarios, particularly when the relative translation scale ratio is below 50. However, in larger-scale settings, unit pixel noise induces rapidly amplified spatial pose perturbations, leading to systematic bias in the front-end PnP solution. The reprojection RMSE values of different frames also become similar, flattening the weighted sampling distribution and weakening its ability to distinguish reliable observations. Hyperparameter studies demonstrate that the method is robust to the confidence level, weighting parameter, and feature-consistency ratio within broad ranges. In Blender simulations with speckle noise, baseline algebraic solvers deviates severely from the ground truth, whereas the proposed method achieves a better balance between accuracy and efficiency. Conclusions This study presents a hierarchical robust hand-eye calibration framework for cooperative UAV relative pose measurement under unknown hand-eye parameters and visual outliers. By reformulating the dual-moving observation geometry as a hand-eye calibration problem and introducing feature-level and frame-level robust mechanisms, the method effectively rejects outlier-contaminated observations. The combination of uncertainty-aware weighted sampling and reprojection-consistency-based inlier selection improves both robustness and computational efficiency compared with uniform sampling. The experimental results suggest that the proposed framework can provide a reliable perception basis for UAV swarm cooperation in GNSS-denied environments. Nevertheless, the current method remains a two-stage pipeline that relies on PnP-derived poses as intermediate variables and does not directly optimize image features with the unknown hand-eye parameters. Future work will investigate end-to-end direct optimization that jointly associates raw image features, inter-UAV relative poses, and hand-eye parameters to reduce accumulated errors and further improve calibration accuracy and robustness.
Due to the extremely limited appearance characteristics and interference from complex backgrounds, the detection of infrared small targets remains a challenge. Single-frame methods rely solely on appearance information, which is insufficient and limits performance in complex background scenes. In contrast, multi-frame methods leverage motion information in the temporal domain simultaneously and have become the focus of infrared small-target detection. Currently, multi-frame algorithms are typically based on convolutional neural network (CNN) or Vision Transformer (ViT). CNN-based methods suffer from a limited local receptive field, while ViT-based methods exhibit high computational complexity. This paper proposed a multi-frame method based on Mamba-like spatio-temporal attention network. By replacing the recurrently computed forget gate in the Mamba architecture with linear attention, the model retains Mamba’s efficient computational capabilities while becoming better suited to non-auto-regressive vision models. This paper utilizes the improved Mamba model to achieve long-range interaction and efficiently fuse spatiotemporal information across sequential images. Moreover, we further introduced an inter-frame self-attention mechanism to extract motion features from sequential images, compensating for the insufficiency of appearance information for small targets. Finally, cross-layer connections are employed to prevent loss of small target features in deep layers caused by pooling operations. Comparative experiments with state-of-the-art algorithms on two typical datasets demonstrate that the proposed algorithm exhibits significant advantages in both effectiveness and computational efficiency.
Multi-camera systems offer rich observation capabilities for visual navigation and 3D scene reconstruction; however, the resulting feature redundancy often compromises computational efficiency. This challenge is particularly pronounced during bundle adjustment, where the non-linear optimization of both system poses and scene points incurs substantial computational overhead. To address this challenge, this paper introduces a pose-only geometric constraint for multi-camera systems and proposes a corresponding pose adjustment algorithm. Specifically, we use generalized camera model to establish a unified representation of the multi-camera system. Building upon this model, we formulate the multi-camera pose-only constraint, which implicitly represents a 3D scene point using two base observations and their associated poses, thereby achieving a pose-only representation of the projection geometry. Subsequently, we introduce a multi-camera pose adjustment algorithm that eliminates 3D points from the parameter space, thereby achieving efficient and focused pose optimization. Experimental results on both synthetic and real-world datasets demonstrate that the proposed algorithm outperforms baseline bundle adjustment methods in computational efficiency, while maintaining or even improving pose estimation accuracy
Conventional dynamics analysis of the human body is often constrained by the need for contact force and torque sensors and controlled laboratory environments. To address this issue, this study proposes an opticalmechanics kinematic-dynamic integrated estimation framework for multibody systems. Specifically, a constrained multibody model is established to describe the system dynamics, while image-measured kinematic quantities are used as non contact inputs for dynamic estimation. The unknown joint torque is then identified through a genetic-algorithm based optimization by minimizing the discrepancy between model-predicted and image-measured kinematic quan tities. Experimental validation on an air-bearing platform showed that the wrist joint torque estimated from image data achieved a mean absolute error of 0.46 Nm compared with sensor measurements. In the forward prediction test, the model-predicted angular velocity achieved a mean absolute error of 0.006 rad/s relative to the image-measured results. This study demonstrates the potential of combining image measurement and mechanical modeling for non-contact dynamic estimation in scenarios where direct force and torque measurement is difficult.
Event cameras are increasingly used for Multiple Object Tracking (MOT), but their asynchronous event output often requires specialized methods. Existing processing methods primarily follow two paradigms, pseudo-frames and event-by-event. The former is the prevailing approach since its data format aligns with images, making image-based techniques applicable. However, it suffers from tracking failures when trajectories overlap or are spatially close on pseudo-frames. Facing this challenge, we propose a multi-view pipeline, Multi-view Tracking (MvT), which preserves the 2D data format to leverage image-based techniques directly while introducing additional spatio-temporal views to resolve tracking ambiguities in a single view. MvT comprises a Multi-view Projection (MvP) module and a Multi-view Fusion (MvF) stage. MvP encodes events into three complementary spatio-temporal views while mitigating the pattern discretization. Within MvF, multi-view results are unified into a 3D coordinate system, and tracklets are associated through an optimization model subject to specific criteria combination. Evaluations on four datasets, including our self-collected Small Objects Dataset (SOD), show that MvT seamlessly integrates image-based methods and outperforms existing non-learning and learning trackers in generalized scenarios, and effectively resolves the single-view tracking ambiguities. Being training-free, MvT is applicable when ground-truth annotation is infeasible, thereby highlighting its practical, data-efficient potential. Code is available at https://github.com/zhazhabiu/MvTracking.
Existing event-based tracking methods typically operate in an online manner, associating tracklets across adjacent batches. However, this limited temporal window often leads to failures in challenging scenarios, such as long-term occlusions or missed detections, where broader temporal context should be incorporated for high-precision tracking. To achieve robust and accurate tracking, this paper concentrates on the event-based multi-dimensional assignment that associates tracklets across multiple batches simultaneously. Unlike image-based detectors, event-based detectors can generate multiple correct yet spatially distinct tracklets for a single object per batch. This violates the one-to-one rule in classic data association, making event-based association an underdetermined many-to-many matching problem. To the best of our knowledge, this is the first work to tackle this NP-hard problem. To address its underdetermined nature, we formulate the assignment problem via fuzzy logic, where pairwise tracklet correlations are quantified through multiple long-range kinematic metrics. This fuzzy design eliminates the need for hard tracking thresholds and enables multiple heterogeneous metrics to be integrated into a unified correlation. By treating tracklets as nodes and their fuzzy correlations as edges, we construct a Hierarchical Sparse Graph (HSG). Building upon the HSG, we pose assignment as a graph partition task, and employ a heuristic solution that measures vertex affinity for partitioning adaptively. The entire pipeline is training-free. Comprehensive experiments on two challenging event datasets, eTram and Ev-UAV, show that our method outperforms existing event-based and most image-based association models. Particularly, it achieves over 2× the performance of OC-SORT across all metrics in the many-to-many association scenario.
As a bio-inspired intelligent sensor, event cameras have introduced a new paradigm in the intelligent perception of spatiotemporal information and visual motion estimation, characterized by their high temporal resolution, low latency, and minimal power consumption. However, their asynchronous data streams present significant challenges to traditional synchronous, frame-based algorithms. To address these challenges, this paper presents a novel framework for full degree of freedom (DoF) egomotion estimation directly from asynchronous optical flow, specifically targeting the joint recovery of angular and linear velocities. We decouple the differential epipolar constraint into distinct angular and linear velocity components, and derive its formulation for asynchronous data. Based on this formulation, an optimization algorithm is developed that enables full-DoF egomotion estimation leveraging at least five points. Furthermore, by applying a first-order approximation to rotational dynamics, we transform the constraint equations into a polynomial form, resulting in the first algebraic minimal 5-point solver for this formulation. To ensure real-time performance in high-speed scenarios, we additionally propose an accelerated solver achieved by truncating high-order angular velocity terms. Extensive evaluations on both synthetic and real-world datasets demonstrate that the asynchronous approach outperforms traditional synchronous methods, particularly in its accuracy and robustness to spatiotemporal noise. We believe that this work establishes a critical foundation for efficient and accurate continuous-time motion estimation in high-speed robotics applications.