Infrared–visible fusion alleviates the limitations of single-modality detection, especially under low illumination and haze. However, most existing methods rely on sequential or parallel attention fusion, which limits representation ability and efficiency. To address these issues, we propose DASYOLO, a cross-modal object detector that combines shallow feature enhancement with dual-attention-guided cross-modal fusion. Specifically, DASYOLO incorporates a BiAttention module to strengthen shallow feature representations through complementary channel and spatial attention, and a Dual Attention Synergy (DAS) module that jointly models channel importance and spatial saliency through multiplicative coupling. This design improves cross-modal feature discrimination while maintaining high inference efficiency. Extensive evaluations on M3FD, LLVIP, and our UGVLQ dataset demonstrate that DASYOLO achieves 83.9
The synergistic integration of infrared and visible video streams enhances scene representation for critical applications such as autonomous navigation and surveillance. However, effectively fusing these cross-modal sequences remains challenging due to their inherent spatiotemporal discrepancies, which often lead to incoherent outputs and degraded performance in downstream video analysis tasks. To address this, we propose TSCP-MF, a novel mimic fusion framework founded on the principle that "differences determine structure, and structure determines outcome." The framework introduces a collaborative time-space perception module that quantitatively captures inter-frame temporal dynamics and intra-frame spatial distributions of cross-modal features. Based on this perception, an intelligent discrimination mechanism adaptively selects between two mimic optimization pathways to dynamically reconfigure the fusion architecture. The entire process is formalized through a correlation-aware joint synthesis rule, which leverages possibility theory to optimize fusion parameters via mimic transformation and multi-projection learning. Extensive experiments demonstrate that TSCP-MF achieves superior spatiotemporal consistency in the fused video, significantly outperforming state-of-the-art methods in preserving thermal saliency for target localization and texture details for scene understanding, while maintaining minimal inter-frame flicker. The proposed framework provides an effective and adaptive solution for cross-modal video fusion, with strong potential for integration into real-time video processing systems.
Existing fusion methods lack the capability to perceive and adaptively adjust fusion structures during critical moments of significant feature variation in infrared and visible videos, resulting in blurring, detail loss, and performance degradation at salient frames. To address this issue, this paper proposes a salient frame-driven video mimic fusion method, termed SFDM-Fusion, which is based on possibility distribution evidence synthesis. The proposed method facilitates structural adaptive adjustment through accurate salient frame detection, thereby enhancing video fusion quality in complex dynamic scenes. First, weighted attributes of feature magnitude and frequency are computed to extract cross-modal intra-frame difference features and single-modal temporal variation features. Second, clustering analysis based on feature uncertainty variations is performed to characterize possibility distributions. A weight distribution matrix and a nonlinear fusion rule are designed to construct a Gaussian-based possibility belief assignment function for effective feature uncertainty quantification. Furthermore, an ordered reliability decision-making method integrating an improved PROMETHEE II and evidence theory is proposed. By establishing a dual-dimensional evaluation criterion and a nonlinear possibility preference function, the net flow ranking result is converted into dynamic weights that measure evidence reliability, thus optimizing the evidence synthesis process and improving salient frame detection accuracy. Finally, the detected salient frames drive the adaptive selection and fusion of mimic fusion variants. Extensive experimental results demonstrate that the proposed SFDM-Fusion not only maintains superior fusion performance at salient frames but also significantly enhances overall video fusion quality, exhibiting remarkable adaptability and robustness in complex dynamic scenarios.
In complex dynamic environments, infrared and visible video sequences exhibit highly variable and unpredictable feature distributions. Existing fusion algorithms with fixed architectures cannot adaptively respond to these dynamic feature changes, resulting in blurred fusion outcomes and the loss of critical detail information. To address this limitation, we propose a salient feature-driven mimic fusion algorithm that continuously monitors feature variations and dynamically reconfigures the fusion architecture to maintain optimized fusion performance. First, we extract amplitude and frequency attributes from infrared and visible video features and perform weighted fusion to calculate single-modality temporal features and cross-modal intra-frame difference features. Second, based on clustering statistical properties of feature distributions, we construct possibility distribution functions to quantify the degree of feature variation, design synthesis rules to derive comprehensive possibility values of feature change, and utilize significantly changing features as driving factors for subsequent mimic variant adjustments. Building upon this foundation, we establish a fusion validity evaluation function by analyzing correlation coefficients between feature changes and fusion quality metrics, and accordingly construct mapping relationships between features and mimic variants. Finally, we determine the optimal mimic variant combination by synthesizing various feature change characteristics to implement mimic fusion. Experimental evaluation demonstrates that our proposed method significantly outperforms existing approaches in adaptive fusion performance in dynamic scenes, with superior preservation of edge and texture details in the fusion results.
To address the challenges in multi-modal trajectory prediction for multi-aircraft interactions within non-towered terminal airspace, including the insufficient extraction of long-range temporal dependencies, neglect of physical separation constraints, and barriers to integrating flight intentions and multi-source environmental context, this paper develops a trajectory prediction model integrated with long-range temporal modeling, physics-aware spatial interaction, and intention context enhancement. A parameter-shared ST-Transformer temporal encoder is established to capture long-term motion patterns of aircraft via global multi-head self-attention, and temporal attention pooling is adopted to mitigate error accumulation in long-term prediction. A physical distance-aware ST-GAT module is designed, which embeds the spatial distance prior between aircraft into the attention calculation and leverages distance masks to reduce interference from distant irrelevant aircraft. Furthermore, an intention-aware context enhancement module (IACEM) is proposed. It identifies the distribution of flight phases and adaptively incorporates meteorological information to construct enhanced features embedded with high-level semantics and environmental priors. Finally, a CVAE-based framework is utilized to generate multiple candidate trajectories satisfying kinematic constraints. Multiple verification experiments are carried out on the TrajAir dataset. The experimental results demonstrate that the proposed model outperforms various baseline models in terms of ADE and FDE. Ablation studies and visual analysis verify that the three core modules produce synergistic improvements. The model achieves higher prediction accuracy under scenarios involving 2D complex maneuvers, 3D climbing turns, and dense multi-aircraft interactions, which proves the effectiveness and superiority of the proposed algorithm for trajectory prediction in non-towered terminal airspace.
Existing end-to-end infrared–visible fusion methods often blur edges, smooth textures and weaken target-to-background contrast. We therefore propose Structure–Detail Constrained Fusion (SDC-Fusion), a frequency-decoupled framework with separate constraints on structure and detail. The proposed method employs the Haar wavelet transform to decompose the source images into low-frequency structural and high-frequency detail components. The high-frequency branch uses a local directional-energy prior to construct the target guidance. Unlike existing Rectified Flow-based approaches that operate on the full image or a generic latent representation, our gated module applies Rectified Flow only to the high-frequency wavelet subbands. It learns a few-step residual trajectory from the visible high-frequency coefficients to the target representation, enhancing infrared target boundaries and visible textures without altering low-frequency structure. In the low-frequency branch, adaptive weighting, four-directional Mamba scanning, and multi-dilation depthwise convolutions are integrated to preserve global luminance and background structure while optimizing local grayscale transitions. Comparative experiments against eleven representative fusion methods on the MSRS and M3FD datasets show that SDC-Fusion ranks first in SSIM, VIF, Qabf, SF, and PSNR on MSRS, and first in SSIM, VIF, Qabf, SD, and PSNR on M3FD, while ranking second in the remaining two metrics on each dataset. Relative to the strongest competing result, the largest improvements reach 10.44% in SF on MSRS and 5.72% in VIF on M3FD. The model contains 0.535 M parameters and requires 67.5 G FLOPs.
Multivariate time series classification finds extensive applications in medical diagnosis, human health monitoring, and other fields. However, existing methods pay insufficient attention to the unique “two-dimensional structure” (variable dimension and time dimension) of multivariate time series, resulting in relatively low classification accuracy. Therefore, this paper proposes a multivariate time series classification model based on a variable-time dual-attention mechanism, which separately extracts features from the variable and temporal dimensions of the sequence to achieve accurate classification. Specifically, the model proposed in this paper primarily consists of two components: a variable-time dual-attention module and a convolutional feature enhancement module. First, the feature enhancement module extracts “state features” at each time step through one-dimensional convolution, and concatenates them with the original input to achieve feature enhancement. Next, the variable-time dual-attention module applies convolutional attention simultaneously across both variable and time dimensions to the enhanced data, extracting key variable features and critical time points while suppressing less significant components. Experimental results on multiple real-world datasets show that the proposed method attains the highest average accuracy of 84.04
Infrared-visible object detection in complex dynamic environments often suffers from weak feature representation and underutilized cross-modal complementarity, leading to missed and false detections. To address these issues, we propose a Dual-modal Enhanced Feature Enhancement and Fusion Network (DEF-Net). To enhance the model's focus on informative features within both infrared and visible modalities, a feature interaction enhancement module is designed to effectively highlight and reinforce salient information. Furthermore, to better exploit the complementary characteristics of the two modalities, a transformer-based fusion architecture incorporating a cross-attention mechanism is introduced, enabling deep inter-modal feature integration. Experiments on SYUGV and LLVIP datasets show that DEF-Net outperforms existing methods in accuracy while maintaining real-time processing speed.
Detection networks based on deep learning mainly adopt a single feature interaction mechanism to capture the deep features of targets. As the quality of infrared or visible images deteriorates moderately or severely, this often results in the insignificance of deep feature. To surmount this deficiency, we present a parallel feature interaction network, termed PFI-Net. This architecture involves dual-branch feature extraction and enhancement, parallel feature interaction and decision fusion detector. With dual-branch feature extraction and enhancement as the premise, we construct a parallel feature interaction module with different interaction mode to avoid mutual interference between features of infrared and visible image. This parallel feature interaction module can ensure the features of infrared and visible are guided into two separate independent channels. Additionally, we devise a weighted detection boxes fusion module to achieve the integration of the parallel detection results. This module integrates the advantages of detection results from different channels to promote detection accuracy and stability. Finally, comprehensive experiments on multiple benchmark models demonstrate that the proposed PFI-Net delivers promising detection performance, outperforming other advanced alternatives.
It is necessary to optimize the fusion strategy according to the difference information between videos to achieve high-precision fusion of multi-source videos. The existing fixed fusion strategy can only fuse the video frame as a whole, and cannot select the appropriate fusion method for different regions in the frame, which is difficult to achieve the global optimum in the frame, resulting in poor fusion effect or even fusion failure. We imitate the multi-mimic characteristics of the mimic octopus, and propose a mimic fusion method driven by difference feature space-moment perception for infrared and visible video, namely DFSMP. Firstly, from the intra-frame region layer, difference feature space-moment is defined and calculated by using difference feature amplitude and the amplitude probability density estimated by the KNN method, and the main difference feature types of each region are determined. Secondly, the method for selecting the size of the meta-image blocks is introduced. Based on the idea of region growing, the intra-block aggregation segmentation method is proposed to adaptively divide the local region of the video frame to be fused. Then, the optimal mimic variables are hierarchically selected for local mimic fusion based on the constructed fusion effectiveness function. Finally, the blocks are spliced and equalized to the method for selecting the size of the meta-image blocks is introduced the global mimic fusion of infrared and visible video. The results show that the proposed method retains more complete useful target information of the source images both globally and locally, and has better fusion performance than other methods in spatial information.
Traditional fusion methods based on deep learning mainly employ convolutional or self-attention operations to model local or global dependencies, which often lead to the oversight of frequency-domain information. To address this deficiency, we introduce a unified frequency adversarial learning network, termed FreqGAN. Our method involves a frequency-compensated generator that employs discrete wavelet transformation to decompose encoded spatial features into multiple frequency bands. Leveraging skip connections, low and high-frequency components are respectively directed into the encoder and decoder, compensating for additional outline and detail. Moreover, we construct a hybrid frequency aggregation module, which enables a progressive optimization of activity levels across multiple scales and makes the various frequency bands correlated. Complementing our generative model, we devise dual frequency-constrained discriminators. These discriminators are tasked with dynamically adjusting weights for each input frequency band, thereby obligating the generator to accurately reconstruct salient frequency information from different modality images. Additionally, a frequency-supervised function is formulated to further safeguard against the loss of frequency information. Our comprehensive experimental evaluations, encompassing a wide range of fusion tasks and subsequent applications, distinctly highlight FreqGAN’s superior performance, establishing it as a frontrunner in comparison to existing state-of-the-art alternatives. The source codes are forthcoming at: https://github.com/Zhishe-Wang/FreqGAN.
The variable and unpredictable distribution of spatiotemporal features in infrared-visible videos under complex dynamic environments is addressed, where traditional fusion algorithms with fixed architectures are found to fail when adapting to continuous dynamic changes in feature distributions, resulting in blurred fusion outcomes and loss of critical detail information. To overcome this bottleneck, a mimic fusion method based on spatiotemporal feature saliency change-driven approach, termed STMFuse, is proposed, whereby spatiotemporal feature variations are continuously monitored and the fusion architecture is dynamically reconfigured, thereby ensuring optimal fusion performance. Specifically, a spatiotemporal dual-domain feature perception module extracts single-modal temporal change features and cross-modal difference features to comprehensively capture dynamic scene characteristics. To address the inadequacy of single threshold settings in measuring dynamic feature variations, a possibility distribution function and synthesis rules is constructed to quantify feature change degree, utilizing significant change features as driving factors for mimic variant adjustments. By investigating the dynamic correlation between feature changes and fusion quality metrics, a mimic variant evaluation function based on correlation coefficient weighting is established, mapping the relationship between features and mimic variants. Finally, a multi-feature collaborative decision mechanism based on energy functions is designed to determine the optimal mimic variant combination by integrating various feature change characteristics. Extensive experimental validation demonstrates that STMFuse significantly outperforms existing methods in terms of adaptive fusion performance and robustness in dynamic scenes, exhibiting superior preservation of edges, texture details, and enhanced visual perception quality.
Infrared and visible image fusion integrates complementary information from multimodal sensors into a single informative representation, thereby enhancing scene perception and promoting applications in computer vision tasks. Despite recent advances, existing generative adversarial network (GAN)-based methods continue to face significant challenges, including entangled holistic feature processing, inefficient frequency separation, and substantial computational overhead. To address these issues, we propose a parallel frequency-invertible adversarial network for infrared and visible image fusion, termed PFI-Fuse. Operating in the wavelet domain, PFI-Fuse employs a divide-and-conquer strategy to process different frequency subbands in parallel, allowing the model to better capture both global structures and fine-grained textures. More importantly, we introduce a frequency-invertible generator, which incorporates invertible neural blocks to ensure that critical frequency information is preserved throughout the fusion process, ensuring minimal information loss. Furthermore, we introduce a wavelet modulation loss, which dynamically adjusts subband contributions during adversarial training. This adaptive loss encourages the network to balance structural preservation, texture detail enhancement, and intensity consistency across all frequency components. Extensive experiments on different benchmarks and downstream applications demonstrate that PFI-Fuse consistently outperforms existing state-of-the-art methods in both quantitative metrics and visual quality, while also achieving superior computational efficiency. The source code will be available at: https://github.com/Zhuoqun-Zhang/PFI-Fuse.
Maize lodging poses a significant challenge to agricultural production, severely constraining yield improvement and mechanized harvesting efficiency. Under modern agricultural practices characterized by high-density planting and multi-variety intercropping, there is an urgent need for precise and efficient monitoring technologies to address lodging issues. This study utilized unmanned aerial vehicle (UAV) light detection and ranging (LiDAR) to acquire high-precision point cloud data of field maize at full maturity. An innovative method was proposed to automatically identify structural differences induced by lodging by analyzing canopy structural similarity across multiple height thresholds through point cloud stratification. This approach enables automated monitoring of maize lodging in complex field environments. The experimental results demonstrate the following: (1) High-precision point cloud data effectively capture canopy structural differences caused by lodging. Based on the structural similarity change curve, the height threshold for lodging can be automatically identified (optimal threshold: 1.76 m), with a deviation of only 2.3% between the calculated lodging area and the manually measured reference (ground truth). (2) Sensitivity analysis of the height threshold shows that when the threshold fluctuates within a +/- 5 cm range (1.71-1.81 m), the calculation deviation of the lodging area remains below 10% (maximum deviation = 8.2%), indicating strong robustness of the automatically selected threshold. (3) Although UAV flight altitude influences point cloud quality (e.g., low altitude: 25 m, high altitude: 80 m), the height threshold derived from low-altitude flights can be extrapolated to high-altitude monitoring to some extent. In this study, the resulting deviation in lodging area calculation was only 5.3%.
Addressing the limitation of existing infrared and visible video fusion models, which fail to dynamically adjust fusion strategies based on video differences, often resulting in suboptimal or failed outcomes, we propose an infrared and visible video fusion algorithm that leverages the autonomous and flexible characteristics of multi-agent systems. First, we analyze the functional architecture of agents and the inherent properties of multi-agent systems to construct a multi-agent fusion model and corresponding fusion agents. Next, we identify regions of interest in each frame of the video sequence, focusing on frames that exhibit significant changes. The multi-agent fusion model then perceives the key distinguishing features between the images to be fused, deploys the appropriate fusion agents, and employs the effectiveness of fusion to infer and determine the fusion algorithms, rules, and parameters, ultimately selecting the optimal fusion strategy. Finally, in the context of a complex fusion process, the multi-agent fusion model performs the fusion task through the collaborative interaction of multiple fusion agents. This approach establishes a multi-layered, dynamically adaptable fusion model, enabling real-time adjustments to the fusion algorithm during the infrared and visible video fusion process. Experimental results demonstrate that our method outperforms existing approaches in preserving key targets in infrared videos and structural details in visible videos. Evaluation metrics indicate that the fusion outcomes obtained using our method achieve optimal values in 66.7% of cases, with sub-optimal and higher values accounting for 80.9%, significantly surpassing the performance of traditional single fusion methods.
Optimizing fusion strategy according to the difference information between videos is a necessary means to realize high-precision fusion of multisource videos. Addressing the limitations of existing video fusion models, which fail to intelligently adapt fusion strategies to dynamic changes in difference features over time, resulting in resource wastage, poor fusion quality, and even failure, we propose a mimic fusion method inspired by the versatile mimicry of the octopus. This method is driven by difference feature time-phase (TP) perception for infrared and visible video, namely DFTP. First, the comprehensive weight function of difference features is constructed by extracting the amplitude and frequency of difference features between bimodal videos, so as to determine the main difference feature types of video frames. Second, the corresponding possibility distribution vector subsets of each difference feature are obtained under the possibility framework, and the weighted synthetic of multiple subsets is realized based on the credibility (CR) and gap value (GA) of each subset. The comprehensive information function of the feature synthesis subset is constructed, and the TP value of the difference feature is calculated. Then, the mimic transformation mode of the current frame is determined by perceiving the TP changes of difference features between frames. Finally, the best combination of mimic variables is selected to realize the mimic fusion of infrared and visible video based on the mapping memory bank and fusion effectiveness function. The experimental results show that the proposed method not only retains the typical infrared targets and visible structural details of the video as a whole but also is significantly better than other single fusion methods in quantitative analysis and qualitative evaluation.
Aiming at the problem that the existing algorithms are difficult to segment efficiently in indoor scenes due to the high similarity of stacked and closely adjacent objects, this paper proposes an indoor point cloud segmentation algorithm based on hierarchical feature fusion. Firstly, according to the results of stratification, the principal component analysis method is adopted to reduce the dimension of each layer point clouds, and then the decision rules on the two-dimensional plane are established based on the structural characteristics of each layer objects. Combined with the analysis of hierarchical connectivity, fusing the features of each layer, so as to obtain the global positioning of the object. Last, for further division, the density segmentation method based on inter cluster constraints are taken into consideration. The experimental results show that the location and segmentation effect of stacked objects is fine, It is worth mentioning that this method has high accuracy, and the average mIoU can reach 0.811, In addition, another advantage of this method is that it takes less time to process point clouds, the average consumption time is about 6.59 seconds. These characteristics represent that this method can better complete the segmentation of indoor scenes.
Path planning is one of the most essential parts of autonomous navigation. Most existing works are based on the strategy of adjusting angles for planning. However, drones are susceptible to collisions in environments with densely distributed and high-speed obstacles, which poses a serious threat to flight safety. To handle this challenge, we propose a new method based on Multiple Strategies for Avoiding Obstacles with High Speed and High Density (MSAO2H). Firstly, we propose to extend the obstacle avoidance decisions of drones into angle adjustment, speed adjustment, and obstacle clearance. Hybrid action space is adopted to model each decision. Secondly, the state space of the obstacle environment is constructed to provide effective features for learning decision parameters. The instant reward and the ultimate reward are designed to balance the learning efficiency of decision parameters and the ability to explore optimal solutions. Finally, we innovatively introduced the interferometric fluid dynamics system into the parameterized deep Q-network to guide the learning of angle parameters. Compared with other algorithms, the proposed model has high success rates and generates high-quality planned paths. It can meet the requirements for autonomously planning high-quality paths in densely dynamic obstacle environments.
Existing deep learning-based methods often follow either image-level or feature-level fusion frameworks to uniformly or separately extract features, ignoring the specialized interactive information learning, which may produce limited fusion performance. To tackle this challenge, we devise a powerful fusion baseline via adaptive interactive Transformer learning, namely AITFuse. Unlike previous methods, our network alternately incorporates local and global relationships through collaborative learning of both CNN and Transformer. In particular, we propose a cascaded token-wise and channel-wise Vision Transformer architecture with different attention mechanisms to model the long-range contexts, and allow feature communication across different tokens and independent channels in an interactive manner. On this basis, the modal-specific feature rectification module employs self-attention operation to revise distinctive features within the same domain for efficient encoding. Meanwhile, the cross-modal feature integration module constructs cross-attention mechanism to fuse complementary characteristics from different domains for multi-level decoding. In addition, we discard the learning position embedding to release our fusion model for the image of arbitrary sizes without splitting operations. Extensive experiments on mainstream datasets and downstream tasks demonstrate the rationality and superiority of our AITFuse. The codes will be available at https://github.com/Zhishe-Wang/AITFuse.
Aiming at the heterogeneous features are often collaboratively optimized for fusion, and the existing feature attributes cannot be targeted to adjust algorithms to drive fusion effectively, resulting in poor fusion. We put forward infrared image fusion algorithm selection based on quality synthesis of intuition possible sets. Firstly, the fusion validity of difference feature of image is calculated. In view of intuition possible set orderings of multiple intervals of each difference feature, the degrees can be divided into three levels, and the corresponding possibility distribution subsets are obtained. Secondly, we introduce credibility and separability to weight the multiple subsets, and calculate negative entropy and credibility of each synthesized subset, in order to determine non-dominated subsets under each fusion algorithm. Thirdly, the score function of multiple fusion algorithms is constructed to select optimal fusion algorithm. Then we employ thirteen evaluation indicators and their comprehensive score to verify the effectiveness of our method.