
Image fusion generates more information-rich integrated images by integrating multi-source image data, and boasts extensive applications in fields such as environmental monitoring and battlefield reconnaissance. Although existing technologies can achieve favorable multi-modal image fusion in non-interfering scenarios and aid in the presentation of scene information, they exhibit significant limitations when confronted with the degradation of low-quality source images induced by extreme scenarios, including low light, haze, rain and snow, motion blur, and strong noise interference. To tackle these challenges, this paper proposes a novel anti-degradation multi-modal image fusion framework guided by supplementary text instructions. Firstly, the Low-rank mechanism is employed to conduct targeted fine-tuning on the CLIP vision-language large model, enabling accurate capture of scene degradation features such as low light, noise, and motion blur, and facilitating the generation of semantic descriptions. Secondly, a heterogeneous mixture-of-experts feature extraction backbone network is designed, which leverages scene degradation semantic descriptions to dynamically regulate heterogeneous expert modules. Meanwhile, a mutual information loss function is introduced to optimize the sparsity of experts during multi-task processing, thereby avoiding redundant computations. Finally, each module undergoes training-testing collaborative optimization to enhance the controllability of the fusion process and improve the generalization capability of the model. Experiments on 4 mainstream benchmark datasets (LLVIP, MSRS, M3FD, ROAD) demonstrate that the proposed text-semantic-guided sparse mixture-of-experts mechanism exhibits significant advantages over state-of-the-art methods in terms of image fusion performance and degradation correction. This research outcome provides new insights for multi-modal image fusion in complex degraded environments and is expected to promote technological advancement in fields such as remote sensing images and medical images.
Owing to the inherent limitation of coarse-grained supervision, contrastive vision-language models such as CLIP typically exhibit suboptimal performance in fine-grained attribute understanding tasks. We propose SubCLIP, a lightweight framework designed to enhance the capability of the CLIP in capturing subtle visual-semantic distinctions. SubCLIP enhances fine-grained attribute understanding by decomposing the positive and negative captions of an image into subtexts containing attribute words and corresponding nouns and introducing a multi-granularity contrastive learning objective. Specifically, we introduce Hierarchical Subtext Augmentation, which leverages a large language model to generate fine-grained, subtexts from existing annotations, thereby providing denser and more informative training signals. To fully exploit these subtexts, we propose Multi-Granularity Contrastive Learning, a unified strategy that jointly optimizes global image-text alignment, intra-sample attribute-level contrast, and subtext-level compositional alignment. This hierarchical design enables SubCLIP to learn both high-level semantic associations and subtle attribute differentiations. On the FG-OVD benchmark, SubCLIP significantly improves fine-grained attribute retrieval performance. When using ROI features for retrieval, it achieves Top-1 accuracy rates of 62.4 % on the hard samples, surpassing previous state-of-the-art methods. Our results demonstrate SubCLIP's effectiveness in enhancing fine-grained vision-language understanding, delivering robust performance while maintaining computational efficiency and domain adaptability. Our code is available at https://github.com/mop134679852/SubCLIP.
Falls pose a serious challenge to the health and safety of the elderly population. In crowded public places, such incidents may also trigger secondary disasters such as stampede, so rapid identification of fall behavior is crucial for maintaining public safety. Traditional fall detection algorithms have significant deficiencies in terms of insufficient feature extraction, single detection method, and poor real-time performance, which make it difficult to cope with complex fall scenarios and real-time detection requirements. In order to overcome the shortcomings of traditional fall detection algorithms, this study improves the PPYOLOE algorithm. We optimize feature extraction using Depthwise Separable Convolution (DSConv) and introduce High-Level Screening-feature Fusion Pyramid Networks (HS-FPN) and Efficient Multi-Scale Attention (EMA) to enhance important features. The improved algorithm significantly improves the detection performance and computational efficiency. Compared with the traditional PPYOLOE algorithm, the fall detection algorithm designed in this paper shows significant advantages in detection accuracy and real-time performance. In the experiment, Our improved algorithm has achieved a 3.5% increase in accuracy on the mAP@0.5 metric, it reaches 94.4% and achieves 24fps on GPU, which realizes the high precision detection of fall events and meets the deployment requirements in low computation scenarios. The real-time and high precision of the algorithm makes it suitable for various monitoring scenarios, which improves the social security level and emergency response capability.
Accurate identification of crop diseases is crucial for enhancing agricultural yields. In the agricultural field, crop disease images vary in resolution and aspect ratio due to different imaging equipment, distances, and environmental conditions, which can affect the accuracy and consistency of disease detection models. We propose a new model, MMCA-ConvNeXt, based on ConvNeXt-T. This model incorporates our novel Multi-scale Multidimensional Collaborative Attention (MMCA) mechanism, which enhances the receptive field of the original model while capturing features at multiple resolutions. This enables better generalization across different crop disease datasets, improving performance and robustness. Additionally, we replaced layer normalization in the downsampling layer and post-global average pooling with Batch Channel Normalization, which allows better exploitation of inter-channel information and enhances gradient flow within the model. Experiments using publicly available datasets for rice, maize, and the 2018 AI Challenger demonstrate the superior performance of our ConvNeXt-based model, achieving accuracies of 99.53% for rice, 99.93% for maize, and 87.57% for the 2018 AI Challenger dataset. It outperforms VGG - 19, ResNet - 101, EfficientNet, and several transformer models with fewer parameters. MMCA shows the best results among different attention mechanisms for plant leaf disease detection, and ablation studies confirm the model’s performance enhancements.
In recent years, the research has been ongoing about crowd simulation. In many situations, crowd are usually organized into several groups and proceed to their respective target location, to accomplish something or perform a task together for a duration, and this process may involve communication and cooperation within the group. We define such crowd as a task-driven group cooperative crowd, and focus on the crowd movement and group behavior during this process. The movement process of the crowd involves multiple stages and exhibits different behaviors at different stage. Within this framework, we consider a problem scenario where crowd within the group orderly gathering in front of objects for a duration to perform a task. We propose a data-driven framework based on Multi-Agent Reinforcement Learning(MARL), including task parameter policy learning and curriculum training strategy. We use a set of task parameters to describe the current task information and integrate them into the iterative optimization process of the policy. Based on this, we design agent observation and a weighted reward function, to make the agent recognize the current task stage and learn the corresponding behaviors. Utilizing curriculum training strategy, the training process is divided into multi steps, gradually increasing the training complexity. We conduct a series of experiments to verify the effectiveness and generalization of the framework and design a museum exhibition scenario to verify the practicality. We provide a demo to supplement the more details, https://github.com/ICSRC/2024ICVRV.
Multiple Object Tracking (MOT) is a fundamental and challenging task in the field of computer vision. Recently, the MOT community has turned its attention from the tracking-by-detection paradigm to a brand new one, namely joint detection and tracking (JDT). In this paper, we address the crowded tracking issue in the JDT paradigm and propose an end-to-end trainable multiple object tracking model based on multi-frame regression, which we term as MFR-tracker. Different from the traditional JDT paradigm which generates frame-level detections, our MFR-tracker takes short-term temporal information into account and directly predicts tracklets in a video clip. Specifically, we approximate a tracklet as an adaptive Bezier curve and learn the object’s arbitrary motion with only several control points. The simple yet effective motion modeling helps a lot in distinguishing different objects, especially when they are occluded somewhere. We conduct all experiments on MOT benchmarks of MOT16, MOT17, and MOT20, demonstrating its remarkable tracking performance.
Existing morphological models of gastropod mollusks often fail to accurately simulate the shell morphology of certain species and overlook morphological variations within the same genus, which are crucial for quantifying biological growth. This study focuses on Conomurex luhuanus, characterized by its highly inflated shell whorls and severely contracted apex, for which current morphological models are inadequate. We introduce a shell aperture opening function to capture its unique morphology. Utilizing a landmark-based geometric morphometrics approach, we extract geometric parameters from landmark points in two-dimensional sample images to construct a three-dimensional shell model. By comparing the parameters of the digital model with those of real samples, we demonstrate that this method accurately simulates individual morphological differences, providing a new analytical tool for quantifying population variations. This research not only enhances our understanding of the growth and development patterns of mollusks but also offers new insights for the morphological analysis of other complex shell species.
Numerous studies have utilized machine learning and deep learning techniques for Alzheimer's disease diagnosis using Structural Magnetic Resonance Imaging (SMRI). However, some networks tend to underutilize the fine-grained spatial information across multiple layers, leading to incomplete utilization of the image's large-scale features. Additionally, overly complex networks increase the risk of overfitting early in the training process, resulting in suboptimal diagnostic accuracy. This paper introduces a novel deep-learning model designed for Alzheimer's disease diagnosis. The proposed model incorporates a Feature Extraction Layer (FEL) with three DeepSearch Units (DSUnits) and a Packet Layer to effectively capture detailed information. Inspired by the ResNet residual structure, the DSUnit employs a combination of convolutional layers with various kernel sizes and ReLU activation functions to extract multi-scale features from structural MRI images. The Packet Layer is composed of multiple feature sub-packets, which effectively reduce mutual interference between neurons, thereby enhancing diagnostic accuracy. The experiments are conducted on the ADNI dataset to compare the proposed model with EfficientNet, DenseNet, ResNet (18, 34, 50), AddFormer, MPS-FFA, and CDFA. The model achieved higher accuracy in binary classifications (CN vs. AD, CN vs. MCI) and exhibited competitive results in the multi-class classification (CN vs. MCI vs. AD) while using fewer parameters than AddFormer, MPS-FFA, and CDFA. Overall, the proposed method offers a more explainable and efficient approach to Alzheimer's diagnosis, demonstrating strong potential for clinical application. The code is available at https://github.com/ourWorks/PktNet and the data is acquired from https://adni.loni.usc.edu.
Weakly-supervised video polyp segmentation is challenging because of no accurate ground truths. This paper proposes a scribble based framework called Voting based Video Polyp Segmentation (VotingVPS), which is based on the statistical advantage of voting in choosing the optimal candidates. For the robust pseudo masks, voting is integrated into the visual foundation model SAM to obtain effective supervision. Considering the initial prediction is also important for the refinement, the framework also designs a morphology based encoder-decoder module for this purpose. Experimental results show the effectiveness of the framework for weakly-supervised video polyp segmentation.
Estimating 3D human pose and shape from monocular videos is a challenging task due to inherent ambiguity and occlusion, which often lead to inaccurate predictions with high uncertainty. Despite significant progress in 3D pose and shape estimation from a single RGB image, achieving accurate and temporally coherent human mesh sequences from monocular videos remains a challenging endeavour. Inspired by the recent success of diffusion models in generating high-quality outputs with low uncertainty through progressively denoising noisy inputs, we propose a novel diffusion-based framework for 3D human pose and shape estimation. This framework formulates the human mesh recovery task as a reverse diffusion process. During training, it diffuses SMPL pose parameters from ground-truth distributions into input-specific distributions and learns to reverse this process. By leveraging the capacity of diffusion models to reduce noise progressively, our method effectively addresses the inherent ambiguity of this monocular task, producing accurate and smooth human mesh sequences from videos. Comprehensive experiments demonstrate that the proposed method significantly outperforms previous video-based methods in both per-frame 3D pose and shape accuracy and temporal coherence on widely used benchmarks, including 3DPW and Human3.6M datasets.
Audio-visual question answering (AVQA) task, which aims to answer questions derived from the original videos, has attracted extensive attention in the fields of multimedia, computer vision, and natural language processing. Although the current studies have achieved promising results, there are still some limitations that need to be further explored. First, most existing methods mainly focus on frame-level visual features while paying little attention to object-level visual features, leading to ambiguous semantic representations of visual objects. Second, these methods ignore the temporal information embedded in the dynamic changes among adjacent time frames, causing the model to be unable to capture the temporal dependence in the video, thereby limiting the model’s capabilities. To address the challenges above, we propose the Graph-based Visual Enhanced Video Network (GVEV-Net), which enables fine-grained question-related visual information queries and guides the fusion of multimodal information through language information. Specifically, we first consider extracting audio-visual cues from object-level, frame-level visual information, and audio information. Then, we highlight the relevant object-level, and frame-level visual information and audio clips by taking the question as the guiding information. Next, graph attention is employed to capture the dynamic change relationship among adjacent frames in the object-level visual information. Finally, question-guided attention is used to fuse all the modal information. Experiments on the MUSIC-AVQA dataset demonstrate that our method outperforms seven existing recent methods. Moreover, the ablation study verifies the necessity and effectiveness of the GVEV-Net proposed in our proposal.
In physical simulation, improving the efficiency and accuracy of real-time simulation of elastic deformable objects becomes paramountly significant. Projective dynamics (PD), as a foundamental method for real-time simulation, holds great necessity in acceleration implementation, such as using Anderson acceleration (AA). However, introducing AA into PD has certain limitations, including insufficient convergence rates or numerical stagnation. To tackle these challenges, we propose the restarted Anderson acceleration with monotonicity control and L1-norm minimization for projective dynamics (RAM-L1) to improve the algorithm’s robustness and global convergence speed, as well as surmount the high computational costs and numerical stagnation. Experiments on triangular and tetrahedral meshes show that the RAM-L1 provides a more efficient and accurate real-time simulation algorithm from both physical and numerical perspectives, proposing a potential framework for designing fast algorithms in physical simulation.
3D Gaussian Splatting has shown exceptional potential in novel view synthesis and 3D reconstruction by leveraging fast rasterization techniques for real-time rendering. Despite its advancements, incorporating semantic information into Gaussians remains challenging due to 2D Object-level Ambiguity and 3D Object-level Ambiguity, significantly impacting the performance of point-level 3D scene understanding and downstream tasks like segmentation and scene editing in 3D space. To address these issues, we propose ConsisGaussian, which ensures all Gaussians corresponding to the same object share a consistent feature. Our method unifies multi-view semantic information, resolving 2D Object-level Ambiguity by classifying objects into a single class and performing class-wise feature fusion. Additionally, we extend the differentiable renderer of 3DGS to create an object-z map, enabling inverse projection to assign consistent features to Gaussians in 3D space and overcome 3D Object-level Ambiguity. Our semantic segmentation and scene editing experiments demonstrate that this approach offers a robust solution for precise 3D scene understanding.
Camera-IMU calibration is essential for achieving precise and dependable data fusion in many augmented reality applications. However, most conventional methods rely on trajectory optimization based on image features, which can be easily affected by other factors. To address this issue, we propose a new calibration framework that utilizes a high-precision external device, such as a robotic arm, to optimize the trajectory of the camera for camera-IMU calibration. The proposed method has two main steps: 1) Optimize camera trajectory estimation using the external device; 2) Calculate camera-IMU parameters with the optimized camera trajectory and pre-processed IMU data. The main advantages of our proposed method are: 1) Utilizing a high-precision external device to reduce image detection errors for improving camera trajectory accuracy; 2) Reducing sensitivity to unexpected factors and increasing the robustness of camera-IMU calibration in various environments. Extensive experiments show that the proposed camera-IMU calibration framework can significantly improve the accuracy and robustness of calibration results by about 35%.
Fusing infrared and visible images is an essential aspect of image processing, designed to integrate both types of data into a single image, thereby improving visual clarity and detail.Most methods for image fusion using deep neural networks often focus on fusion at the pixel or feature level, overlooking the semantic information embedded in the images at a more abstract level. To meet the needs of complex vision tasks and enhance image fusion performance in these scenarios, this paper introduces a multi-scale semantic-focused fusion network for infrared and visible images, named MSDFusion. This network integrates a semantic segmentation module with an image fusion module, allowing it to preserve more semantic information. Experimental results demonstrate that MSDFusion achieves superior visual effects and richer semantic information relative to other leading-edge algorithms
Traditional deep learning semantic segmentation networks can already achieve high segmentation accuracy, but they generally suffer from large parameter size and slow inference speed, making it challenging to deploy them on mobile embedded devices for real-time segmentation tasks. However, the real-time segmentation networks that have emerged in recent years have proved their effectiveness and efficiency, with major advancements in lightweight encoders or decoders and multi-branch architectures. Taking their various advantages, we propose a real-time semantic segmentation network based on a general lightweight backbone and multi-scale feature fusion, which uses sharing shallow features, including semantic, spatial, and boundary three-branch decoder and a global boundary attention fusion module-GBAU. We use different backbone models to conduct experiments for verifying its outstanding performance. Experimental results show that when using STDC1 as the backbone and an input size of 512*1024, this model achieves 72.95% mIoU on the Cityscapes dataset. The model has 8.17M parameters and the inference speed can reach 194.7 FPS, so achieve good balance between segmentation accuracy, inference speed and parameter size.
The popularity of intelligent mobile devices and location-based social networks have promoted research on point-of-interest (POI) recommendation. Most studies have used complex models to extract multi-dimensional user check-in features, which have improved recommendation performance but significantly increased parameter size and computational complexity. In this study, we propose a lightweight (LS-LNR) method for next POI recommendation to achieve the optimal balance between performance and computational cost. To improve the recommendation performance, user check-in behavior learning is divided into long-term and short-term patterns, and multi-dimensional user check-in sequence features are integrated in LS-LNR. To reduce computational complexity, only the feedforward layer in the transformer mechanism is used to extract user features. Moreover, we propose a lighter feedforward layer consisting of linear and convolutional layers, and adopt the GELU activation function. Extensive experiments are conducted on a large-scale real-life dataset collected from Gowalla, and the results demonstrate that LS-LNR has significantly improvements in recall and normalized discounted cumulative gain even with lower operational parameters compared with baseline approaches.
Considering the challenges posed by low background contrast, varying target scales, high computational complexity of the detection model, and real-time detection limitations in marine environments, this paper proposes an algorithm for naval target detection based on an enhanced YOLOv8. Firstly, we introduce a novel feature-focused diffusion pyramid network (FDN) to enhance the representation of feature maps by integrating features from different layers and employing diffusion mechanisms to address the issues of low background contrast and diverse target scales. Secondly, we incorporate the AIFI attention mechanism to facilitate feature interaction, reducing computational redundancy and enhancing detection performance. Additionally, by drawing inspiration from StarNet’s approach, we have further enhanced the C2f module to effectively address issues pertaining to gradient vanishing. Finally, we propose a task-align dynamic detection head (TADH) for classification and location branching tasks that learns task interaction features from multiple convolutional layers. This enables us to obtain joint features while reducing model complexity and addressing target scale inconsistencies.
Accurate defect segmentation is a tricky challenge due to large differences in defect sizes and appearances and blurred boundaries. In this paper, we propose a Dual-Branch Network with Semantics and Details Aggregation for Defect Segmentation, abbreviated as DSDNet. Specifically, DSDNet first applies the bilateral branches to encode semantic and detail information in defect images, respectively. In the semantic branch, we adopt the decoder network to reconstruct the encoder features layer by layer to mitigate the issue of feature distortion. To produce more accurate decoder features, we propose a Strip Context Fusion Module (SCFM) using various strip convolutions and channel attention through skip connections to fully extract context information. Simultaneously, we devise a Selective Semantics Aggregation Module (SSAM) via adaptively integrating multi-scale semantic information in the decoder stage with fine-grained boundary features to provide sufficient semantic guidance to the detail branch. Finally, to alleviate feature misalignment between the two branches, a Bilateral Attention Modulation Unit (BAMU) is introduced to facilitate the efficient feature fusion from semantic and detail branches through symmetric attention guidance. Experimental results on publicly available defect datasets, NEU-Seg, MT-Defect, and MSD, demonstrate that DSDNet can effectively improve segmentation accuracy compared with state-of-the-art algorithms.
Vision-based hand gesture recognition methods enable a natural and efficient human-robot interaction process. However, gesture images tend to have interfering factors such as complex environmental backgrounds, varying light intensities, and skin-like pixels, which makes it difficult for existing gesture recognition networks to weigh the relationship between recognition accuracy and computational cost. To this end, a lightweight cross-dimensional feature adaptive enhancement network (CFAENet) is proposed in this paper. The network efficiently integrates convolutional and attention modules from the perspective of information transfer flow, and proposes a space and channel information consistency enhancement (SCICE) module. In addition, this paper proposes a lightweight spatial attention guidance (SAG) and a channel attention guidance (CAG) to achieve cross-dimensional feature aggregation within a module, and employs differential processing of spatial information to construct an efficient multi-view feature aggregation (MFA) bottleneck. Compared with some gesture recognition networks on multiple gesture benchmark datasets, the proposed CFAENet shows better performance in terms of recognition accuracy and computational cost, and achieves a good interaction experience in human-robot interaction experiments.