To tackle the frequent missed and false detection issues arising from the tiny scale of objects and strong background clutter in UAV aerial photography scenarios, this paper proposes a novel algorithm named SODet-YOLO for UAV aerial imagery. First, to effectively extract the features of aerial objects and alleviate background interference, we integrate a high-resolution detection head denoted as P2 into the YOLO11n, which is connected to the feature layer from the second downsampling stage of the Backbone and Neck networks, we design the Fine-Grained Aggregation-Asymptotic Feature Pyramid Network (FGA-AFPN) to realize adequate fusion of feature information at different levels. Second, we redesign the original C3k2 module by embedding the Inception Depthwise Convolution (IDC). This design effectively expands the receptive field, enriches multi-scale contextual feature extraction, and mitigates adverse interference from complex background clutter. In addition, a novel IoU loss function named MPDInterpIoU is proposed by combining InterpIoU with MPDIoU. This function promotes faster convergence at the early learning stage and optimizes detection-related performance. Finally, the Parallelized Patch-aware Attention (PPA) is incorporated before the downsampling module to preserve the key features of small objects throughout multiple downsampling steps. The experimental findings validate that SODet-YOLO achieves an mAP@0.5 score of 41.487% on the VisDrone2019 object detection dataset, representing an 8.92% performance enhancement relative to the baseline YOLO11n model. However, the computational cost increases moderately, with the number of parameters increasing by 1.08 M, the computational complexity increasing by 26.1 GFLOPs, and the average inference time growing by 34.7 ms.
In response to the difficulty of detecting small-sized targets in drone aerial scenes, an improved YOLOv11n small target detection algorithm based on progressive feature pyramid network is proposed. Firstly, a detection head is added to the high-resolution feature layer of the backbone network, and feature information between non adjacent layers is fused through an asymptotic feature pyramid network (AFPN) to alleviate the problem of feature information loss caused by downsampling and reduce conflicts between cross level feature information. Secondly, an improved SPPF based on spatial channel attention mechanism is used to replace the original SPPF, highlighting important features and suppressing irrelevant feature information through spatial and channel attention mechanisms, further enhancing the model's performance. Finally, combining MPDIoU and InnerIoU improves detection accuracy. The experiment shows that the improved algorithm achieves a performance of 38.537 on the Visdrone2019 dataset mAP@0.5 Compared to the benchmark model, it has increased by 5.97%.
In order to solve the problems of low detection accuracy and poor real-time performance caused by small size, dense distribution and complex background of UAV aerial vehicle detection in UAV aerial vehicle detection scenes, this paper proposes an improved UAV aerial vehicle detection algorithm based on YOLOv11n. Firstly, a detection head was added to the high-resolution feature layer of the YOlOv11n backbone network to reduce the problem of small target information loss caused by the reduction of resolution after downsampling, and improve the detection accuracy of small target. Secondly, the Inner-IoU Loss is used to replace the traditional IoU Loss, and the auxiliary box is used to accelerate the convergence of the regression process. In addition, the Multi-scale Attention Aggregation Module (MSAA) is introduced to fuse multi-scale feature information by using spatial and channel attention mechanisms, which can improve the effect of multi-scale spatial and channel fusion while reducing background interference. Experiments show that the improved algorithm achieves a of 38.302
Semantic segmentation is an essential task in polarimetric synthetic aperture radar (PolSAR) image interpretation. To address the issue of insufficient measurement ability of single-view similarity, an unsupervised semantic segmentation method for PolSAR images based on multiview similarity is proposed to estimate the number-of-classes (NoC) and perform classification. NoC estimation is commonly neglected in semantic segmentation methods, to ensure rationality, a NoC estimation method before clustering is proposed based on multiview polarimetric rotation domain features and the visual assessment of tendency method for PolSAR images without supervision. Then, the norm distance, geodesic distance, maximum likelihood distance, and generalized likelihood ratio test distance based on the statistical characteristics of PolSAR images are comprehensively analyzed. Various advantages of different distances are integrated to combine the multiview vector information and scattering information to construct six multikernel similarity matrices. Subsequently, the consensus similarity network fusion method is utilized to further strengthen the discriminative ability of the similarity matrices. In addition, efficient superpixel segmentation is also adopted to reduce the speckle noise. Finally, based on the estimated NoC and the fused similarity matrix, spectral clustering is utilized to obtain semantic segmentation results. Extensive experiments are conducted on two AIRSAR datasets and one Gaofen-3 dataset demonstrate that the proposed method can effectively combine the spatial neighborhood similarity information and achieve higher semantic segmentation accuracy.
The complex Wishart distribution is a widely used statistical model for multilook polarimetric synthetic aperture radar (PolSAR) image data, of which the equivalent number of looks (ENL) is a critical parameter. Over the past decades, various estimators have been developed to estimate the ENL of complex Wishart distribution, of which the maximum likelihood (ML) estimator is important since it is asymptotically unbiased and has a small variance. However, this estimator is very time-consuming since it has no analytical solution and is usually solved numerically. To address this problem, this letter proposes an efficient ML estimator of ENL by deriving an approximate closed-form solution. Moreover, to estimate the ENL map of a PolSAR image, we also develop an efficient way to compute the local sample statistics parallelly. The experimental results on two PolSAR images show that our method yields highly approximate ENL values as the traditional ML estimator while is much more efficient. It costs less than 0.8 s on a general laptop to estimate the ENL map of a PolSAR image with 900 $\times $ 1024 pixels.
Recently, convolutional neural networks (CNNs) have shown significant advantages in the tasks of image classification; however, these usually require a large number of labeled samples for training. In practice, it is difficult and costly to obtain sufficient labeled samples of polarimetric synthetic aperture radar (PolSAR) images. To address this problem, we propose a novel semi-supervised classification method for PolSAR images in this paper, using the co-training of CNN and a support vector machine (SVM). In our co-training method, an eight-layer CNN with residual network (ResNet) architecture is designed as the primary classifier, and an SVM is used as the auxiliary classifier. In particular, the SVM is used to enhance the performance of our algorithm in the case of limited labeled samples. In our method, more and more pseudo-labeled samples are iteratively yielded for training through a two-stage co-training of CNN and SVM, which gradually improves the performance of the two classifiers. The trained CNN is employed as the final classifier due to its strong classification capability with enough samples. We carried out experiments on two C-band airborne PolSAR images acquired by the AIRSAR systems and an L-band spaceborne PolSAR image acquired by the GaoFen-3 system. The experimental results demonstrate that the proposed method can effectively integrate the complementary advantages of SVM and CNN, providing overall classification accuracy of more than 97%, 96% and 93% with limited labeled samples (10 samples per class) for the above three images, respectively, which is superior to the state-of-the-art semi-supervised methods for PolSAR image classification.
Superpixel generation of polarimetric synthetic aperture radar (PolSAR) images is widely used for intelligent interpretation due to its feasibility and efficiency. However, the initial superpixel size setting is commonly neglected, and empirical values are utilized. When prior information is missing, a smaller value will increase the computational burden, while a higher value may result in inferior boundary adherence. Additionally, existing similarity metrics are time-consuming and cannot achieve better segmentation results. To address these issues, a novel strategy is proposed in this article for the first time to construct the function relationship between the initial superpixel size (number of pixels contained in the initial superpixel) and the structural complexity of PolSAR images; additionally, the determinant ratio test (DRT) distance, which is exactly a second form of Wilks' lambda distribution, is adopted for local clustering to achieve a lower computational burden and competitive accuracy for superpixel generation. Moreover, a hexagonal distribution is exploited to initialize the PolSAR image based on the estimated initial superpixel size, which can further reduce the complexity of locating pixels for relabeling. Extensive experiments conducted on five real-world data sets demonstrate the reliability and generalization of adaptive size estimation, and the proposed superpixel generation method exhibits higher computational efficiency and better-preserved details in heterogeneous regions compared to six other state-of-the-art approaches.
行人重识别旨在多个视频传感器条件下,从图像库中出检索特定的行人目标,具有重要的实际应用价值.针对以往对局部特征利用不足的情况,创新一种基于注意力引导的局部特征关系融合方法,使在对局部特征分别计算的同时,通过注意力引导,探索各局部特征之间的内部关系.首先将图像通过残差网络ResNet-50获取特征,然后对特征进行水平分割获取局部特征后,通过注意力引导的局部特征关系融合网络,最后使用难采样三元组损失函数和交叉熵损失函数对模型进行训练.实验表明,该算法在行人重识别公开数据集Market-1501上mAP值达到86.4%,Rank-1达到94.7%.
Clustering-based methods of polarimetric synthetic aperture radar (PolSAR) image superpixel generation are popular due to their feasibility and parameter controllability. However, these methods pay more attention to improving boundary adherence and are usually time-consuming to generate satisfactory superpixels. To address this issue, a novel cross-iteration strategy is proposed to integrate various advantages of different distances with higher computational efficiency for the first time. Therefore, the revised Wishart distance (RWD), which has better boundary adherence but is time-consuming, is first integrated with the geodesic distance (GD), which has higher efficiency and more regular shape, to form a comprehensive similarity measure via the cross-iteration strategy. This similarity measure is then utilized alternately in the local clustering process according to the difference between two consecutive ratios of the current number of unstable pixels to the total number of unstable pixels, to achieve a lower computational burden and competitive accuracy for superpixel generation. Furthermore, hexagonal initialization is adopted to further reduce the complexity of searching pixels for relabelling in the local regions. Extensive experiments conducted on the AIRSAR, RADARSAT-2 and simulated data sets demonstrate that the proposed method exhibits higher computational efficiency and a more regular shape, resulting in a smooth representation of land cover in homogeneous regions and better-preserved details in heterogeneous regions.
Distance measure plays a critical role in various applications of polarimetric synthetic aperture radar (PolSAR) image data. In recent decades, plenty of distance measures have been developed for PolSAR image data from different perspectives, which, however, have not been well analyzed and summarized. In order to make better use of these distance measures in algorithm design, this paper provides a systematic survey of them and analyzes their relations in detail. We divide these distance measures into five main categories (i.e., the norm distances, geodesic distances, maximum likelihood (ML) distances, generalized likelihood ratio test (GLRT) distances, stochastics distances) and two other categories (i.e., the inter-patch distances and those based on metric learning). Furthermore, we analyze the relations between different distance measures and visualize them with graphs to make them clearer. Moreover, some properties of the main distance measures are discussed, and some advice for choosing distances in algorithm design is also provided. This survey can serve as a reference for researchers in PolSAR image processing, analysis, and related fields.
针对极化SAR图像分类中卷积神经网络(CNN)方法训练时间长、收敛速度慢,原始Softmax函数无法对极化SAR图像的类内差异有效应对的问题,提出一种基于模型微调与加性边际Softmax(AM-Soft-max)的极化SAR图像分类方法.该方法通过预训练网络的整体微调,来改进CNN模型的效率和分类准确率,然后以AM-Softmax替代Softmax,以解决SAR图像中类内变化较大的问题,进一步提升分类精度.实验表明该方法具有快收敛的优势并且能够较好解决极化SAR图像类内差异较大的问题,模型的分类总体精度达到96%以上.
The distance measure plays a crucial role in the PolSAR image superpixel segmentation. In most cases, the commonly used simple weighting is adopted to combine multiple distance measures to calculate the similarity, thus leading to large computational burden and low segmentation performance. To solve this problem, this paper proposes a novel PolSAR image superpixel segmentation method based on a novel cross iteration strategy to incorporate the advantages of the geodesic distance and the revised Wishart distance. First, the PolSAR image is initialized as hexagonal distribution and all pixels are set as unstable pixels. Second, the revised Wishart distance and geodesic distance are adopted by the cross iteration strategy to relabel all unstable pixels. Finally, the postprocessing procedure is used to generate the final superpixels. Extensive experiments conducted on the AirSAR dataset demonstrate that the proposed method exhibits higher computational efficiency and more regular shape, resulting in smooth representation of the land covers in homogeneous regions, and better preserved details in heterogeneous regions.
The scattering feature plays a crucial role in the polarimetric synthetic aperture radar (PolSAR) image classification. In most cases, the commonly used feature concatenation is adopted to construct the similarity matrix, thus leading to the discriminability loss of some single-view features. To solve this problem, this paper proposes an unsupervised PolSAR image terrain classification method based on the cross-view tensor product graph (CV-TPG) diffusion. Multi-view learning can effectively integrate the data sets from different views, moreover, the diffusion based on the CV-TPG is capable of mining the intrinsic affinity along the manifold structure of data. First, the PolSAR image is over-segmented into many superpixels. Second, five feature vectors are extracted from the PolSAR image via superpixels to form three high-dimensional feature vectors, resulting in three corresponding similarity matrices with the Gaussian kernel. Third, multiple CV-TPGs are obtained by the tensor product operation based on three similarity matrices, then CV-TPGs are linearly fused for applying the diffusion process to achieve a more discriminative similarity matrix. Finally, spectral clustering based on the diffused similarity matrix is adopted to perform terrain classification. Extensive experiments conducted on the Oberpfaffenhofen data set demonstrate that CV-TPG diffusion can effectively combine the characteristic information of multiple views and achieve higher classification accuracy, compared to five other competitive state-of-the-art methods.
Considering the lack of similarity capabilities of the distance metric used in the traditional Polarimetric Synthetic Aperture Radar (PolSAR) image superpixel segmentation algorithm, a novel PolSAR image superpixel segmentation algorithm based on geodesic distance is proposed in this paper. First, the PolSAR image is initialized as a hexagonal distribution, and all pixels are initialized as unstable pixels. Thereafter, the geodesic distance between two real symmetric Kennaugh matrices is used to measure the similarity between the current unstable point and another cluster point in the search region to more accurately assign labels to unstable points, thereby effectively reducing the number of unstable points. Finally, the postprocessing procedure is used to remove small, isolated regions and generate the final superpixels. To verify the effectiveness of the initialization method and the high efficiency of the geodesic distance, extensive experiments are conducted using simulated PolSAR images. Moreover, the proposed algorithm is analyzed and compared with four other algorithms using simulated and real-world images. Experimental results show that the superpixels generated using the proposed method exhibit higher computational efficiency and a more regular shape that can more accurately fit the edges of real objects compared with those using the four other algorithms.
This article proposes a robust distributed receding horizon control (RDRHC) synthesis approach for the simultaneous tracking, regulation and formation of multiple perturbed wheeled vehicles with collision avoidance. By successfully extending a tube‐based RHC approach proposed in our previous work and elaborately designing the collision avoidance and compatibility constraints, a nominal control optimization problem is constructed for each vehicle with recursive feasibility guarantee, and the associated RDRHC algorithms with and without on‐line optimization are presented for implementation. By applying the presented algorithms, the multiple perturbed wheeled vehicles can be steered to achieve the desired tracking, regulation and formation objective with satisfying the pre‐specified constraints and avoiding collision. Both theoretical properties and practical effectiveness of the proposed approach are verified through a simulation example.
Recently, convolutional neural networks (CNNs) have been successfully developed and used in the classification of polarimetric synthetic aperture radar (PolSAR) images. However, they often suffer from some problems, such as time-consuming, unsatisfactory detail-preservation, and bad effectiveness given limited training samples. Focusing on these problems, we propose a complex-valued CNN (CV-CNN)-based algorithm for PolSAR image classification in this article. On the one hand, a superpixel-oriented (SPO) scheme is employed to reduce the computational cost of the algorithm and preserve image details simultaneously, which takes superpixels instead of single pixels as classification units. In particular, to meet the input requirement of CV-CNN, three alternative methods of superpixel regularization are designed and compared. On the other hand, considering that both measured data (MD) and manually designed polarimetric features (PFs) have their own advantages, the hybrid data (HD) combining them is employed to drive CV-CNN, which is helpful to improve the effectiveness of the algorithm. We perform experiments on three actual PolSAR image data sets acquired by AIRSAR and Radarsat-2 systems as well as a semisimulated data set. The experimental results demonstrate that, compared to conventional pixel-oriented methods, the proposed SPO scheme is much more time-efficient and is also beneficial to detail preservation. Moreover, the CV-CNN driven by HD generally obtains consistently better classification results than that driven by pure MD or manually designed PFs.
Building extraction technology in urban areas has been a hot topic in recent years, but how to accurately distinguish vegetation, buildings, and man-made objects and improve classification accuracy has always been a difficult point. Aiming at the problem of low classification accuracy, we propose a point cloud classification algorithm based on random forest. First, the improved cloth filtering algorithm is used to perform ground filtering on the point cloud data. And a decision tree is constructed and the correlation analysis based on the largest mutual information coefficient is performed to select the decision tree with the smallest correlation coefficient and the highest accuracy to obtain a weakly correlated random forest model. The decision results arc processed by weighted voting, and finally a point cloud classification algorithm combining cloth filtering and weighted weakly correlated random forest is obtained. Compared with the traditional random forest classification algorithm, the algorithm is verified by the Vaihingen urban dataset, and the classification accuracy is improved by 4.2%.
In this paper, we proposed a structure oriented descriptor (SOD) for local image feature description. Different from the traditional methods, we explored the structure information via hierarchical strategy and structure coding elements. Firstly, the support region is partitioned into hierarchical sub-regions according to the intensity orders. Secondly, a pre-designed structure coding image is explored to pooling the features according to the hierarchical sub-regions. The final descriptor is a conjunction of the feature of each sub-region. We evaluated the proposed descriptor on the public dataset and compared it with the state-of-the-art works. The experimental results indicate that the proposed descriptor is robust to image appearance change and outperform the state-of-the-art works in most cases.
针对当前计算机类课程课堂教学过程中存在的理论性强、过程枯燥等问题,系统梳理总结"起承转合"式教学法的基本理念和实施步骤,并以数字图像处理课程中直方图均衡化的课堂教学为例,详细分析该教学方法在课堂教学中的具体运用,最后说明"起承转合"式教学法在课堂教学中需要注意的问题.
In this study, a weakly supervised classification method is proposed to classify the Polarimetric Synthetic Aperture Radar (PolSAR) images based on sample refinement using a Complex-Valued Convolutional Neural Network (CV-CNN) to solve the problem that the bounding-box labeled samples contain many heterogeneous components. First, CV-CNN is used for iteratively refining the bounding-box labeled samples, and the CV-CNN that can be used for direct classification is trained simultaneously. Then, the given PolSAR image is classified using the trained CV-CNN. The experimental results obtained using three actual PolSAR images demonstrate that the heterogeneous components can be effectively eliminated using the proposed method, obtaining significantly better classification results when compared with those obtained using the traditional fully supervised classification method in which original bounding-box labeled samples are used. Furthermore, the proposed method with CV-CNN is superior to those in which the classical Support Vector Machine(SVM) and Wishart classifier are used.