In remote sensing building detection tasks, data acquisition remains a critical bottleneck that limits both model performance and large-scale deployment. Due to the high cost of manual annotation, limited geographic coverage, and constraints of image acquisition conditions, obtaining large-scale, high-quality labeled datasets remains a significant challenge. To address this issue, this study proposes an automatic semantic labeling framework for remote sensing imagery. The framework leverages geospatial vector data provided by OpenStreetMap, precisely aligns it with high-resolution satellite imagery from Bing Maps through projection transformation, and incorporates a quality-aware sample filtering strategy to automatically generate accurate annotations for building detection. The resulting dataset comprises 36,647 samples, covering buildings in both urban and suburban areas across multiple cities. To evaluate its effectiveness, we selected three publicly available datasets—WHU, INRIA, and DZU—and conducted three types of experiments using the following four representative object detection models: SSD, Faster R-CNN, DETR, and YOLOv11s. The experiments include benchmark performance evaluation, input perturbation robustness testing, and cross-dataset generalization analysis. Results show that our dataset achieved a mAP at 0.5 intersection over union of up to 93.2%, with a precision of 89.4% and a recall of 90.6%, outperforming the open-source benchmarks across all four models. Furthermore, when simulating real-world noise in satellite image acquisition—such as motion blur and brightness variation—our dataset maintained a mean average precision of 90.4% under the most severe perturbation, indicating strong robustness. In addition, it demonstrated superior cross-dataset stability compared to the benchmarks. Finally, comparative experiments conducted on public test areas further validated the effectiveness and reliability of the proposed annotation framework.
This study experimentally explored the effects of equivalence ratio settings on ethanol fuel combustion oscillations with a laboratory-scale combustor. A contrary flame equivalence ratio adjusting trend was selected to investigate the dynamic characteristics of an ethanol atomization burner. Research findings denote that optimizing the equivalence ratio settings can prevent the occurrence of combustion instability in ethanol burners. In the combustion chamber, the sound pressure amplitude increased from 138 Pa to 171 Pa and eventually dropped to 38 Pa, as the equivalence ratio increased from 0.45 to 0.90. However, the sound pressure amplitude increased from 35 Pa to 199 Pa and eventually dropped to 162 Pa, as the equivalence ratio decreased from 0.90 to 0.45. The oscillation frequency of the ethanol atomization burner presents a migration characteristic; this is mainly due to thermal effects associated with changes in the equivalence ratio that increase/decrease the speed of sound in burnt gases, leading to increased/decreased oscillation frequencies. The trend of the change in flame heat release rate is basically like that of sound pressure, but the time-series signal of the flame heat release rate is different from that of sound pressure. It can be concluded that the reversible change in equivalence ratio will bring significant changes to the amplitude of combustion oscillations. At the same time, the macroscopic morphology of the flame will also undergo significant changes. The flame front length decreased from 25 cm to 18 cm, and the flame frontal angle increased from 23 to 42 degrees when the equivalence ratio increased. A strange phenomenon has been observed, which is that there is also sound pressure fluctuation inside the atomized air pipeline, and it presents a special square waveform. This study explored the equivalence ratio adjusting trends on ethanol combustion instability, which will provide the theoretical basis for the design of ethanol atomization burners.
Ground-based Interferometric Synthetic Aperture Radar (GB-InSAR) is a valuable technique for monitoring deformation of landslides. GB-InSAR can construct high-resolution 2D images in a range-doppler plane. However, it is difficult for geo-engineers who are unfamiliar with radar monitoring geometry to interpret the results. Geometric mapping method has been applied to the co-registration of terrain models and radar images for the deformation zonation, but the mismatching correction with GB-InSAR and terrain model is not fully investigated by existing researches. To address this issue, this paper proposes a method exploiting an optimum linear transformation based on ground control points (GCPs). The proposed method was evaluated on simulated data and a field campaign in an open pit mine. The deformation mapping result of the method was verified by a collapse event. The method proposed in this paper has an RMSE value ranging from 0.502 to 0.720m (@700m average monitoring range) when the average density of the terrain point cloud is 0.5m. Although this method cannot accurately evaluate the antenna footprint vector deflection parameters, it can meet the needs of coarse matching calibration in relatively flat slope sub-target areas of open-pit mines for engineering applications.
工程应用中的手势识别需要较高的实时性和准确性,而现场环境通常无法提供足够的计算能力,采用轻量化神经网络在解决了上述问题的同时,还能达到与深度神经网络相当的识别效果;为此,提出一种基于改进轻量化神经网络的手势识别方法;该方法改进用于手部关键点检测的ReXNet网络结构,以改善骨骼点的局部关注;同时将关键点检测损失函数MSE替换为Huber loss,以提升离群点的抗干扰性;实验环境搭建基于普通单目镜头捕获图像后,经YOLO v3手部识别模型和改进的ReX-Net 关键点检测模型,并根据约束手部骨骼关键点的向量角而定义的不同手势,最后达到实时检测的效果;改进模型在RWTH公开数据集上的测试结果表明,改进后的手势识别方法的检测准确度较改进前整体提升2.62%,达到了 96.18%,且收敛速度更快.
近年来,人工智能及其相关专业复合型人才的培养需求逐步增大.立足于应用型本科院校人工智能类人才培养需求,教学团队从教学理念、信息化教学平台与资源建设、线上下教学方案设计、项目化教学设计、过程化教学评价等多个维度出发,凝练出一套适用于应用型本科院校人工智能类课程的线上线下信息化教学思路,并提出了以师—生为主体、以培养一流工科应用人才为目标的创新教学理论与实践方法.所提出的理论和方法可以为高校人工智能类人才培养提供支撑,亦可为人工智能类相关课程或其他课程教学提供参考.
The determination to develop high-speed, efficient and versatile micro/nano photonic systems has inspired vast studies on photonic circuits. We demonstrate here a near-ultraviolet (NUV) monolithic multicomponent integrated circuit on Si substrate, including two non-suspended multiple quantum wells diodes (MQW-diodes) and an arc-shape waveguide. The two MQW-diodes which are fabricated by the same process can function as emitter and detector, respectively, and their roles can be switched with each other, because the InGaN/GaN multiple quantum wells, which are employed in the emitter to produce near-ultraviolet light, are also utilized by the detector for photodetection. The arc-shape waveguide which serves as a communication channel between the emitter and the detector can change the light propagation direction by 90°. Due to light confinement structure of the wafer, silicon removal and back GaN etching is avoided, which makes the devices more robust, as well as simplifying the processing flow. Two MQW-diodes can communicate with each other at 100 Mbps via the arc-shape waveguide, and the light signals of 200 Mbps from the emitter can also be detected by a commercial photodiode module via free space, which forms a three-dimensional NUV light communication system. This work paves the way toward comprehensive photonic integration for wide variety of potential applications.
针对风机设备油液渗漏影响风机正常运行亟需解决的对风机设备油污的识别问题,提出了一种基于改进深度学习的风机油污检测方法;基于深度学习在目标检测中的应用特点,对目标检测网络YOLOv5n(You Only Look Once v5n)进行改进,将原网络中的非极大抑制(NMS,non maximum suppression)替换为Soft-NMS,降低了网络的误检率,添加CA(Coordinate At-tention)注意力机制,增强了模型对目标的定位能力,改进原网络损失函数为α-IoU(Alpha-Intersection over Union)损失函数,提高了边界框检测的准确度;实验结果表明:模型平均精度提升了 8.1%,查全率提高了 19.1%,网络推理速度提高了 28.6%;改进后的模型能准确检测风机油污,有效解决了风机实际运行中油液渗漏所带来的问题.
为了提高图像的特征质量,保证最后提取到的特征高度精炼,提出了一种新的方法;该方法首先将低分辨率图像经过小波变换分解成高频分量和低频分量,并结合插值法进行插值,最后通过小波逆变换得到高分辨率图像来为后续的特征提取提供高质量的图片输入;接着,选取ResNet-50网络作为基础网络,将ECA模块与ResNet残差结构结合形成一个全新的ECA-Res-Net50模块,ECA模块具有的通道级的注意力机制,可以让整个网络更加专注于提取显著特征;经实验测试,该方法对于图像特征提取的质量有着明显的提升,均方误差下降可达6.65;结果表明,该方法可行有效,具有良好的工程应用前景.
发电厂厂区内违规吸烟易导致火灾、爆炸等事故,会带来巨大损失;针对电厂内人员违规吸烟行为检测精度不高的问题,提出一种基于改进YOLOv5s(You Only Look Once v5s)的电厂内人员违规吸烟检测方法;该方法以YOLOv5s网络为基础,将YOLOv5s网络C3模块Bottleneck中的3×3卷积替换为多头自注意力层以提高算法的学习能力;接着在网络中添加ECA(Efficient Channel Attention)注意力模块,让网络更加关注待检测目标;同时将YOLOv5s网络的损失函数替换为SIoU(Scylla Intersection over Union),进一步提高算法的检测精度;最后采用加权双向特征金字塔网络(BiFPN,Bidirectional Feature Pyramid Network)代替原先YOLOv5s的特征金字塔网络,快速进行多尺度特征融合;实验结果表明,改进后算法吸烟行为的检测精度为89.3%,与改进前算法相比平均精度均值(mAP,mean Average Precision)提高了 2.2%,检测效果显著提升,具有较高应用价值.
Action recognition is a challenging task of modeling both spatial and temporal context. Numerous works focus on architectures modality and successfully make worthy progress on this task. While due to the redundancy in time and the limit of computation resources, several works focus on the efficiency study like frame sampling, some for untrimmed videos, and some for trimmed videos. With the intent of improving the effectiveness of action recognition, we propose a novel Computational Spatiotemporal Selector (CSS) to refine and reinforce the key frames with discriminative information in video. Specifically, CSS includes two modules: Temporal Adaptive Sampling (TAS) module and Spatial Frame Resolution (SFR) module. The former can refine the key frames in the temporal space for capturing the key motion information, while the latter can further zoom out some refined frames in the spatial space for eliminating the discrimination-irrelevant structural information. The proposed CSS is flexible to be embedded into most representative action recognition models. Experiments on two challenging action recognition benchmarks, i.e., ActivityNet1.3 and UCF101, show that the proposed CSS improves the performance over most existing models, not only on trimmed videos but also untrimmed videos.
扇翼机兼具固定翼飞机和直升机的优越性能,通过在厚机翼前缘嵌入横流风扇代替传统固定机翼,可同时提供升力和推力,并会附加一定的低头力矩,使得扇翼机纵向高度控制和姿态控制存在强耦合.独特的动力结构使扇翼机具有较高的静稳定性,但扇翼机机载空间有限,基于此提出基于角速率陀螺的扇翼机简化配置控制技术研究,并分析了控制系统的频域和时域特性.通过非线性仿真对比分析,验证了简化配置控制方案的有效性,且纵向简化配置相较于全配置更适合扇翼机纵向的飞行控制.
针对输电线路横跨地域广,输电通道中隐患目标多的问题,提出了输电线路通道可视化分级预警模型.首先改进深度残差网络提取输入图像的多光谱信息,通过软阈值化来减少噪声影响,提高输电线路通道场景分析模型的准确度;然后利用YOLOv3目标检测算法构建输电线路通道隐患目标识别模型,针对隐患中的烟雾、施工车辆目标小的问题,采用难负样本挖掘策略,减少图片背景的影响,再根据输电线路通道的分级预警结构构建分级预警模型.研究结果表明,结合场景分析的输电线路通道可视化分级预警模型能够科学、准确地反映出输电线路通道的隐患预警状态,为输电线路运行维护工作提供指导.
针对现阶段实时监测发电厂内人员存在违规行为活动较为困难的问题,提出基于改进YOLOv5s网络的枪球联动行为边缘计算系统.该系统利用枪球联动提高监控画面中小目标的成像质量,降低检测算法对小目标的检测难度,提高检测能力;同时针对检测网络计算复杂的问题,对YOLOv5s检测模型进行轻量化改进,提出YOLOv5s-light轻量型检测网络,修改预置检测anchors以及BottleNeck模块,提出BN-SiLU-weight特征提取结构,降低模型计算难度,提高模型检测速度,优化边缘计算部署能力.实验结果表明,YOLOv5s-light模型在mAP仅下降0.5%的前提下,实现了模型参数下降30%,推理时间缩短21.4%,结合枪球联动可以实现mAP提高1.5%,满足实时快速的边缘检测需求.
This paper addresses the autonomous landing control of FanWing. FanWing has good slow flight capability that shares certain fundamental features of both a fixed-wing aircraft and a rotary aircraft results in short take-off and landing (STOL) capability and no stall with a large angle of attack. FanWing is an aircraft configuration that uses a simple cross-flow fan mounted in the wing to provide distributed propulsion and augmented wing lift at very low flying speeds, leading to the high coupling of the attitude control and the altitude control. A control solution based on the pitch angle rate is proposed to obtain the pitch attitude stability control. And the combination control of the elevator and the fan wing is proposed to actualize the stability control of the longitudinal attitude and height control.
针对UWB定位中的标签容量限制和通信冲突的问题,提出了一种基于洪泛机制准同步的改进定位方法;该方法通过设置主从基站,利用洪泛机制逐级传递准同步报文实现基站和标签间的时钟准同步,满足了DS-TWR方法下TOA算法对通信中时钟误差的需求,从而提高了测距精度;同时通过多Hash运算为标签基站对分配唯一时隙,有效提高了多标签情况下的通信成功率;实验结果表明,改进后的系统多标签条件下平均通信成功率提高了16%,标签能耗降低了30%,单位时间内获取数据量提高了27.8%,在提高系统通信效率的同时降低了能耗,具有较高的工程应用价值.
针对山火烟雾的检测存在由于监控范围广、发生频率不固定等造成的高成本问题,在边缘计算思维的启发下,提出了一个基于YOLOv5改进的适用于前端布设的轻量级识别网络.该方法针对YOLOv5模型过大的缺陷,通过修改网络结构,将融合了通道注意力机制CoordAttention的Ghostbottleneck模块与YOLOv5结合,提出一种改进型卷积神经网络CG-yolo识别网络.实验结果表明,CG-yolo相对于YOLOv5s算法速度提高了9.5%,查全率提升了1.8%,查准率仅损失1.7%,部署在NVIDIA的Jetson Nano边缘计算平台上时运行速度可以达到13fps,更好地满足了隐患监测的工程实际需求.
Metro Vehicle Door System (MVDS) is one of the most frequently used parts in rail transit. In this work, we propose an effective and visualized state detection technique of MVDS by utilizing feature extraction and matching, which is helpful to improve the safety and reliability of the door system. Five states of MVDS, the normal opening, opened, closing and closed state, as well as the anti-extrusion (also named as anti-drag) state, are to be detected in our work. In particular, we design an improved Speeded Up Robust Features (SURF) method to reduce the mismatching pairs by exploiting local space distance, which is crucial for practical usage. Furthermore, we present a scheme to detect the five states of MVDS mentioned above via our improved SURF. Finally, we conduct experiments in subjective and objective, and the resultant performance shows the effectiveness of our method.
图像的特征选择需要筛除大量噪声节点,在像素较高的图像内处理效率低下,为减少图像特征选择的处理时间,设计基于适用性骨干粒子群优化算法的图像特征选择方法.提取图像视觉特征参数,包括颜色参数、纹理参数以及形状参数.获取分类面的线性判别函数,建立最优超平面,得到满足约束条件的目标函数以及特征选择的适应度函数,基于适用性骨干粒子群引入粒子位置与速度更新机制,设计特征选择算法,得到一个新的图像特征选择处理方法.获取不同阈值以及不同学习效率下的最优特征数量,分别测试四种数据集内特征选择算法的运行时间,实验结果显示,在四种数据集内,适用性骨干粒子群优化算法的运行时间均小于其他算法,可见该算法为相同图像相同参数下的最优算法.
为了对物体表面温度实时在线测量并能实时获得红外热图像,论文设计了一种低成本的红外热成像系统.该系统以微处理器为控制核心,通过IIC接口接收MLX90640所采集到的温度数据,并实时显示测量区域的红外热图像,针对红外传感器分辨率低的问题,提出了一种基于边缘保护的新插值算法,可有效提高红外热图像的分辨率.实验结果表明,该系统可以实时获得高分辨率、高质量的热力图,准确定位发热点位置,具有良好的工程应用前景.
针对发电厂存在大型设备遮挡、跟踪算法易出现目标ID切换引起跟踪失败的问题,提出基于DeepSORT的改进算法.该算法采用YOLO v5s+DeepSORT实现目标跟踪,针对目标遮挡问题加入边缘撞线机制,设计遮挡补偿函数,缓解目标因遮挡转为删除态的问题.对目标特征保存提出基于置信度权重和关键帧的保存方法,优化特征库质量,提高算法的重识别能力.试验结果表明,改进后的跟踪算法IDS平均下降40.4%、IDF1平均提高16.1%、MOTA平均提高11.6%,MOTP平均提高2.1%,遮挡发生时ID切换次数明显减少,提高了跟踪系统的抗遮挡能力.