In recent years, 3D Gaussian Splatting (3DGS) has garnered significant attention for its superior rendering quality and real-time performance. However, the inefficient utilization of Gaussians in 3DGS necessitates the use of millions of Gaussian primitives to adapt to the geometry and appearance of 3D scenes, leading to significant redundancy. To address this issue, we propose an efficient adaptive density control strategy that incorporates Cross-Section-Oriented splitting and Heterogeneous cloning operations. These modifications prevent the proliferation of redundant Gaussians and improve Gaussian utilization. Furthermore, we introduce opacity adaptive pruning, adaptive thresholds, and Gaussian importance weights to refine the Gaussian selection process. Our post-processing Gaussian refinement pruning further eliminates small-scale and low-opacity Gaussians. Experimental results on various challenging datasets demonstrate that our method achieves state-of-the-art rendering quality while consuming less storage space, reducing the number of Gaussians by up to 42% compared to 3DGS. The code is available at: https://github.com/zhiyu-cv/EGU.
Pedestrian behavior and trajectory prediction in highly dynamic and interactive scenes have emerged as among the most daunting challenges in the realm of autonomous driving. In addressing the modeling of pedestrian interaction and the generation of multimodal trajectories for pedestrian trajectory prediction, we present a novel approach: a context-based conditional variational generative adversarial network (Context-CVGN). This network is capable of capturing the physical environment, pedestrian interactions, and other scene elements by representing them as a bird's-eye view (BEV) semantic map. It can then infer various potential pedestrian trajectories in the future. By training and evaluating our model on the ETH&UCY dataset, we demonstrate superior performance compared to several state-of-the-art methods, particularly in terms of the final displacement error (FDE). These results substantiate the efficacy of our model in accurately predicting future pedestrian trajectories.
Pedestrian trajectory prediction in dynamic and strongly interactive scenes has become one of the most challenging problems in fields such as automated driving. In this paper, we propose Goal-CurveNet, a multimodal trajectory prediction network combining heterogeneous graph attention goal prediction and curve fitting. The model addresses the problems of pedestrian interaction modeling, multimodal trajectory prediction, and performance of predicted trajectories in pedestrian trajectory prediction. Goal-CurveNet can better model the historical trajectories and interaction behaviors in the scene systematically based on heterogeneous graph attention. It predicts the complete trajectories by curve fitting, which effectively improves the quality of predicted trajectories. The model architecture of “goal first and then trajectory” and targeted training paradigm also enhance the final performance. Through detailed training and testing on ETH & UCY datasets, we validate the effectiveness of each contribution of Goal-CurveNet. Compared to many state-of-the-art models, Goal-CurveNet achieves performance improvements in key metrics and effective prediction of pedestrian trajectories.
Compared to traditional traffic signal control methods, the method driven by Deep Reinforcement Learning (DRL) has shown better performance. But the problem of low sample utilization in reinforcement learning also arises. To deal with the problem, this paper presents a novel Twin Delayed Deep Deterministic Policy Gradient with Dual Buffer (TD3_DB) for traffic signal control. In the proposed framework, two experience buffers are used to store important samples and normal samples, separately, and the proportion of the two buffers is adjusted adaptively. In addition, lane pressure, describing the dynamic feature of lane traffic flow, is used for the state design of the TD3 agent, which enhances the perception of the agent toward intersections. Comprehensive experiments on different traffic flow modes has shown, the dual experience replay scheme can improve the sample utilization, and the proposed TD3_DB performs better than other methods such as original TD3, Proximal Policy Optimization (PPO), etc., effectively reducing vehicle queue length and waiting time.
The traffic scenes are complex and ever-changing, and traditional signal control methods have limitations. DRL based methods rely on a large amount of interactive data between intelligent agents and the environment, which leads to a long learning process and high computational costs. Based on this, this paper introduces imitative learning into reinforcement learning for traffic signal control. At the same time, it improves the design of reward function for the problem of frequent phase switching and too long red-light in the flow imbalance scene, and increases the cost of phase switching and penalty factors. Experiments verify the effectiveness of the algorithm in this paper.
In view of the complexity of UAV maneuvering decision calculation, the game decision algorithm of UAV maneuvering is studied based on receding horizon control. The time domain of UAV air combat is segmented, the maneuver decision problem is transformed into the optimal control sequence solution problem under receding horizon control. The state equation of the UAV system is determined to select the control variables in the maneuver process. The index function of air combat maneuvering decision is constructed and the situation superiority value of UAV is calculated to obtain the game matrix of UAV maneuvering situation and solve the Nash equilibrium solution of UAV maneuvering decision, realizing the maneuvering game decision of UAV. The feasibility and effectiveness of the constructed model is verified by simulation examples.
深度强化学习(deep reinforcement learning,DRL)可广泛应用于城市交通信号控制领域,但在现有研究中,绝大多数的DRL智能体仅使用当前的交通状态进行决策,在交通流变化较大的情况下控制效果有限.提出一种结合状态预测的DRL信号控制算法.首先,利用独热编码设计简洁且高效的交通状态;然后,使用长短期记忆网络(long short-term memory,LSTM)预测未来的交通状态;最后,智能体根据当前状态和预测状态进行最优决策.在SUMO(simulation of urban mobility)仿真平台上的实验结果表明,在单交叉口、多交叉口的多种交通流量条件下,与三种典型的信号控制算法相比,所提算法在平均等待时间、行驶时间、燃油消耗、C02排放等指标上都具有最好的性能.
特征点提取是图像处理领域的一个重要方向,在视觉导航、图像匹配、三维重建等领域具有广泛的应用价值;基于卷积神经网络的特征点提取方法是目前的主流方法,但由于传统卷积层的感受野大小不变、采样区域的几何结构固定,在尺度、视角和光照变化较大的情况下,特征点提取的精度和鲁棒性较差;为解决以上问题提出了一种结合多尺度与可变形卷积的自监督特征点提取网络;以L2-NET为网络骨干,在深层网络中引入多尺度卷积核,增强网络的多尺度特征提取能力,获得细粒度尺度信息的特征图;使用单应矩阵约束的可变形卷积以提取不规则的特征区域,同时降低运算量,并采用归一化约束单应矩阵的求解,均衡不同采样点对结果的影响,配合在网络中增加的卷积注意力机制和坐标注意力机制,提升网络的特征提取能力;文章在HPatches数据集上进行了对比试验和消融实验,与R2D2等7种主流方法进行对比,文章方法的特征点提取效果最好,相比于次优数据,特征点重复度指标(Rep)提升了约1%,匹配分数(M.s.)提升了约1.3%,平均匹配精度(MMA)提高了约0.4%;文章提出的方法充分利用了可变形卷积提供的深层信息,融合了不同尺度的特征,使特征点提取结果更加准确和鲁棒.
To further improve retinal vessel segmentation accuracy, we propose a deformable convolutional neural network based on cascade U-Net for retinal vessel segmentation: DCU-Net. The overall structure of DCU-Net is composed of two U-Net. We introduce deformable convolution to build a feature extraction module, which enhances the modeling ability of the model for vessel deformation. For improving the efficiency of information transfer between U-Net models, we use a residual channel attention module to connect U-Net. DCUNet achieves excellent results on public datasets. On DRIVE and CHASE_DB1 datasets, the Acc reaches 0.9568, 0.9664, respectively, the AUC reaches 0.9810, and 0.9872, respectively. From the experimental results, the residual channel attention module and residual deformable convolution module greatly improve the retinal vessel segmentation accuracy. The comprehensive performance of our method is better than that of some state-of-the-art methods.
For the deep learning based super-resolution (SR) reconstruction method, researchers try to expand the receptive field to improve the reconstruction quality. With the increase of network depth, the consumption of computing resources is impressive. Therefore, SR technology is challenging to apply to small mobile clients. To deal with this defect, we propose a lightweight multi-stage residual distillation network (MRDN) for the SR task in this paper. The model has made the following two main improvements: First, we design a multi-stage residual distillation block (MRDB). It combines the channel separation and the skip connection to reduce the parameter number and guarantee the network's performance by residual learning. Secondly, we propose efficient pixel attention (EPA) module, which weights different channels according to their importance so that the network can learn more details of pictures. Experiments show that the proposed MRDN is superior to the state-of-the-art models in subjective visual effects and objective evaluation criteria such as PSNR, SSIM, and computational complexity. Taking the basic data set manga109 as an example, the PSNR of this algorithm on scale x4 reaches 30.66, and the average inference time is 0.021 s. The performance surpasses RFDN, the champion of AIM 2020 Challenge on Efficient Super-Resolution.
Abstract Early detection of retinal vessel lesions by fundus image examination is essential to prevent and screen the related diseases. However, due to retinal vessels’ complex structure, it is time-consuming and laborious to extract retinal vessels manually by professionals. Therefore, retinal vessel automatic segmentation has great application value in modern medicine. In recent years, deep learning technology has achieved remarkable results in medical image segmentation, and the segmentation accuracy and speed are significantly improved compared with traditional methods. In this paper, we propose an efficient retinal vessel segmentation model ECU-Net. In this model, a feature extraction module combining edge attention module and efficient channel attention mechanism is designed to improve the vessel structure’s recognition ability. To improve the fusion efficiency of the texture information in the shallow layers and the semantic information in the deep layers, ECU-Net constructs an up-sampling fusion module. In addition, we design a multi-scale channel attention module to connect U-Net’s subnetwork to achieve efficient feature extraction. Compared with the state-of-the-art models, our ECU-Net has reached the highest level in some classical evaluation criteria and has a certain application value in the medical field.
To better extract feature maps from low-resolution (LR) images and recover high-frequency information in the high-resolution (HR) images in image super-resolution (SR), we propose in this paper a new SR algorithm based on a deep convolutional neural network (CNN). The network structure is composed of the feature extraction part and the reconstruction part. The extraction network extracts the feature maps of LR images and uses the sub-pixel convolutional neural network as the up-sampling operator. Skip connection, densely connected neural networks and feature map fusion are used to extract information from hierarchical feature maps at the end of the network, which can effectively reduce the dimension of the feature maps. In the reconstruction network, we add a 3×3 convolution layer based on the original sub-pixel convolution layer, which can allow the reconstruction network to have better nonlinear mapping ability. The experiments show that the algorithm results in a significant improvement in PSNR, SSIM, and human visual effects as compared with some state-of-the-art algorithms based on deep learning.
At present, the deep learning super-resolution (SR) method has achieved excellent results, but it also faces problems such as large models, high computational cost, a large amounts of training data, and poor interpretability. However, traditional machine learning-based methods still have room for improvement in feature extraction and model structure. This paper constructs a gradient embedding cascade forest structure on the basis of random forest and proposes a limit gradient embedding cascaded forest SR (LGECFSR) model. In feature construction, we not only adopt the first-order gradient, the second-order gradient, and other features of the image but also fuse the information of the original LR image. In addition, image blocks of different sizes are used for training, which increases the model’s generalization ability. Compared with the state-of-the-art machine learning-based methods, our method achieves the best performance and the second-best computational speed. In addition, compared with some deep learning-based methods, our model has a similar reconstruction effect and the best computational speed. In detail, for some reconstruction tasks, the Multi-Adds of LGECFSR is one-tenth to one-4000th of that of some current models. However, the SR performance of LGECFSR is the same or slightly better than that of some current classical algorithms.
In this paper, we propose a multi-scale convolution adaptive fusion super-resolution reconstruction network. Firstly, the input is passed through three convolution kernels of different sizes, and then the results are added and fused. Then, after pooling and full connection, the output results of the convolution layer with different sizes are weighted and added by the Softmax weighting mechanism to get the fusion feature map. Since the weights of the different branches with different convolution kernel sizes can be adaptively changed with the input information, the SR reconstruction is effectively improved. The detailed comparative experiments on the public datasets show that the SR reconstruction effect of our model is better than that of some state-of-the-art networks in objective criteria PSNR, SSIM, and subjective visual effect.
As a member of low-level visual tasks, image super-resolution (SR) is now mostly implemented by deep learning. Although the deeper convolution neural network can bring larger receptive field, it will increase the amount of calculation, make the training difficult and reduce efficiency. In addition, the feature information obtained by each channel plays a different and important role in the detail recovery during the SR process. To settle the above problems and improve the performance, we develop a multi-branch attention SR model. The main network contains multiple residual bodies, which are composed of several residual units. Additionally, we construct a multi-branch attention mechanism, which divides all channels into equal parts. Then, the network learns the relationship between channels and focuses more on the high-frequency feature channels. Experimental results show that the proposed algorithm is superior to the state-of-the-art algorithms in terms of subjective visual quality and objective evaluation criteria.
针对行人检测算法在交通场景下应用时的遮挡问题,提出一种结合双重注意力机制的遮挡感知行人检测算法.以RetinaNet作为基础框架,在回归和分类支路分别添加空间注意力和通道注意力子网络,增强网络对于行人可见区域的关注;同时引入行人可见边界框信息对传统的回归损失函数进行优化,使其能够随着遮挡程度自适应地调节预测框贡献的权重.在Caltech和CityPerson数据集上的实验结果表明:相较于RetinaNet等8种先进算法,该方法具有较好的鲁棒性和检测精度,尤其是严重遮挡情况下,该算法的对数平均漏检率仅为45.69%,小于其他算法12%以上;此外,该算法能够实现准实时检测,在Caltech和CityPerson上的检测速度分别为11.8帧/s和10.0帧/s.所提出的双重注意力机制和遮挡感知回归损失函数的检测方法具有可行性和有效性,对于遮挡行人的处理有显著优势.
In recent years, to improve the nonlinear feature mapping ability of the image super-resolution network, the depth of the convolutional neural network is getting deeper and deeper. In the existing residual network, the the residual block’s output and input are added directly through the skip connection to deepen the nonlinear mapping layer. However, it can not be proved that every addition is useful to improve the network’s performance. In this paper, based on Dirac convolution, an improved Dirac residual block is proposed, which uses the trainable parameters to adaptively control the balance of the convolution and the skip connection to increase the nonlinear mapping ability of the model. The main body network uses multiple Dirac residual blocks to learn the nonlinear mapping of high-frequency information between LR and HR images. In addition, the global skip connection is realized by sub-pixel convolution, which can learn to use linear mapping of low-frequency features of input LR image. In the training stage, the model uses Adam optimizer for network training and L1 as the loss function. The experiments compare our algorithm with some other state-of-the-art models in PSNR, SSIM, IFC, and visual effect on five different benchmark datasets. The results show that the proposed model has excellent performance both in subjective and objective evaluation.
目的 无监督单目图像深度估计是3维重建领域的一个重要方向,在视觉导航和障碍物检测等领域具有广泛的应用价值.针对目前主流方法存在的局部可微性问题,提出了一种基于局部平面参数预测的方法.方法 将深度估计问题转化为局部平面参数估计问题,使用局部平面参数预测模块代替多尺度估计中上采样及生成深度图的过程.在每个尺度的深度图预测中根据局部平面参数恢复至标准尺度,然后依据针孔相机模型得到标准尺度深度图,以避免使用双线性插值带来的局部可微性,从而有效规避陷入局部极小值,配合在网络跳层连接中引入的串联注意力机制,提升网络的特征提取能力.结果 在KITTI(Karlsruhe Institute of Technology and Toyota Technolog-ical Institute at Chicago)自动驾驶数据集上进行了对比实验以及消融实验,与现存无监督方法和部分有监督方法进行对比,相比于最优数据,误差性指标降低了10%~20%,准确性指标提升了2%左右,同时,得到的稠密深度估计图具有清晰的边缘轮廓以及对反射区域更优的鲁棒性.结论 本文提出的基于局部平面参数预测的深度估计方法,充分利用卷积特征信息,避免了训练过程中陷入局部极小值,同时对网络添加几何约束,使测试指标及视觉效果更加优秀.
针对现有的基于卷积神经网络的行人重识别方法所提取的特征辨识力不足的问题,提出了一种基于多尺度多粒度特征的行人重识别方法。在训练阶段,该方法在卷积神经网络的不同尺度提取特征;然后对获得的多尺度特征图进行分块和池化,从而得到不同尺度的全局特征和局部特征的多粒度特征,使用不确定性权重调节Softmax损失和三元组损失来对特征向量进行监督训练。在推理阶段,对所获得的多尺度多粒度的特征进行融合,使用融合特征在图像库中进行相似度匹配。在Market-1501和DukeMTMC-ReID数据集上的实验表明,所提方法相比基准网络ResNet-50在Rank-1评价指标上分别提升了4.3%和3.6%,在mAP评价指标上分别提升了6.2%和6.6%。实验结果表明,所提方法能够增强提取特征的辨识力,提高行人重识别的性能。
At present, the main super-resolution (SR) method based on convolutional neural network (CNN) is to increase the layer number of the network by skip connection so as to improve the nonlinear expression ability of the model. However, the network also becomes difficult to be trained and converge. In order to train a smaller but better performance SR model, this paper constructs a novel image SR network of multiple attention mechanism (MAMSR), which includes channel attention mechanism and spatial attention mechanism. By learning the relationship between the channels of the feature map and the relationship between the pixels in each position of the feature map, the network can enhance the ability of feature expression and make the reconstructed image more close to the real image. Experiments on public datasets show that our network surpasses some current state -of-the-art algorithms in PSNR, SSIM, and visual effects.