To solve the problems in current co-saliency detection algorithms, a novel co-saliency detection algorithm is proposed which applies fully convolution neural network and global optimization model. First, a fully convolution saliency detection network is built based on VGG16Net. The network can simulate the human visual attention mechanism and extract the saliency region in an image from the semantic level. Second, based on the traditional saliency optimization model, the global co-saliency optimization model is constructed, which realizes the transmission and sharing of the current superpixel saliency value in inter-images and intra-image through superpixel matching, making the final saliency map has better co-saliency value. Third, the inter-image saliency value propagation constraint parameter is innovatively introduced to overcome the disadvantages of superpixel mismatching. Experimental results on public test datasets show that the proposed algorithm is superior over current state-of-the-art methods in terms of detection accuracy and detection efficiency, and has strong robustness.
针对目前基于稀疏表示的显著性检测算法中存在的边界显著性检测不足、字典表达能力不够等问题,提出一种基于稀疏恢复与优化的检测算法。首先对图像进行滤波平滑和超像素分割,并从边界与内部超像素中挑选可靠的背景种子构建稀疏字典;然后基于该字典对整幅图像进行稀疏恢复,根据稀疏恢复误差生成初始显著图;再运用改进的基于聚类的二次优化模型对初始显著图进行优化;最后经过多尺度融合得到最终显著图。在三大公开测试数据集上的实验结果表明,所提算法能够保持高效快速、无训练等优点,同时性能优于目前主流的非训练类算法,在处理边界显著性方面表现优异,具有较强的鲁棒性。
The microcomputer interfacing technology is an important course to communication engineering students. A project-centric flipped classroom learning model is proposed to reform the course teaching paradigm. A Proteus project framework is introduced to organize and integrate lecture contents, and is implemented by the flipped classroom learning model which is composed of students’ activity loop, teacher’s activity loop, and problem driven loop in the classroom. Practical experiences show that the new learning model has improved the students’ performance.
This paper proposes a bottom-up saliency detection algorithm based on multi-dictionary sparse recovery. Firstly, the SLIC algorithm is used to segment the image into superpixels in multilevel and atoms with a high background possibility are selected from the boundary superpixels to construct the multidictionary. Secondly, sparse recovery of the entire image is achieved using multi-dictionary to get subsaliency maps from the perspective of sparse recovery errors. The final saliency map is generated in a weighted fusion manner. Experimental results on three public datasets demonstrate the effectiveness of our model.
In order to solve the fusion of space-time information and excessive detection area in pedestrian detection,a pedestrian detection method was proposed based on objectness and space-time covariance features.Firstly,binarized normed gradients algorithm is used for a test image to get objectness evaluations,and a pedestrian detection candidate area is formed.Secondly,the spatial and temporal features are extracted.Finally,a space-time detector based on cova-riance information was proposed to improve the accuracy.Experimental results on the INRIA and Caltech demonstrate that the proposed method outperforms the state-of-art pedestrian detectors in accuracy.
Due to its ability to model image corruptions explicitly, sparse representation (SR) attracts much attention from the community of visual tracking in recent years. However, the existing tracking approaches generally employ the ℓ 1 norm regularization method to achieve sparse recovery. As a result, they have to set the regularization parameter to an appropriate value, which is actually a crucial but difficult task in practice. To avoid this difficulty, we develop a new algorithm to solve the resultant SR problem under the framework of variational Bayesian inference. The algorithm can simultaneously estimate the sparse coefficients and other unknown parameters in an automatic manner. Consequently, using this algorithm to achieve SR involves little user intervention. In addition, we employ a new observation likelihood function to allow for the incorporation of a particle screening mechanism into the tracking process. Based on the above modifications, we develop a new algorithm for visual tracking based on sparse prototypes. Experimental results over various video sequences demonstrate the effectiveness and robustness of the algorithm.
Feature learning and metric learning are two important components in person re-identification (re-id). In this paper, we utilize both aspects to refresh the current State-Of-The-Arts (SOTA). Our solution is based on a classification network with label smoothing regularization (LSR) and multi-branch tree structure. The insight is that some middle network layers are found surprisingly better than the last layers on the re-id task. A Hierarchical Deep Learning Feature (HDLF) is thus proposed by combining such useful middle layers. To learn the best metric for the high-dimensional HDLF, an efficient eXQDA metric is proposed to deal with the large-scale big-data scenarios. The proposed HDLF and eXQDA are evaluated with current SOTA methods on five benchmark datasets. Our methods achieve very high re-id results, which are far beyond state-of-the-art solutions. For example, our approach reaches 81.6%, 96.1% and 95.6% Rank-1 accuracies on the ILIDS-VID, PRID2011 and Market-1501 datasets. Besides, the code and related materials (lists of over 1800 re-id papers and 170 top conference re-id papers) are released for research purposes.
In this paper we consider recovering non-negative sparse signals under heterogeneous noise from a Bayesian inference perspective. To induce sparsity and non-negativity simultaneously, we assign a rectified Gaussian scale mixture prior to the signal of interest. With such a prior, the signal posterior is analytically intractable. To handle this, we employ an approximate approach to simplify the inference process and obtain the marginal posteriors approximately. Moreover, to reduce the high computational cost, we use a conjugate gradient based scheme to implement the above process. Based on these efforts, we develop a novel recovery algorithm for the problem of interest. Results of numerical experiments demonstrate that the algorithm can achieve high recovery accuracy as well as low computational cost.
In order to improve the accuracy of pedestrian detection, we proposed an algorithm of multi-channel feature detection based on Discrete Cosine Transform (DCT). We use a two-layer convolution network for arranging the image information after DCT to build a new channel in the frequency domain. The channel can describe complex textures about pedestrian. Combined with the features of the histogram of gradient and the color space as well as the DCT frequency domain, a low cost multi-channel pedestrian detector has been trained based on the Adaboost algorithm. Experimental results on two datasets demonstrate that our model outperforms state-of-the-art methods and the effect is remarkable in low false positive per image.
The video stream is encapsulated to form a packet after coding, and transported to the receiving end through the network.The quality of the video sequence during transmission is affected by the network state.It's unavoidable to loss the packet when the network is violent jitter and unstable, thus resulting the damage of the video quality.The subjective-oriented perception video quality evaluation index is used to analyze the importance of the frames of the video sequence, so as to define the important level of different types of frames.From the result, it finds that P frames are more important than I frames, while I frames are more important than B frames for subjective-oriented perception.The resulting important level can provide the basis for unequal error protection and frame dropping strategy.
In order to solve the problem that the whole reference video is required for calculation in model PDMOSL, we focus on the impact of coding parameters and network conditions on the model, especially the influence of the QP, frame rate and packet loss rate. Based on our findings, we proposed a novel way to calculate PDMOSL without the need of reference video, and only two parameters are need. The proposed way can reduce the time complexity and is applicable to practical systems, especially for network nodes.
In view of the detection error caused by the target on the image boundary, this paper proposes a boundary saliency algorithm.We firstly conduct a multi-scale image segmentation at super-pixel level and compute the boundary discriminations to estimate its boundary saliency.Then reliable saliency seeds are obtained by pattern mining with the boundary saliency.Finally, the saliency maps are obtained by saliency propagation.We compare our model to other 18 saliency detection algorithms and extensive experiments on three datasets.The results show that the proposed model outperforms state-of-the-art methods.
针对现有视频检测算法在空间和时间显著度上一致性不足,提出了空时一致性模型.首先构造梯度流场,整合空间上的颜色对比度与时间上的目标运动信息.而后基于空时梯度流场构造全局对比度,综合局部对比度和全局对比度,得到初始检测结果.最后通过马尔可夫随机场,对其进行空时一致性优化,得到最终显著图.在3个公开数据集上的大量实验表明,所提算法检测性能较好,并且具有较强的鲁棒性.
In pedestrian detection,multi-channel feature detection has the defect of incomplete using features.In this paper,an algorithm of multi-channel feature detection based on discrete cosine transform (DCT) was proposed.We used a two-layer convolution network for arranging the image information after DCT to build a new channel in the frequency domain.This channel can describe complex textures about pedestrian.Combined with the features of the histogram of gradient and the color space as well as the DCT frequency domain,a low cost multi-channel pedestrian detector has been trained based on the Adaboost algorithm.Experiments on popular pedestrian databases show that the proposed method improves the accuracy of detection,and the effect is remarkable in low false positive per image.
This paper proposes a new spatio-temporal appearance feature named Phasic Maximal and Local Maximal Occurrence (PM-LOMO) representation for video-based person re-identification. To perform temporal alignment of the sequence, we selected the optimal period of walking cycle and divide frames into several phases based on the extreme points of the sequence's Flow Energy Profile (FEP). To describe the appearance of the video, we averaged the local maximal occurrence representations of all the frames in a same phase, and then we chose maximal features among the phases. Extensive experiments on public datasets demonstrate the better effectiveness of the proposed method compared with the state-of-art approaches.
This paper proposes a no-reference video quality assessment model by reducing the complexity of the human visual system(HVS).The characteristics of spatial domain and temporal domain of the videos are firstly extracted.Then multiweight convergence is conducted by simulating visual perception according to different granularity from fine-gained to coarsegrained of video local block,video frame,video segment,etc.Finally the feature vector of the whole video is achieved.The support vector regression(SVR) is taken as quality assessment tool in this algorithm.The quality assessment of the unknown video is obtained without reference after supervised training.The experiments we have done show that the algorithm is not only superior to all of the other no-reference quality assessment algorithms,but also can be compared to part-reference algorithms.
针对通信专业本科《计算机网络》课程实验教学中存在的问题,提出了融合式实践教学改革方案.利 用翻转课堂把分层渐进的实践任务与整个理论教学过程相融合,促进基本原理知识的理解吸收.利用开放实验和创新课题等综合性网络实践突破课程界限,着重培养网络思维与实践能力.近年的实施效果表明该方案为学生打好理论基础、提高综合应用能力起到了十分积极地作用.
To address the misjudgment caused by all boundaries of an image being equally and artificially selected as background in most of state-of-the-art models using background prior,this paper proposes an algorithm called weighted contrast optimization based on discriminative background.Firstly,a metric is constructed to roughly but objectively estimate a saliency map,which is used to choose a better background map.Based on this metric,a reliable background detection model is constructed through geodesic distance transformation after discriminating each boundary via Hausdorff distance.Then,the only background weighted contrast is improved into fore-background weighted contrast.Last,the final saliency map is obtained through weighted optimization framework.Extensive experiments on five public datasets demonstrate that the proposed algorithm outperforms state-of-the-art methods.
本文针对汇编语言教学效果不佳的现状,探讨了开展翻转课堂教学实践的可行性.首先结合C语言课程的先修内容,部署了学生自主学习阶段的内容.然后,结合计算机系统的运转机理,阐述了如何在课堂内化阶段与学生展开深入探讨和研究.最后,总结了本次教学改革实践的一些经验和教训.