Low-light image enhancement is a very challenging subject in the field of computer vision such as visual surveillance, driving behavior analysis, and medical imaging . It has a large number of degradation problems such as accumulated noise, artifacts, and color distortion. Therefore, how to solve the degradation problems and obtain clear images with high visual quality has become an important issue. It can effectively improve the performance of high-level computer vision tasks. In this study, we propose a new two-stage low-light enhancement network with a progressive attention fusion strategy, and the two hallmarks of this method are the use of global feature fusion (GFF) and local detail restoration (LDR), which can enrich the global content of the image and restore local details. Experimental results on the LOL dataset show that the proposed model can achieve good enhancement effects. Moreover, on the benchmark dataset without reference images, the proposed model also obtains a better NIQE score, which outperforms most existing state-of-the-art methods in both quantitative and qualitative evaluations. All these verify the effectiveness and superiority of the proposed method.
Crowd counting has been widely applied in various fields including social security, urban planning, and intelligent monitoring. A series of excellent fully-supervised crowd counting methods spring up and achieve great performance. Nevertheless, all of the fully-supervised methods deeply depend on large quantities of annotated crowd density maps. Collecting and annotating crowd images is time-consuming and expensive especially for highly dense crowds. In contrast, unlabeled crowd images can be acquired without having to make a great effort. However, it is challenging to effectively exploit unlabeled data for crowd counting. To this end, we propose a semi-supervised crowd counting method that aims to optimize the crowd counting models via exploiting large amounts of unlabeled crowd images. Firstly, we design an effective proxy task based on image patch counts statistics. Then, we present an end-to-end iterative learning strategy to train our semisupervised framework. To prove the effectiveness of our semisupervised method, we conducted various experiments on three benchmark crowd counting datasets. Experimental results demonstrate that our semi-supervised algorithm achieves competitive performance compared with the the-state-of-art semi-supervised crowd counting approaches. Furthermore, experimental results show that our method performs well on cross-dataset.
Crowd counting has received increasing attention due to its important roles in multiple fields, such as social security, commercial applications, epidemic prevention and control. To this end, we explore two critical issues that seriously affect the performance of crowd counting including nonuniform crowd density distribution and cross-domain problems. Aiming at the nonuniform crowd density distribution issue, we propose a density rectifying network (DRNet) that consists of several dual-layer pyramid fusion modules (DPFM) and a density rectification map (DRmap) auxiliary learning module. The proposed DPFM is embedded into DRNet to integrate multi-scale crowd density features through dual-layer pyramid fusion. The devised DRmap auxiliary learning module further rectifies the incorrect crowd density estimation by adaptively weighting the initial crowd density maps. With respect to the cross-domain issue, we develop a domain adaptation method of randomly cutting mixed dual-domain images, which learns domain-invariance features and decreases the domain gap between the source domain and the target domain from global and local perspectives. Experimental results indicate that the devised DRNet achieves the best mean absolute error (MAE) and competitive mean squared error (MSE) compared with other excellent methods on four benchmark datasets. Additionally, a series of cross-domain experiments are conducted to demonstrate the effectiveness of the proposed domain adaption method. Significantly, when the A and B parts of the Shanghaitech dataset are the source domain and target domain respectively, the proposed domain adaption method decreases the MAE of DRNet by 47.6% .
Due to the underexposure, the lack of details and the noise issues, Low-light images always have a high degree of degradation. In this paper, we thoroughly study the degradation mechanism of low-light images and design a pre-denoising 3D multi-scale fusion attention network (P3DMFE) with Retinex decomposition theory. This work is divided into three modules, firstly, The proposed three-branch decomposition module decouples the original space into three sub-spaces: reflection decomposition, illumination decomposition and noise decomposition, where the noise decomposition allows us to obtain the higher-quality reflection map and illumination map. Secondly, the 3D multi-scale fusion improvement module removes the noise map and performs image reshaping, structure restoration and detail restoration on the combined reflection map and illumination map. Thirdly, the Illumination improvement module provides a suitable illumination map. The experimental results show that the proposed P3DMFE can not only enrich the details and improve the brightness and contrast of low-light images, but also have a good denoising effect. Specifically, the proposed method can achieve 22.04 PSNR, 0.84 SSIM, 1250.4 LOE and 5.03 NIQE on LOL dataset, which are the best performance compared with some state-of-the-art methods. The experiments on common low-light datasets such as NPE, VV, MIT5K, MEF, LIME, and DICM also verify the good generalization ability and superiority of the proposed method.
The assessment of pathological Images plays a crucial role in cancer cure and research. The automatic evaluation method based on artificial intelligence can provide assistance for doctors and effectively improve the efficiency and accuracy of doctors’ diagnosis. However, the current methods based on artificial intelligence have the problem of low quality of collected pathological images, which has a negative impact on the prediction of intelligent algorithms. Due to the influence of the illumination deviation, the brightness distribution of the collected pathological images is uneven, and the noise generated by the uneven illumination will be mixed with the useful signal in the image, and some details are partially blurred or completely obscured. To tackle these problems, this paper proposes a LinkNet-based light field correction method for polarization correction of pathological images. This method takes advantage of Gaussian image characteristics to generate a large number of images with different center points and different brightness diffusion speeds to simulate the light field distribution map, and uses WSI (Whole-slides Image) as a standard unbiased pathological image to train a polarization correction model. Compared with traditional image enhancement methods, the method proposed in this paper is not only the best in terms of visual effect contrast, but also has very desirable effects in image light field correction and detail enhancement. The experimental results show that the method proposed in this article and other evaluation methods we use are the best in optional various index evaluations.
Abstract Crowd counting has become a noteworthy vision task due to the needs of numerous practical applications, but it remains challenging. State‐of‐the‐art methods generally estimate the density map of the crowd image with the high‐level semantic features of various deep convolutional networks. However, the absence of low‐level spatial information may result in counting errors in the local details of the density map. To this end, a novel framework named Multi‐level Feature Fusion Network (MFFN) for single image crowd counting is proposed. The proposed MFFN, which is constructed in an encoder–decoder fashion, incorporates semantic and spatial information for generating high‐resolution density maps of input crowd images. Skip connections are developed between the encoder and the decoder so that low‐level spatial information and high‐level semantic features can be combined by element‐wise addition. In addition, a dense dilated convolution block is placed behind the encoder, extracting multi‐scale context features to guide feature fusion by a channel attention mechanism. The model is trained by multi‐task learning; semantic segmentation supervision is introduced to enhance feature representation. Extensive experiments are conducted on three crowd counting datasets (ShanghaiTech, UCF_CC_50, UCF‐QNRF), and the results show that MFFN outperforms state‐of‐the‐art methods. In addition, sufficient ablation studies are performed to verify the effectiveness of each component in our proposed method.
Crowd counting plays a significant role in crowd monitoring and management, which suffers from various challenges, especially in crowd-scale variations and background interference issues. Therefore, we propose a method named depth and edge auxiliary learning for still image crowd density estimation to cope with crowd-scale variations and background interference problems simultaneously. The proposed multi-task framework contains three sub-tasks including the crowd head edge regression, the crowd density map regression and the relative depth map regression. The crowd head edge regression task outputs distinctive crowd head edge features to distinguish crowd from complex background. The relative depth map regression task perceives crowd-scale variations and outputs multi-scale crowd features. Moreover, we design an efficient fusion strategy to fuse the above information and make the crowd density map regression generate high-quality crowd density maps. Various experiments were conducted on four main-stream datasets to verify the effectiveness and portability of our method. Experimental results indicate that our method can achieve competitive performance compared with other superior approaches. In addition, our proposed method improves the counting accuracy of the baseline network by 15.6%.
In order to solve the problem of crowd occlusion and scale change in a single image,this paper proposes a crowd counting algorithm based on multi-column convolution neural network.The algorithm uses Convolutional Neural Network(CNN) with receptive fields of different sizes and the feature attention module to adaptively extract multi-scale crowd features.The deformable convolution is introduced to enhance the learning ability of spatial geometric deformation of the network and optimize the feature map,so as to generate a high quality density map.Experimental results on the Shanghai Tech and UCF_CC_50 datasets show that the algorithm can learn the mapping relationship between input images and crowd density maps,and has high counting accuracy and robustness.
Crowd counting plays an important role in crowd analysis and monitoring. To this end, we propose a novel method called Adaptive Weighted Crowd Receptive Field Network (AWRFN) for crowd counting to estimate the number of people and the spatial distribution of input crowd images. The proposed AWRFN is composed of four modules: backbone, crowd receptive field block (CRFB), recurrent block (RB), and channel attention block (CAB). Backbone utilizes the first ten layers of VGG16 to extract base features of input images. CRFB is a multi-branch architecture simulating a real human visual system for further obtaining refined and discriminative crowd features. RB generates strong semantic and global information by recurrently stacking convolutional layers with the same parameters. CAB outputs appropriate weights to supervise each channel of the feature maps output from CRFB, which uses the outputs of RB as guidance. Different from previous works using Euclidean Loss, we employ L1_Smooth Loss to train our network in an end-to-end fashion. To demonstrate the effectiveness of our proposed method, we implement AWRFN on two representative datasets including the ShanghaiTech dataset and the UCF_CC_50 dataset. The experimental results prove that our method is both effective and robust compared with the state-of-the-art approaches.
Single image crowd counting remains challenging primarily due to various issues, such as large scale variations, perspective and non-uniform crowd distribution. In this paper, we propose a novel architecture referred to Second-Order Convolutional Network (SOCN) to deal with this task from the perspective of improving the feature transformation capability of the network. The proposed SOCN applies a convolutional neural network as the backbone. We introduce three cascaded second-order blocks located behind the backbone to augment the family of transformation operations and increase the nonlinearity of the network, which can extract multi-scale and discriminative features. Furthermore, we design a context attention module (CAM) including dilated convolutions to assign weights to the score map of each second-order block for the purpose that the features which contribute to counting can be highlighted. We conduct various experiments on ShanghaiTeach1 and UCF_CC_502 datasets, and the results demonstrate the effectiveness of our method.
目的在视频监控和人群模式行为理解的重要应用中,识别分割场景中的集体行为仍然是一个极具挑战性的问题。在这项研究中,提出一种基于流形密度的集体聚类算法,能够识别具有任意形状和不同密度条件下的集体行为的局部和全局模式。方法受群体运动行为的流形拓扑结构启发,首先提出一种新的流形距离度量方式用于挖掘群体运动的深层行为模式。进一步定义了集体聚集密度的概念,并通过基于聚集密度的聚类算法识别具有局部一致性行为的群组,这种策略更适用于识别具有任意形状的聚类。同时考虑到子群组之间的复杂交互作用,引入层次聚集合并算法得到全局集体行为模式,可以有效地表征全局一致性关系。结果针对不同情况下的复杂场景,本文算法在集体视频监控数据集下的实验结果表明了其有效性和鲁棒性,相比于传统的聚类方法和标准经典算法,以平均误差(AD)和方差(VAR)作为评价指标来评价算法性能,本文方法将识别分割聚集行为群组的误差率结果控制在了0. 81和0. 99以内,相比许多经典方法有较大提升。同时在具有复杂流形结构及任意密度条件下的人群场景中能够取得精确有效的识别结果,解决了经典方法在该特殊场景下存在的缺点。结论本文针对已有方法在流形结构场景识别集体行为流向缺乏精确性和稳定性的描述和分析这一问题,提出了基于流形密度的群组聚集聚类识别算法,在多个复杂真实视频数据集中进行实验,证明了所提方法的有效性,并相比于已有方法具有更高的识别精度。
To estimate the crowd density map and count the crowd from a single image accurately is always a challenging task. With arbitrary perspective and random crowd density, occlusions, appearance variations and perspective distortions may occur. Some of current crowd counting methods are based on image cropping. And some popular deep learning models are difficult to optimize. In this paper, we propose a Dilated Multi-column Convolutional Neural Network architecture for crowd density estimation in still images improved from the MCNN model [1]. We also use the dilated layer and optimize the loss function to get better accuracy. The DMCNN model is lightweight, easy to train and has better fitting ability. Meanwhile the architecture is an end-to-end system and robust for images with different perspective or crowd density. Furthermore, the ground truth (density map) is generated based on our Perspective-Adaptive Gaussian Kernels which can better represent the heads of pedestrians. We conduct experiments on the WorldExpo'10 dataset, the ShanghaiTech dataset, the UCF_CC_50 dataset, and the mall dataset. The results show that our method achieves better estimation and is convenient to utilize. Our DMCNN model has a good practical application prospect.
Due to the fast-growing industry of intelligent vehicles the advanced driver assistance system (ADAS) has engrossed a lot of attention of the scholars. One of the biggest hurdles for new autonomous vehicles is to detect curvy lanes, multiple lanes, and lanes with a lot of discontinuity and noise. The purpose of this paper is to analyze the possibilities of image processing techniques for a computer vision application focusing on the problem of lane detection to enable traffic safety and driving comfort. The proposed algorithm is a combination of two sub-algorithms. The first sub-algorithm called Fuzzy Noise Reduction Filter (FNPF); removes the noise and smoothen the sequences of images received by the camera. While the second sub-algorithm aims to detect lane in normal as well as challenging scenarios by applying the concept of Hough Transform (HT) with a capable region of interest. The novelty of the proposed research study is the tracking of the lanes under inclement weather and challenging lightening conditions with improved computational time. The result achieved through our proposed algorithm is satisfactory in video sequences captured on several road types and under very challenging lighting and weather conditions.
The rapid development of convolutional neural networks (CNNs) is usually accompanied by an increase in model volume and computational cost. In this paper, we propose an entropy-based filter pruning (EFP) method to learn more efficient CNNs. Different from many existing filter pruning approaches, our proposed method prunes unimportant filters based on the amount of information carried by their corresponding feature maps. We employ entropy to measure the information contained in the feature maps and design features selection module to formulate pruning strategies. Pruning and fine-tuning are iterated several times, yielding thin and more compact models with comparable accuracy. We empirically demonstrate the effectiveness of our method with many advanced CNNs on several benchmark datasets. Notably, for VGG-16 on CIFAR-10, our EFP method prunes 92.9% parameters and reduces 76% float-point-operations (FLOPs) without accuracy loss, which has advanced the state-of-the-art.
Crowd counting is a challenging vision task which aims to accurately estimate the crowd count from a single image. To this end, we propose a novel architecture called De-background Detail Convolutional Network (DDCN) to learn a mapping from the input image to the corresponding crowd density map. DDCN focuses on removing the interference of background from crowds and reducing the mapping range from input to output. Such design optimizes the learning process to a large extent. The proposed DDCN is composed of three components: a decomposer, a feature extraction CNN and a regression CNN. Specifically, the decomposer produces a detail layer by subtracting the background interference from the crowd image. Feature extraction CNN works for extracting high level features and regression CNN is used to estimate the density map. In addition, a weighted Euclidean loss is designed to calculate the Euclidean distances of the crowd and the background separately with different loss weights, which further improves the counting performance. Extensive experiments were conducted on three crowd counting datasets to validate the performance of DCNN. And experimental results demonstrate that DDCN achieves performance improvements compared with the state-of-the-art.
Detecting coherent motion remains a challenging problem with important applications for the video surveillance and understanding of crowds. In this study, we propose the Density-based Manifold Collective Clustering approach to recognize both local and global coherent motion having arbitrary shapes and varying densities. Firstly, a new manifold distance metric is developed to reveal the underlying patterns with topological manifold structure. Based on the novel definition of collective density, the Density-based collective clustering algorithm is further presented to recognize the local consistency, where its strategy is more adaptive to recognize clusters with arbitrary shapes. Finally, considering the complex interaction among subgroups, a hierarchical collectiveness merging algorithm is introduced to fully characterize the global consistency. Experiments on several challenging video datasets demonstrate the effectiveness of our approach for coherent motion detection, and the comparisons show its superior performance against state-of-the-art competitors.
In recent years, crowd counting in still images has attracted many research interests due to its applications in public safety. However, it remains a challenging task for reasons of perspective and scale variations. In this paper, we propose an effective Skip-connection Convolutional Neural Network (SCNN) for crowd counting to overcome the issue of scale variations. The proposed SCNN architecture consists of several multi-scale units to extract multi-scale features. Each multi-scale unit including three convolutional layers builds connections between the input and each convolutional layer. In addition, we propose a scale-related training method to improve the accuracy and robustness of crowd counting. We evaluate our method on three crowd counting benchmarks. Experimental results verify the efficiency of the proposed method, and it achieves superior performance compared with other methods.
Crowd counting on still images is very challenging due to heavy occlusions and scale variations. In this paper, we aim to develop a method that can accurately estimate the crowd count from a still image. Recently, convolutional neural networks have been shown effective in many computer vision tasks including crowd counting. To this end, we propose a fully convolutional network (FCN) architecture to map the input image of arbitrary size or resolution to its density map. In order to address the perspective and scale variation issues, Inception-like modules with multiple kernel size filters are used to capture multi-scale features, which is necessary for higher crowd counting performance. We test our model on challenging ShanghaiTech dataset, the results show that our method outperforms the state-of-the-art methods.