
In case of a crime or terrorist attack, nowadays much video footage is available from surveillance and mobile cameras recorded by witnesses. While immediate results can be crucial for the prevention of further incidents, the investigation of such events is typically very costly due to the human resources and time that are needed to process the mass data for an investigation. In this paper, we present an approach that creates a 4D reconstruction from mass data, which is a spatio-temporal reconstruction computed from all available images and video footage. The resulting 4D reconstruction gives investigators an intuitive overview of all camera locations and their viewing directions. It provides investigators the ability to view the original video or image footage at any specific point in time. Combined with an innovative 4D interface, our resulting 4D reconstruction enables investigators to view a crime scene in a way that is similar to watching a video where one can freely navigate in space and time. Furthermore, our approach augments the scene with automatic detections and their trajectories and enrich the crime scene with annotations serving as clues.
Accurate and fast face detection is a crucial step in forensic applications such as surveillance, facial fugitives recognition, and child sexual abuse detection. Several deep-learning-based methods addressed the face detection problem with high accuracy but they require a large computation power and processing time. Although GPUs have been used to speed-up computations, nowadays multiple generations of GPU architectures are available (e.g. Turing, Pascal or Keppler) making difficult to choose the most appropriate for face detection. In this work, we evaluate the speed-accuracy trade-off of three deep-learning-based face detectors in various CPUs/GPUs considering images reduced to different proportions as input. We successfully used this image resizing strategy in one previous work to improve the performance of face detection with Child Sexual Exploitation Material (CSEM). The results showed that the best speed-accuracy trade-off is achieved using the Pascal and Turing GPUs with images reduced to 50% of its original size.
Biometric systems can be subverted using presentation attack artefacts. This work presents a way to deal with the vulnerability to such spoofing attacks. In this work we propose the use of pupillary movements to detect such presentation attacks. The pupillary movements were stimulated by presentation of a moving visual challenge to ensure that some pupillary motion can be captured. Photo, 2D mask and 3D mask attack artefacts were evaluated based on data captured from 80 volunteers performing genuine attempts and spoofing attempts. The results indicate the effectiveness of the proposed pupillary movement feature to stop presentation attacks.
Human Interactive Proofs (HIPs) are an important tool in eCommerce, science and engineering, since they enable secure separation or identification of humans and computer bots. Although, many HIP techniques have been introduced and deployed in practical applications over the last two decades, the challenge of finding a waterproof method remains open. In this paper, we introduce a novel strategy to produce HIPs exploiting Brauer configuration algebras theory. The proposed approach is based on the generation of images of shapes that can be identified by humans but hard to be recognised by a computer program. Several experimental results are reported to demonstrate the robustness and feasibility of the proposed approach.
Emotion recognition has been one of the stimulating issues over the years due to the irregularities in the complexity of models and unpredictability between expression categories. So many Emotion detection algorithms have developed in the last two decades and still facing problems in accuracy, complexity and real-world implementation. In this paper, we propose two feature extraction techniques: Mouth region-based feature extraction and Maximally Stable Extremal Regions (MSER) method. In Mouth based feature extraction method mouth area is calculated and based on that value the emotions are classified. In the MSER method, the features are extracted by using connecting components and then the extracted features are given to a simple ANN for classification. Experimental results shows that the Mouth area based feature extraction method gives 86% accuracy and MSER based feature extraction method outperforms it by achieving 89% accuracy on DEAP. Thus, it can be concluded that the proposed methods can be effectively used for emotion detection.
An investigation on the spreading of blood droplets in fabric, as well as finger formation around the rim of the spreading stain immediately after droplet impact was done. Blood droplets were release perpendicular to the fabric substrate, at varying impact velocities (0.44–4.4 m/s), and different substrate backing material (air, leather, aluminum). A camera was used to record the impact spreading, and formation of fingers. The maximum spreading diameter and the number of fingers formed scaled with the Reynolds and Weber number, respectively. Both increased with increasing impact velocity, but slightly decreased with harder backing material. It was also observed, that the number of fingers approaches a constant maximum value so that it remains constant even if impact velocity is further increased, deviating slightly from predictive models.
Urban vitality is closely related to the built environments such as street architectures in the city. We propose a vitality evaluation system which calculates the amount of human activities under the use of mobile phone signaling Origin-Destination (OD) data. Meanwhile, because of the strong correlation between urban safety and urban vitality, the potential security risks of street architectures of different categories at different moments in a day are able to be appraised by our method. Experiments on the database of the street view images and mobile OD data in Nanjing, China validate the practicality and accuracy of our method. Moreover, we study the distribution of Points of Interest (POI) nearby street architectures of different vitality. A series of heuristic conclusion which will help optimizing the planning of street facilities are drawn from our study.
The pursuit to develop an effective people management system has widened over the years to manage the enormous increase in population.Any management system includes identification, verification and recognition stages.Iris recognition has become notable biometrics to support the management system due to its versatility and non-invasive approach.These systems help to identify the individual with the texture information distributed around the iris region.Many classification algorithms are available to help in iris recognition.But those are very sophisticated and require heavy computation.In this paper, an improved Kohonen selforganizing neural network (KSONN) is used to boost the performance of existing KSONN.This improvement is brought by the introduction of optimization technique into the learning phase of the KSONN.The proposed method shows improved accuracy of the recognition.Moreover, it also reduces the iterations required to train the network.From the experimental results, it is observed that the proposed method achieves a maximum accuracy of 98% in 85 iterations.
The extensive coverage of surveillance camera networks has supported the ever-growing research of vehicle re-identification (re-ID) due to their significant applications in matching and tracking vehicles-of-interest. The inherent challenging characteristics such as intra-class variance and inter-class similarity make the re-identification one of the most difficult tasks in computer vision. In this paper, we proposed a novel approach for vehicle re-id based on multi-block features. It implements the idea of information fusion from intermediate levels of representation and multi-stage supervision into a fully convolutional neural network. To demonstrate the effectiveness and superiority of our approach we perform extensive experiments and analysis on two standard vehicle re-id benchmarks.
The everyday life of a prison is not only shaped by the objective of resocialisation but also, as a rule, by aspects of security, which have the highest priority. A particular challenge is the prompt detection of potentially dangerous behavior. In the face of scarce human resources and psychologically limited attention, the use of video live streams is not an optimal solution. An alternative are 3D sensors, which enable the automatic and robust detection of such behavior in detention rooms and other areas such as hospital wards or workshops, while respecting the privacy of the monitored people. In this paper we present our research on this matter, based on realistic data that was acquired in an Austrian prison over 3.5 months. We discuss the recording setup and resulting dataset, and present algorithms for detecting selected behaviors. The experimental results show that these behaviors can be detected reliably, demonstrating that automatic behavior analysis in 3D data is a promising means for supporting the security personnel.
Use of video surveillance cameras in public space is the recent solution to control vandalism acts and emergency incidents. Such type of incidents requires an urgent and an appropriate action by security personnel. Key problem for security personnel is to manage and monitor high volume of visual data. Since last decade, human detection for video surveillance systems is an emerging research area. There are several crucial factors that effect the performance of human/object detection, such as illumination changes, background clutter, dynamic background, occlusion, and camera orientation etc. In this paper, a hybrid approach is presented for the localization and detection of person inside a moving train with challenging environment. Our proposed framework contains two modules i.e. human localization and human detection. We proposed a hybrid approach using GMM background modeling for foreground extraction followed by head and face detection to be used as a clue for human detection. Along with head and face detection, Histogram of Oriented Gradient (HOG) feature representation is used for human localization. For detection, ensemble classifier outperforms SVM and KNN classifiers on BOSS Dataset (On Board Wireless Secure Video Surveillance) and achieved 90% accuracy.
In this paper, we present a method for identifying and tracking illegally parked vehicles. This approach is based on deep learning for vehicles detection and hand crafted descriptors for the tracking which are designed to cope with occlusions. The tracking of the parked vehicle is achieved by key-point extraction of the detected vehicles and feature point matching. For each frame, a bounding box was generated to represent the vehicle and feature points extracted in that area. All parked vehicles have a unique ID which was generated by the Hungarian algorithm and Kalman filter, and the parked vehicle with the same ID was matched frame by frame. Based on this matching result, the stationary vehicles in the forbidden area can be tracked. Our approach tested efficiency and robustness on a public database and is shown to produce state of the art results.
Biometric person recognition using EEG signals has received considerable attention in recent years. This paper proposes a new feature based on the co-activation of EEG sensors. A visual representation of this co-activation feature is used to illustrate the identity-bearing nature of the proposed feature. The DEAP database was used to evaluate the proposed feature which was presented in the form of a visual signature indicating the spatial correlations around the scalp of EEG signals for an individual. The results show a high identification accuracy irrespective of the emotional state of the data subjects.
Text recognition can be used to retrieve textual information embedded in images. This task can be complex due to the low-resolution of the images and the orientation of the text, which are problems commonly found in Tor darknet images. In this work, we combine three different super-resolution algorithms, together with a rectification network, to increase the performance in the text recognition step. We evaluated these combinations in four state-of-the-art datasets, and in TOICO-1K, a Tor-based image dataset which was semi-automatically labelled for the task of Text Spotting in Tor darknet. We obtained the highest performance increase in ICDAR 2015 dataset, with an improvement of 3.77% when combining Residual Dense and the rectification networks. In TOICO-1K, we achieved a 3.41% of improvement when we combined Deep CNN and the rectification network. Our conclusion is that rectification performs slightly better than super-resolution when they are applied standalone, while their combination obtains the best results in the datasets evaluated.
We propose a semi-automatic progressive enhancement method for latent fingerprints. This method requires three inputs; a latent image, a manual segmentation, and an initial block location. The method starts enhancing the initial block using a matched filter in the frequency domain, then the enhanced block is padded back on the input latent image. The padded image is fed back as an input image and the surrounding blocks of the enhanced block are progressively enhanced. The proposed method performs iterative enhancement and feedback until the enhanced blocks fill the entire segmentation. The proposed method is benchmarked against several state-of-the-art methods with the NIST SD27 database. The experimental results show that the proposed method outperforms other methods in terms of identification accuracy.
Footwear impressions are commonly found at crime scenes and are therefore a valuable source of evidence for criminal investigations. Forensic experts can show that a footwear impression was made by a specific shoe or impressions at different crime scenes were made by the same suspect by comparing individual characteristics of the impressions. However, this process is very time consuming, therefore automated solutions are desired. Yet, testing and training such methods requires datasets that on the one hand reflect real data from criminal cases and on the other hand provide ground truth information. To solve this, we created an acquisition line and captured footwear impressions of 300 different pairs of shoes under varying conditions with the help of the Austrian police. In this work the creation of this dataset and the dataset itself is described in detail.
Solving Single-Shot Person Re-Identification (Re-Id) by training Deep Convolutional Neural Networks is a daunting challenge, due to the lack of training data, since only two images per person are available. This causes the overfitting of the models, leading to degenerated performance. This paper formulates the Triplet Permutation method to generate multiple training sets, from a certain re-id dataset. This is a novel strategy for feeding triplet networks, which reduces the overfitting of the Single-Shot Re-Id model. The improved performance has been demonstrated over one of the most challenging Re-Id datasets, PRID2011, proving the effectiveness of the method.
Person re-identification (re-ID) is a very active area of research in computer vision, due to the role it plays in video surveillance. Currently, most methods only address the task of matching between colour images. However, in poorly-lit environments CCTV cameras switch to infrared imaging, hence developing a system which can correctly perform matching between infrared and colour images is a necessity. In this paper, we propose a part-feature extraction network to better focus on subtle, unique signatures on the person which are visible across both infrared and colour modalities. To train the model we propose a novel variant of the domain adversarial feature-learning framework. Through extensive experimentation, we show that our approach outperforms state-of-the-art methods.
In this paper we study 3D pedestrian motion direction estimation using a projective geometry based method. The idea for estimation is based on image sequence analysis, where the detection of both head and toe points will provide measurements for vanishing point estimation. Knowing that once the walk direction is fixed, these two points move along 3D parallel lines. This estimation assumes that there is a binary segmentation of the pedestrian. For each frame, the top and the bottom points of the segmentation are extracted. One line is fitted to the collection of top points and the other to the collection of the bottom points. Given these two lines, the vanishing point is estimated and thence the direction of the pedestrian. Our approach either assumes two intrinsic parameters cameras that are known or estimated at a learning stage using two previously known walking pedestrian directions. Experiments using a publicly available database show an accurate estimation of the direction of the body with less than 3 degrees error.
With the growing amount of pornography content over Internet and cases of Child Sex Abuse (CSA) material possession and distribution, there is a rising demand for automatic detection of such content especially in certain environments such as educational or work places. The contribution of this paper is three fold. First, we present a critical review of automatic pornography and CSA detection in images and videos. Second, we provide an empirical evaluation of five selected pornography detection approaches representing traditional skin detection based as well as more recent deep learning based methods. The evaluations are performed under common criteria using two publicly available pornographic databases. Finally, we assess these methods on a dataset of real-world CSA material provided by Spanish Police Forces. This study observes that for pornography or CSA detection, the methods involving multiple features perform better than those using simple features like skin color or single image descriptor. It is also found that deep learning based methods outperform all of the other methods and report current state-of-the-art.