
Automatic white balance (AWB) is a crucially important part of digital still camera. It keeps constant color of an image by eliminating the color cast caused by non-canonical illuminant. A dynamic threshold is used to remove outliers in $$ C_{b} $$ and $$ C_{r} $$ components and detect the near-white region in an image. And we also describe a technique using both the internal illumination and all pixels in the near-white region to estimate the illuminant. The results show that the proposed technique is superior or comparable to the existing AWB algorithms. The algorithm is attractive for practical applications because of the low complexity.
Aiming at the problem that the object detection accuracy of SSD algorithm in warehouse environment is not too high, a warehouse object detection algorithm based on improved SSD that fusion of DenseNet and SSD is proposed (Dense-SSD). Firstly, a large number of images containing cargos, trays and forklifts in real warehouse environment are collected through the camera, and the collected images are labeled to building the warehouse object dataset. Further the based network of the improved algorithm based on SSD pipeline is adapted with DenseNet, the Dense-SSD is trained from scratch on the PASCAL VOC and self-built warehouse object dataset, respectively. Finally, the trained models are tested on the above two datasets respectively. Experimental results show that the proposed method can reach 77.62
Common target detection is usually based on single frame images, which is vulnerable to affected by the similar targets in the image and not applicable to video images. In this paper, anchor mask is proposed to add the prior knowledge for target detection and an anchor mask net is designed to improve the RPN performance for single target detection. Tested in the VOT2016, the model perform better.
This paper describes a teaching aid called ANIMA, which integrates the whole flow of animation design, and uses a virtual role and speech recognition technology to guide students to make animation. This paper examines the learning performance of two groups of participants: experimental group (ANIMA aids) and control group (traditional manufacturing tools). A quasi-experimental study was conducted on 44 college students. The learning performance of different production tools was evaluated, including learning effect and design thinking. The results showed that compared with the traditional animation tools, ANIMA can improve students’ learning effect and cultivate their complete design thinking.
In this paper, a multi-image encryption method is presented by using the Random Discrete Fractional Mellin Transform (RDFrMT) and the cascaded phase retrieval algorithm (PRA). Firstly, we define the RDFrMT and discuss its properties including nonlinear, multi-parameter, and randomness. Then, based on the RDFrMT, a cascaded multiple-image encryption scheme is proposed. In our algorithm, an image generated by logistic map is transformed to the RDFrMT domain, and the resulting image is fed into the PRA. This nonlinear process makes our cryptosystem robust for potential attacks and ensures its security. Thereafter, the output is cut based on the Hamming distance and the keys are related to multiple input images, which makes our cryptosystem has distinct security level for different images. We repeat the above steps and get a cascaded encryption scheme. Numerical simulations demonstrate the security and feasibility of the proposed scheme.
The Planar Interferometric Imaging System samples the object visibility in the Fourier domain and then digitally reconstructs an image. Firstly, the principle of SPIDER imaging is studied, the principle of AWG is analyzed, and the influence of different bandwidth on the visibility of interference fringes is simulated. The results show that increasing the number of spectral channels of AWG can improve the visibility of interference fringes, effectively increase the spatial frequency coverage and improve the imaging quality.
The geometric link model of integrated remote sensing imaging of star sensor and camera is proposed in order to improve the geometric positioning accuracy of remote sensing satellite image, reduce the error between satellite coordinate system and other coordinate systems, and increase the installation stability between star sensor and satellite platform. Based on the model, the beam adjustment algorithm of the star-sensing camera integrated remote sensing imaging geometric link model with additional distance constraints is proposed. The simulation results, using the initial adjustment parameters of ZY-3, show that the beam adjustment algorithm can effectively eliminate the systematic error component in the geometric positioning error and improve the geometric positioning accuracy of the remote sensing satellite image.
Background modeling method is one of the most commonly methods for target detection. Gaussian mixture model (GMM) is a widely used background modeling method which can get good performance in video of surveillance scenes. However, when the GMM is directly applied to the detection of small moving targets, it may cause problems such as incomplete and missing detection of the target contour. Therefore, we propose a new background modeling method by using skew normal mixture model (SNMM). A skew normal mixture model is established at each pixel position in frames of video. After updating the frames of video, the parameters of the background model SNMM are updated, and the detection of small moving target is performed. Experimental results show that the SNMM can obtain better contour of the small moving targets in videos than the GMM.
With the rapid development of augmented reality technology in the field of education, the availability of learning methods based on mobile AR technology has been verified. However, there is little research on the impact of mobile AR on Chinese character learning. This paper designed and developed a mobile AR Chinese characters sandbox game. Compared with the traditional Chinese characters learning applications, the game is more interesting and effective. The game combines 3D touch interaction, 2D interface interaction, image recognition-based AR interaction and so on. The game can intelligently recommend suitable learning content to learners. In the game, participants write Chinese characters by interacting with the real environment, while changing the virtual environment. The results of this study show that the game has a positive impact on learners’ learning of Chinese characters. It can greatly improve learners’ learning interest and motivation.
3D human pose estimation is a fundamental task in computer vision. However, most of the related works focus on recovering human joint positions, which provides sparse and insufficient pose information for many applications like 3D avatar animation. Therefore, this paper presents a deep network for recovering joint angles from 3D joint positions, which learns the prior dependence between them. We test the validity and robustness of our method. We also discuss some details in designing and training the network. Our method is simple, effective and extensive. It can be combined with work of 3D human pose estimation that predict 3D joint positions from image or depth data to produce more detailed and natural poses. It builds a map between two joint sets with different numbers of joints, which provides a framework to unify multiple datasets for human pose estimation with different annotation formats.
In recent years, deep learning has become a popular method for 3D object recognition with point clouds. In this work, we introduce a multi-resolution feature fusion convolution neural network using point cloud data for 3D object recognition. Experiments are conducted on ModelNet40 dataset. It achieves better accuracy with 86.9% for 3D object recognition on point cloud data through four different feature fusion. Experimental results have demonstrated the superior performance of the proposed multi-resolution feature fusion network.
With the growth in the elderly population, fall detection methods for the elderly are of great significance. In this paper, we propose a deep learning-based method for real-time fall detection continuous depth maps with Residual Fall Detection Network (RFD-Net). Our method incorporates feature extraction with fall detection. In feature extraction part, seven important features that accurately represent the body posture are extracted from the depth maps to reduce the computation load. In the fall detection part, a novel RFD-Net is proposed to recognize body posture for fall detection. Meanwhile, two other networks are developed to compare with RFD-Net. The experimental results show that the extracted features are good representative of the body posture, and our method delivers performance with a fall detection accuracy of 98.51%, which is higher than other related methods.
A new local feature extraction method (BSPL) is proposed and applied to heterogeneous image matching to solve the problem that the traditional SIFT features have poor matching performance in heterogeneous image matching. A number of improvements have been made to ensure that common features of heterogeneous images can be extracted efficiently. The gradient histogram-equalized image is used as the input matching image; The bilateral filtering is used to construct the scale space pyramid to replace the Gaussian filtering of the traditional SIFT, which can make the details such as the edges of the image better preserved; PCA-based LDB descriptor is used as feature expression to improve the robustness of feature expression. Experimental results show that the proposed local feature descriptor has rotation and scale invariance, and effectively improves the number of matching points, matching accuracy, matching precision and matching adaptability, which is an effective infrared and visible image matching method.
For the demand of domestic digital film industry to improve the quality and efficiency of special effects film and television works, This paper studies the transmission quality of virtual and real fusion and the fusion efficiency, and proposes a better video transmission method combined with virtual and real fusion technology, which made the digital video industry more complete. Firstly, the Hadoop platform is built as the basis of distributed transmission, and then the parallel distributed transmission method and the virtual and real fusion technology proposed in this paper are combined to complete the experiment. In this paper, the efficiency of this method is compared with other existing advanced transmission methods. The experimental results show that the efficiency and transmission quality are improved compared with other methods.
In automatically recognizing human faces, it is an important problem how to extract the effective features from the corrupted face. This paper propose a new face recognition algorithm based on fusion of global and local Gaussian-Hermite moments (GHMs). Firstly, in order to solve the interference of noise on features, we use the GHMs of face image as facial feature. Second, we construct the face image spatial pyramid to extract the global and local features of the face, and then we compute scatter-ratio to seclect highly discriminative feature. Lastly we use sparse representation classifier to improve the robust of algorithm. Experiments on ORL, FERET and Yale A face databases reveal that the accuracy of proposed algorithm is better than traditional algorithm, especially when the face images are corrupted by salt&pepper noise.
Image change detection is a process that analyzes images of the same scene taken at different times in order to identify changes that may have occurred between the multitemporal images. This letter proposes a remote sensing image change detection algorithm based on BM3D and PCANet. Firstly, the BM3D algorithm is utilized to remove the noise in the log-ratio image, then the gray level co-occurrence matrix (GLCM) and FCM algorithm are utilized to select the image patches which are used to train the PCANet model. Finally the pixels in the multitemporal images are classified by the trained PCANet model, the changed and unchanged pixels are combined to form the final change map. The experimental results obtained in this letter verify the effectiveness of the proposed algorithm.
We present MMRPet, a modular mixed reality pet system based on passive props. In addition to superimposing virtual pets onto pet entities to take advantages of physical interactions provided by pet entities and personalized appearance and rich expressional capabilities provided by virtual pets, the key idea behind MMRPet is the modular design of pet entities. The user can reconfigure limited modules to construct pet entities of various forms and structures. These modular pet entities can provide flexible haptic feedback and support the system to render virtual pets of personalized form and structure. By integrating tracking information from the head and hands of the user, as well as each module of pet entities, MMRPet can infer rich interaction intents and support rich human-pet interactions when the user touches, moves, rotates or gazes each module. We explore the design space for the construction of modular pet entities and the design space of the human-pet interaction enabled by MMRPet. Furthermore, a series of prototypes demonstrate the advantages of using modular entities in a mixed reality pet system.
The combination of global and partial features has been an effective method to improve the precision for Person Re-identification. However, illumination, camera angle and pedestrian pose, etc. still have adverse effects on the retrieval results. In particular, a lot of background and other redundant information is contained in the boundingbox. Meanwhile, the part-based solutions are imprecise on account of unbalanced partitioning. In order to minimize the impact of these factors on the retrieval results, we introduced the pixel, channel attention modules and middle layer supervision into the ReID system to aggregate person features. In this paper, we propose a novel architecture for Person Re-Identification, with the pixel and channel attention modules that are beneficial for feature extraction. Comprehensive experiments results on the mainstream datasets including Market-1501, DukeMTMC-ReId, CUHK03-labeled and CUHK03-detected show that our method achieves better results.
Audio-visual language is the language of film and television animation art. The scene scheduling in the audio-visual language of animation includes action design, the design of characters’ moving routes in space, and the arrangement and design between the roles and the background. Traditional teaching methods are not enough to provide virtual scenes for learners to practice. Therefore, this paper develops an immersive gamification teaching environment for the practical teaching of scene scheduling content in audio-visual animation language, which includes three aspects: role scene scheduling, shot scene scheduling, role scheduling and shot scheduling. Given this teaching system we have carried on the teaching pilot experiment, the experimental results showed that the immersive teaching environment could improve students’ learning effect, promote the construction of action program schemata, and promote the formation and maintenance of skills.
In the design of the traditional CNN model, there is always a balance between the spatial dimension and the number of channels. The high-dimensional spatial resolution is to preserve more detailed local information, while the large number of channels ensures more complex feature representation. For the current Light-field depth estimation algorithm, the designers usually choose to maintain a high spatial dimension in the network to improve the accuracy of the depth estimation, which results in a situation where the model has a large size and will take up huge amount of computing resources in the depth estimation process. In this paper, we introduce an effective pooling method: spectral pooling, to improve the problems of the original Light-field depth estimation network mentioned above. We transfer the light field feature map to the frequency domain through Fourier transform and reasonably reduce the size of the feature map in the frequency domain to an arbitrary size. The method can be used to reduce the complexity of the network and speed up training with similar performance to the original algorithm. More importantly, the new model with less memory demand can be better applied to light mobile device.