[Abstract] This paper has proposed an algorithm that detecting for dense small vehicle in large image efficiently. It is consisted of two Ensemble Deep-Learning Network algorithms based on Coarse to Fine method. The system can detect vehicle exactly on selected sub image. In the Coarse step, it can make Voting Space using the result of various Deep-Learning Network individually. To select sub-region, it makes Voting Map by to combine each Voting Space. In the Fine step, the sub-region selected in the Coarse step is transferred to final Deep-Learning Network. The sub-region can be defined by using dynamic windows. In this paper, pre-defined mapping table has used to define dynamic windows for perspective road image. Identity judgment of vehicle moving on each sub-region is determined by closest center point of bottom of the detected vehicle's box information. And it is tracked by vehicle's box information on the continuous images. The proposed algorithm has evaluated for performance of detection and cost in real time using day and night images captured by CCTV on the road.
Minutiae used in most fingerprint recognition devices is robust to presentation attack, but generates a high false match rate. Thus, it is applied along with orientation map or skeleton images. There has been plenty of research on security vulnerability of minutiae, whereas few research has been conducted on orientation map or skeleton images. This study analyzes vulnerability of presentation attack for skeleton images. For this purpose, it proposes a new algorithm of recovering fingerprints with the use of machine learning and skeleton image features of fingerprints. In the proposed method, we suggest the new machine learning Pix2Pix model to generate more natural images. The suggested model is developed in the way of adding a latent vector to the conventional image-to-image translation model Pix2Pix. In the experiment, fingerprints were recovered with the use of the proposed Pix2Pix model, and it was found that a fingerprint recognition device which recognized the recovered fingerprints had a high success rate of recognition. Therefore, it was proved that a fingerprint recognition device using skeleton images as well was vulnerable to presentation attack. It is expected that the algorithm proposed in this study will be very useful to many different application areas related to image processing, including biometrics, fingerprint recognition and recovery, and image surveillance.
This paper proposes a new method for estimating the symmetric axis of a pottery from its small fragment using surface geometry. Also, it provides a scheme for grouping such fragments into shape categories using distribution of surface curvature. For automatic assembly of pot from broken sherds, axis estimation is an important task and when a fragment is small, it is difficult to estimate axis orientation since it looks like a patch of a sphere and conventional methods mostly fail. But the proposed method provides fast and robust axis estimation by using multiple constraints. The computational cost is also too lowered. To estimate the symmetric axis, the proposed algorithm uses three constraints: (1) The curvature is constant on a circumference CH. (2) The curvature is invariant in any scale. (3) Also the principal curvatures does not vary on CH. CH is a planar circle which is one of all the possible circumferences of a pottery or sherd. A hypothesis test for axis is performed using maximum likelihood. The variance of curvature, multi-scale curvature and principal curvatures is computed in the likelihood function. We also show that the principal curvatures can be used for grouping of sherds. The grouping of sherds will reduce the computation significantly by omitting impossible configurations in broken pottery assembly process.
As communication medium of information, speech is not only used a lot, but also is the most comfortable. When we have conversation by speech, transmission of the information, which wanted to be delivered, is affected by the noise level. In speech signal processing, speech enhancement is using to improve speech signal corrupted by noise. Usually noise estimation algorithm need flexibility for variable environment and it can only apply on silence region to avoid effects of speech signal. So we have to preprocess finding voiced region before noise estimation. we proposed SNR estimation method for speech signal without silence region. For unvoiced speech signal, vocal track characteristic is reflected by noise, so we can estimate SNR by using spectral distance between spectrum of received signal and estimated vocal track. The proposed estimation method on voiced speech and the method by using voiced/unvoiced region energy are operated with simple logic as time domain method. And the estimation method on unvoiced region is possible to estimated noise level for narrow-band speech signal by using vocal track properties. It can be applied to rate decision of vocoder and used for pre-processing to decide threshold of noise reduction.
Voice conversion is a method that aims to transform the input speech signal such that the output signal will be perceived as produced by another speaker. Speech synthesizers using voice conversion technologies allow developers to create more voices from a single database and users to personalize the synthesizer to speak with any desired voice after a training period. In this paper, we present the method that converts time and pitch scaling using spectral mapping and PSOLA technique with OLA. This new synthesis scheme allows very flexible modifications of the pitch-scale, the time-scale and the spectral envelope characteristics while producing high-quality speech output. This synthesis scheme is thus well suited to voice conversion. Further work will be conducted on a matching method to correspond well with each phonetic information, and larger corpora to assess the robustness of the method.
The guidance system proposed in this paper aims to complement the white cane by monitoring road conditions in a medium range for blind pedestrians in real time. The system prototype employs only one webcam fixed at the waist of user. One of the main difficulties of using a single camera in outdoor obstacle detection is the discrimination of obstacles from a complex background. To solve this problem, this paper re-formulates top-view mapping as an inhomogeneous re-sampling process, so that background edges are sub-sampled while obstacle edges are oversampled in the top-view domain. Morphology filters are then used to enhance obstacle edges as edge-blobs, which are further represented using a directional ellipse as a new model for obstacle classification. Based on the identified obstacles, safe walking area is estimated by tracking a polar edge-blob histogram. To transfer the information obtained from image domain to language domain, this paper proposes a verbal message generation scheme based on fuzzy logic. The efficiency of the system is confirmed by testing the system with visually impaired people on outdoor pedestrian paths.
This paper presents a travel-aid system for sight-impaired pedestrians in outdoor environment.Unlike many existing travel-aid systems that rely on stereo-vision,the proposed system aims to detect generic obstacles in a cluttered road environment by using just the single camera attached to user’s body.To achieve this goal,the original image is re-sampled and mapped to a bird-eye view virtual plane,on which background edges are sub-sampled while obstacle edges are oversampled.Morphology filters with connected component analysis are used to enhance obstacle edges as edge-blobs with larger size,whereas sparse edges from background are filtered out.Based on the identified obstacles,safe walking areas are estimated by calculating a polar edge histogram.The algorithm is tested in different sidewalk scenes with complex pavements,and its efficiency has been confirmed.
This paper presents a lane-mark detection method which can extract lane-mark positions accurately in a crowded road environment. Many existing lane detection methods rely on the fitting of lane model among a great deal of outlier feature-points, while the proposed method aims at removing those outliers as much as possible at feature extraction stage. To achieve this goal, a multi-channel Haar-like filter is firstly used to get possible lane-mark patterns on a top-view image of the road, and then a directional feature-point tracing algorithm is proposed to get feature-point components, which are further classified as lane-mark or non lane-mark components via geometric properties. Based on the extracted lane-mark components, a tangent circle model is fitted by means of circle template matching. Experimental results show that the proposed method is effective in identifying lane-mark patterns among heavy clutters in crowded road environment.
Active appearance model(AAM)is an efficient method for the localization of facial feature points,which is also useful for the subsequent work such as face detection and facial expression recognition.In this paper,we mainly discuss the AAMs based on principal component analysis(PCA).We also propose an efficient facial fitting algorithm,which is named inverse compositional image alignment(ICIA),to eliminate a considerable amount of computation resulting from traditional gradient descent fitting algorithm.Finally,3D facial curvature is used to initialize the location of facial feature,which helps select the parameters of initial state for the improved AAM.
This paper proposes a new method for estimating the symmetric axis of a pottery from its small fragment using surface geometry. For the automatic assembly of broken sherds, the axis estimation is an important measure [2]. When a fragment is small, it is difficult to estimate axis orientation since it looks like a patch of a sphere and conventional methods mostly fail, but the proposed method provides reliable axis estimation by using multiple constraints. The computational cost is also much lowered. To estimate the symmetric axis, the proposed algorithm uses three constraints: (1) the curvature is constant on a circumference C H . (2) the curvature is invariant in any scale. (3) also the principal curvatures does not vary on C H . C H is a planar circle which is one of all the possible circumferences of a pottery or sherd. A hypothesis test for axis is performed using maximum likelihood. The variance of curvature, multi-scale curvature and principal curvatures are computed in the likelihood function. We also show that the principal curvatures can be used for grouping of sherds. The grouping of sherds will reduce the computation significantly by omitting impossible configurations in pottery assembly.
본 논문은 비마커 증강현실(Marker-less Augmented Reality)을 위한 색상 및 깊이 정보를 융합한 Mean-Shift 추적 알고리즘 기반 손 자세의 추정 기법을 제안한다. 기존 비마커 증강현실의 연구는 손을 검출하기 위해 단순한 실험 배경에서 피부색상 기반으로 손 영역을 검출한다. 그리고 손가락의 특징점을 검출하여 손의 자세를 추정하므로 카메라에서 검출할 수 있는 손 자세에 많은 제약이 따른다. 하지만, 본 논문은 3D 센서의 색상 및 깊이 정보를 융합한 Mean-Shift 추적 기법을 사용함으로써 복잡한 배경에서 손을 검출할 수 있으며 손 자세를 크게 제약하지 않고 손 영역의 중심점과 임의의 2점의 깊이 값만으로 정확한 손 자세를 추정한다. 제안하는 Mean Shift 추적 기법은 피부 색상정보만 사용하는 방법보다 약 50픽셀 이하의 거리 오차를 보였다. 그리고 증강실험에서 제안하는 손 자세 추정 방법은 복잡한 실험환경에서도 마커 기반 방법과 유사한 성능의 실험결과를 보였다. This paper proposes a new method of estimating the hand pose through the Mean-Shift tracking algorithm using the fusion of color and depth information for marker-less augmented reality. On marker-less augmented reality, the most of previous studies detect the hand region using the skin color from simple experimental background. Because finger features should be detected on the hand, the hand pose that can be measured from cameras is restricted considerably. However, the proposed method can easily detect the hand pose from complex background through the new Mean-Shift tracking method using the fusion of the color and depth information from 3D sensor. The proposed method of estimating the hand pose uses the gravity point and two random points on the hand without largely constraints. The proposed Mean-Shift tracking method has about 50 pixels error less than general tracking method just using color value. The augmented reality experiment of the proposed method shows results of its performance being as good as marker based one on the complex background.
A new method is proposed for stereo vision where solution to disparity map is presented in terms of Line constraint and Reliability space -the first constraint proposes a progressive framework for stereo matching which applies local area pixel-values from corresponding lines in the left and right image pairs. The second states that reliability space based on corresponding points records the disparity and then we are able to apply the median filter in order to reduce the noises which occur in the process. A coarse to fine result is presented after the median filtering, which improves the final result qualitatively. Experiment is evaluated by rectified stereo matching images pairs from Middlebury datasets and has proved that those two adopted strategies yield good matching quantitative results in terms of fast running speed.
본 논문에서는 단일 카메라를 이용하여 차량의 위치를 검출하고 연속적인 프레임에서의 차량의 움직임을 추적하는 알고리즘을 제안한다. 차량의 특징을 검출하기 위해 Haar-like 에지 검출기를 사용하고, 카메라의 캘리브레이션 정보를 이용하여 차량의 위치를 추정한다. 신뢰도를 높이기 위해 k 개의 연속적인 프레임에서의 누적된 차량 정보를 추출한다. 최종 검출된 차량을 템플릿으로 지정하고 SURF (Speeded Up Robust Features) 알고리즘을 통해 연속적으로 입력되는 프레임에서 동일한 차량을 추출한다. 이를 통해 동일 차량으로 추출된 차량 정보를 새로운 템플릿으로 업데이트 한다. 비교 검출을 위한 수행 시간을 줄이기 위해 이전 프레임에서 검출된 차량의 범위를 확장한 영역만을 관심 영역으로 지정한다. 이 과정은 공통된 대응점을 찾지 못할 때까지 검출과 추적 과정을 반복하여 진행한다. 실 도로 상에서 얻어진 영상에 대해 적용함으로써 제안된 알고리즘의 효율성을 보였다. This paper proposes vehicle detection and tracking algorithm using a CCD camera. The proposed algorithm uses Haar-like wavelet edge detector to detect features of vehicle and estimates vehicle's location using calibration information of an image. After that, extract accumulated vehicle information in continuous k images to improve reliability. Finally, obtained vehicle region becomes a template image to find same object in the next continuous image using SURF(Speeded Up Robust Features). The template image is updated in the every frame. In order to reduce SURF processing time, ROI(Region of Interesting) region is limited on expended area of detected vehicle location in the previous frame image. This algorithm repeats detection and tracking progress until no corresponding points are found. The experimental result shows efficiency of proposed algorithm using images obtained on the road.
Power management is an important issue in mobile wireless network devices with multiple network communication interfaces, as most mobile wireless network devices are battery-operated. Therefore, it is important to reduce the energy consumption to maximize the lifetime. In this paper, we analyze the energy consumption of network devices according to the variation of the network detection interval. We obtain the optimal network detection interval using a probabilistic analysis according to the state transition rate of networks and the energy consumption of network interfaces. Finally, we propose a heuristic network detection scheme that dynamically adjusts the network detection interval to reduce the amount of energy consumed. With the proposed scheme, we can reduce the energy consumption of network devices. We evaluated our scheme in a simulation. The simulation results show that the energy consumption of our scheme is nearly identical to of an optimal scheme under certain conditions.
We present an approach to 3D vehicle class recognition (which of SUV, mini-van, sedan, pickup truck) with one or more fixed video-cameras in arbitrary positions with respect to a road. The vehicle motion is assumed to be straight. We propose an efficient method of Structure from Motion (SfM) for camera calibration and 3D reconstruction. 3D geometry such as vehicle and cabin length, width, height, and functions of these are computed and become features for use in a classifier. Classification is done by a minimum probability of error recognizer. Finally, when additional video clips taken elsewhere are available, we design classifiers based on two or more video clips, and this results in significant classification-error reduction.
본 논문은 야외 환경에서 하나의 카메라를 이용한 시각 장애인을 위한 보행 안내 시스템을 제안한다. 기존의 스테레오 비전을 이용한 보행 지원 시스템과는 다르게 제안된 시스템은 사용자의 허리에 고정된 하나의 카메라를 이용하여 꼭 필요한 정보만을 얻는 것을 목표로 하는 시스템이다. 제안하는 시스템은 먼저 탑-뷰 영상을 생성하고, 생성된 탑-뷰 영상 내 지역적인 코너 극점을 검출한다. 검출된 극점에서 방사형의 히스토그램을 분석하여 장애물을 검출한다. 그리고 사용자 움직임은 사용자에 가까운 지역 안에서 옵티컬 플로우를 사용하여 추정한다. 이렇게 영상으로부터 추출된 정보들을 기반으로 음성 메시지 생성 모듈은 보행 지시 정보를 합성된 음성을 통해 시각 장애인에게 전달한다. 다양한실험 영상들을 사용하여 제안한 보행 안내 시스템이 일반 인도에서 유용한 안내 지시를 제공하는 것이 가능함을 보인다. This paper presents a walking guidance system for blind pedestrians in an outdoor environment using just one single camera. Unlike many existing travel-aid systems that rely on stereo-vision, the proposed system aims to get necessary information of the road environment by using just single camera fixed at the belly of the user. To achieve this goal, a top-view image of the road is used, on which obstacles are detected by first extracting local extreme points and then verified by the polar edge histogram. Meanwhile, user motion is estimated by using optical flow in an area close to the user. Based on these information extracted from image domain, an audio message generation scheme is proposed to deliver guidance instructions via synthetic voice to the blind user. Experiments with several sidewalk video-clips show that the proposed walking guidance system is able to provide useful guidance instructions under certain sidewalk environments.
This paper presents a vision-based navigation method for blind people in an outdoor sidewalk environment. Unlike many existing navigation systems that use stereo-vision based methods, the proposed method is able to get obstacle position as well as user motion information by using just single camera fixed at the belly of the user. To achieve this goal, a top-view image of the road is used for obstacle detection and user motion estimation, based on which a grid map is generated for navigation. Obstacles are detected by using a beam-ray model, while user motion is estimated by using optical flow in a user surrounding area. For navigation part, a step score is calculated on the grid map for evaluating the safety of next moving step. Experiments with several sidewalk video-clips show that the proposed navigation method is able to provide useful guidance instructions under certain sidewalk environments.