
Entropy-based image thresholding has received considerable interest in recent years. Two types of entropy are generally used as thresholding criteria: Shannon's entropy and relative entropy, also known as Kullback-Leibler information distance, where the former measures uncertainty in an information source with an optimal threshold obtained by maximising Shannon's entropy, whereas the latter measures the information discrepancy between two different sources with an optimal threshold obtained by minimising relative entropy. Many thresholding methods have been developed for both criteria and reported in the literature. These two entropy-based thresholding criteria have been investigated and the relationship among entropy and relative entropy thresholding methods has been explored. In particular, a survey and comparative analysis is conducted among several widely used methods that include Pun and Kapur's maximum entropy, Kittler and Illingworth's minimum error thresholding, Pal and Pal's entropy thresholding and Chang et al.'s relative entropy thresholding methods. In order to objectively assess these methods, two measures, uniformity and shape, are used for performance evaluation
A new multi-spectral laser radar (ladar) system based on the time-correlated single photon counting, time-of-flight technique has been designed to detect and characterise distributed targets at ranges of several kilometres. The system uses six separated laser channels in the visible and near infrared part of the electromagnetic spectrum. The authors present a method to detect the numbers, positions, heights and shape parameters of returns from this system, used for range profiling and target classification. The algorithm has two principal stages: non-parametric bump hunting based on an analysis of the smoothed derivatives of the photon count histogram in scale space, and maximum likelihood estimation using Poisson statistics. The approach is demonstrated on simulated and real data from a multi-spectral ladar system, showing that the return parameters can be estimated to a high degree of accuracy.
3D graphics performance is increasing faster than any other computing application. Almost all PC systems now include 3D graphics accelerators for games, computer aided design or visualisation applications. This article investigates the suitability of field programmable gate array devices as an accelerator for implementing 3D affine transformations. Proposed solution based on processing large matrix multiplication have been implemented, for large 3D models, on the RC1000 Celoxica board based development platform using Handel-C. Outstanding results have been obtained for the acceleration of 3D transformations using fixed and floating-point arithmetic.
A new analytic method for estimating the pose of a calibrated camera from a single view of a planar target with a known geometry is presented. A set of fiducial points is extracted from an image of this target and metrically rectified through a 2D homography. The author's contribution consists in deriving from the coefficients of this transformation the analytic expressions of the 3D rigid motion parameters defining the pose of the camera. Experiments with synthetic data show the effectiveness of the proposed method against another gold-standard method for pose estimation and its robustness against perturbations of the fiducial points. These features make the algorithm a perfect candidate for all those applications where pose estimation has to be performed reliably and in real time.
A novel unsupervised strategy for content-based image retrieval is presented. It is based on a meaningful segmentation procedure that can provide proper distributions for matching via the earth mover's distance as a similarity metric. The segmentation procedure is based on a hierarchical watershed-driven algorithm that extracts meaningful regions automatically. In this framework, the proposed robust feature extraction and the many-to-many region matching along with the novel region weighting for enhancing feature discrimination play a major role. Experimental results demonstrate the performance of the proposed strategy.
The authors demonstrate the potential of femtosecond laser pulsed illumination and picosecond detection in biomedical imaging via an efficient and simple Monte Carlo laser pulse diffusion simulation, based on realistic biological parameters. The evolution of the contrast of a sinusoidal grid in a diffuse medium as a function of the width of a temporal gate is shown.
The author presents a CORDIC-based split-radix fast Fourier transform (FFT)/inverse FFT (IFFT) processor dedicated to the computation of 2048/4096/8192-point discrete Fourier transforms (DFTs). The arithmetic unit of a butterfly processor and a twiddle factor generator are based on a CORDIC algorithm. An efficient implementation of the CORDIC-based split-radix FFT algorithm is demonstrated. The chip of 2048/4096/8192-point FFT/IFFT core processor is fabricated in a 0.18 mu m CMOS technology. The core size is 4860 x 7883 mu m(2) and contains about 200 822 gates for logic and memory, and the power dissipation is 350 mW with a clock rate of 150 MHz at 1.8 V. All control signals are generated internally on-chip. The processor performs 8192-point FFT/IFFT every 138 mu s and 2048-point FFT/IFFT every 34.5 mu s, respectively, which exceeds orthogonal frequency division multiplexer symbol rates. The modified-pipelining CORDIC arithmetic unit is employed for complex multiplication. A CORDIC twiddle factor generator is proposed and implemented for reducing the size of ROM required for storing the twiddle factors. Compared with conventional FFT implementations, the power consumption is reduced by 25%.
The paper presents a novel method and software platform for remote and interactive browsing of a summary of long video sequences as well as revealing the semantic links between shots and scenes in their temporal context. The solution is based on interactive navigation in a scalable mega image resulting from a JPEG 2000 coded key-frame-based video summary. Each key-frame could represent an automatically detected shot, event or scene, which is then properly annotated using some semi-automatic tools or learning methods. The presented system is compliant with the new JPEG 2000 Part 9 'JPIP - JPEG 2000 interactivity, API and protocols,' which lends itself to working under varying transmission channel conditions such as GPRS or 3G wireless networks. While keeping the advantages of a single 2D video summary, like the limited storage cost, the flexibility offered by JPEG 2000 allows the application to highlight interactively key-frames corresponding to the desired content first within a low-quality and low-resolution version of the full video summary. It then offers fine grain scalability for a user to navigate and zoom into particular scenes or events represented by the key-frames. This possibility of visualising key-frames of interest and playing back the corresponding video shots within the context of the whole sequence (e.g. an episode of a media file) enables the user to understand the temporal relations between semantically related events/actions/physical settings, providing a new way to present and search for contents in video sequences.
An improved sinusoidal modelling method based on perceptual matching pursuits computed in the Bark scale with application to parametric audio coding is proposed. Complex exponentials compose the overcomplete dictionary for matching pursuits. The main contribution is the minimisation of a perceptual distortion measure defined in the Bark scale to select the optimum atom at each iteration of the pursuits. In addition, a psychoacoustic stopping criterion for the pursuits is presented. The proposed sinusoidal modelling method is suitable to be integrated into a parametric audio coder based on the three-part model of sines, transients and noise, as appreciated in experimental results. The method provides significant advantages over previously proposed methods mainly because it operates in the Bark scale rather than in the frequency scale.
The authors present the use of visible colour difference in a new quantitative evaluation scheme for colour segmentation. To avoid directly evaluating the subjectively perceived quality of colour segmentation, two objective visual quantities, the quantity of missing boundaries and the quantity of fake boundaries, are considered. To explore how missing boundaries and fake boundaries affect the perceived quality of colour segmentation, a few visual rating experiments are made. On the other hand, to fit for humans' visual perception on colour difference, the visible colour difference is defined. Based on the experiments and the visible colour difference, two measures, named intra-region visual error and inter-region visual error, are designed to estimate the degrees of missing boundaries and fake boundaries, respectively. With these two measures, a complete scheme for the evaluation of colour segmentation is proposed. The simulation results demonstrate that this new scheme may evaluate segmentation results without any ground truth, and could help the automatic selection of parameter settings for a given segmentation algorithm
Digital data hiding applications have been developed mainly with a view for sending secret information safely through the Internet or other computer networks. Most of the existing vector quantisation (VQ)-based data hiding methods, however, are capable of embedding only a limited quantity of secret data into the medium, usually a secret bit in every block at most. To improve the embedding capacity without seriously sacrificing qualities of carrier media, an adaptive VQ-based data hiding scheme based on a codeword clustering technique is presented. After the adaptive clustering process, the proposed scheme embeds secret data into the VQ-index table by performing codeword-order-cycle permutation. With the help of the cycle technique, more options can be offered for proper codeword substitution, which adds more flexibility to the whole data hiding scheme and therefore improves the quality of the embedded image. As the experimental results demonstrate, the proposed scheme is capable of providing better image quality and embedding capacity than the currently existing VQ-based hiding schemes. In fact, the stego-image quality is so good that it is visually indistinguishable from the VQ-compressed image
A new approach to local Wiener filtering in the presence of Gaussian noise is presented. The two-step denoising algorithm is formulated. In the first step, standard local Wiener filtering is applied. The output of the first step is used for the construction of the matched filter, which enables us to better estimate the signal energy. The signal energy estimate is then used for the construction of the improved local Wiener filter. The performance of this method, using the dual-tree complex wavelet transform, is evaluated. The results obtained are comparable with the best in the literature but at a significant smaller computational cost.
A study of the classic line spectral frequency (LSF) extraction methods along with the assumptions made during their estimation is presented. LSF extraction is investigated from an over-sampling and decimation perspective and LSFs are shown to contain high-frequency variations that led to spectral overlapping problems. An anti-aliasing filter prior to decimation, with a cut-off frequency dependent on the final LSF vector transmission rate, is proposed to alleviate the aliasing problem of the classic extraction methods. The proposed method shows a clear advantage because it produces the same quantisation distortion as the classic methods at a lower bit requirement, with significant reduction in 2 and 4 dB outliers that greatly affect synthesised speech quality. This was also confirmed through a listening test
A new technique, called inter-band compensated prediction, for coding colour and multispectral images is presented. It is suitable to use for coding any spectral domain and can code colour and multispectral images with any number of bands. This technique is based on the same principles as the very efficient motion compensated prediction largely used in video coding. Thus, each band is predicted in the spectral direction by compensating the differences in the neighbouring bands and then coding the prediction error spatially by another method. This is a forward adaptive prediction and the information used for compensation is coded as side information with prediction error. The comparison of the coding results with the state-of-the-art coding algorithms, based on spectral transformations, proves that this technique is very efficient and can even outperform them. In addition, compensation can be combined with any spatial coder that allows lossless, lossy and scalable coding of any spectral content of the image. It has also the advantages of being simple to implement and to use with parallel architectures.
The IIR digital integrator is designed by using the Simpson integration rule and fractional delay filter. To improve the design accuracy of a conventional Simpson IIR integrator at high frequency, the sampling interval is reduced from T to 0.5T. As a result, a fractional delay filter needed to be designed in the proposed Simpson integrator. However, this problem can be solved easily by applying well-documented design techniques of the FIR and all-pass fractional delay filters. Several design examples are illustrated to demonstrate the effectiveness of the proposed method.
A voice-over-Internet protocol technique with a new hierarchical data security protection (HDSP) scheme using a secret chaotic bit sequence has been recently proposed. Some insecure properties of the HDSP scheme are pointed out and then used to develop known/chosen-plaintext attacks. The main findings are: given n known plaintexts, about (100 - (50/2(n)))% of secret chaotic bits can be uniquely determined; given only one specially-chosen plaintext, all secret chaotic bits can be uniquely derived; and the secret key can be derived with practically small computational complexity when only one plaintext is known (or chosen). These facts reveal that HDSP is very weak against known/chosen-plaintext attacks. Experiments are given to show the feasibility of the proposed attacks. It is also found that the security of HDSP against the brute-force attack is not practically strong. Some countermeasures are discussed for enhancing the security of HDSP and several basic principles are suggested for the design of a secure encryption scheme.
An encoder/decoder system for 3D volumetric medical data is proposed. The system allows fast access to any 2D image by decoding only the relevant information from each subband image and thus provides minimum decoding time. This will be of immense use for the medical community because most of the computerised tomography, magnetic resonance imaging and positron emission tomography modalities produce volumetric data. As a fully-fledged 3D wavelet transform is used for compression, the advantage of good compression ratio is preserved. Preprocessing is carried out prior to wavelet transform, to enable easier identification of coefficients from each subband image. Inclusion of special characters in the bit stream (markers) facilitates access to corresponding information from the encoded data. Experiments are carried out by performing Daub4 filter along x (row) and y (column) directions and Haar filter along the z (slice) direction to account for the difference between interslice and intraslice resolution. The performance of the system has been evaluated on four sets of volumetric data and the results are compared to other 3D encoding/2D decoding schemes. Results show that for slice spacing of 3-10 mm, there is substantial improvement in decoding time. The speedup is found to be similar to 2.
A novel adaptive discriminative vector quantisation technique for speaker identification (ADVQSI) is introduced. In the training mode of ADVQSI, for each speaker, the speech feature vector space is divided into a number of subspaces. The feature space segmentation is based on the difference between the probability distribution of the speech feature vectors from each speaker and that from all speakers in the speaker identification (SI) group. Then, an optimal discriminative weight, which represents the subspace's role in SI, is calculated for each subspace of each speaker by employing adaptive techniques. The largest template differences between speakers in the SI group are achieved by using optimal discriminative weights. In the testing mode of ADVQSI, discriminative weighted average vector quantisation (VQ) distortions are used for SI decisions. The performance of ADVQSI is analysed and tested experimentally. The experimental results confirm the performance improvement employing the proposed technique in comparison with existing VQ techniques for SI and recently reported discriminative VQ techniques for SI (DVQSI).
Image classification is usually based on various visual descriptors extracted from the images. The descriptors characterising, for example, image colours or textures are often high dimensional and their scaling varies significantly. In the case of natural images, the feature distributions are often non-homogenous and the image classes are also overlapping in the feature space. This can be problematic, if all the descriptors are combined into a single feature vector in the classification. A method is presented for combining different visual descriptors in rock image classification. In our approach, k-nearest neighbour classification is first carried out for each descriptor separately. After that, the final decision is made by combining the nearest neighbours in each base classification. The total numbers of the neighbours representing each class are used as votes in the final classification. The experimental results with rock image classification indicate that the proposed classifier combination method is more accurate than the conventional plurality voting.
A new efficient technique for the classification of signals, in the form of earthquake-induced ground-acceleration time histories, according to the damage that they cause in buildings, is presented for the first time. A training set of real seismic accelerograms with well-known damage effects is utilised and fuzzy representations of prototype signals are extracted. These prototypes are selected with respect to the architectural and structural damage caused by the seismic-acceleration time histories. The classification of the unknown accelerograms takes place through a fuzzy comparison with the prototypes and each is classified to the most similar prototype. Real, seismic time-acceleration records were used for testing the algorithm and the high percentage of the correctly recognised signals prove the effectiveness of the algorithm. Correct classification rates of up to 84% are achieved.