Adaptive beamformers at the element level usually require a great number of training samples and the computational cost for calculating the weight vector for large phased array antennas is very high, which make it difficult for real-time applications. To address this problem, a two-dimensional (2-D) adaptive beamformer applicable to large planar array antennas that have low computational complexity and low training sample requirement is proposed. In the proposed method, the weight matrix is first reconstructed as a matrix that has the same or close columns and rows by utilising the special Kronecker property of the array steering matrix. Then, the weight vector is determined by adopting a bi-quadratic cost function and a bi-iterative algorithm. Experimental results show that the proposed method can achieve fairly good performance even when the training samples are small.
Effective ground stationary operational target detection is the precondition of subsequent target identification for helicopter-borne fire-controlled radar. More abundant target information could be obtained when adopting wide-band radar than traditional narrow-band one. While there exit such problem of high false alarm and poor adaptability when applying traditional detection methods. Thus, a novel ground stationary target detection method for airborne wide-band radar based on statistical characteristic is presented in this paper. By considering statistical distribution property in both range and azimuth directions, this method could distinguish the target from strong ground clutter background adaptively and effectively. Experiment results show that our algorithm not only can improve detection performance significantly but also could enhance processing efficiency.
There exists periodic modulation problem in radar echoes due to the main rotor blades periodic blockage in helicopter-borne fire-control radar which is mounted atop the main rotor mast of the helicopter. Such modulation echo induces a set of ghosts in synthetic aperture radar (SAR) image, further resulting in a poor performance of subsequent tracking and striking. To address this problem, this article proposes a method on rotor blades blockage modulation suppression for helicopter-borne SAR. By decomposing the blocked echo into the form of Fourier series in the azimuth direction, a reference function could be constructed to suppress the modulation directly by adopting an iterative approximation strategy. This method effectively avoids the complex blocked data recovery methods, and thus can be used to suppress various kinds of periodic modulation components without requiring certain distribution models. Both simulated and real-measured data are processed to demonstrate the effectiveness of the proposed algorithm.
We train fully convolutional neural networks with no recurrent layers for the end-to-end phoneme recognition task, using the Connectionist Temporal Classification (CTC) loss function. The adopted network, U-Net, was introduced initially for semantic image segmentation tasks, and is often applied to segmenting features in medical imaging and remote sensing. The similarities between CTC-based automatic speech recognition and semantic segmentation problems are discussed. We extend the encoder-decoder architecture of U-Net and show it is capable of good performance in the acoustic modelling of a speech recognition system. We investigate the importance of the concatenation step in the design of U-net, and report results using the core test set of the TIMIT corpus.
We investigate the problem of acoustic scene classification, using a deep residual network applied to log-mel spectrograms complemented by log-mel deltas and delta-deltas. We design the network to take into account that the temporal and frequency axes in spectrograms represent fundamentally different information. In particular, we use two pathways in the residual network: one for high frequencies and one for low frequencies, that were fused just two convolutional layers prior to the network output. We conduct experiments using two public 2019 DCASE datasets for acoustic scene classification; the first with binaural audio inputs recorded by a single device, and the second with single-channel audio inputs recorded through various devices. We show the performance of our models are significantly enhanced by the use of log-mel deltas, and that overall our approach is capable of training strong single models, without use of any supplementary data from outside the official challenge dataset, with excellent generalization to unknown devices. In particular, our approach achieved second place in 2019 DCASE Task 1b (0.4% behind the winning entry), and the best Task 1B evaluation results (by a large margin of over 5%) on test data from a device not used to record any training data.
Automatic transcription of ancient handwritten manuscripts can be a challenging task when compared with a transcription of contemporary handwriting. Characters and words can have unusual and varying shapes, with significant variation between writers, and sufficient labelled data from which to train machine learning algorithms can be difficult to access. This paper describes ancient Thai handwriting transcription on block-based from archive manuscripts, using a hybrid deep neural network with both convolutional (CNN) and recurrent (RNN) layers, trained using Connectionist Temporal Classification (CTC) loss. Six architecture variations are compared. Data augmentation was applied to synthetically increase the number of training samples, resulting in improved learning. Thai archive manuscripts were collected from the Thai National Library. The character error rate (CER) in the best architecture was found to be 11.9 percent.
In order to improve the angle accuracy and meet the requirement of side-lobe blanking, the fire-control radar system usually adopts sum beam, azimuth difference beam, elevation difference beam and several protecting antennas simultaneously. A ΣΔG-STAP method based on sum, difference and protecting channels is proposed to improve the performance of ground moving target detection. Moreover, the target is located on the Doppler beam sharpening image by the monopulse angle estimation technique. The effectiveness of the proposed method is verified using the simulated data and measured data.
A novel bi-iterative dimension-reduced space–time adaptive processing (STAP) algorithm for clutter suppression and moving target detection in airborne radar system is proposed. A dimension-reduced processing in both space and time is firstly performed. Then a bi-quadratic cost function is constructed to estimate the dimension-reduced spatiotemporal filter coefficients, which are used for suppressing the clutter and keeping the target signal. An optimal solution to the cost function can be obtained efficiently by a bi-iteration algorithm. Theoretical analysis and computer simulation results illustrate that the algorithm has fast convergence, low computation load and sampling requirements. The results of experiment by using the measured data show the validity of the proposed algorithm.
This technical report describes our approach to Tasks 1a, 1b and 1c in the 2019 DCASE acoustic scene classification challenge. Our focus was on developing strong single models, without use of any supplementary data. We investigated the use of a deep residual network applied to log-mel spectrograms complemented by log-mel deltas and delta-deltas. We designed the network to take into account that the temporal and frequency axes in spectrograms represent fundamentally different information. In particular, we used two pathways in the residual network: one for high frequencies and one for low frequencies, that were fused just two convolutional layers prior to the network output.
The introduction of skip connections used for summing feature maps in deep residual networks (ResNets) were crucially important for overcoming gradient degradation in very deep convolutional neural networks (CNNs). Due to the strong results of ResNets, it is a natural choice to use features that it produces at various layers in transfer learning or for other feature extraction tasks. In order to analyse how the gradient degradation problem is solved by ResNets, we empirically investigate how discriminability changes as inputs propagate through the intermediate layers of two CNN variants: all-convolutional CNNs and ResNets. We found that the feature maps produced by residual-sum layers exhibit increasing discriminability with layer-distance from the input, but that feature maps produced by convolutional layers do not. We also studied how discriminability varies with training duration and the placement of convolutional layers. Our method suggests a way to determine whether adding extra layers will improve performance and show how gradient degradation impacts on which layers contribute increased discriminability.
Video Synthetic Aperture Radar (ViSAR) is a novel technique to continuously monitor moving targets in the interested area. From the video echo provided by the ViSAR system, the motion trail of the moving target can be easily obtained. Therefore, the echo of the moving target should be carefully studied, which is important in practical application for moving targets tracking in ViSAR system.