Disruptive events within a country can have global repercussions, creating a need for the anticipation and planning of these events. Crystal Cube (CC) is a novel approach to forecasting disruptive political events at least one month into the future. The system uses a recurrent neural network and a novel measure of event similarity between past and current events. We also introduce the innovative Thermometer of Irregular Leadership Change (ILC). We present an evaluation of CC in predicting ILC for 167 countries and show promising results in forecasting events one to twelve months in advance. We compare CC results with results using a random forest as well as previous work.
We propose an ensemble approach for multi-target binary classification, where the target class breaks down into a disparate set of pre-defined target-types. The system goal is to maximize the probability of alerting on targets from any type while excluding background clutter. The agent-classifiers that make up the ensemble are binary classifiers trained to classify between one of the target-types vs. clutter. The agent ensemble approach offers several benefits for multi-target classification including straightforward in-situ tuning of the ensemble to drift in the target population and the ability to give an indication to a human operator of which target-type causes an alert. We propose a combination strategy that sums weighted likelihood ratios of the individual agent-classifiers, where the likelihood ratio is between the target-type for the agent vs. clutter. We show that this combination strategy is optimal under a conditionally non-discriminative assumption. We compare this combiner to the common strategy of selecting the maximum of the normalized agent-scores as the combiner score. We show experimentally that the proposed combiner gives excellent performance on the multi-target binary classification problems of pin-less verification of human faces and vehicle classification using acoustic signatures.
The goal of Crystal Cube is to create an automated capability for the prediction of disruptive events. In this paper we present initial prediction results on six prediction categories previously shown to be of interest in the literature. In particular, we compare the performance of static classification models, often used in previous work for these prediction tasks, with a gated recurrent unit sequence model that has the ability to retain information over long periods of time for the classification of sequence data. Our results show that the sequence model is comparable in performance to the best performing static model (the random forest), and that more work is needed to classify highly dynamic prediction categories with high probability.
The any-combiner is a classifier combination approach for target classification problems in which the target class can be naturally decomposed into multiple subclasses. This kind of classification problem can often occur in sensor-based system applications, such as biometric user verification, biosurveillance or underwater mine detection, in which the system goal is to identify a test exemplar as belonging to a category of objects of interest to the exclusion of all other exemplars (clutter). We propose an approach to the target classification problem in which an ensemble of classifier agents are trained to distinguish individual target subclasses from clutter. The any-combiner is then trained by optimizing the multi-agent ensemble for maximum recognition performance across all target subclasses over a range of acceptable operating points. Once deployed, the any-combiner classifies a test example as a target if any of the agents indicates a true positive classification for its target subclass. Experiments show that the any-combiner yields excellent performance on the tasks of biometric verification using face images and underwater object classification using acoustic features.
We consider the problem of classifying a test sample given incomplete information. This problem arises naturally when data about a test sample is collected over time, or when costs must be incurred to compute the classification features. For example, in a distributed sensor network only a fraction of the sensors may have reported measurements at a certain time, and additional time, power, and bandwidth is needed to collect the complete data to classify. A practical goal is to assign a class label as soon as enough data is available to make a good decision. We formalize this goal through the notion of reliability--the probability that a label assigned given incomplete data would be the same as the label assigned given the complete data, and we propose a method to classify incomplete data only if some reliability threshold is met. Our approach models the complete data as a random variable whose distribution is dependent on the current incomplete data and the (complete) training data. The method differs from standard imputation strategies in that our focus is on determining the reliability of the classification decision, rather than just the class label. We show that the method provides useful reliability estimates of the correctness of the imputed class labels on a set of experiments on time-series data sets, where the goal is to classify the time-series as early as possible while still guaranteeing that the reliability threshold is met.
Early classification of time series is important in time-sensitive applications. An approach is presented for early classification using generative classifiers with the dual objectives of providing a class label as early as possible while guaranteeing with high probability that the early class matches the class that would be assigned to a longer time series. We give a specific algorithm for early quadratic discriminant analysis (QDA), and demonstrate that this classifier meets the requirement of reliable early classification.
the availability of integrated tools to explore, analyze and understand the data warehoused in these archives is lagging far behind the ability to gain access to the same data. In particular, locating and identifying patterns of interest in numerical time series data is an increasingly important problem for which there are few available techniques. Temporal pattern recognition poses many interesting problems in classification, segmentation, prediction, diagnosis and anomaly detection. This research focuses on the problem of classification or characterization of numerical time series data. Highway vehicles and their drivers are examples of complex dynamic systems (CDS) which are being used by transportation agencies for field testing to generate large-scale time series datasets. Tools for effective analysis of numerical time series in databases generated by highway vehicle systems are not yet available, or have not been adapted to the target problem domain. However, analysis tools from similar domains may be adapted to the problem of classification of numerical time series data.
We present local discriminative Gaussian (LDG) dimensionality reduction, a supervised dimensionality reduction technique for classification. The LDG objective function is an approximation to the leave-one-out training error of a local quadratic discriminant analysis classifier, and thus acts locally to each training point in order to find a mapping where similar data can be discriminated from dissimilar data. While other state-of-the-art linear dimensionality reduction methods require gradient descent or iterative solution approaches, LDG is solved with a single eigen-decomposition. Thus, it scales better for datasets with a large number of feature dimensions or training examples. We also adapt LDG to the transfer learning setting, and show that it achieves good performance when the test data distribution differs from that of the training data.
We consider the problem of classifying a signal that is the output of a linear, time-invariant channel in the presence of additive noise, given two distinct sets of labeled data: one dataset of examples of the signals input to the channel, and a second dataset of example signals corrupted by the channel. We propose a distribution-based Bayesian quadratic discriminant analysis classifier that uses the input examples along with a model for the channel to form a prior for the likelihood of the output examples. Preliminary experiments with this proposed transfer BDA classifier show that it effectively uses both sets of data and is also robust to errors in channel modeling.
In many signal processing applications, a signal to be classified has been corrupted by a channel and additive noise. A standard approach is to estimate the clean signal, then classify it. We consider two robust approaches that account for the estimation procedure. The first approach is an application of the MAP rule for noisy features, and the second is an approach for discriminative classifiers that treats that training points as random. An experiment confirms that the robust approaches offer performance gains.
We present a robust probabilistic method to classify targets based on their tracks. As is customary in supervised learning problems, it is assumed that example tracks from various classes are available to train a classifier. We present an optimal but computationally intensive sequential solution, and show that a computationally feasible naive Bayes approximation works better than ignoring sequential information. We show how to take into account the uncertainty of the track, as quantified by the error covariance matrix from a Kalman tracker, using the recently proposed expected maximum likelihood rule coupled with a robust local Bayesian discriminant analysis classifier. In addition, we propose an expected maximum a posterior rule to take test sample uncertainty into account for classifiers that model the posterior, and use it to define a robust kernel classifier. Simulations with a Kalman tracker show significantly improved performance by taking into account the tracked state covariance.
The impact of bottom sediment type in relation to acoustic communications via orthogonal frequency division multiplexing (OFDM) is shown via experimental results and simulation. Experimental data from Lake Washington, Seattle with a “silty clay” bottom show that the multipath delay spread is longer at 250 m than at 4 km. This results in better OFDM performance at the longer range. Similar results are shown via simulation using a channel model developed from Bellhop, a Gaussian Ray tracing tool [M. Porter, “Bellhop Gaussian beam/finite element beam code,” Available: http://oalib.hlsresearch.com/Rays/index.html (2007)]. Through simulation, results are also shown under similar conditions to the experiment but with varying bottom type. The results show that the performance of OFDM signaling is dependent on the bottom type as well as specific source/receiver geometry. [Work supported by NASA ESTO.]
Challenges in developing high performance underwater acoustic modems can be summarized as: 1. Lack of maturity in underwater acoustic communications technology at all layers of the communication stack . This stems in part from the complex and dynamic nature of the underwater channel. 2. Spatially and temporally dynamic nature of the underwater acoustic channel . Theoretically "high performance" communications may not be achievable using the largely channel-specific (channel-static) processing techniques similar to those which have evolved for radio frequency and optical communications. 3. Resources . Without the expectation that the processing being implemented will reach a point of high maturity and meet a high-priority near-term operational need, it can be difficult to justify such an application specific investment.
We propose an OFDM receiver capable of estimating and correcting, on a symbol-by-symbol basis, the subcarrier dependent Doppler shifting due to the movement of source and receiver in an underwater acoustic network. We propose two methods of estimation: one of which is based upon the marginal maximum likelihood principle, and one of which is ad-hoc. We compare the performance of both estimators to the Cramer-Rao lower bound. We show through simulation that the proposed receiver design performs well for a source that is accelerating at 0.29 m/s2.
We address several inter-related aspects of underwater network design within the context of a cross-layer approach. We first highlight the impact of key characteristics of the acoustic propagation medium on the choice of link layer parameters; in turn, the consequences of these choices on design of a suitable MAC protocol and its performance are investigated. Specifically, the paper makes contributions on the following fronts: a) Based on accepted acoustic channel models, the pointto- point (link) capacity is numerically calculated, quantifying sensitivities to factors such as the sound speed profile, power spectral density of the (colored) additive background noise and the impact of boundary (surface) conditions for the acoustic channel; b) It provides an analysis of the Micromodem-like linklayer based on FH-FSK modulation; and finally c) it undertakes performance evaluation of a simple MAC protocol based on ALOHA with Random Backoff, that is shown to be particularly suitable for small underwater networks.
Orthogonal Frequency Division Multiplexing (OFDM) is a wideband modulation scheme that has recently gained attention for underwater acoustic communications. The benefits of OFDM include its ability to overcome long channel delay spreads through the use of a guard interval, its ability to transform a frequency selective channel into multiple frequency non-selective channels, and its relatively easy implementation through the use of the fast fourier transform and its inverse. However, OFDM has several distinct challenges in the underwater channel environment. The underwater channel is known to be highly frequency selective with large delay and Doppler spreading as well as fast time variance. The frequency selectivity requires accurate estimation of the channel transfer function in order to recover the transmitted symbols. Large delay spreading requires a longer guard interval, decreasing the data rate. Finally, Doppler spreading and fast time variations will introduce added noise at the demodulator due to inter-chip interference (ICI) and inaccurate channel estimation. In this work in progress poster, we will show results of an experiment using zero padded (ZP) OFDM where good performance was achieved at ranges of 250m to 4km. The results will show that, counter-intuitively, the worst performance was achieved at 250m. The results highlight the sensitivity of the OFDM parameters to the specific conditions of the underwater channel.
In many areas of Earth science, including climate change research, there is a need for near real-time integration of data from heterogeneous and spatially distributed sensors, in particular in-situ and space-based sensors. The data integration, as provided by a smart sensor web, enables numerous improvements, namely, 1) adaptive sampling for more efficient use of expensive space-based sensing assets, 2) higher fidelity information gathering from data sources through integration of complementary data sets, and 3) improved sensor calibration. The specific purpose of the smart sensor web to be demonstrated as part of the development presented here is to provide for adaptive sampling and calibration of space-based data via in-situ data. Our ocean-observing smart sensor web presented herein is composed of both mobile and fixed underwater in-situ ocean sensing assets and Earth Observing System (EOS) satellite sensors providing larger-scale sensing. An acoustic communications network forms a critical link in the web between the in-situ and space-based sensors and facilitates adaptive sampling and calibration. After an overview of primary design challenges, we report on the development of various elements of the smart sensor web. These include (a) a cable-connected mooring system with a profiler under real-time control with inductive battery charging; (b) a glider with integrated acoustic communications and broadband receiving capability; (c) satellite sensor elements; (d) an integrated acoustic navigation and communication network; and (e) a predictive model via the Regional Ocean Modeling System (ROMS). Results from field experiments as well as simulation and theoretical studies on acoustic communication system performance, link capacity computation, and development of a media access control (MAC) layer protocol for underwater networking, are described. Plans for future adaptive sampling demonstrations using the smart sensor web are also presented.
Signals transmitted through underwater channels experience attenuation due to dissipation of acoustic energy by spreading as well as by absorption. The path loss due to absorption is found to be highly dependent upon the frequency. Ambient noise, which also greatly affects accurate reception of the signal, is also highly dependent upon frequency. For these reasons, the received SNR cannot be assumed to be constant over wideband acoustic signaling schemes. In this paper we determine the (signaling) rate vs. range curves for a FH-FSK modem considering the frequency selective nature of signal attenuation in an acoustic medium.
This paper discusses a methodology for predicting underwater acoustic communications performance using high fidelity acoustic time series simulation and acoustic modem processing emulation. Multiple source/receiver combinations can be simultaneously simulated, so that aspects of a complete underwater network can be studied. Here, the fundamental modeling and emulation capability will be described, with examples of the propagation modeling, time series simulation, and modem processing over multiple realizations of example communications channels. The results show the dependence of source and receiver location in the water column with respect to the sound speed profile on communications performance. The utility of such simulations for ad hoc network design in the presence of moving communications nodes will be discussed.