NOTE: A good introduction to various machine learning models. NOTE: The theory is explained here with all the equations. [4] Vladimir N. Vapnik. The nature of statistical learning theory. Springer, second edition, 1995. NOTE: A good introduction to the theory, not much equations. NOTE: Very good paper proposing a series of tricks to make neural networks really working.
NOTE: A good introduction to various machine learning models. NOTE: The theory is explained here with all the equations. [4] Vladimir N. Vapnik. The nature of statistical learning theory. Springer, second edition, 1995. NOTE: A good introduction to the theory, not much equations. NOTE: Very good paper proposing a series of tricks to make neural networks really working.
A network model of disparity estimation was developed based on disparity-selective neurons, such as those found in the early stages of processing in the visual cortex. The model accurately estimated multiple disparities in regions, which may be caused by transparency or occlusion. The selective integration of reliable local estimates enabled the network to generate accurate disparity estimates on normal and transparent random-dot stereograms. The model was consistent with human psychophysical results on the effects of spatial-frequency filtering on disparity sensitivity. The responses of neurons in macaque area V2 to random-dot stereograms are consistent with the prediction of the model that a subset of neurons responsible for disparity selection should be sensitive to disparity gradients.
Local disparity information is often sparse and noisy, which creates two conflicting demands when estimating disparity in an image region: the need to spatially average to get an accurate estimate, and the problem of not averaging over discontinuities. We have developed a network model of disparity estimation based on disparity-selective neurons, such as those found in the early stages of processing in visual cortex. The model can accurately estimate multiple disparities in a region, which may be caused by transparency or occlusion, in real images and random-dot stereograms. The use of a selection mechanism to selectively integrate reliable local disparity estimates results in superior performance compared to standard back-propagation and cross-correlation approaches. In addition, the representations learned with this selection mechanism are consistent with recent neurophysiological results of von der Heydt, Zhou, Friedman, and Poggio [8] for cells in cortical visual area V2. Combining multi-scale biologically-plausible image processing with the power of the mixture-of-experts learning algorithm represents a promising approach that yields both high performance and new insights into visual system function.
Motion perception requires the visual system to satisfy two conflicting demands: first, spatial integration of signals from neighboring regions of the visual field to overcome noisy signals, and second, sensitivity to small velocity differences to segment regions corresponding to different objects (Braddick, 1993). We have developed a computational model for the visual processing of motion in area MT that accounts for these conflicting demands. The model has two types of units, similar to those found in area MT. One type of unit in the model integrates information about the direction motion to estimate the local velocity; these local velocity units compete among themselves to determine the most likely local velocity. A second type of unit selects regions of the visual field where the velocity estimates are most reliable; these selection units have nonclassical receptive field surrounds by virtue of competition with pools of similar units across the visual field. The output of the model is a distributed segmentation of the image into patches that support distinct objects moving with a common velocity. T h e processing of motion in the primate's visual cortex begins in area V1, where cells with reliable selec-tivity for direction of motion are found (Maunsell and Newsome, 1987); however, these cells d o not detect true velocity but instead are tuned to a limited range of spatiotemporal frequencies and exhibit spatially restricted receptive fields so that they can report only the perpendicular component of the velocity for straight edges. This so-called aperture problem is illustrated in figure 27.1. T o overcome these limitations and compute true local velocity measurements, it is necessary to integrate motion responses from cells with a variety of directions and spatiotemporal frequency tunings over a Neurons that respond selectively to velocity over a wide range of spatial frequencies are found in visual area M T , which receives a direct projection from area A class of cells in M T , the "pattern cells" of Movshon and colleagues (1985), respond to the direction of overall motion of plaid patterns composed of two differently oriented gratings rather than to the direction of the individual components. Psychophysical studies suggest that the perceived velocity of such patterns generally is close to the velocity that is uniquely consistent with the constraints imposed by the individual component's motions (Adelson and Movshon, 1982), although other possibilities have been suggested (Wilson et al., 1992; Rubin and Hochstein, 1993). In …
We describe a system that can track a hand in a sequence of video frames and recognize hand gestures in a user-independent manner. The system locates the hand in each video frame and determines if the hand is open or closed. The tracking system is able to track the hand to within ±10 pixels of its correct location in 99.7% of the frames from a test set containing video sequences from 18 different individuals captured in 18 different room environments. The gesture recognition network correctly determines if the hand being tracked is open or closed in 99.1% of the frames in this test set. The system has been designed to operate in real time with existing hardware.
We present a new approach to computing from image sequences the two-dimensional velocities of moving objects that are occluded and transparent. The new motion model does not attempt to provide an accurate representation of the velocity flow field at fine resolutions but coarsely segments an image into regions of coherent motion, provides an estimate of velocity in each region, and actively selects the most reliable estimates. The model uses motion-energy filters in the first stage of processing and computes, in parallel, two different sets of retinotopically organized spatial arrays of unit responses: one set of units estimates the local velocity, and the second set selects from these local estimates those that support global velocities. Only the subset of local-velocity measurements that are the most reliable is included in estimation of the velocity of objects. The model is in agreement with many of the constraints imposed by the physiological response properties of cells in primate visual cortex, and its performance is similar to that of primates on motion transparency.
We describe an extension to the Mixture of Experts architecture for modelling and controlling dynamical systems which exhibit multiple modes of behavior. This extension is based on a Markov process model, and suggests a recurrent network for gating a set of linear or non-linear controllers. The new architecture is demonstrated to be capable of learning effective control strategies for jump linear and non-linear plants with multiple modes of behavior.
We propose a new adaptation algorithm for equalizers operating on very distorted channels. The algorithm is based on the idea of adjusting the equalizer tap gains to maximize the likelihood that the equalizer outputs would be generated by a mixture of two gaussians with known means. The familiar decision-directed least mean square (LMS) algorithm is shown to be an approximation to maximizing the likelihood that the equalizer outputs come from such an i.i.d. source. The algorithm is developed in the context of a binary PAM channel and simulations demonstrate that the new algorithm converges in channels for which the decision-directed LMS algorithm does not converge.
Geoffrey E. Hinton Department of Computer Science . U ni versi ty of Toran to Toronto, Canada M5S lA4 One way of simplifying neural networks so they generalize better is to add an extra t.erm 10 the error fUll c tion that will penalize complexit.y. \Ve propose a new penalt.y t.erm in which the dist rihution of weight values is modelled as a mixture of multiple gaussians . C nder this model, a set of weights is simple if the weights can be clustered into subsets so that the weights in each cluster have similar values . We allow the parameters of the mixture model to adapt at t.he same time as t.he network learns. Simulations demonstrate that this complexity term is more effective than previous complexity terms.
We present a local learning rule in which Hebbian learning is conditional on an incorrect prediction of a reinforcement signal. We propose a biological interpretation of such a framework and display its utility through examples in which the reinforcement signal is cast as the delivery of a neuromodulator to its target. Three examples are presented which illustrate how this framework can be applied to the development of the oculomotor system.
Neurons in area MT of primate visual cortex encode the velocity of moving objects. We present a model of how MT cells aggregate responses from VI to form such a velocity representation. Two different sets of units, with local receptive fields, receive inputs from motion energy filters. One set of units forms estimates of local motion, while the second set computes the utility of these estimates. Outputs from this second set of units "gate" the outputs from the first set through a gain control mechanism. This active process for selecting only a subset of local motion responses to integrate into more global responses distinguishes our model from previous models of velocity estimation. The model yields accurate velocity estimates in synthetic images containing multiple moving targets of varying size, luminance, and spatial frequency profile and deals well with a number of transparency phenomena.
One way of simplifying neural networks so they generalize better is to add an extra term to the error function that will penalize complexity. Simple versions of this approach include penalizing the sum of the squares of the weights or penalizing the number of nonzero weights. We propose a more complicated penalty term in which the distribution of weight values is modeled as a mixture of multiple gaussians. A set of weights is simple if the weights have high probability density under the mixture model. This can be achieved by clustering the weights into subsets with the weights in each cluster having very similar values. Since we do not know the appropriate means or variances of the clusters in advance, we allow the parameters of the mixture model to adapt at the same time as the network learns. Simulations on two different problems demonstrate that this complexity term is more effective than previous complexity terms.
We present a new supervised learning procedure for systems composed of many separate networks, each of which learns to handle a subset of the complete set of training cases. The new procedure can be viewed either as a modular version of a multilayer supervised network, or as an associative version of competitive learning. It therefore provides a new link between these two apparently different approaches. We demonstrate that the learning procedure divides up a vowel discrimination task into appropriate subtasks, each of which can be solved by a very simple expert network.
We compare the performance of the modular architecture, composed of competing expert networks, suggested by Jacobs, Jordan, Nowlan and Hinton (1991) to the performance of a single back-propagation network on a complex, but low-dimensional, vowel recognition task. Simulations reveal that this system is capable of uncovering interesting decompositions in a complex task. The type of decomposition is strongly influenced by the nature of the input to the gating network that decides which expert to use for each case. The modular architecture also exhibits consistently better generalization on many variations of the task.
An algorithm that is widely used for adaptive equalization in current modems is the “bootstrap” or “decision-directed” version of the Widrow-Hoff rule. We show that this algorithm can be viewed as an unsupervised clustering algorithm in which the data points are transformed so that they form two clusters that are as tight as possible. The standard algorithm performs gradient ascent in a crude model of the log likelihood of generating the transformed data points from two gaussian distributions with fixed centers. Better convergence is achieved by using the exact gradient of the log likelihood.
One popular class of unsupervised algorithms are competitive algorithms. In the traditional view of competition, only one competitor, the winner, adapts for any given case. I propose to view competitive adaptation as attempting to fit a blend of simple probability generators (such as gaussians) to a set of data-points. The maximum likelihood fit of a model of this type suggests a softer form of competition, in which all competitors adapt in proportion to the relative probability that the input came from each competitor. I investigate one application of the soft competitive model, placement of radial basis function centers for function interpolation, and show that the soft model can give better performance with little additional computational cost.
Bruce Porter合作论文数Department of Computer Science The University of Texas at Austin1
John Moody合作论文数Computational Finance Lab;International Computer Science Institute;Berkeley & Portland1