Nanowire networks (NWNs) mimic the brain’s neurosynaptic connectivity and emergent dynamics. Consequently, NWNs may also emulate the synaptic processes that enable higher-order cognitive functions such as learning and memory. A quintessential cognitive task used to measure human working memory is the n -back task. In this study, task variations inspired by the n -back task are implemented in a NWN device, and external feedback is applied to emulate brain-like supervised and reinforcement learning. NWNs are found to retain information in working memory to at least n = 7 steps back, remarkably similar to the originally proposed “seven plus or minus two” rule for human subjects. Simulations elucidate how synapse-like NWN junction plasticity depends on previous synaptic modifications, analogous to “synaptic metaplasticity” in the brain, and how memory is consolidated via strengthening and pruning of synaptic conductance pathways.
We present multiplexed gradient descent (MGD), a gradient descent framework designed to easily train analog or digital neural networks in hardware. MGD utilizes zero-order optimization techniques for online training of hardware neural networks. We demonstrate its ability to train neural networks on modern machine learning datasets, including CIFAR-10 and Fashion-MNIST, and compare its performance to backpropagation. Assuming realistic timescales and hardware parameters, our results indicate that these optimization techniques can train a network on emerging hardware platforms orders of magnitude faster than the wall-clock time of training via backpropagation on a standard GPU, even in the presence of imperfect weight updates or device-to-device variations in the hardware. We additionally describe how it can be applied to existing hardware as part of chip-in-the-loop training or integrated directly at the hardware level. Crucially, because the MGD framework is model-free it can be applied to nearly any hardware platform with tunable parameters, and its gradient descent process can be optimized to compensate for specific hardware limitations, such as slow parameter-update speeds or limited input bandwidth.
In this short paper, we will introduce a simple model for quantifying philosophical vagueness. There is growing interest in this endeavor to quantify vague concepts of consciousness, agency, etc. We will then discuss some of the implications of this model including the conditions under which the quantification of `nifty' leads to pan-nifty-ism. Understanding this leads to an interesting insight - the reason a framework to quantify consciousness like Integrated Information Theory implies (forms of) panpsychism is because there is favorable structure already implicitly encoded in the construction of the quantification metric.
In this paper, we take a brief look at the advantages and disadvantages of dominant frameworks in consciousness studies -- functionalist and causal structure theories, and use it to motivate a new non-equilibrium thermodynamic framework of consciousness. The main hypothesis in this paper will be two thermodynamic conditions obtained from the non-equilibrium fluctuation theorems -- TCC 1 and 2, that the author proposes as necessary conditions that a system will have to satisfy in order to be 'conscious'. These descriptions will look to specify the functions achieved by a conscious system and restrict the physical structures that achieve them without presupposing either of the two. These represent an attempt to integrate consciousness into established physical law (without invoking untested novel frameworks in quantum mechanics and/or general relativity). We will also discuss it's implications on a wide range of existing questions, including a stance on the hard problem. The paper will also explore why this framework might offer a serious path forward to understanding consciousness (and perhaps even realizing it in artificial systems) as well as laying out some problems and challenges that lie ahead.
As the compute demands for machine learning and artificial intelligence applications continue to grow, neuromorphic hardware has been touted as a potential solution. New emerging devices like memristors, spintronics, atomic switches, etc have shown tremendous potential to replace CMOS-based circuits but have been hindered by multiple challenges with respect to device variability, stochastic behavior and scalability. In this paper we will introduce a Description ↔ Design framework to analyze past successes in computing, understand current problems and identify a path moving forward. Engineering systems with these emerging devices might require the modification of both the type of descriptions of learning that we will design for, and the design methodologies we employ in order to realize these new descriptions. We will explore ideas from complexity engineering and analyze the advantages and challenges they offer over traditional approaches to neuromorphic design with novel computing fabrics. A reservoir computing example is used to understand the specific changes that would accompany in moving towards a complexity engineering approach. The time is ideal for a fundamental rethink of our design methodologies and success will represent a significant shift in how neuromorphic hardware is designed and pave the way for a new paradigm.
The hardware and software foundations laid in the first half of the 20th Century enabled the computing technologies that have transformed the world, but these foundations are now under siege. The current computing paradigm, which is the foundation of much of the current standards of living that we now enjoy, faces fundamental limitations that are evident from several perspectives. In terms of hardware, devices have become so small that we are struggling to eliminate the effects of thermodynamic fluctuations, which are unavoidable at the nanometer scale. In terms of software, our ability to imagine and program effective computational abstractions and implementations are clearly challenged in complex domains. In terms of systems, currently five percent of the power generated in the US is used to run computing systems - this astonishing figure is neither ecologically sustainable nor economically scalable. Economically, the cost of building next-generation semiconductor fabrication plants has soared past 10 billion. All of these difficulties - device scaling, software complexity, adaptability, energy consumption, and fabrication economics - indicate that the current computing paradigm has matured and that continued improvements along this path will be limited. If technological progress is to continue and corresponding social and economic benefits are to continue to accrue, computing must become much more capable, energy efficient, and affordable. We propose that progress in computing can continue under a united, physically grounded, computational paradigm centered on thermodynamics. Herein we propose a research agenda to extend these thermodynamic foundations into complex, non-equilibrium, self-organizing systems and apply them holistically to future computing systems that will harness nature's innate computational capacity. We call this type of computing "Thermodynamic Computing" or TC.
There is a significant amount of interest in the field of big data and machine learning right now. This has been driven by use of sophisticated learning algorithms along with large datasets and powerful computing hardware to achieve extraordinary success in narrow tasks. Such success has been classified as narrow artificial intelligence (AI), in order to distinguish it from general intelligence, which continues to be the holy grail of computing. If we are to progress from narrow to general AI, it is important to have a better understanding of what intelligence is and what it entails. As we seek to reboot computing across the stack, this is an important question to address, to help us identify the optimal devices, architectures and design techniques that will allow us to build the intelligent systems of the future. In this paper, I will review the fundamental ideas and assumptions that have allowed us to achieve computing in artificial systems over the years. Building off these ideas, I will discuss the important distinction between a good example of a system with general intelligence i.e. ourselves, and the intelligence achieved through our current computational approaches. Following this, I will use recent results to explore a new framework - a physically grounded theory of thermodynamic intelligence, and discuss the design paradigm that seeks to achieve such intelligence in systems.
We present the fundamental lower bound on dissipation in feedforward neural networks associated with the combined cost of the training and testing phases. Finite state automata descriptions of output generation and the weight updates during training, are used to derive the corresponding lower bounds in a physically grounded manner. The results are illustrated using a simple perceptron learning the AND classification task. The effects of the learning rate parameter and input probability distribution on the cost of dissipation are studied. Derivation of neural network learning algorithms that minimize the total dissipation cost of training are explored.
The focus of the computing industry continues to shift towards designing and building intelligent systems that can handle and learn from large amounts of data. The availability of powerful processing hardware like GPUs and TPUs, has powered the tremendous success of many sophisticated resource intensive machine learning algorithms. However as device scaling and energy dissipation fast approach the physical limits, we have started to look away from traditional CMOS devices and the von Neumann architecture, and towards neuromorphic and other bio-inspired computing systems built using novel devices such as memristors and photonics. Biological systems like the human brain are a product of millions of years of bottom-up self- organizing evolutionary processes and often operate near the limits of energy efficiency. Improved understanding of these systems, their capabilities and the self- organization processes that produce them from a computational perspective will go a long way in helping us build better learning machines. The use of non- equilibrium thermodynamics and information theory to understand such biological processes have gained significant momentum in recent years. In this paper, based on a prediction focused definition of intelligence, results characterizing the thermodynamic constraints on finite state automata models of intelligent physical systems will be presented. As we seek to reboot computing, these constraints can prove be to extremely beneficial in guiding the design and fabrication of energy efficient intelligent systems.
At roughly kT energy dissipation per operation, the thermodynamic energy efficiency “limits” of Moore's Law were unimaginably far off in the 1960s. However, current computers operate at only 100-10,000 times this limit, forming an argument that historical rates of efficiency scaling must soon slow. This paper reviews the justification for the ~kT per operation limit in the context of processors for von Neumann-class computer architectures of the 1960s. We then reapply the fundamental arguments to contemporary applications and identify a new direction for future computing in which the ultimate efficiency limits would be much further out. New nanodevices with high-level functions that aggregate the functionality of several logic gates and some local memory may be the right building blocks for much more energy efficient execution of emerging applications-such as neural networks.
A hierarchical methodology for the determination of fundamental lower bounds on energy dissipation in nanoprocessors is described. The methodology aims to bridge computational description of nanoprocessors at the instruction-set-architecture level to their physical description at the level of dynamical laws and entropic inequalities. The ultimate objective is hierarchical sets of energy dissipation bounds for nanoprocessors that have the character and predictive force of thermodynamic laws and can be used to understand and evaluate the ultimate performance limits and resource requirements of future nanocomputing systems. The methodology is applied to a simple processor to demonstrate instruction- and architecture-level dissipation analyses.
Performance-power-area tradeoffs for on-chip error correction with unreliable decoders are considered from a physical-information-theoretic perspective. Fundamental upper bounds on information throughput under decoding power and heat removal constraints are obtained as a function of decoder efficacy for (n, k) linear block codes. Evaluation of these bounds is demonstrated for implementation of a (7,4) Hamming code in an illustrative example involving a classical bit-flip channel and a noisy decoder, and extension to codes with memory is briefly discussed. These bounds depend only on the statistical decoder input characteristics, the efficacy with which the decoder implements its intended function, and the decoder area; they are not derived from technology-specific assumptions about the decoder circuitry. As such, they can be used to explore tradeoffs between ultimate capabilities and resource requirements in the kinds of error correction scenarios that may be encountered in post-CMOS nanocomputing systems, where the decoding hardware designed to enhance reliability may itself be unreliable.
Irreversibility and dissipation in finite-state automata (FSA) are considered from a physical-information-theoretic perspective. A quantitative measure for the computational irreversibility of finite automata is introduced, and a fundamental lower bound on the average energy dissipated per state transition is obtained and expressed in terms of FSA irreversibility. The irreversibility measure and energy bound are germane to any realization of a deterministic automaton that faithfully registers abstract FSA states in distinguishable states of a physical system coupled to a thermal environment, and that evolves via a sequence of interactions with an external system holding a physical instantiation of a random input string. The central result, which is shown to follow from quantum dynamics and entropic inequalities alone, can be regarded as a generalization of Landauerʼs Principle applicable to FSAs and tailorable to specified automata. Application to a simple FSA is illustrated.
Christof Teuscher合作论文数Los Alamos National Laboratory1