Spinal codes are a recently proposed capacity-achieving rateless code. While hardware encoding of spinal codes is straightforward, the design of an efficient, high-speed hardware decoder poses signif-icant challenges. We present the first such decoder. By relaxing data dependencies inherent in the classic M-algorithm decoder, we obtain area and throughput competitive with 3GPP turbo codes as well as greatly reduced latency and complexity. The enabling architectural feature is a novel “ α - β ” incremental approximate selection algorithm. We also present a method for obtaining hints which anticipate successful or failed decoding, permitting early termination and/or feedback-driven adaptation of the decoding parameters. We have validated our implementation in FPGA with on-air testing. Provisional hardware synthesis suggests that a near-capacity implementation of spinal codes can achieve a throughput of 12.5 Mbps in a 65 nm technology while using substantially less area than competitive 3GPP turbo code implementations.
Systems, methods, and apparatuses relating to debugging a configurable spatial accelerator are described. In one embodiment, a processor includes a plurality of processing elements and an interconnect network between the plurality of processing elements to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is to be overlaid into the interconnect network and the plurality of processing elements with each node represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements are to perform an operation by a respective, incoming operand set arriving at each of the dataflow operators of the plurality of processing elements. At least a first of the plurality of processing elements is to enter a halted state in response to being represented as a first of the plurality of dataflow operators.
Systems, methods and apparatus relating to a configurable spatial accelerator are described. In one embodiment, a processor includes a synchronizer circuit coupled between a first tile interconnect network and a second tile interconnect network and including memory for storing the data to be sent between the first tile interconnect network and the second tile interconnect network wherein the synchronizer circuit converts the data from the memory between a first voltage or a first frequency of the first tile and a second voltage or frequency of the second tile to generate converted data and the converted data between the interconnection network of the first tile and the interconnect network of the second tile.
Systems, methods, and apparatuses relating to conditional queues in a configurable spatial accelerator are described. In one embodiment, a configurable spatial accelerator includes a first output buffer of a first processing element coupled to a first input buffer of a second processing element and a second input buffer of a third processing element via a data path that is to send a dataflow token to the first input buffer of the second processing element and the second input buffer of the third processing element when the dataflow token is received in the first output buffer of the first processing element; a first backpressure path from the first input buffer of the second processing element to the first processing element to indicate to the first processing element when storage is not available in the first input buffer of the second processing element; a second backpressure path from the second input buffer of the third processing element to the first processing element to indicate to the first processing element when storage is not available in the second input buffer of the third processing element; and a scheduler of the second processing element to cause storage of the dataflow token from the data path into the first input buffer of the second processing element when both the first backpressure path indicates storage is available in the first input buffer of the second processing element and a conditional token received in a conditional queue of the second processing element from another processing element is a true conditional token.
An integrated circuit includes a memory interface, coupled to a memory to store data corresponding to instructions, and an operations queue to buffer memory operations corresponding to the instructions. The integrated circuit may include acceleration hardware to execute a sub-program corresponding to the instructions. A set of input queues may include an address queue to receive, from the acceleration hardware, an address of the memory associated with a second memory operation of the memory operations, and a dependency queue to receive, from the acceleration hardware, a dependency token associated with the address. The dependency token indicates a dependency on data generated by a first memory operation of the memory operations. A scheduler circuit may schedule issuance of the second memory operation to the memory in response to the dependency queue receiving the dependency token and the address queue receiving the address.
Methods and apparatuses relating to distributed memory hazard detection and error recovery are described. In one embodiment, a memory circuit includes a memory interface circuit to service memory requests from a spatial array of processing elements for data stored in a plurality of cache banks; and a hazard detection circuit in each of the plurality of cache banks, wherein a first hazard detection circuit for a speculative memory load request from the memory interface circuit, that is marked with a potential dynamic data dependency, to an address within a first cache bank of the first hazard detection circuit, is to mark the address for tracking of other memory requests to the address, store data from the address in speculative completion storage, and send the data from the speculative completion storage to the spatial array of processing elements when a memory dependency token is received for the speculative memory load request.
Verfahren und Einrichtungen im Zusammenhang mit bevorzugter Auslegung in raumlichen Arrays werden beschrieben. In einer Ausfuhrungsform umfasst ein Prozessor Verarbeitungselemente; ein Verbindungsnetzwerk zwischen den Verarbeitungselementen; und eine Auslegungssteuerung, die mit einer ersten und einer zweiten, unterschiedlichen Teilmenge der mehreren Verarbeitungselemente gekoppelt ist, wobei die erste Teilmenge eine Ausgabe hat, der mit einer Eingabe der zweiten, unterschiedlichen Teilmenge gekoppelt ist, wobei die Auslegungssteuerung dazu dient, das Verbindungsnetzwerk zwischen der ersten Teilmenge und der zweiten, unterschiedlichen Teilmenge der mehreren Verarbeitungselemente auszulegen, um Kommunikation auf dem Verbindungsnetzwerk zwischen der ersten Teilmenge und der zweiten, unterschiedlichen Teilmenge nicht zu erlauben, wenn ein Bevorzugungsbit auf einen ersten Wert gesetzt ist, und Kommunikation auf dem Verbindungsnetzwerk zwischen der ersten Teilmenge und der zweiten, unterschiedlichen Teilmenge zu erlauben, wenn das Bevorzugungsbit auf einen zweiten Wert gesetzt ist.
Es werden Systeme, Verfahren und Vorrichtungen in Bezug auf konfigurierbare netzwerkbasierte Datenflussoperator-Schaltungen beschrieben. In einer Ausfuhrungsform enthalt ein Prozessor ein dreidimensionales Array von Verarbeitungselementen und ein paketvermittelndes Kommunikationsnetzwerk, um Daten innerhalb des dreidimensionales Arrays zwischen Verarbeitungselementen entsprechend einem Datenflussgraphen zu routen, um eine erste Datenflussoperation des Datenflussgraphen durchzufuhren, wobei das paketvermittelnde Kommunikationsnetzwerk des Weiteren mehrere Netzwerk-Datenflussendpunktschaltungen umfasst, um eine zweite Datenflussoperation des Datenflussgraphen durchzufuhren.
Systems, methods, and apparatuses relating to integrated performance monitoring in a configurable spatial accelerator are described. In one embodiment, a configurable spatial accelerator includes a first performance monitoring circuit coupled to a first proper subset of processing elements by a network to receive at least one monitoring value from each of the first plurality of the processing elements, generate a first aggregated monitoring value based on the at least one monitoring value from each of the first plurality of the processing elements, and send the first aggregated monitoring value to a performance manager circuit on a different network when a first threshold value is exceeded by the first aggregated monitoring value; and the performance manager circuit is to perform an action based on the first aggregated monitoring value.
Verfahren und Gerate in Zusammenhang mit konfigurierbarem Clock-Gating in raumlichen Arrays sind beschrieben. Bei einer Ausfuhrungsform weist ein Prozessor Verarbeitungselemente; ein Interconnect-Netzwerk zwischen den Verarbeitungselementen; und einen Konfigurationscontroller auf, der mit einem ersten Verarbeitungselement und einem zweiten Verarbeitungselement der Vielzahl von Verarbeitungselementen gekoppelt ist, und wobei das erste Verarbeitungselement einen Ausgang mit einem Eingang des zweiten Verarbeitungselements gekoppelt hat, um das zweite Verarbeitungselement zu konfigurieren, um mindestens ein getaktetes Bauelement des zweiten Verarbeitungselements zu clock-gaten, und das erste Verarbeitungselement zu konfigurieren, um ein Wiederaktivierungssignal auf dem Interconnect-Netzwerk zu dem zweiten Verarbeitungselement zu senden, um das mindestens eine getaktete Bauelement des zweiten Verarbeitungselements wieder zu aktivieren, wenn Daten von dem ersten Verarbeitungselement zu dem zweiten Verarbeitungselement zu senden sind.
Es werden Systeme, Verfahren und Vorrichtungen bezuglich eines konfigurierbaren raumlichen Beschleunigers beschrieben. In einer Ausfuhrungsform weist ein Prozessor mehrere Verarbeitungselemente auf; und ein Zwischenverbindungsnetz zwischen den mehreren Verarbeitungselementen zum Empfangen einer Eingabe eines Datenflussgraphen, der mehrere Knoten umfasst, wobei der Datenflussgraph in das Zwischenverbindungsnetz und die mehreren Verarbeitungselemente zu uberlagern, wobei jeder Knoten als ein Datenflussoperator in den mehreren Verarbeitungselementen reprasentiert ist, und die mehreren Verarbeitungselemente eine atomare Operation durchzufuhren haben, wenn ein eingehender Operand bei den mehreren Verarbeitungselementen eingeht.
Systems, methods, and apparatuses relating to operations in a configurable spatial accelerator are described. In one embodiment, a configurable spatial accelerator includes a first processing element that includes a configuration register within the first processing element to store a configuration value that causes the first processing element to perform an operation according to the configuration value, a plurality of input queues, an input controller to control enqueue and dequeue of values into the plurality of input queues according to the configuration value, a plurality of output queues, and an output controller to control enqueue and dequeue of values into the plurality of output queues according to the configuration value.
Es werden Systeme, Verfahren und Vorrichtungen bezuglich eines konfigurierbaren raumlichen Beschleunigers beschrieben. In einer Ausfuhrungsform weist ein Prozessor mehrere Verarbeitungselemente auf; und ein Zwischenverbindungsnetz zwischen den mehreren Verarbeitungselementen zum Empfangen einer Eingabe eines Datenflussgraphen, der mehrere Knoten umfasst, wobei der Datenflussgraph in das Zwischenverbindungsnetz und die mehreren Verarbeitungselemente zu uberlagern, wobei jeder Knoten als ein Datenflussoperator in den mehreren Verarbeitungselementen reprasentiert ist, und die mehreren Verarbeitungselemente eine Operation durchzufuhren haben, wenn ein eingehender Operand bei den mehreren Verarbeitungselementen eingeht. Mindestens eines der mehreren Verarbeitungselemente weist mehrere Steuereingaben auf.
Systems, methods, and apparatuses relating to a configurable spatial accelerator are described. In one embodiment, a processor includes a plurality of processing elements; and an interconnect network between the plurality of processing elements to receive an input of two dataflow graphs each comprising a plurality of nodes, wherein a first dataflow graph and a second dataflow graph are be overlaid into a first and second portion, respectively, of the interconnect network and a first and second subset, respectively, of the plurality of processing elements with each node represented as a dataflow operator in the plurality of processing elements, and the first and second subsets of the plurality of processing elements are to perform a first and second operation, respectively, when incoming first and second, respectively, operand sets arrive at the plurality of processing elements.
Described are systems, methods, and apparatus related to a sequencer dataflow operator of a configurable spatial accelerator. In one embodiment, an interconnect network between a plurality of processing elements receives an input of a data flow graph comprising a plurality of nodes forming a loop construct superimposing the data flow graph into the interconnect network and the plurality of processing elements, each node being represented as a data flow operator in the plurality of processing elements at least one data flow operator is controlled by a sequencer data flow operator of the plurality of processing elements and the plurality of processing elements perform an operation when an incoming operand set arrives at the plurality of processing elements, and the sequencer data flow operator generates control signals for the at least one data flow operator in the plurality of processing elements.
Systems, methods, and apparatuses relating to a configurable spatial accelerator are described. In one embodiment, a processor includes a plurality of processing elements; and an interconnect network between the plurality of processing elements to receive an input of a dataflow graph comprising a plurality of nodes, wherein the dataflow graph is to be overlaid into the interconnect network and the plurality of processing elements with each node represented as a dataflow operator in the plurality of processing elements, and the plurality of processing elements is to perform an operation when an incoming operand set arrives at the plurality of processing elements. The processor also includes a streamer element to prefetch the incoming operand set from two or more levels of a memory system.
An integrated circuit includes a processor to execute instructions and to interact with memory, and acceleration hardware, to execute a sub-program corresponding to instructions. A set of input queues includes a store address queue to receive, from the acceleration hardware, a first address of the memory, the first address associated with a store operation and a store data queue to receive, from the acceleration hardware, first data to be stored at the first address of the memory. The set of input queues also includes a completion queue to buffer response data for a load operation. A disambiguator circuit, coupled to the set of input queues and the memory, is to, responsive to determining the load operation, which succeeds the store operation, has an address conflict with the first address, copy the first data from the store data queue into the completion queue for the load operation.
Systems, methods, and apparatuses relating to remote memory access in a configurable spatial accelerator are described. In one embodiment, a configurable spatial accelerator includes a first memory interface circuit coupled to a first processing element and a cache, the first memory interface circuit to issue a memory request to the cache, the memory request comprising a field to identify a second memory interface circuit as a receiver of data for the memory request; and the second memory interface circuit coupled to a second processing element and the cache, the second memory interface circuit to send a credit return value to the first memory interface circuit, to cause the first memory interface circuit to mark the memory request as complete, when the data for the memory request arrives at the second memory interface circuit and a completion configuration register of the second memory interface circuit is set to a remote response value.
Es werden Systeme, Verfahren und Einrichtungen beschrieben, die sich auf Multicast in einem konfigurierbaren raumlichen Beschleuniger beziehen. In einer Ausfuhrungsform enthalt ein Beschleuniger einen ersten Ausgabepuffer eines ersten Verarbeitungselements, der mit einem ersten Eingabepuffer eines zweiten Verarbeitungselements und einem zweiten Eingabepuffer eines dritten Verarbeitungselements gekoppelt ist; und das erste Verarbeitungselement bestimmt, dass es fahig war, eine Ubertragung in einem vorhergehenden Zyklus fertigzustellen, wenn das erste Verarbeitungselement fur sowohl das zweite Verarbeitungselement als auch das dritte Verarbeitungselement beobachtet hat, dass entweder ein Spekulationswert auf einen Wert eingestellt wurde, um anzugeben, dass ein Datenfluss-Token in seinem Puffer gespeichert wurde (wie z. B. durch einen Empfangswert (z. B. Bit) angegeben ist), oder ein „Backpressure“-Wert auf einen Wert eingestellt wurde, um anzugeben, dass Speicher in seinem Eingabepuffer verfugbar sein soll, bevor das Datenfluss-Token aus der Warteschlange des ersten Ausgabepuffer entfernt wird.
George A. Constantinides合作论文数Early Career Researcher Institute, Imperial College London3
Tryggve Fossum合作论文数1