Open Coherent Accelerator Processor Interface (OpenCAPI) is a new industry-standard device interface that enables the development of host-agnostic devices that can coherently connect to any host platform that supports the OpenCAPI standard. This in turn allows such devices to coherently cache host memory to facilitate accelerator execution, perform direct memory access and atomics to host memory, send messages and interrupts to the host, and act as a host memory home agent. OpenCAPI utilizes high-frequency differential signaling technology while providing the high bandwidth and low latency needed by advanced accelerators. OpenCAPI encapsulates the serializing cache access and address translation constructs in high-speed host silicon technology to minimize overhead and design complexity in attached silicon such as field-programmable gate arrays and application-specific integrated circuits. Finally, OpenCAPI architecturally ties together transaction layer, link layer, and physical layer attributes to optimally align to high serializer/deserializer (SerDes) ratios and enable high-bandwidth, highly parallel exploitation of attached silicon.
IBM POWER9 is a family of processor chips designed to serve a diverse set of workloads. New features have been added to POWER9 to address emerging workloads such as cognitive and artificial intelligence applications. POWER9 also further enhances features introduced in IBM POWER8 for big data and cloud applications. Distinct chips using common intellectual property building blocks are provided to enable enterprise applications requiring large symmetric multiprocessor servers with large memory footprints, as well as one to two socket industry form-factor servers. In this paper, we describe new POWER9 features for both system types. Several highly differentiated new features are described in other papers in this issue of the IBM Journal, and they provide a more in-depth description of their unique design aspects.
In this paper we will present new adaptive routing algorithms for faulty processor arrays. Past research has shown that packet switched based communication performance in mesh connected networks is significantly degraded by the presence of faulty processors. Nondeterministic routing algorithms have been developed based on transport modeling of packet flow in disordered arrays. By utilizing nondeterministic routing strategies, based on biased random walkers, we can implement deadlock free routing, at the expense of not following the shortest path. These algorithms will be shown to be capable of increasing network bandwidth in the presence of faulty processors and interconnects. These algorithms offer an alternative to conventional adaptive routing techniques by utilizing a computationally simple algorithm based on local (nearest neighbor) information. Although we concentrate our efforts on 2 dimensional processor arrays, the algorithm are also suitable for higher dimensional topologies such as hypercubes.
PETSC2.0 is a software toolkit for portable, parallel (and serial) numerical solution of partial differential equations and minimization problems. It includes software for the solution of linear and nonlinear systems of equations. These codes are written in a data-structure-neutral manner to enable easy reuse and flexibility.
I. [6] The Work-Time (W-T) presentation of EREW sequence reduction (Algorithm 2 in PRAM handout) has work complexity WW(nn) = Ο(nn) and step complexity SS(nn) = Ο(lgnn). Following the strategy of Brent’s theorem, the translation of this algorithm will yield a pp processor EREW PRAM program with running time TTCC(nn,pp) = Ο(nn pp ⁄ + lgnn) (a) Construct an alternate sequence reduction algorithm directly for the bare bones EREW PRAM with running time TTCC(nn,pp) = Ο(nn pp ⁄ + lg pp). (b) Explain why your solution to (a) cannot be expressed in the W-T model.
Cevdet Aykanat合作论文数Computer Engineering Department of Bilkent University54