The rise of GPUs as accelerators is reasoned by the claim for power efficient architectures. Despite their remarkable performance while being still power efficient, the main bottleneck of distributed GPU computing remains communication. A hybrid model is required, because the CPU is used to handle communication. Besides performance aspects, this also adds complexity and makes programming even more difficult. Furthermore this affects energy consumption negatively, as either communication models don’t match the GPU architecture, or communication is off-loaded to the CPU which prevents it from entering power