Optimizing Selective Protection for CNN Resilience

Abdulrahman Mahmoud,Siva Kumar Sastry Hari,Christopher W. Fletcher,Sarita V. Adve,Charbel Sakr,Naresh R. Shanbhag,Pavlo Molchanov,Michael B. Sullivan,Timothy Tsai,Stephen W. Keckler

2021 IEEE 32nd International Symposium on Software Reliability Engineering (ISSRE)（2021）

引用 11|浏览40

暂无评分

摘要

As CNNs are being extensively employed in high performance and safety-critical applications that demand high reliability, it is important to ensure that they are resilient to transient hardware errors. Traditional full redundancy solutions provide high error coverage, but the associated overheads are often prohibitively high for resource-constrained systems. In this work, we propose software-directed selective protection techniques to target the most vulnerable work in a CNN, providing a low-cost solution. We propose and evaluate two domain-specific selective protection techniques for CNNs that target different granularities. First, we develop a feature-map level resilience technique (FLR), which identifies and statically protects the most vulnerable feature maps in a CNN. Second, we develop an inference level resilience technique (ILR), which selectively reruns vulnerable inferences by analyzing their output. Third, we show that the combination of both techniques (FILR) is highly efficient, achieving nearly full error coverage (99.78% on average) for quantized inferences via selective protection. Our tunable approach enables developers to evaluate CNN resilience to hardware errors before deployment using MAC operations as overhead for quicker trade-off analysis. For example, targeting 100% error coverage on ResNet50 with FILR requires 20.8% additional MACs, while measurements on a Jetson Xavier GPU shows 4.6% runtime overhead.

查看译文

关键词

Reliability,Vulnerability,Errors,Silent Data Corruptions (SDC),Software directed,Convolutional Neural Networks (CNNs)

AI 理解论文

溯源树

样例

生成溯源树，研究论文发展脉络

Chat Paper

正在生成论文摘要