Chrome Extension
WeChat Mini Program
Use on ChatGLM

Improving Alignment and Robustness with Circuit Breakers

Andy Zou, Long Phan, Justin Wang, Derek Duenas, Maxwell Lin,Maksym Andriushchenko,J Zico Kolter,Matt Fredrikson,Dan Hendrycks

NeurIPS 2024(2024)

Cited 0|Views21
Key words
alignment,adversarial robustness,adversarial attacks,harmfulness,security,reliability,ML safety,AI safety
AI Read Science
Must-Reading Tree
Example
Generate MRT to find the research sequence of this paper
Chat Paper
Summary is being generated by the instructions you defined