2025 IEEE International Conference on Data Mining Workshops (ICDMW)(2025)
University of Southern California
被引用1|浏览0
摘要
Machine Learning as a Service (MLaaS) systems face a fundamental tension between regulatory demands for explainability and security vulnerabilities from model extraction attacks. While global regulations increasingly mandate transparent AI systems, emerging evidence suggests increasing model explainability can increase the vulnerability to model extraction attacks. This paper advocates for investigating the underexplored relationship between model extractability and explainability. Like how bulletproof glass provides visibility without vulnerability, we envision theoretical frameworks enabling MLaaS systems to achieve human-level explainability without enabling extraction attacks. Impeding this vision, we identify three critical technical gaps: the absence of general frameworks to predict extraction risk, standardized extractability metrics, and predictive security approaches. Institutional fragmentation across model owners, security teams, explainability researchers, and regulators further blocks coordinated solutions. The urgency of this agenda stems from rapid MLaaS adoption, expanding regulatory mandates, and increasing AI-related security incidents. Without an explainability-extractability framework, organizations must choose not to comply with regulatory transparency requirements or face exploitation through model theft and training data extraction. Success in this area can extend to complex systems beyond AI, from financial markets to government operations, wherever transparency is desirable but potentially dangerous.
更多
查看译文
关键词
Machine learning as a service,model extraction attacks,explainable artificial intelligence,proactive security