Machine Learning and Knowledge Discovery in Databases Research Track(2026)
Kyushu University
被引用0|浏览0
摘要
Many large-scale recommender systems adopt a two-stage architecture to balance computational efficiency and performance. Off-policy evaluation (OPE) of candidate generators in such two-stage recommender systems is challenging because action-set changes can violate the overlap assumption required by standard importance-weighting estimators. Existing proxy-based methods address this by aggregating actions into equivalence classes, but their reliance on heuristics can incur substantial approximation bias. In this work, we propose Latent-Proxy Alignment IPS (LPAIPS), which addresses this limitation by learning reward-aware latent proxies. LPAIPS learns a mapping from context–action pairs to discrete latent classes and promotes latent-space overlap through distribution alignment. As a theoretical contribution, we formalize sufficient conditions for latent overlap control and characterize the bias–variance trade-off in terms of mean squared error (MSE). Through experiments on both synthetic data and a real-world public dataset (Open Bandit Dataset), we show that LPAIPS achieves substantially lower MSE than IPS, IIPS, and Proxy IPS in severe support-mismatch settings, with up to a 9.2-fold reduction in MSE relative to Proxy IPS under severe support mismatch and up to a 74