Machine Learning and Knowledge Discovery in Databases Research Track(2026)
Toronto Metropolitan University
被引用0|浏览0
摘要
We investigate the challenge of promoting diversity in offline reinforcement learning (RL), where agents must develop diverse strategies despite being trained on homogeneous datasets with limited behavioral variation. Existing offline RL approaches, including those leveraging expectation-maximization algorithms for unsupervised clustering, often struggle with either insufficient diversity or performance degradation in such settings. To overcome these limitations, we introduce a novel Unique Behavior objective function that can be directly computed to quantify the distinctiveness between agents, eliminating the need for additional estimators and reducing potential estimation errors. By maximizing uniqueness, our approach encourages agents to learn diverse behaviors effectively, even when the training data lacks variety. Extensive experiments on standard and diverse D4RL benchmarks, together with Atari evaluations, demonstrate that our method consistently achieves stronger quality-diversity trade-offs than DIVEOFF, CLUE, and SORL while maintaining competitive task performance across homogeneous and heterogeneous datasets.
更多
查看译文
关键词
Offline reinforcement learning,Behavior diversity,Homogeneous data