Multi-modal Offline Decomposition for Expertise: A Structural Approach to Robust Offline Meta-RL | AMiner
Multi-modal Offline Decomposition for Expertise: A Structural Approach to Robust Offline Meta-RL
Chun-Yu Lin,Pei-Yuan Wu
Machine Learning and Knowledge Discovery in Databases Research Track(2026)
National Taiwan University
被引用0|浏览0
摘要
Offline Meta-Reinforcement Learning (OMRL) enables agents to generalize to unseen tasks using pre-collected static datasets. However, a critical challenge arises when these datasets are of mixed quality, originating from multi-modal behavior policies that blend expert demonstrations with sub-optimal explorations. While current approaches prioritize behavior-invariant purification, this strategy does not necessarily imply that the learned representation preserves the minimal sufficient statistics of task-defining transition and reward dynamics, resulting in an under-specified latent space that fails to resolve unseen task structures during inference. To overcome this, we propose Multi-modal Offline Decomposition for Expertise (MODE), a framework to address the challenge of learning optimal meta-policy from mixed-quality offline dataset. MODE consists of three components: (1) a Recurrent GM-VAE context model that incorporates Recurrent Neural Network (RNN) and Gaussian Mixture Model (GMM) on top of Variational Auto-encoder (VAE) architecture to robustly capture task representations from temporal features; (2) a GMM-based behavior decomposition process that distinguishes between varying quality modes within the dataset; and (3) a Lower Confidence Bound (LCB) selection mechanism that guides a hyper-policy to selectively combine the most promising Gaussian components based on value estimation and uncertainty. Experimental results demonstrate that our method consistently outperforms state-of-the-art baselines on continuous control benchmarks in MuJoCo, particularly in scenarios involving mixed-qualities. Our method not only achieves higher asymptotic returns, but also exhibits significantly lower variance, highlighting its robustness to task variations and data noise.
更多
查看译文
关键词
Offline Meta Reinforcement Learning,Gaussian Mixture Model,Lower Confidence Bound