Machine Learning and Knowledge Discovery in Databases Research Track(2026)
Stockholm University
被引用0|浏览0
摘要
We study stochastic bandits with ordered actions, unimodal rewards, and resource constraints, motivated by treatment selection problems where intervention intensity improves outcomes up to a peak while incurring increasingly higher costs. Unlike existing constrained bandit approaches, whose regret bounds typically scale with the number of actions because feasibility and optimality must be explored across an unstructured action set, our setting incorporates two key forms of structure: unimodality of rewards, which reduces optimality learning to local exploration around the empirical peak, and monotonicity of costs, which implies that the feasible region is a contiguous prefix that can be identified through a single threshold. We propose F-OSUB, an algorithm that interleaves feasibility identification with unimodal leader–neighbor exploration and show that it achieves logarithmic regret and logarithmic budget violation with high probability, with constants depending only on local reward gaps rather than the total number of actions. These results demonstrate that exploiting structural properties enables substantially more efficient and safer learning in resource-constrained decision problems such as treatment intensity selection.