A pattern-matrix learning algorithm for adaptive MDPs: the regularly communicating case (不確実な状況における意思決定の理論と応用--RIMS研究集会報告集) | AMiner