The classical Multi-armed Bandit (MAB) framework relies on immediate access to reward feedback, yet in many practical settings rewards are inaccessible or severely delayed. To address this gap, inverse MAB seeks to recover the latent reward function by observing the actions of one or more demonstrators. In this work, we study the problem of learning from multiple noisy demonstrations in a stochastic MAB environment, in the absence of reward information. Four algorithms for learning from noisily optimal demonstrators, that is, demonstrators who choose the optimal arm with fixed probability and a suboptimal one otherwise, are considered; three adapted from existing literature and one newly proposed. Their performance is evaluated both theoretically, through regret bounds, and empirically, via numerical tests.