As social robots increasingly engage in collaborative activities, they must express intentions through interpretable nonverbal cues. In human-robot interaction, intention inference may involve both identifying the target directly communicated by a robot and inferring the robot's underlying intention behind that communicative behavior. This study examined how humans infer a robot's intention when gaze direction and torso orientation are inconsistent. In a collaborative decoration task, participants interacted with a social robot under one of four experimental groups based on three gaze-torso consistency conditions: consistency, head-gaze inconsistency, and eye-gaze inconsistency. Results suggested that target inference remained relatively stable across conditions, whereas inference of underlying intention tended to be higher under inconsistency, especially when the mismatch was produced through eye-gaze. At the session level, eye-gaze inconsistency also increased negative mental state attribution without reducing social acceptance or likeability. These findings support a two-layer view of robot intention inference and suggest that eye-gaze and head-gaze play different roles in robot intention expression.