Simultaneous Vision-Language Knowledge Transfer for Zero-Shot Human-Object Interaction Detection. | AMiner