Multimodal interaction in transactional systems raises a mixed-initiative design question: should users or AI systems decide whether users click, speak, or use chat at each step of the workflow? A controlled experiment was conducted using a multi-step gifting checkout flow to examine the effects of AI governing the input modality in multimodal interfaces on participants’ performance, usability, and cognitive experience. Participants completed identical gifting workflows under three conditions: (A) traditional on-screen interaction (B) user-selected input modality (choosing between traditional, speech, or chat input per step), and (C) AI-selected modality (system-assigned input modality for each step). When the LLM governed modality selection, participants took significantly longer to complete tasks ( η ^2 = 0.416) and reported lower ratings of usability, satisfaction, trust, and perceived control than in the user-selectable condition. When participants governed their own modality choice, task completion performance was comparable to that of traditional on-screen interaction; trust ratings and perceived control remained high; and perceived workload did not differ across conditions. Notably, despite the availability of multimodal options, participants relied exclusively on standard on-screen input, without adopting voice or chat modalities. These results demonstrate a gap between modality availability and actual usage. Finally, AI-governed interaction is not inherently beneficial when decision authority over input modality transfers without constraint.
更多
查看译文
关键词
Multimodal interaction,Mixed-initiative interaction,Human-AI interaction,Modality governance,Voice interfaces,Conversational interfaces,E-commerce checkout,Trust in automation