Online retailers increasingly offer reusable express packaging, but adoption remains difficult because sustainable options can conflict with the convenience expectations that define online retailing. This article examines reusable packaging adoption as a circular retail service conversion problem. Across three studies with adult online shoppers, we test how digital service nudges and responsibility framing influence adoption through perceived service friction, perceived environmental efficacy, and sustainable service trust, and how these effects depend on return convenience. Study 1 uses a quota-based, post-stratification-weighted survey to show that decision-relevant service nudges reduce perceived service friction and increase adoption intention, while responsibility framing increases perceived environmental efficacy and sustainable service trust. Study 2 uses a discrete choice experiment to estimate trade-offs among nudge timing, return convenience, deposit/refund rules, delivery delay, environmental benefit visibility, and responsibility framing. Study 3 uses an incentivized simulated checkout task to test behaviorally consequential reusable-packaging choice at the decision point. The results show that return convenience and deposit clarity are stronger adoption drivers than generic environmental claims, and that consumers penalize delivery delay more strongly than they reward verified environmental-benefit information. Decision-point nudges increase reusable-packaging checkout choice more than post-checkout reminders, especially when return convenience is high. The findings contribute to retail service convenience, digital nudging, and circular retailing research by showing that sustainable retail adoption depends on service-touchpoint design, not only on environmental attitudes. We position the model as a middle-range service-conversion framework grounded in service convenience, limited-attention decision-making, and sustainable-consumption efficacy and trust.
This study investigates how multi-attribute health-related tasks shape users’ observed choices among single-modal and multimodal information combinations. It examines modality configuration as a task-triggered, micro-level instantiation of task–technology fit in personalized health information systems. Grounded in task facet classification theory, the study conceptualizes health tasks along three dimensions: task type, task complexity, and topic group. These effects are examined through a quasi-experimental design based on simulated work task scenarios. A total of 120 participants were recruited, yielding 25,000 valid health information records selected and saved in a researcher-supervised controlled live-web SWTS setting. Distributional differences in modality choices across task configurations are examined, and the findings are triangulated using descriptive odds ratios and multinomial logit models to assess the consistency of effect direction and relative strength in task–modality relationships. The results show that users’ observed modality choices are highly concentrated in a small number of dominant text-centered single-modal and multimodal formats, indicating a clear distributional concentration pattern rather than indiscriminate multimodal use. Systematic distributional differences are observed across task configurations in the selection proportions of key modality combinations, including text-only, text–voice, text–image, text–image–video, and text–video formats, suggesting that task attributes are associated with observed modality choices in multimodal information system use. By translating task attributes into implementable modality configuration logic, this study operationalizes modality configuration as a task-triggered, micro-level instantiation of task–technology fit and provides actionable principles for adaptive and risk-sensitive modality design in personalized and generative AI–enabled health information systems.
With the continuous development of artificial intelligence and autonomous driving technology, cars are transitioning from traditional transportation to user-centered intelligent mobile spaces. In order to enhance drivers' road safety awareness and entertainment experience, this study explores the design of multi-modal games in smart cockpits in the context of L3-L4 autonomous driving based on the theory of Situation Awareness (SA). First, using SA and KANO models, key scenarios and design elements are identified, and a method combining contextawareness and avatar design is proposed. Second, typical driving scenarios are identified through user interviews and evaluation scales, and experimental studies are conducted to analyze and refine the multi-modal game interaction information in the smart cockpit environment. Finally, based on Pokemon-style, an AR game design solution suitable for a smart cockpit was developed. The 2 (with and without game interaction) x 4 (vision, vision + hearing, vision + touch, vision + hearing + touch) mixed factorial experimental design was designed, and the proposed design scheme was evaluated using a driving simulator. The experimental results demonstrate that the game interaction design for L3-L4 autonomous driving conditions significantly enhances driver perception of road and environmental information, reduces reaction time in emergencies, improves user experience, and increases trust in the autonomous driving system compared to the non-game interaction design.
As automotive cockpits evolve into proactive interaction environments supported by End-to-End (E2E) Large Language Models (LLMs), their opaque black-box reasoning processes may create a sensing-design disconnect, contributing to automation surprise and reduced human-AI trust. To address this transparency challenge, this study repositions Explainable AI (XAI) from a post-hoc diagnostic tool toward a design-oriented methodology. We propose the Explainable End-to-End Cockpit Agent (EECA) framework, which integrates multi-sensor fusion, LLM-based reasoning, and structured graph logic. An Explainable Interaction Design Knowledge Graph further maps multimodal driver states to experience goals, design attributes, explanation cues, and evaluation indicators, thereby supporting the generation of transparent and explainable multimodal interaction prompts. A hardware-software prototype was evaluated across 200 complex cockpit scenarios. Compared with selected baseline models, EECA achieved an F1-score of 0.892 and an information gain of 8.3 under edge-side inference conditions. Furthermore, a simulator-based study with 30 participants provided preliminary mixed-method evidence that EECA was associated with shorter driver takeover reaction time and higher trust-related user-experience ratings under the tested cockpit scenarios.
Although social media extensively captures user preferences and aesthetic intentions, its highly unstructured and emotionally expressive semantics have long been regarded as noise that is difficult to translate into reliable generative cues. Here we show that social media semantics contain stable patterns that can be systematically abstracted and validated, and that these patterns exert interpretable effects on the structural quality and semantic consistency of three-dimensional (3D) generation. Using 60,500 real-world social media posts, we validate this finding through CoX-3D, an explainable generative framework that introduces an Explainable Mediation Layer to bridge macroscopic social computing with microscopic physical rendering. Incorporating these stable semantic patterns reduces structural reconstruction error by over 50% and improves semantic consistency, measured by CLIP similarity, by 13.1%. Semantic activation and attribution analyses further reveal differentiated roles of semantic factors in geometric and stylistic decisions. These results demonstrate that social media semantics constitute a systematically modelable and interpretable cognitive resource, establishing a generalizable paradigm for socially informed generative 3D design.
To explore the potential knowledge in natural language and achieve trusted content generation in human-computer collaboration environments, this paper proposes a product design generation and explainable design optimisation (GEDO) method that fuses cross-modal semantic associations between text and images. Firstly, big data is collected from social media to construct a text-image pair dataset, and a multi-layer perceptron (MLP) is used to realise cross-modal feature alignment. Second, based on feature modelling, the gradient-weighted class activation map (Grad-CAM) method is further introduced to visualise and explain the model decision-making process, to enhance the transparency and controllability of the design generation process. Finally, an interpretable network with a ResNet-Grad-CAM mechanism is constructed for optimising image generation and semantic alignment effects. The results show that GEDO achieves 93.71% accuracy as well as a 0.68 confidence level in the graphic alignment task, which is better than the baseline model. Among them, the accuracy is improved by 1.2% and the confidence level is improved by 15.3% compared with the optimal baseline model. The method in this paper not only improves the accuracy and transparency of the design process but also provides an effective path for trusted content generation under human-computer collaboration.
Amid platform-based collaboration, expanding digital government interfaces, and the commoditisation of data, corporate digital innovation increasingly depends on the interoperability and coordination of cross-organisational data and information systems. However, governance interface mismatches and interoperability failures often generate systemic conversion losses between data restructuring and innovation outputs. The organisational mechanisms through which such losses are mitigated remain insufficiently understood. Drawing on boundary-spanning theory and extending it to the information systems context, this study conceptualises the Chief Information Officer (CIO) as an institutionalised information boundary-spanning capability grounded in information systems orchestration and governance. By identifying, integrating, and orchestrating data resources, information systems, and cross-boundary interfaces, CIO-enabled IS orchestration enhances the governability and reusability of digital resources, thereby improving digital innovation efficiency. Using firm-year panel data from Chinese A-share listed firms between 2014 and 2024 (approximately 50,000 observations), we treat a firm’s first CIO appointment as a staggered institutional event and employ propensity score matching combined with a staggered difference-in-differences approach for causal identification. The results show that the initial appointment of a CIO significantly increases digital innovation output. Mechanism analyses indicate that information systems orchestration capability serves as a key mediating channel. Further analyses reveal that open innovation ecosystem maturity and governance interface configurations shaped by public and firm-level digital governance significantly strengthen CIOs’ boundary-spanning effects. From an information management perspective, this study demonstrates that digital innovation returns in open innovation contexts are not technologically self-generating, but depend critically on the alignment between governance interfaces and information systems orchestration.
In autonomous driving, environmental perception increasingly demands such informatics-driven integration to transform raw sensory data into interpretable and actionable knowledge. However, monocular vision-based perception pipelines remain sensitive to environmental disturbances, including illumination variation, noise, and occlusion, which limits robustness and constrains system-level interpretability. This paper proposes an engineering informatics-oriented perception framework, denoted as MAD, that systematically couples visual perception, semantic modeling, and signal-level analysis within a unified information processing architecture. The framework is designed around three complementary layers of information abstraction. At the object level, dynamic vehicle perception is performed to establish stable target representations from monocular imagery. At the semantic level, scene understanding is formalized through semantic segmentation-driven spatial representations, enabling the construction of interpretable occupancy grids and vehicle cost maps that explicitly encode drivable regions and spatial constraints. At the signal level, multi-Q-factor Gabor wavelet analysis is applied to tri-axial vibration signals to enhance time-frequency resolution and provide auxiliary perceptual cues that improve robustness under adverse sensing conditions. In addition, projection-based spatial reasoning strategies are incorporated to ensure consistent information alignment across heterogeneous representations. Experimental evaluation on public driving datasets demonstrates that the proposed framework achieves an F1-score of 87.6% and an intersection-over-union value of 72.4%, outperforming representative baseline methods while maintaining computational efficiency suitable for embedded deployment. The results indicate that MAD exemplifies how engineering informatics principles can bridge visual semantics and temporal signal analysis, delivering an interpretable, scalable, and system-oriented perception solution for autonomous driving applications.
Passenger trust is key to acceptance and reliance in autonomous eVTOL aircraft, especially during low-altitude risk scenarios such as obstacle avoidance. Grounded in situation awareness theory and human-automation trust research, this study examines passenger trust formation, focusing on explanatory information and its spatial layout. A passenger-centered framework is proposed in which trust is mediated by visual attention, cognitive workload, perceived usability, and situation awareness. A simulated experiment compared three interface configurations: full head-down display, full head-up display, and split-screen. Subjective measures and eye-tracking data capture perceptual and cognitive processes. Results indicate that full head-up display enhances situation awareness, usability, and trust by supporting efficient attention allocation and cognitive coherence. Split-screen layouts increase integration demands and limit trust development. Findings clarify passenger trust mechanisms and highlight the importance of integrated, low-friction interface design under dynamic low-altitude risk conditions.
User-generated content has emerged as a valuable source for identifying product and service needs. To address the challenges of accurately extracting user intent from social media data and mitigating the misinterpretation of intentions by large models, this paper proposes DM-CAM. This product design method integrates diffusion models with class activation mapping. A dataset of 24,003 user posts is collected from the Weibo platform, and user demand keywords are extracted through community detection using the Leiden algorithm. These keywords are transformed into semantic input vectors for the generation process. The image generation is guided by a SqueezeNet-based pre-trained CNN, with CAM employed to visualise salient regions and refine design outcomes. To evaluate the effectiveness and adaptability of the proposed method, a series of comparative and ablation experiments are conducted across typical product categories, including conceptual robots, vacuum cleaners, and cars. Results from standard image quality metrics indicate that DM-CAM achieves a prediction confidence of 87% during generation, outperforming LIME, Grad-CAM, and OS by 11%, 22%, and 26%, respectively. The proposed method demonstrates notable improvements in visual clarity, detail enhancement, and design naturalness, offering a promising solution for explainable, data-driven design workflows.
Wilderness search and rescue (WiSAR) often involves asymmetric information accessibility and role responsibilities that shape decision-making pathways. This study investigates how information-sharing strategies shape human-drone collaboration under asymmetric information in WiSAR. A controlled indoor laboratory simulation with 26 dyads compared no sharing, full sharing, and adaptive sharing. Compared with no sharing, both full and adaptive sharing improved subjective experience, whereas no consistent reductions in task completion time were observed. Differences between full and adaptive sharing were generally nonsignificant across most subjective measures. Behavioral, subjective, and eye-tracking data suggest that both sharing strategies support coordination under asymmetry by facilitating information alignment and reducing uncertainty. Adaptive sharing introduced more structured interaction patterns but did not consistently outperform full sharing. These findings suggest that information sharing primarily mitigates asymmetric constraints rather than demonstrating the superiority of adaptive mechanisms, highlighting the need to balance informativeness and coordination support in human-AI collaboration design.
During the lunar base construction phase, manned lunar rovers are expected to be used frequently for long-distance surface travel and multi-site task coordination. In this context, rover drivers must continuously perform navigation and obstacle avoidance in complex and high-risk lunar environments, resulting in high cognitive demands. Vehicle-mounted augmented reality head-up displays (AR-HUDs) offer opportunities to integrate navigation and environmental information directly into the driver’s field of view; however, methodological guidance on how to systematically organize AR-HUD interface information around task constraints remains limited. To address this gap, this study applies Work Domain Analysis (WDA), a core component of Cognitive Work Analysis (CWA), to model the goals, functional constraints, and physical environment of manned lunar rover navigation and obstacle avoidance tasks. Based on the WDA results, two alternative guidance-line presentation schemes were developed and compared in a controlled navigation experiment, focusing on drivers’ subjective workload and information comprehension. The results indicate differences between the two designs in subjective workload and information understanding, while participants generally reported positive evaluations of the WDA-derived interfaces. Overall, the findings suggest that WDA can serve as a structured analytical and design support framework for organizing AR-HUD interface information in complex driving contexts. Although this study does not aim to demonstrate direct performance improvements, it provides methodological insights for task-constraint-based interface design in manned lunar rovers and other high-risk scenarios.
Although prior research has explored public trust in fully autonomous vehicles (FAVs), the pedestrian perspective remains underexplored. Unlike drivers and passengers, pedestrians do not directly use FAVs, so the mechanisms underlying their trust may differ. Grounded in the Stimulus-Organism-Response and Risk Perception Attitude frameworks, this study examines how FAV design characteristics (anthropomorphism, transparency, and interactivity) influence pedestrian trust through perceived risk and self-efficacy. Based on 168 valid samples, the results show that all three characteristics were significantly associated with pedestrian trust, with interactivity showing the strongest association ((3 = 0.230), followed by anthropomorphism ((3 = 0.193) and transparency ((3 = 0.171). Perceived risk was negatively associated with trust ((3 =-0.154), whereas self-efficacy was positively associated with trust ((3 = 0.278). Mediation analysis shows that self-efficacy partially mediates the relationships of anthropomorphism ((3 = 0.065) and transparency ((3 = 0.098) with trust. FsQCA identifies three configurations associated with high pedestrian trust. Under low perceived risk, S1 (anthropomorphism, transparency, and self-efficacy) and S2 (anthropomorphism, interactivity, and self-efficacy) show the highest explanatory power (raw coverage = 0.667 for both; consistency = 0.975 and 0.967), suggesting that transparency and interactivity can substitute for each other when other conditions are present. When perceived risk is present, S3 (interactivity, transparency, and self-efficacy) remains associated with high trust (raw coverage = 0.291; consistency = 0.917), although its applicability is narrower. These findings suggest that pedestrian trust in FAVs is related to the joint configuration of design characteristics and psychological conditions across contexts.
To address the limitations of conventional target tracking methods-such as reliance on external devices, poor robustness, and low accuracy in motion intent detection under complex scenarios-this study proposes a novel color marker-based target tracking and motion intent detection method (CMTT). Specifically, the target object is first annotated using color markers. Then, color segmentation and occlusion modeling techniques are applied for image preprocessing, enabling preliminary localization of the target region. Building on this, the method integrates GoogLeNet with gradient-weighted class activation mapping (Grad-CAM) for fine-grained binary classification within the identified region, enhancing the detection of key target areas. A Kalman filter is subsequently employed to perform dynamic system state estimation, achieving real-time tracking performance. To further improve robustness under varying lighting conditions, an image reconstruction strategy based on average grayscale values is introduced, effectively mitigating the impact of illumination interference on tracking accuracy. Experimental results demonstrate that the proposed CMTT method achieves substantial performance gains across several benchmark datasets. On the visual object tracking (VOT) dataset, compared to the baseline model, adaptive target classification and interaction strategy (ATCAIS), CMTT achieves a 34.3 % increase in expected average overlap (EAO), an 8.0 % improvement in accuracy (ACC), and a 23.8 % enhancement in Robustness (ROB). On the depth-based tracking benchmark (DepthTrack) dataset, CMTT achieves 31.9 %, 38.7 %, and 25.6 % in EAO, ACC, and ROB, respectively, highlighting its superior stability and adaptability in complex environments. An empirical study with 15 participants confirmed the consistency between the proposed method and data from inertial measurement units (IMU).
With the widespread application of mixed reality (MR) technology, systems now demand higher standards for interaction response speed and target detection accuracy. To address the issues of response latency and insufficient gesture recognition accuracy in existing MR controllers, this paper proposes an interaction response enhancement method that integrates target detection with explainable artificial intelligence-TDXAI. This method introduces the InceptionConv module into the lightweight YOLOv5s model to enhance feature extraction capabilities and detection accuracy. Additionally, it combines Grad-CAM technology to enhance the spatial perception and explainability of the human-machine interface, thereby improving users' understanding and trust in system behavior. In experiments using automotive parts assembly as the application scenario, TDXAI achieved an accuracy rate of 98.33 % and a recall rate of 98.38 % on a self-built three-class dataset. Through five sets of comparative experiments (including whether to use object detection, attention heat map, and the integration of both), TDXAI improved task completion accuracy to 92 % and reduced the average response time to 7 s, significantly enhancing user satisfaction and trust. The experimental results demonstrate that the proposed method significantly improves the interaction performance of MR systems and possesses good practicality and potential for widespread application.
In-vehicle agents (IVAs) have been applied to intelligent vehicles to enhance driver experience in human-vehicle interaction (HVI). Anthropomorphic virtual agents are currently the mainstream in vehicles. However, it remains unclear how the anthropomorphic appearances of virtual agents affect the driver experiences in the HVI context. This research divides the anthropomorphism of virtual IVA appearances into two dimensions: head-to-body ratio and morphological completeness. We investigate their impacts on the driver experiences through two online surveys in a Chinese driver sample (n = 257) regarding Pleasure, Fear, Trust, Comprehensibility, and Acceptance. Results indicate that virtual IVA appearances with a medium head-to-body ratio and high morphological completeness provided better experiences in the five dimensions. Based on these findings, a driver visual preference map and four design suggestions are proposed for virtual in-vehicle agent anthropomorphic design. This work enhances understanding of the driver experiences caused by anthropomorphic appearance and supports future anthropomorphic design.
A complexity evaluation method based on interpretable deep learning models is proposed to assess the complexity of Human-Machine Interfaces (HMI) in automotive intelligent cockpits. This method aims to optimize the HMI design in intelligent cockpits, enhancing driver attention while allowing for evaluation during low-fidelity design stages or in no-code development environments, thus reducing design time and personnel costs. The evaluation process includes both static HMI image analysis and post-software development assessments. It involves forward model inference, selecting relevant layers, calculating gradients, and generating heatmaps to provide a comprehensive evaluation and visual feedback of the interface. Examining the heatmaps, we analyze whether the model focuses on the correct regions, detecting biases or misunderstandings. This helps understand model behavior, improve model performance, and enhance interpretability. Results show that the generated heatmaps effectively predict webpage complexity and are closely related to user-perceived complexity. Ultimately, comprehensive evaluations using the NASA-TLX scale validate the method's effectiveness in reducing driver cognitive load and improving driving safety, significantly enhancing user experience and aiding developers in optimizing HMI interfaces.
Large Language Models (LLMs) have demonstrated remarkable generative capabilities across a wide range of natural language processing tasks. However, the frequent occurrence of hallucinations-outputs that appear plausible but are factually incorrect or logically inconsistent-poses a significant challenge to the reliability and practical utility of these models. This paper proposes a novel Emotion-Augmented Inference (EAI) method based on the Wheel of Emotions, aiming to mitigate hallucinations in multimodal generation tasks involving LLMs. EAI integrates two core mechanisms: visual-contrastive decoding and affective textual symbolization, which jointly enable the perception, regulation, and reconstruction of emotional signals during generation. These mechanisms enhance emotional coherence and semantic reliability in the model's outputs. Experimental results on two multimodal datasets, MSCOCO and GQA, show that EAI achieves improvements of 4%-8 % over baseline models in terms of key metrics such as accuracy, precision, recall, and F1-score. Additionally, under three emotional contexts-neutral (S1), positive (S2), and negative (S3)-EAI demonstrates particularly strong performance in hallucination suppression. In the S3 condition, accuracy improves by 5.48% and 2.23% compared to S1 and S2, respectively. These findings also indicate that EAI enhances the ability to manage emotion and maintain textual coherence. In summary, EAI not only stabilizes hallucination suppression in multimodal generation but also provides a new perspective for interpreting the emotional states embedded in LLM outputs. The proposed method offers a promising direction for building more trustworthy, controllable, and human-centered AI systems.