Rapid advancement of artificial intelligence and immersive technologies is revolutionizing various sectors, notably product design. However, the traditional personalized design process, which depends on predefined elements with limited user input, often results in products that do not fully align with individual preferences and lack substantial user engagement. To fill this gap, the emergence of artificial intelligence (AI)-generated content (AIGC) presents a significant opportunity for mass personalization through natural language interactions. Inspired by this paradigm, this article proposes an AIGC-enabled personalized product design approach, which integrates a configuration retrieval model with a fine-tuned text-to-3D generative model (TAPS3D model), enabling users to create personalized products within an immersive environment. While the current system requires approximately two minutes for 3D shape generation, this level of responsiveness is considered suitable for concept exploration in early-stage design workflows, where rapid iteration is prioritized over instantaneous feedback. Furthermore, a case study is conducted focusing on the design of personalized steering wheels to demonstrate the feasibility of this methodology. Furthermore, the effectiveness of the proposed approach in improving user experience is evaluated using a comparative experiment with the traditional configuration system. The findings indicate that our proposed AIGC-enabled personalized design system effectively enhances personalization, facilitates user engagement, improves the interaction experience, and increases user satisfaction.
While mass personalization manufacturing paradigm increasingly requires robots to handle complex and variable tasks, traditional robot-centric programming methods remain constrained by their expert-dependent nature and lack of adaptability. To address these limitations, this research proposes a scene-centric robot programming approach using MR-assisted interactive 3D segmentation, where operators naturally manipulate the digital twin (DT) of real-world objects to control the robot, rather than considering cumbersome end-effector programming. This framework combines Segment Anything Model (SAM) and 3D Gaussian Splatting (3DGS) for cost-effective, zero-shot, and flexible scene reconstruction and segmentation. Scale consistency and multi-coordinate calibration ensure seamless MR-driven interaction and robot execution. Finally, experimental results verify improved segmentation accuracy and computational efficiency, particularly in cluttered industrial environments, while case studies validate the method's feasibility for real-world implementation. This research illustrates a promising human-robot collaborative manufacturing paradigm where virtual scene editing directly informs robot actions, demonstrating a novel MR-assisted interaction method beyond low-level robot movement control.
Human-robot collaboration enhances efficiency by enabling robots to work alongside human operators in shared tasks. Accurately understanding human intentions is critical for achieving a high level of collaboration. Existing methods heavily rely on case-specific data and face challenges with new tasks and unseen categories, while often limited data is available under real-world conditions. To bolster the proactive cognitive abilities of collaborative robots, this work introduces a Visual-Language-Temporal approach, conceptualizing intent recognition as a multimodal learning problem with HRC-oriented prompts. A large model with prior knowledge is fine-tuned to acquire industrial domain expertise, then enables efficient rapid transfer through few-shot learning in data-scarce scenarios. Comparisons with state-of-the-art methods across various datasets demonstrate the proposed approach achieves new benchmarks. Ablation studies confirm the efficacy of the multimodal framework, and few-shot experiments further underscore meta-perceptual potential. This work addresses the challenges of perceptual data and training costs, building a human-robot bridge (H2R Bridge) for semantic communication, and is expected to facilitate proactive HRC and further integration of large models in industrial applications.
Recent rapid developments of dexterous robotic hands have greatly enhanced the manipulative capabilities of robots, enabling them to perform industrial tasks in human-like dexterity. These advancements not only enhance operational efficiency but also liberate human operators from monotonous tasks, allowing them to focus on creative and intellectually demanding. Despite the considerable attention robotic hands have garnered, existing reviews tend to focus on isolated topics, failing to provide a comprehensive perspective of the manufacturing sector. To empower robotic hands in human-centric smart manufacturing, this paper explores the latest research on holistic perception and dexterous skill learning of robotic hands. Specifically, the perceptual challenges in dexterous manipulation concerning different entities are investigated, including human hand perception, object inside-hand and outside-hand perception based on vision or tactility, and hand-object interactions, which help robots accurately understand environmental information. Furthermore, learning-based control methods are discussed, enhancing the execution capabilities of robotic hands through learning from scratch and learning from human demonstrations. Lastly, this paper identifies current challenges and offers several promising directions for future developments.
Robot learning has attracted an ever-increasing attention by automating complex tasks, reducing errors, and increasing production speed and flexibility, which leads to significant advancements in manufacturing intelligence. However, its low training efficiency, limited real-time feedback, and challenges in adapting to untrained scenarios hinder its applications in smart manufacturing. Introducing a human role in the training loop, a practice known as human-in-the-loop (HITL) robot learning, can improve the performance of robots by leveraging human prior knowledge. Nonetheless, the exploration of HITL robot learning within the context of human-centric smart manufacturing remains in its infancy. This study provides a holistic literature review for understanding HITL robot learning within an industrial context from a human-centric perspective. A united structure is presented to encompass different aspects of human intelligence in HITL robot learning, highlighting perception, cognition, behavior, and notably, empathy. Then, the typical applications in manufacturing scenarios are analyzed to expand the research landscape for smart manufacturing. Finally, it introduces the empirical challenges and future directions for HITL robot learning in the next industrial revolution era.
In Industry 5.0, where human ingenuity is combined with cutting-edge technologies such as artificial intelligence (AI) and robotics to revolutionize manufacturing with a focus on sustainability and human well-being, Digital Twins (DT) have become essential to real-time optimization. However, the complexity of managing DT for large-scale systems poses challenges in terms of data transmission, analytics, and advanced applications, which can be potentially addressed by Large Language Model (LLM). This research firstly performs a literature review to study the roles and functions of LLM in DT in the context of Industry 5.0. Subsequently, we propose a framework named Interactive-DT for LLM-DT integration that reveals the technical pathway for how LLM can be effectively integrated and function within DT environments. Within this framework, the roles and functionalities of LLM at the edge layer, DT layer, and service layer are elaborated upon. Finally, the identified research gaps and prospects for the integration of LLM and DT are outlined and discussed. The research outcomes of this paper highlight the potential of LLM to augment DT capabilities through improved construction and operation, enhanced cloud-edge collaboration, and sophisticated data analytics, ultimately promoting industrial practices that are both efficient and aligned with human-centric and sustainability principles in Industry 5.0.
Semantic expertise remains a reliable foundation for industrial decision-making, while Large Language Models (LLMs) can augment the often limited empirical knowledge by generating domain-specific insights, though the quality of this generative knowledge is uncertain. Integrating LLMs with the collective wisdom of multiple stakeholders could enhance the quality and scale of knowledge, yet this integration might inadvertently raise privacy concerns for stakeholders. In response to this challenge, Federated Learning (FL) is harnessed to improve the knowledge base quality by cryptically leveraging other stakeholders' knowledge, where knowledge base is represented in Knowledge Graph (KG) form. Initially, a multi-field hyperbolic (MFH) graph embedding method vectorizes entities, furnishing mathematical representations in lieu of solely semantic meanings. The FL framework subsequently encrypted identifies and fuses common entities, whereby the updated entities' embedding can refine other private entities' embedding locally, thus enhancing the overall KG quality. Finally, the KG complement method refines and clarifies triplets to improve the overall quality of the KG. An experiment assesses the proposed approach across different industrial KGs, confirming its effectiveness as a viable solution for collaborative KG creation, all while maintaining data security.
In smart manufacturing, autonomous mobile robots play an indispensable role in conducting inspection and material handling operations, yet they face significant limitations regarding adaptability and resilience within unstructured environments. Vision and language navigation (VLN), a human-guided navigation paradigm, emerges as a compelling solution to these challenges. Nevertheless, VLN’s practical implementation is constrained by limited task generalization capabilities, inadequate response to diverse linguistic commands, and insufficient consideration of sensor-induced noise in environmental perception. This research addresses these limitations by introducing an innovative vision-language model (VLM)-based human-guided mobile robot navigation approach in an unstructured environment for human-centric smart manufacturing (HSM). This approach encompasses robust Three-dimensional (3D) scene reconstruction through advanced point cloud techniques, zero-shot semantic segmentation via a VLM, and natural language processing through a large language model (LLM) to interpret instructions and generate control code for navigation. The system’s efficacy is validated through extensive experiments in an unstructured manufacturing setup.
human-robot collaboration (HRC) is set to transform the manufacturing paradigm by leveraging the strengths of human flexibility and robot precision. The recent breakthrough of Large Language Models (LLMs) and Vision-Language Models (VLMs) has motivated the preliminary explorations and adoptions of these models in the smart manufacturing field. However, despite the considerable amount of effort, existing research mainly focused on individual components without a comprehensive perspective to address the full potential of VLMs, especially for HRC in smart manufacturing scenarios. To fill the gap, this work offers a systematic review of the latest advancements and applications of VLMs in HRC for smart manufacturing, which covers the fundamental architectures and pretraining methodologies of LLMs and VLMs, their applications in robotic task planning, navigation, and manipulation, and role in enhancing human–robot skill transfer through multimodal data integration. Lastly, the paper discusses current limitations and future research directions in VLM-based HRC, highlighting the trend in fully realizing the potential of these technologies for smart manufacturing.
Human-centric assembly is emerging as a promising paradigm for achieving mass personalization in the context of Industry 5.0, as it fully capitalizes on the advantages of human flexibility with robot assistance. However, in small-batch and highly customized assembly tasks, frequently changes in production procedures pose significant cognition challenges. To address this, leveraging computer vision technology to enhance human cognition becomes a feasible solution. Therefore, this review aims to explore the cognitive characteristics of human beings and classify existing computer vision technologies in a manner that discusses the future development of cognition-augmented human-centric assembly. The concept of cognition-augmented assembly is first proposed based on the brain's functional structure - the frontal, parietal, temporal, and occipital lobes. Corresponding to these brain regions, cognitive issues in spatiality, memory, knowledge, and decision-making are summarized. Recent studies conducted between 2014 and 2023 on visual computation of assembly are categorized into four groups: position registration, multi-layer recognition, contextual perception, and mixed-reality fusion, all aimed at addressing these cognitive challenges. The applications and limitations of current computer vision technology are discussed. Furthermore, considering the rapidly evolving technologies such as the metaverse, cloud services, large language models, and brain-computer interfaces, future trends on computer vision are prospected to augment human cognition corresponding to the cognitive issues.
Human–robot collaboration (HRC) has been recognized as a potent pathway towards mass personalization in the manufacturing sector, by leveraging the synergy of human creativity and robotic precision. Previous approaches rely heavily on visual perception to autonomously comprehend the HRC environment. However, the inherent ambiguity in human–robot communication cannot be consistently neutralized by relying solely on visual cues. With the recently soaring popularity of large language models (LLMs), the consideration of language data as a complementary information source has increasingly drawn research attention, while the application of such large models, particularly within the context of HRC in manufacturing, remains largely under-explored. In response to this gap, a vision-language reasoning approach is proposed to mitigate the communication ambiguity prevalent in human–robot collaborative manufacturing scenarios. A referred object retrieval model is first designed to alleviate the object–reference ambiguity in the human language command. This model is then seamlessly integrated into an LLM-based robotic action planner to achieve an improved HRC performance. The effectiveness of the proposed approach is demonstrated empirically through a series of experiments conducted on the object retrieval model and its application in a human–robot collaborative assembly case.
Human-Robot Collaboration (HRC) has emerged as a pivot in contemporary human-centric smart manufacturing scenarios. However, the fulfilment of HRC tasks in unstructured scenes brings many challenges to be overcome. In this work, mixed reality head-mounted display is modelled as an effective data collection, communication, and state representation interface/tool for HRC task settings. By integrating vision-language cues with large language model, a vision-language-guided HRC task planning approach is firstly proposed. Then, a deep reinforcement learning-enabled mobile manipulator motion control policy is generated to fulfil HRC task primitives. Its feasibility is demonstrated in several HRC unstructured manufacturing tasks with comparative results.
Human-robot collaborative disassembly (HRCD) has gained much interest in the disassembly tasks of end-of-life products, integrating both robot’s high efficiency in repetitive works and human’s flexibility with higher cognition. Explicit human-object perceptions are significant but remain little reported in the literature for adaptive robot decision-makings, especially in the close proximity co-work with partial occlusions. Aiming to bridge this gap, this study proposes a vision-based 3D dense hand-object pose estimation approach for HRCD. First, a mask-guided attentive module is proposed to better attend to hand and object areas, respectively. Meanwhile, explicit consideration of the occluded area in the input image is introduced to mitigate the performance degradation caused by visual occlusion, which is inevitable during HRCD hand-object interactions. In addition, a 3D hand-object pose dataset is collected for a lithium-ion battery disassembly scenario in the lab environment with comparative experiments carried out, to demonstrate the effectiveness of the proposed method. Note to Practitioners —This work aims to overcome the challenge of joint hand-object pose estimation in a human-robot collaborative disassembly scenario, of which can also be applied to many other close-range human-robot/machine collaboration cases with practical values. The ability to accurately perceive the pose of the human hand and workpiece under partial occlusion is crucial for the collaborative robot to successfully carry out co-manipulation with human operators. This paper proposes an approach that can jointly estimate the 3D pose of the hand and object in an integrated model. An explicit prediction of the occlusion area is then introduced as a regularization term during model training. This can make the model more robust to partial occlusion between the hand and object. The comparative experiments suggest that the proposed approach outperforms many existing hand-object estimation ones. Nevertheless, the dependency on manually labeled training data can limit its application. In the future, we will consider semi-supervised or unsupervised training to address this issue and achieve faster adaptation to different industrial scenarios.
In contemporary digital landscape, the demand for personalized design is significantly influenced by rising consumer expectations. Enabling active user engagement and catering to individual preferences has become essential for creating personalized experiences, improve customer satisfaction and foster brand loyalty. These factors are fundamental to the success of products and services. However, traditional personalized design methodologies often lack active user active participation and are typically implemented through configurations predefined by experts. In this study, we introduce DesignGemini, a highly personalized design approach that integrates generative models with human digital twin. This methodology facilitates the generation of personalized designs through natural language interactions with generative models, while simultaneously enabling online preference recognition and ergonomic analysis via the establishment of human digital twins. Additionally, we present a case study on the personalized design of vehicle seats to demonstrate the feasibility of the proposed approach. This case study effectively showcases product generation utilizing TAPS3D, a text-to-3D generative model, along with rapid ergonomic analysis employing a vision-based human digital twin modeling technique.
The evolution of smart vehicle cockpits is transitioning from serving as mere driving tools to becoming intimate partners that significantly enhance user experiences through advanced technologies. This research addresses the growing demand for personalized design in smart vehicle cockpits by proposing a framework, CockpitGemini. This framework integrates generative model-based multi-agent systems and human digital twins, enabling tailored designs and services based on user preferences and real-time status. The capabilities of the proposed framework are illustrated through four dimensions: personalized product design, personalized interactive interface design, user state monitoring and personalized regulation, and personalized driving strategy recommendations. A case study on the design of personalized vehicle seats demonstrates the feasibility and usability of the CockpitGemini framework, highlighting its potential to enhance user satisfaction in smart vehicle cockpits.
Industry 5.0 prioritizes Human-centric Smart Manufacturing (HSM), aiming to enhance human operators’ well-being and needs. This necessitates collaborative robots with advanced natural interaction capabilities and improved perception, cognition, and action intelligence. The Large Language Model (LLM) exhibits strong reasoning abilities and generalization capabilities, which can significantly advance the development of HSM once integrated into human–robot interaction and collaboration. Accordingly, this paper explores part of LLM’s ability in the context of smart manufacturing, focusing on addressing interruptions in the manufacturing process caused by repetitive tool fetching. To alleviate this issue, a vision and language cobot navigation approach is innovatively adopted in the manufacturing environment that could be further used to assist operators in retrieving tools. Specifically, a real Human–Robot Collaboration (HRC) manufacturing scene is first reconstructed and annotated using Three-Dimensional (3D) point cloud techniques. Then the LLM is utilized for the Automated Guided Vehicle (AGV) to comprehend natural language commands and generate Python code to triger navigation actions. Finally, the Pathfinder algorithm is applied for corresponding path planning. The framework is implemented in the Artificial Intelligence (AI) Habitat simulator, and the case studies demonstrate that the AGV can accurately comprehend complex language instructions, empowering human operators to complete manufacturing tasks efficiently.
Segmentation of intracranial aneurysm(IA)from computed tomography angiography(CTA)images is of sig-nificant importance for quantitative assessment of IA and further surgical treatment.Manual segmentation of IA is a la-bor-intensive,time-consuming job and suffers from inter-and intra-observer variabilities.Training deep neural networks usually requires a large amount of labeled data,while annotating data is very time-consuming for the IA segmentation task.This paper presents a novel weight-perceptual self-ensembling model for semi-supervised IA segmentation,which em-ploys unlabeled data by encouraging the predictions of given perturbed input samples to be consistent.Considering that the quality of consistency targets is not comparable to each other,we introduce a novel sample weight perception module to quantify the quality of different consistency targets.Our proposed module can be used to evaluate the contributions of unlabeled samples during training to force the network to focus on those well-predicted samples.We have conducted both horizontal and vertical comparisons on the clinical intracranial aneurysm CTA image dataset.Experimental results show that our proposed method can improve at least 3%Dice coefficient over the fully-supervised baseline,and at least 1.7%over other state-of-the-art semi-supervised methods.
Recognizing sitting posture is significant to prevent the development of work-related musculoskeletal disorders for office workers. Multimodal data, i.e., infrared map and pressure map, have been leveraged to achieve accurate recognition while preserving privacy and being unobtrusive for daily use. Existing studies in sitting posture recognition utilize handcrafted features with machine learning models for multimodal data fusion, which significantly relies on domain knowledge. Therefore, a deep learning model is proposed to fuse the multimodal data and recognize the sitting posture. This model contains modality-specific backbones, a cross-modal self-attention module, and multi-task learning-based classification. Experiments are conducted to verify the effectiveness of the proposed model using 20 participants’ data, achieving a 93.08% F1-score. The high-performance result indicates that the proposed model is promising for sitting posture-related applications.
Human-Robot Collaboration (HRC) has played a pivotal role in today's human-centric smart manufacturing scenarios. Nevertheless, limited concerns have been given to HRC uncertainties. By integrating both human and artificial intelligence, this paper proposes a Collaborative Intelligence (CI)-based approach for handling three major types of HRC uncertainties (i.e., human, robot and task uncertainties). A fine-grained human digital twin modelling method is introduced to address human uncertainties with better robotic assistance. Meanwhile, a learning from demonstration approach is offered to handle robotic task uncertainties with human intelligence. Lastly, the feasibility of the proposed CI has been demonstrated in an illustrative HRC assembly task.
Human-robot collaboration (HRC) has been identified as a highly promising paradigm for human-centric smart manufacturing in the context of Industry 5.0. In order to enhance both human well-being and robotic flexibility within HRC, numerous research efforts have been dedicated to the exploration of human body perception, but many of these studies have focused only on specific facets of human recognition, lacking a holistic perspective of the human operator. A novel approach to addressing this challenge is the construction of a human digital twin (HDT), which serves as a centralized digital representation of various human data for seamless integration into the cyber-physical production system. By leveraging HDT, performance and efficiency optimization can be further achieved in an HRC system. However, the implementation of visual perception-based HDT remains underreported, particularly within the HRC realm. To this end, this study proposes an exemplary vision-based HDT model for highly dynamic HRC applications. The model mainly consists of a convolutional neural network that can simultaneously model the hierarchical human status including 3D human posture, action intention, and ergonomic risk. Then, on the basis of the constructed HDT, a robotic motion planning strategy is further introduced with the aim of adaptively optimizing the robotic motion trajectory. Further experiments and case studies are conducted in an HRC scenario to demonstrate the effectiveness of our approach.