At present, the dissemination of Tai Chi is predominantly facilitated through offline instructional methods combined with video practice, lacking effective feedback on movement poses and exhibiting relatively low efficiency. This article proposes a novel algorithm named TC-YOLO for pose estimation based on YOLOv8. TC-YOLO enhances the efficiency through utilizing real-time detection of key points in Tai Chi practitioners' movements to provide more accurate and intuitive feedback for instructional evaluation and pose correction. Focusing on the demonstration of the Essential Eighteen Movements of Chen-style Tai Chi as the research subject, Tai Chi movement dataset comprising 3,688 images is constructed. To enhance model efficiency, the backbone network is reparameterized through the Reparametrized C2f (RC2f) module, which optimizes feature extraction process by allowing more effective information flow and reducing computational complexity. Furthermore, a simplified neck network structure is designed to provide effective information transmission and multiscale feature fusion, thereby enhancing detection accuracy. Experimental results show that the TC-YOLO algorithm achieves a mAP of 97.2% and a recognition speed of 146.1 FPS with lower parameters and computational cost on the self-built dataset, which is better than YOLO-Pose and other models. TC-YOLO could make an instructive contribution to the field of sports science through the analysis of Tai Chi movement poses, promoting the inheritance and global dissemination of Tai Chi.
The distilled liquor brewing technique is one of China’s national-level intangible cultural heritages, and it has a rich history. In order to enhance public awareness and understanding of this technique, previous researchers have utilized digital technologies such as virtual reality to create virtual museums for disseminating and showcasing its cultural significance. However, existing virtual museums often employ passive roaming displays, lacking interactivity, immersion, and engagement, which makes it challenging to attract the younger generation. This paper introduces a multimodal interactive virtual liquor culture museum system. By incorporating sensory feedback such as visual, auditory, tactile, and temperature cues, participants can learn about and experience the art of distilled liquor brewing. Within the museum, three virtual avatars (visitor, guide, poet) are implemented and powered by advanced language models, enabling continuous dialogue and interactive gameplay with participants and enhancing immersion and interest. The study, validated through evaluations involving 30 participants, demonstrates the usability and effectiveness of the System.
This article explores whether there are significant differences in the creativity level of future teachers in teaching design. In virtual reality (VR) and mixed reality (MR) experimental teaching environments, the K-DOCS creative power scale, the flow state scale, and the expert group assessment scale were used to investigate the mediation effect of flow, attention, and meditation on the creativity level. The results indicate that in the MR experimental teaching environment, the three factors have a moderate mediating effect on the creativity level of future teachers (n = 65), and meditation has a certain regulating effect on heart flow (0.175*); In the VR experimental teaching environment (n = 67), only the factor of heart flow has a moderate promotion effect on the creativity level of future teachers (0.822***). The study results indicate that the quality of future teachers' creative instructional design in the MR environment is higher than that in the VR environment. In addition more MR experimental teaching environments should be provided to cultivate the creativity level of future teachers in teaching design.
The objective of this paper is to investigate fine-grained 3D face reconstruction. Recently, many methods based on 3DMM and mixed multiple low-level losses for unsupervised or weak supervised learning have achieved some results. Based on this, we propose a more flexible framework that can reserve original structure and directly learn detail information in 3D space, which associates 3DMM parameterized model with an end-to-end self-supervised system. To encourage high-quality reconstruction, residual learning is introduced. Evaluations on popular benchmarks show that our approach can attain comparable state-of-the-art performance. In addition, our framework is applied to a reconstruction system for interaction.
In this paper, we study the task of referring semantic segmentation in a highly practical setting, in which labeled visual data with corresponding text descriptions are available in the source, but only unlabeled visual data (without text descriptions) are available in the target. It is a challenging task that has many difficulties: (1) how to obtain proper queries for the target domain; (2) how to adapt visual-text joint distribution shifts; (3) how to maintain the original segmentation performance. Thus, we propose a cycle-consistent vision-language matching network to narrow down the domain gap and ease adaptation difficulty. Our model has significant practical applications since they are capable generalising to new data sources without requiring corresponding text annotations. First, a pseudo-text selector is devised to handle the missing modality, through the pre-trained clip model to measure the gap between query features of the source and visual features of the target. Next, a cross-domain segmentation predictor is adopted, which prompts the joint representations to be domain invariant and minimize the discrepancy between two domains. Then, we present a cycle-consistent query matcher to learn discriminative features via reconstructing visual features from masks. Instead of doing the textual comparison, we match the visual features to the pseudo queries. Extensive experiments show the effectiveness of our method.
眼动人机交互利用眼动特点可以增强用户的沉浸感和提高舒适度,在虚拟现实(VR)系统中融入眼动交互技术对VR系统的普及起到至关重要的作用,已成为近年来的研究热点.对VR眼动交互技术的原理和类别进行阐述,分析了将VR系统与眼动交互技术结合的优势,归纳了目前市面上主流VR头显设备及典型的应用场景.在对有关VR眼动追踪相关实验分析的基础上,总结了VR眼动的研究热点问题,包括微型化设备、屈光度矫正、优质内容的匮乏、晕屏与眼球图像失真、定位精度、近眼显示系统,并针对相关的热点问题展望相应的解决方案.
Unlike the traditional environment, VR and MR environments provide novel, natural, lifelike 3D user interfaces, which can stimulate the imagination of users. Whether VR and MR environments can influence the creativity and its impact factors (flow, attention, and relaxation) of future teachers' scene expansion in instruction design are the focus of the study. Moreover, the differences between the impacts of VR and MR environments on creativity and its impact factors are also questions worth exploring. In this study, we developed VR and MR experiment environments as creativity support systems to stimulate future teachers' creativity before they carry out the instruction design of experimental courses. The results show that the creativity, flow, and attention of future teachers which use the VR and MR environments were higher than in the traditional environment. Moreover, the future teachers' creativity of scene expansion in MR environment was higher than VR environment, the future teachers' attention of VR environment was higher than MR environment, and the future teachers.' relaxation of VR environment was lower than the traditional environment. These findings provide a useful inspiration, i.e. too high concentration or too high relaxation is not conducive to the production of creativity. The MR creativity support system that provides moderate and balanced attention and relaxation can good stimulate the creativity of future teachers in the instruction design. We expect these findings to inspire the design of creativity support systems and foster future teachers' creativity level in instructional design.
Mixed reality (MR) is the merging of real and virtual worlds to produce new environments and visualizations, where virtual and physical objects can interact with each other in real time. The traditional fixed single camera desktop AR platform provides a new display and interaction mode for experimental teaching in middle schools. However, to observe the experimental phenomenon, we have to use a fixed-position screen with a fixed viewing angle which makes the experiment operation unnatural. The use of traditional Hamming code mark can not let people know what kind of AR object it is. In this paper, we propose a virtual experiment platform that can conduct various experiments in middle school. In this virtual platform, users observe the experimental phenomena with a pair of movable MR glasses. Meanwhile, the mark images required for interaction are customized and indicates what the AR object is. In this way, users can select the AR objects they need easily during the experiment instead of memorizing the AR objects corresponding to different marks in advance. However, the value and importance of different modules are different. What criteria do we use to judge the value of different modules? Kano's Theory provides the guidance of requirements classification attributes for the research of virtual experimental platform. Therefore, we evaluate five requirements classification attributes, including the function of camera capture, custom mark, visualization of magnetic line, voice interaction and multi-camera tracking. Nineteen pedagogical volunteers participated in our experiment. Experimental results prove that our platform stimulates the users’ interest in learning. Further, most participants believe that it would promote the development of experiment education effectively in middle schools.
Sign Language Production (SLP) aims to translate a spoken language description to its corresponding continuous sign language sequence. A prevailing solution for this problem is in a two-staged manner: it formulates SLP as two sub-tasks, i.e., Text to Gloss (T2G) translation and Gloss to Pose (G2P) animation, with gloss annotations as pivots. Although two-staged approaches achieve better performance than their direct translation counterparts, the requirement of gloss intermediaries causes a parallel data bottleneck. In this paper, to reduce reliance on gloss annotations in two-staged approaches, we propose DualSign, a semi-supervised two-staged SLP framework, which can effectively utilize partially gloss-annotated text-pose pairs and monolingual gloss data. The key component of DualSign is a novel Balanced Multi-Modal Multi-Task Dual Transformation (BM3T-DT) method, where two well-designed models, i.e., a Multi-Modal T2G model (MM-T2G) and a Multi-Task G2P model (MT-G2P), are jointly trained by leveraging their task duality and unlabeled data. After applying BM3T-DT, we derive the expected uni-modal T2G model from the well-trained MM-T2G with knowledge distillation. Considering that the MM-T2G may suffer from modality imbalance when decoding with multiple input modalities, we devise a cross-modal balancing loss, further boosting the system's overall performance. Extensive experiments conducted on the PHOENIX14T dataset show the effectiveness of our approach in the semi-supervised setting. By training with additionally collected unlabeled data, DualSign substantially improves previous state-of-the-art SLP methods.
We present a simulation microscope device called GiantScope, which combines virtual microscope, cloud computing, and embedded technologies. Users can complete most of the microscope‐based experiments in biology courses by operating our device, while learning the operating skills of microscopes at the same time. Our device supports most of the operation functions of optical microscopes, including quasifocus screw adjustment, slide movement recognition and so on, and also has auxiliary enhancement functions including manual measurement, annotation and so on. In addition, we have built a cloud‐based digital slide database, which enables users to select experimental observations through digital slides, including static cell specimens or dynamic cell activities. After user study, we found that using GiantScope for biological experiments has better learning efficiency and user experience than traditional microscopes.
Nowadays, researches on mixed reality (MR) have made a lot of exploration in the aspects of user experience, hardware devices, interaction technologies, application systems, etc. However, more research is still needed to explore how to improve the experience in the shared MR Environment and design appropriate collaborative interaction models. The traditional VR handle can realize tool simulation and 3D interaction, while smartphones can realize 2D interactions such as fast 2D gestures, text input and handwriting. Using the mobile phone as the handle of the MR headset can realize the complementary advantages of the two. In this study, we design a cross-device collaborative system sharing a hybrid reality environment. In the MR Environment, the smartphone will realize the functions of the controller for remote selection and manipulation of objects, and its advantages in 2D interaction can be well applied to some special tasks. We design the user interface with three levels of immersion: high, medium and low. High immersion: We provide a hybrid user interface combining a HoloLens2 and a smartphone. Medium immersion: We provide a 3D user interface represented by HoloLens2. Low immersion: We provide a multi-touch user interface represented by tablet. User studies show that the hybrid user interfaces can bring users satisfactory immersion and interactive experience, but it also needs to design accurate and efficient input methods according to the interactive tasks.
Surface flattening plays an important role in the whole process of garment design. We proposed a novel method by using three-dimensional triangle mesh flattening in this study. First, the three-dimensional triangle mesh is flattened to a two-dimensional plane to approximate the original surface. The initial flattening results are then used as preliminary guesses for subsequent optimizations. Considering that the deformation energy in the real woven fabric is related to tensile or shear deformation, a simplified fabric deformation model based on energy is proposed to update the energy distribution to determine the best two-dimensional pattern. An innovative unified axis system process is proposed to obtain the deformation energy, and energy relaxation in local flattening is proposed to release the distortion of flattening. Finally, the experimental results show that complex surfaces such as garments could achieve better flattening results. Compared with other energy-based methods in garment design, our proposed methods are more flexible and practical.
Augmented reality technology has been widely used in experimental education. Augmented reality provides virtual-real integration, real-time interaction and three-dimensional immersion, which provides a new development direction for simulating the teaching environment and promoting learning interaction. We have developed an MR experiment of diluting concentrated sulfuric acid, which helps students learn experimental manipulations and scientific concepts before conducting real chemical experiments. Virtual experiments can avoid the risks of real chemical experiments and reduce the waste of chemical materials. At the same time, we correctly render the occlusion relationship between the hand, the beaker and the virtual liquid, providing a realistic experimental effect. The tracking of multiple cameras allows students to interact more naturally. We compared the availability and students' attitudes of three virtual experiments about diluting concentrated sulfuric acid. The results show that the usability of the MR experimental system reaches the level of user satisfaction. The realistic visual effects and natural interaction method of the MR experiment have been recognized by the students.
Metaverse is the next generation gaming Internet, and virtual humans play an important role in Metaverse. The simultaneous representation of motions and emotions of virtual humans attracts more attention in academics and industry, which significantly improves user experience with the vivid continuous simulation of virtual humans. Different from existing work which only focuses on either the expression of facial expressions or body motions, this paper presents a novel and real-time virtual human prototyping system, which enables a simultaneous real-time expression of motions and emotions of virtual humans (short for SimuMan). SimuMan not only enables users to generate personalized virtual humans in the metaverse world, but also enables them to naturally and simultaneously present six facial expressions and ten limb motions, and continuously generate various facial expressions by setting parameters. We evaluate SimuMan objectively and subjectively to demonstrate its fidelity, naturalness, and real-time. The experimental results show that the SimuMan system is characterized by low latency, good interactivity, easy operation, good robustness, and wide application.
针对如何在中学物理实验中,在增强现实(AR)的虚拟环境中模拟出符合物理规律的磁感线、电场线等曲线的问题,研究了一套在三维空间中拟合出符合磁感线性质的磁感线算法,并将其应用到多模态自然交互的AR实验系统中.该算法使用四阶龙格库塔方法生成磁感线,并在必要时使用能量最小化的方法进行修正.该AR系统使用基于实物套件的增强现实并利用多相机协同AR三维注册来克服传统二维MARK跟踪失效问题.最终以中学教学中常见的电生磁实验为例,测试了该磁感线生成算法和实验系统.结果表明,该磁感线符合物理定律,并可以较好地服务于电磁相关的物理实验中,具有解决实验现象不明显,使学生理解更为直观透彻的实际意义.
Real chemical experiments may be dangerous or pollute the environment; meanwhile, the preparation of drugs and reagents is time-consuming. Due to the above-mentioned reasons, few experiments can be actually operated by students, which is not conducive to the chemistry learning and the phenomena principle understanding. Recently, due to the impact of Covid-19, many schools adopt online teaching, which is even more detrimental to students’ learning of chemistry. Fortunately, MR(mixed reality) technology provides us with the possibility of solving the safety issues and breaking the space-time constraints, while the theory of human needs (Maslow’s hierarchical needs) provides us with a way to design a comfortable and stimulant MR system with realistic visual presentation and interaction. The paper combines with the theory of human needs to propose a new needs model for virtual experiment. Based on this needs model, we design and develop a comprehensive MR system called MagicChem, which offers a robust 6-DoF interactive and illumination consistent experimental space with virtual-real occlusion, supporting realistic visual interaction, tangible interaction, gesture interaction with touching, voice interaction, temperature interaction, olfactory interaction and virtual human interaction. User study shows that MagicChem satisfies the needs model better than other MR experimental environments that partially meet the needs model. In addition, we explore the application of the needs model in VR environment.
技术的革新带动了多维空间的融合和多模态数据的生成,MR技术启动了课堂数学新一轮的交互式革命.人工智能的快速发展与教育应用,也带来了教学中的人机协同.在人机协同以及具身认知等相关理论的支撑下,结合中学数理实验教学案例,建构了基于MR实验的"多模态+人机协同"教学方式,它是一种融合视觉、听觉、触觉等多模态数据,实现师生与智能设备互联,围绕固定教学目标而协同共进的教学.这一教学方式不仅具有数据多模态、人机协同共进等特点,而且应用于MR实验教学中,相对于传统实验和VR实验教学而言,能够更好地激发并培养学生的创新性思维与动手能力.随着"AI+教与学"应用的不断发展,未来"多模态+人机协同"教学的价值与趋势主要在于:依托多情境自由转换,助力学习者创新能力的提升;融合"视-听-触"多感官通道,提升交互过程的具身认知;建立多模态数据采集、分析和反馈机制,推动反思性观察;建立多阶段数据建模跟踪,实现拓展性迁移等.
Due to the reasons such as the complicated preparation of reagents, danger, pollution, and lack of teacher resources, there are few chemical experiments that students can actually operate during the chemistry learning, which is not conducive to the chemistry learning and the phenomena principle understanding. Fortunately, MR technology provides us with the possibility of solving the safety issues and the space-time constraints, while the theory of human needs provides us with a way to think about designing a comfortable and stimulant MR system with realistic visual presentation and interaction. This study combines with the theory of human needs to propose a new needs model for virtual experiment. Based on this needs model, we design and develop a comprehensive MR system called MagicChem to verify the needs model. User study shows MagicChem that satisfies the needs model is better than the MR experimental environment that partially meet the needs model. In addition, we explore the application of the needs model in VR environment.
In terms of visual presentation, this study studies a set of multi-camera collaborative technology, which can conduct three-dimensional registration of MR without dead Angle. We also study an accurate and real-time virtual-real occlusion algorithm based on depth calculation. In terms of human-computer interaction, we support three kinds of multi-modal interaction: gesture interaction with touching, tangible interaction and voice interaction. Taking potassium permanganate oxygen production experiment as an example, a MR experimental prototype system was developed. After testing and feedback, the system was highly praised by the experimental subjects.
The modeling of hair is too difficult to simulate because of its number and shape, as well as the texture of the hair itself. The traditional way of constructing hair based on physics and geometry requires complex calculations and various parameters. In recent years, hair modeling methods based on single images, based on multiple images, and based on videos have begun to develop. The advantage is that modeling is fast. At present, the geometry of the hair is mainly represented by lines of three-dimensional points. In this paper, a three-dimensional multi-strip is used to represent the geometry of the hair. Through deep learning, the position and type of the hair in a single image are obtained, and the similar hairstyle model in the database is matched. The selected hair model and head model are connected and fixed by further fitting. Then we simulate dynamic hair by setting gravity, friction, collision detection, and other more. The model preserves the image appearance of the image as much as possible and can be used to simulate common hair geometry.