Visual-language foundation models (VLFMs) are rapidly emerging as key enablers of interaction in virtual and augmented reality (VR/AR). By integrating advanced visual perception with language-based reasoning, they overcome the limitations of rule-driven systems and enable more natural, adaptive, and context-aware engagement. This paper provides a systematic survey of the evolving roles of VLFMs in immersive environments. We introduce the Interaction-Role Taxonomy, a novel framework that categorizes VLFMs into three complementary functions: the Semantic Interpreter, which perceives and interprets environments; the Embodied Agent, which facilitates collaborative interaction; and the Generative Engine, which creates dynamic content. Building on this framework, we review representative applications, analyze implementation pipelines and evaluation protocols, and critically discuss technical and ethical challenges. Finally, we identify open research directions to enhance the capability, efficiency, and trustworthiness of VLFMs, paving the way for next-generation human–computer interaction in immersive VR/AR ecosystems.
Football coaching is undergoing a paradigm shift with the introduction of immersive technologies, making it critical to evaluate how virtual reality (VR) can reshape training and performance development. This survey examines the transformative potential of VR in football coaching, with a focus on its capacity to enhance technical skill development, tactical comprehension, and performance evaluation. We review foundational VR technologies—such as high-fidelity scene reconstruction, accurate motion capture, and multi-sensory feedback—that collectively enable immersive and photorealistic training environments. Beyond replicating conventional training tasks, VR facilitates personalized instruction, fosters collaborative tactical learning, and supports rehabilitation through data-driven feedback. At the same time, critical challenges remain, including realistic motion simulation, synchronous multi-user interaction, mitigation of cybersickness, and the effective transfer of acquired skills to real-world contexts. We also highlight emerging research directions, including AI-driven adaptive training, athlete-specific digital twin modeling, hybrid VR/AR integration, and scalable cloud-based platforms. By systematically synthesizing opportunities, limitations, and future trajectories, this survey provides a comprehensive reference for coaches, researchers, and developers aiming to advance football training and reimagine coaching methodologies.
Accurate prediction of canopy light interception is essential for advancing precision management in planted forests. However, conventional approaches that rely on manual pruning strategies are limited in scalability and fail to effectively capture the dynamic interactions between canopy structure and light distribution. We propose a novel Canopy Light Interception Prediction with Transformer-LSTM Network(CLIP-TLNet) that combines 3D spatial complexity quantification with temporal modeling to predict light distribution within individual tree canopies over time. Leveraging multi-sensor UAV-derived LiDAR point clouds and synchronized photometric measurements from triploid poplar plantations, we develop a regional canopy complexity algorithm based on multi-scale fractal dimension analysis, enabling precise quantification of structural heterogeneity. This complexity metric enhances model interpretability and improves learning performance by characterizing how canopy architecture influences light penetration. In parallel, the spatio-temporally decoupled Transformer-LSTM architecture effectively captures temporal trends while maintaining spatial awareness, yielding a Mean Absolute Percentage Error (MAPE) of 6.8% and outperforming the second-best CNN-LSTM model by a 33.6% reduction in Root Mean Square Error (RMSE).Ablation studies confirm that incorporating canopy complexity as a structural prior caused a 20.6% increase in MAPE upon its removal, underscoring its critical role in predictive performance. By enabling data-driven, complexity-aware pruning strategies and temporally optimized interventions, this framework bridges static structural assessment with dynamic environmental response modeling. It offers a powerful tool for precision silviculture and intelligent canopy management in plantation forestry.
With the advancement of 3D digital dentistry, accurate 3D tooth segmentation has become increasingly important in orthodontics and computer-aided diagnosis. However, existing supervised approaches heavily rely on exhaustive face-wise annotations and often exhibit limited generalization across complex clinical meshes. Although self-supervised learning offers a promising alternative to alleviate annotation costs, current paradigms remain challenged by sensitivity to data augmentations, suboptimal representation learning in pure masking schemes, and the complex structural characteristics of dental geometry. To address these limitations, we propose Dental-CMAE, a graph-enhanced hierarchical Contrastive masked AutoEncoder framework tailored for 3D tooth segmentation. The framework incorporates a dual-branch masking strategy that leverages graph-based structural priors to generate distinct corrupted views while preserving intrinsic mesh topology, thereby facilitating robust reconstruction. This is integrated with a feature-level contrastive objective designed to enforce semantic consistency between co-masked regions, which enhances representation discriminability without the requirement for negative sample queues. Additionally, the architecture utilizes a hierarchical multi-scale attention mechanism that partitions feature channels into parallel streams, enabling the simultaneous capture of fine-grained morphological variations and the overarching global dental arch context. Extensive experiments demonstrate that our Dental-CMAE consistently outperforms state-of-the-art fully supervised and self-supervised methods across multiple evaluation metrics. Specifically, our framework achieves an Overall Accuracy (OA) of 95.57%, a mean Intersection-over-Union (mIoU) of 88.14%, and a mean Accuracy (mAcc) of 90.85%. Supported by these quantitative findings, our method validates its effectiveness for robust 3D tooth segmentation, highlighting its strong potential to alleviate annotation bottlenecks and improve the reliability of automated 3D digital dental workflows.
The accurate point cloud completion of individual tree crowns is critical for quantifying crown complexity and advancing precision forestry, yet it remains challenging in dense plantations due to canopy occlusion and LiDAR limitations. In this study, we extended the scope of conventional point cloud completion techniques to artificial planted forests by introducing a novel approach called Multi−feature Fusion Completion of Populus (MFCPopulus). Specifically designed for Populus Tomentosa plantations with uniform spacing, this method utilized a dataset of 1050 manually segmented trees with expert−validated trunk−canopy separation. Key innovations include the following: (1) a hierarchical adversarial framework that integrates multi−scale feature extraction (via Farthest Point Sampling at varying rates) and biologically informed normalization to address trunk−canopy density disparities; (2) a structural characteristics split−collocation (SCS−SCC) strategy that prioritizes crown reconstruction through adaptive sampling ratios, achieving a 94.5% canopy coverage in outputs; (3) a cross−layer feature integration enabling the simultaneous recovery of global contours and a fine−grained branch topology. Compared to state−of−the−art methods, MFCPopulus reduced the Chamfer distance variance by 23% and structural complexity discrepancies (ΔDb) by 33% (mean, 0.12), while preserving species−specific morphological patterns. Octree analysis demonstrated an 89−94% spatial alignment with ground truth across height ratios (HR = 1.25−5.0). Although initially developed for artificial planted forests, the framework generalizes well to diverse species, accurately reconstructing 3D crown structures for both broadleaf (Fagus sylvatica, Acer campestre) and coniferous species (Pinus sylvestris) across public datasets, providing a precise and generalizable solution for cross−species trees’ phenotypic studies.
In forestry data management and analysis, data integrity and analytical accuracy are of critical importance. However, existing techniques face a dual challenge: first, sensor failures, data transmission interruptions, and human errors lead to the prevalence of missing data in forestry datasets; second, the multidimensional heterogeneity and environmental complexity of forestry systems not only increase the difficulty of missing value estimation, but also significantly affect the accuracy of resolving the potential correlations among data. In order to solve the above problems, we proposed the L2 model using the aspen woodland as the experimental object. The L2 model consists of a complementary model and a predictive model. The L2 complementary model integrates low tensor tensor kernel norm minimisation (LRTC-TNN) to capture global consistency and local trends, and combines long and short-term memory and convolutional neural network (LSTM-CNN) to extract temporal and spatial features, which is effective in accurately reconstructing the missing values in forestry time-series data. We also optimised the LRTC-TNN model to handle multi-class data and incorporated a self-attention mechanism into the LSTM-CNN framework to improve performance in the case of complex missing data. The L2 prediction model adopts a dual attention mechanism (temporal attention mechanism and feature attention mechanism) based on LSTM to construct a stem diameter prediction model, which achieves high-precision prediction of stem diameter variation. Then we further analyzed the effects of various factors on stem diameter using SHAP (Shapley Additive Explanations).Experimental results demonstrate that our L2 significantly improves data completion accuracy while preserving the original structure and key characteristics of the data. Moreover, it enables a more precise analysis of the factors affecting stem diameter, providing a robust foundation for advanced forestry data analysis and informed decision making.
A method for identifying pine wood nematode-infected wood using terahertz time-domain spectroscopy (THz-TDS) was developed. The disease caused by this nematode is highly pathogenic and spreads rapidly. Current quarantine methods are time-consuming, so alternative techniques are needed. THz-TDS shows potential for wood identification. This study collected terahertz spectral data from different sections of healthy and infected Pinus massoniana Lamb, and 420 spectral datasets were obtained. The absorption coefficient and refractive index were analysed, and principal component analysis (PCA) was applied to extract features within the 0.2-2 THz range. A support vector machine (SVM) model was developed for classification. Results show differences in terahertz spectra between infected and healthy samples, with classification performance varying across sampling positions. The top section showed the highest accuracy (88%) in detecting infected wood. Furthermore, the performance of four neural network models - BP neural network, learning vector quantisation (LVQ), extreme learning machine (ELM), and SVM - was compared. The SVM model based on absorption coefficient data outperformed the others, achieving 87.5% overall accuracy. These findings highlight the potential of THz-TDS combined with machine learning for rapid and non-destructive identification of pine wood nematode infections.
Video analysis of soccer matches is vital for enhancing training methodologies and performance assessment. However, accurately tracking players and detecting the ball in real-time poses significant challenges due to occlusion, rapid movements, and variable lighting conditions. In this study, we introduce SoccerEDMF, an Enhanced Dual-Model Framework specifically designed for precise player tracking and ball detection in soccer broadcasts. By employing separate modules optimized for each task, our system leverages unique jersey color patterns for player identification and applies interpolation techniques to predict ball positions during occlusions. Experimental results on custom datasets demonstrate that SoccerEDMF outperforms state-of-the-art methods, achieving a mean average precision of 97.7 https://github.com/JianglangKang/SoccerEDMF .
Image coloring has always been a challenging problem. Current coloring methods cannot eliminate the ambiguity of bright material surfaces from specular highlight images, as the presence of specular highlights introduces complex brightness and color changes. To tackle this issue, we propose a novel multi-stage colorization network (MSC-Net) to identify and eliminate color ambiguity caused by specular highlights accurately. Firstly, we design a new method for detecting and removing specular highlights in images, including a specular highlight segmentation network(SHSNet) and a highlight-eliminating module, in order to accurately detect and eliminate specular highlight areas for subsequent coloring. Subsequently, we give a deep learning-based coloring network(ColorNet) and integrated highlight features to color the image. After coloring, the completed image will be combined with the original specular highlight information to achieve a comprehensive restoration of the highlighted area. Experiments on multiple public datasets are conducted, validating that our MSC-Net achieves favorable results in specular highlight segmentation and performs excellently in coloring images with specular highlights. Our MSC-Net provides a mature solution for image coloring with specular highlights, which has significant theoretical and practical implications. Our code is available at: https://github.com/ymm0304/MSC-Net .
We propose a comprehensive soccer match video analysis pipeline tailored for broadcast footage, which encompasses three pivotal stages: soccer field localization, player tracking, and soccer ball detection. Firstly, we introduce sports camera calibration to seamlessly map soccer field images from match videos onto a standardized two-dimensional soccer field template. This addresses the challenge of consistent analysis across video frames amid continuous camera angle changes. Secondly, given challenges such as occlusions, high-speed movements, and dynamic camera perspectives, obtaining accurate position data for players and the soccer ball is non-trivial. To mitigate this, we curate a large-scale, high-precision soccer ball detection dataset and devise a robust detection model, which achieved the mAP50-95$$ mA{P}_{50-95} $$ of 80.9%. Additionally, we develop a high-speed, efficient, and lightweight tracking model to ensure precise player tracking. Through the integration of these modules, our pipeline focuses on real-time analysis of the current camera lens content during matches, facilitating rapid and accurate computation and analysis while offering intuitive visualizations.
Soccer is a popular sport, and there is a growing need for automated analysis of soccer videos, while the detection and tracking of the players is the indispensable prerequisite. In this paper, we first introduce and classify multi-object tracking and then present two mostly used multi-object tracking methods, DeepSort and TrackFormer. When multi-object tracking is applied to soccer scenarios, some preprocessing and post-processing are generally performed, with preprocessing including processing of the video, such as splicing and background removing, and post-processing including further applications, such as player mapping for a 2D stadium. By directly employing the two methods above, we test the real scene and train TrackFormer to get further results. Meanwhile, in order to facilitate researchers who are interested in multi-object tracking as well as in the direction of player tracking, recent advances in preprocessing and processing methods for soccer player tracking are given and future research directions are suggested.
In order to generate a highly personalized, interactive and time-independent layout of artifacts in museums, to solve the limitations of real-world museums in terms of places, layouts and exhibition forms, to enrich the functions of existing digital museums, and to use the layout in digital museums to better guide the layout in real museums, a framework for personalized layout of museum exhibition halls is proposed. The framework consists of five modules: information preparation, personalized recommendation, random placement, optimization, and user interaction. The personalized recommendation module uses neural networks to train personalized recommendation models; the random placement module includes the layout of exhibit cases in the exhibition hall and the random filling of exhibit cases with artifacts, and proposes the Placement depends on Integrated path (PIP) algorithm. The optimization module includes efficiency optimization and rendering optimization; the user interaction module includes personal information collection, roaming, map navigation, and exhibit interaction functions. The experimental results show that the layout generated by this framework has a high degree of realism, scene loading speed and smoothness of online viewing, and the efficiency of artifact screening is improved to 7 times compared with that before optimization; the user tuning results show that more than 85% of people give realistic or very realistic evaluation to the layout, and the framework and algorithm in the paper can realize the personalized recommendation of artifacts and exhibition hall layout in a realistic, effective, real-time and interactive way.
This paper presents a visualization algorithm for wood fracture simulation based on wood science and wood internal structure reconstruction. The algorithm can simulate a reasonable and realistic wood fracture effect. First, the 3D point-cloud data of the bark structure are obtained using a laser scanner, and the cross-section of the branch is obtained by voxelization of the surface mesh model. Then, the outer contour of the cross-section is shrunk inward to reconstruct the annual rings and wood fiber bundles, and reasonable internal structures of branch 3D models are generated. The internal structure consists of a hierarchical model composed of several ring-like annual rings, and each annual ring is divided into a series of continuous fan rings. On the basis of the reconstruction results, the wood fracture surface model generated by the parameter control can be mapped to the irregularly shaped 3D branch model. In this research, the internal structure of branches and the shape of annual rings on the fracture surface of branches are analyzed to provide a reliable fracture model for different branch fractures of trees. In addition, the realistic fractured tree branch model generated by this algorithm can be widely applied in fields such as animation film special effects, game scene simulation, virtual reality scene construction, and mechanical research on broken tree branches.
To solve the problem of high difficulty factor of ski jumping and the difficulty of extracting data of this sport due to the danger of invasive devices such as wearable sensors and high price, a monocular video-based data extraction method for the flight phase of ski jumping is proposed. Firstly, we pre-process the video for the problem of distortion and background clutter, correct the distortion by calibrating the camera parameters, and use the inter-frame differential method to remove the background. Then, using the OpenPose to initially identify the joint position of the athlete, and obtained the 2D pixel coordinates of each joint-point in each frame. An iterative fitting algorithm is proposed to correct the nodes with errors in recognition by combining the pose characteristics. Finally, the athletes’ motion features are extracted and calculated according to the modified joint-points, including the extraction of athletes’ 2D motion trajectories, the calculation of spatial and temporal features, and the generation of SMPL 3D human model, and the human models are applied to compare with the athletes’ pose. The experimental results show that the iterative fitting algorithm improves the accuracy and precision of joint-points recognition, and the SMPL generated after the correction of joint-points is more suitable for the reality, which also proves the effectiveness of the algorithm for joint-points correction.
Flight simulation of catkins using computer technology helps their prevention and control. However, this is a challenging task due to the complex characteristics, and irregular shapes of catkins, while existing methods mainly focus on rain and snow, which are not suitable for catkins. In this paper, we propose a physics-based algorithm for the dynamic simulation of fluttering catkins. Our approach includes an L-system based 3D modeling method for simulating the natural phenomena of the catkin. We consider the motion of wind, free fall of catkins, and the dynamics of catkins under the joint action of attraction between them, while adhering to the physical motion law of catkins. To provide wind force, we first establish a three-dimensional wind field based on Boltzmann’s equation. We then use the kernel function idea to calculate the attraction force between catkins and finally update the position of the catkin. We incorporate the phenomena of collision and adhesion, attraction, and accumulation of catkins while simulating motion states depending on the adjusted wall height and ground humidity parameters. Our approach overcomes limitations of previous models by achieving good simulation while using relatively less code to simulate various realistic motion states. According to our users’ study, more than 71% of users found the simulation results to be acceptable, authentic, and realistic, confirming the authenticity of our simulation. Our method can generate highly realistic effects, significantly improving efficiency by several orders of magnitude compared to manual modeling. In addition, it can effectively simulate the dynamics of catkins in different scales, providing a decision-making reference for catkin control.
Sand painting is a significant form of artistic expression. It is necessary to convert real photographs into sand painting style automatically. However, current methods for creating artwork mostly concentrate on oil painting, watercolor painting, and other fields, with little involvement in sand painting. In this paper, we propose a sand painting conversion algorithm that aims to create a unique art style while preserving details. Our Sand Painting Conversion (SPC) method consists of two primary processes: coarse waving for the whole image and detail preservation for specific parts. The former process involves parameter optimization to render the entire image and obtain preliminary sand painting results. The latter process comprises three parts: detail painting based on sand reduction, refined waving for inhomogeneous areas, and edge detail processing based on multi-technique selection, in order to achieve detail preservation from a source image to the sand painting style. Our SPC can generate a series of realistic sand painting works without any user intervention. Comparative experiments and user feedback have revealed the effectiveness and superiority of our algorithm in converting natural images into high-quality and realistic sand painting style images.
Plantation forests, cultivated through artificial seeding and planting methods, are of great significance to human society. However, most experimental sites for these forests are located in remote areas. Therefore, in-depth studies on remote forest management and off-site experiments can better meet the experimental and management needs of researchers. Based on an experimental plantation forest of Triploid Populus Tomentosa, this paper proposes a digital twin architecture for a virtual poplar plantation forest system. The framework includes the modeling of virtual plantation and data analysis. Regarding this system architecture, this paper theoretically analyzes the three main entities of the physical world, digital world, and researchers contained in it, as well as their interaction mechanisms. For virtual plantation modeling, a tree modeling method based on LiDAR point cloud data was adopted. The transitional particle flow method was proposed to combine with AdTree method for tree construction, followed by integration with other models and optimization. For plantation data analysis, a database based on forest monitoring data was established. Tree growth equations were derived by fitting the tree diameter at breast height data, which were then used to predict and simulate trends in diameter-related data that are difficult to measure. The experimental result shows that a preliminary digital twin-oriented poplar plantation system can be constructed based on the proposed framework. The system consists of 2160 trees and simulations of 10 types of monitored or predicted data, which provides a new practical basis for the application of digital twin technology in the forestry field. The optimized tree model consumes over 67% less memory, while the R-2 of the tree growth equation with more than 100 data items could reach more than 87%, which greatly improves the performance and accuracy of the system. Thus, utilizing forestry information networking and digitization to support plantation forest experimentation and management contributes to advancing the digital transformation of forestry and the realization of a smart management model for forests.
病害仿真是农业信息化的重要研究内容之一,通过模拟病害发生过程,可方便直观有效地宣传作物病害有关知识.病害模拟在教学、影视及游戏等领域也有重大应用价值.为实现作物病害过程模拟,以小麦为研究对象,在温度与湿度变化环境下,对小麦条锈病、叶锈病、秆锈病3种锈病类型进行基于纹理特征的动态仿真模拟.用噪声模拟分布在小麦茎叶上的孢子堆,通过调整噪声属性模拟不同的锈病类型,从而达到逼真的模拟效果;通过改变遮罩贴图中红绿蓝(RGB)3种颜色分量的分布,指定模拟发病位置,区分小麦锈病的发病类型;通过着色器调整指定位置作物模型表面贴图,以径向为扩散路径,完成植株从完好到完全发病的可视化过程.该方法模拟效果较好,具有一定的研究意义与应用价值,可为其他作物的多病害模拟提供参考.
Plant disease visualization simulation belongs to an important research area at the intersection of computer application technology and plant pathology. However, due to the variety of plant diseases and their complex causes, how to achieve realistic, flexible, and universal plant disease simulation is still a problem to be explored in depth. Based on the principles of plant disease prediction, a time-varying generic model of diseases affected by common environmental factors was established, and interactive environmental parameters such as temperature, humidity, and time were set to express the plant disease spread and color change processes through a unified calculation. Using the apparent symptoms as the basis for plant disease classification, simulation algorithms for different symptom types were propose. The composition of disease spots was deconstructed from a computer simulation perspective, and the simulation of plant diseases with symptoms such as discoloration, powdery mildew, ring pattern, rust spot, and scatter was realized based on the combined application of visualization techniques such as image processing, noise optimization and texture synthesis. To verify the effectiveness of the algorithm, a simulation similarity test method based on deep learning was proposed to test the similarity with the recognition accuracy of symptom types, and the overall accuracy reaches 87%. The experimental results showed that the algorithm in this paper can realistically and effectively simulate five common plant disease forms. It provided a useful reference for the popularization of plant disease knowledge and visualization teaching, and also had certain research value and application value in the fields of film and television advertising, games, and entertainment.
Since forest and fruit wood borer insects are very harmful, and the formed galleries are complex and not easy to observe, the 3D reconstruction and visual prediction simulation of their galleries are of great importance in agricultural and forestry research. A single image-based 3D reconstruction and visualization method is proposed. The method is divided into two steps: (1) photographing the complete insect galleries on different sample wood segments, correcting the images to obtain the complete insect tract outline, and then redefining the height of model expansion based on the distance from the outline to the midline of the outline via the sketch-based reconstruction method to reconstruct the 3D geometric model of insect tracts; (2) setting the influencing factors, such as forest and fruit wood borer pest species, host plants and insect population density, and simultaneously judging the newly added sample points and updating the original skeleton points according to the category of sample points and the comprehensive consideration of influencing factors, so as to obtain the changes of insect gallery structure under different conditions and achieve the predictive simulation of insect tract structure. We found that modeling 3D wood borer galleries by different pests on different host plants can be achieved. Compared to the hand drawing method, our method can obtain 3D models in a very short time, and the experimental models are all reconstructed within 1.5 s. The predicted variation in the range of insect tracts indicate that it was inversely proportional to the population density and positively proportional to the moth-eating ability of the pests, indicating that the method reflects the relationship between the range of insect tracts and the influencing factors. The proposed method provides a new approach to the study and control of wood borer galleries in the forest and fruit industry. In conclusion, we provide a method to reconstruct and predict the wood borer galleries in three dimensions.