Camera calibration is crucial to accurate vision tasks for its establishment of the mapping between 3D space and 2D image space. Traditional calibration methods often suffer from limited accuracy and time-consuming processes. In this study, we propose an automatic guidance system leveraging an improved particle swarm optimization algorithm to achieve fast and high-precision camera calibration. Our system dynamically recommends optimal camera poses for the next calibration image, effectively reducing calibration uncertainty and enhancing accuracy. Furthermore, for wide-angle cameras, we introduce a pre-estimation of distortion coefficients to guide the calibration process, significantly improving the calibration of distortion parameters. Experimental results demonstrate that our method outperforms existing guidance systems, achieving higher calibration accuracy with fewer images and shorter calculation time. The results of camera parameters calibrated by the system are applied to the reconstruction based on point clouds, and can achieve desirable reconstruction effect. The proposed system holds promise for applications in film and television shooting, promoting the development of the industry by reducing calibration errors and equipment debugging time.
Industrial robots, particularly six-axis serial manipulators, have been widely deployed in manufacturing workflows including assembly, welding, material handling, inspection and precision machining. As the demand for higher end-effector positioning accuracy and trajectory tracking performance grows, end-position errors induced during manipulator operation—stemming from geometric deviations, joint friction, load fluctuations, current surges, as well as variations in velocity and acceleration—have emerged as a critical bottleneck limiting high-precision applications. Conventional error compensation approaches mostly rely on geometric calibration, empirical formulas or fixed regression algorithms, which struggle to adequately characterize error trends featuring strong temporal dependencies, nonlinearity and multi-factor coupling. To address the aforementioned limitations, this paper takes the UR5 industrial manipulator as the research object. Leveraging the NIST-released dataset for manipulator positional accuracy degradation monitoring, this study develops and implements a physics-aware Transformer-based compensation framework that integrates a physics-consistent constraint loss and a nonlinear exponential error amplification strategy with a standard Transformer encoder for end-effector positional accuracy degradation. Multiple variables including target joint position, velocity, acceleration, torque, motor current and control current are selected to construct time-window input vectors, which are used to train the Transformer regression model to capture the correlation between historical motion states and real-time end-effector positional accuracy degradation. Experimental results demonstrate that the proposed Transformer model can fully capture temporal contextual correlations and multi-feature fusion information embedded within manipulator kinematic data, delivering superior error compensation performance for the six-dimensional end-effector pose error prediction task. The self-attention-based time-series modeling framework is well-suited to the nonlinear, coupled and time-varying characteristics of manipulator operational errors. This work provides valuable references for accuracy enhancement of industrial robots and the design of intelligent error compensation schemes. This work provides valuable references for accuracy enhancement of industrial robots and the design of intelligent error compensation schemes, with the proposed physics-aware strategies being model-agnostic and potentially extensible to other regression architectures.
Although geometric reconstruction of general objects from images has made remarkable progress in recent years, slender structures remain largely underexplored, despite their critical importance in engineering, biomedical, and agricultural applications. To bridge this gap, we propose a dedicated 2DGS-based geometric reconstruction framework tailored for slender structures, achieving accurate and faithful geometry recovery. Our method first addresses the challenge that most slender objects are texture-less, which hinders reliable feature matching and pose estimation in traditional SfM pipelines. By leveraging the curve-like nature of slender structures, we perform a curve-guided SfM process that provides robust camera poses and accurate 3D curve initialization for Gaussian primitives. To ensure SfM reliability, we introduce a high-precision mask extraction strategy that integrates geometric priors with a segmentation network, effectively handling self-occlusion and thin geometry. Furthermore, to enhance fine geometric recovery, we incorporate a differentiable Poisson reconstruction module to extract an initial mesh during training, which is then refined via image-space iterative optimization using differentiable mesh rasterization. In contrast to conventional approaches that rely on differentiable Gaussian rasterization followed by TSDF-based mesh extraction, our method avoids the additional geometric errors and artifacts introduced during the intermediate TSDF conversion, thereby improving the overall reconstruction quality. Comprehensive experiments on both synthetic and real-world datasets validate that our method achieves superior reconstruction quality compared to state-of-the-art approaches.
Dressing animations have broad applications in film and animation, but current methods require retraining networks for new styles, which is resource-intensive. While 2D pattern parameters are used for static 3D clothing modeling and editing, applying them to control style variations in dynamic clothing deformation is challenging due to the uncertainty introduced by multiple parameters. Ensuring consistent and stable style features across time-varying deformation sequences is difficult. Therefore, we present StyleGarNet, a novel approach for style-parameter-controlled dressing animation generation and editing. We employ a conditional variational autoencoder architecture to build our network, with style parameters serving as constraint information. This allows us to learn the probabilistic distribution model of deformations under style constraints, crucial for enhancing the robust representation of styles during the deformation process. Simultaneously, considering the temporal nature of motion, we introduce a transformer layer to capture the temporal dependencies of both motion and clothing deformation, thereby enhancing the stability of clothing deformation. Ultimately, our approach enables flexible manipulation of dressing animation generation through inputting style and motion features. Evaluation results demonstrate the efficacy of our approach and show that StyleGarNet outperforms existing methods in terms of prediction speed, accuracy, and stability of deformation sequences.
We propose Self-Adapting NeRF for high-quality novel view synthesis based on non-ideal video input. We first using Lie algebra to encode camera poses, which were dynamically controlled by parameters from the precomputed Fourier transform encoding in the NeRF network input, achieving joint optimization of poses and the model. Subsequently, by monitoring intermediate training results, we supplement areas with poor performance, implementing a training strategy based on keyframe supplementation and gradient prioritization. This addresses the challenge of achieving high-quality novel view synthesis with NeRF-series in non-ideal input scenarios. Finally, we employ a strategy of geometry and material information separation, along with a reflection lighting model, to address issues in scenes with specular reflections.
Virtual character dressing deformation simulation is widely used in digital filmmaking, 3D gaming, animation, and metaverse construction to generate realistic dressing deformations and animations based on human body shapes and poses. Data-driven methods, compared to physically driven ones, offer advantages such as ease of control, speed, and data reusability, making them increasingly mainstream. However, they are often time-consuming and expensive, unsuitable for rapid iteration in current dressing animation. We present an unsupervised clothing deformation prediction model suitable for various body shapes and poses. Our method enables network training without a clothing deformation dataset by converting physical constraints into optimization objectives. Using a variational autoencoder with an encoder-decoder structure, we map body parameters (pose and shape) to clothing deformation. The reparameterisation module learns the latent space conditional probability distribution model from body features to clothing deformation. The feature-deformation transformation space is then learned to convert encoded vectors of different body features into corresponding clothing deformation vertex sets. Experimental results show that our model can be trained quickly without a clothing deformation dataset, even on a CPU, and can rapidly synthesize realistic clothing animation effects based on given body parameters, excelling in prediction speed and minimizing penetration loss.
This paper presents a multilayered garment animation generation method. Generating realistic dynamics in 3D garment animations is a challenging task due to the complex nature of multilayered garments and the variety of outer forces involved. Existing data-driven approaches have mainly focused on the study of static draping deformation of multilayer garments, with less consideration for the temporal deformation of garments, such as the time-varying motion behaviors of individual layers and their continuous interactions during motion. In addition, these methods require a substantial amount of high-quality paired garment datasets for network training, leading to a costly data acquisition and annotation process. To address these challenges, we propose a multilayered garment animation generation method that explicitly models different garment layers as separate meshes, and employs a combination of unsupervised and temporally supervised learning strategies to analyze and model the behavior of individual garment layers and their interactions. Our primary contribution lies in introducing a two-stage network architecture for layered garment processing, which decomposes multilayer garment deformation prediction into single-layer garment generation and interlayer garment interaction deformation. We focus more on generating two-layered clothing animations. Of course, our two-layered approach can be used iteratively to support more layers by using the current outer layer as the inner layer for the next iteration. This approach achieves dynamic simulation of multilayer garments, and experimental results demonstrate that our method can generate realistic multilayer garment deformation effects, outperforming existing methods both visually and in terms of evaluation metrics.
The fish motion skeleton serves as the foundation for 3D fish motion modeling, enabling the manipulation of fish posture deformations and movements, while also providing a robust framework for analyzing fish behavior to assess their health status and overall performance. However, the joints within the fish motion skeleton, responsible for driving the fish’s movements, are not always stable, which undergo changes as the fish grows. The unstable topology of the skeleton poses a challenge when attempting to simulate a lifelike fish skeleton. In this paper, we present a novel method for generating a 3D fish skeleton based on fish posture data. Our approach establishes an initial motion skeleton including the spine and fins. We then determine its parameters, encompassing joint positions and the number of joints, through iterative optimization, employing collected data from fish with various shapes and five common postures as constraints. Furthermore, the skeletons generated through this optimization process are utilized as sample data for training the FishSkeletonNet network, a framework introduced in this paper for predicting fish motion skeletons of input 3D fish bodies. To validate the effectiveness of our approach, we introduce a new dataset of grass carp postures, on which we carry out experiments and conduct both quantitative and qualitative evaluations. The experiments illustrate that our method generates fish motion skeletons that closely emulate the actual motion skeleton structure of fish, demonstrating a higher level of biological plausibility compared to existing methods.
Achieving a high grasping success rate in a stacked environment is the core of the robot’s grasping task. Most methods achieve a high grasping success rate by training the network on a dataset containing a large number of grasping annotations which requires a lot of manpower and material resources. Therefore, achieving a high grasping success rate for stacked scenes without grasping annotations is a challenging task. To address this, we propose a No-Grasp annotation grasp detection network for stacked scenes (NG-Net). Our network consists of two modules: an object selection module and a grasp generation module. Specifically, the object selection module performs instance segmentation on the raw point cloud to select the object with the highest score as the object to be grasped, and the grasp generation module uses mathematical methods to analyze the geometric features of the point cloud surface to achieve grasping pose generation without grasping annotations. Experiments show that on the modified IPA-Binpicking dataset G, NG-Net has an average grasp success rate of 97
Objective In industrial production,influenced by the complex environment during manufacturing and produc-tion processes,surface defects on products are difficult to avoid.These defects not only destroy the integrity of the products but also affect their quality,posing potential threats to the health and safety of individuals.Thus,defect detection on the surface of industrial products is an important part that cannot be ignored in production.In defect detection tasks,the tar-gets must be accurately classified to determine whether they should be subjected to recycling treatment.At the same time,the detection results must be presented in the form of bounding boxes to assist enterprises in analyzing the causes of defects and improving the production process.The traditional method of surface defect detection is the manual inspection method.However,in practice,manual inspection often has large limitations.In recent years,the performance of computers has improved by leaps and bounds,and traditional machine vision technology has been widely tested in various production fields.These methods rely on image processing and feature engineering,and in specific scenarios,they can reach a level close to manual detection,truly realizing the productivity replacement of machines for some manual labor.However,the shortcoming is the difficulty in extracting features from complex backgrounds,often resulting in inaccurate detection.Therefore,it is hardly reused in other types of workpiece inspection tasks.Deep learning has played an increasingly impor-tant role in the field of computer vision in recent years.Deep learning-based defect detection methods learn the features of numerous defect samples and utilize the defect sample features to achieve classification and localization.With high detec-tion accuracy and applicability,they have addressed the complexity and uncertainty associated with manual feature extrac-tion in traditional image processing,achieving remarkable results in industrial product surface defect detection.However,given the complex background of some industrial product surfaces,the high similarity between some surface defects and the background,and the small difference between different defects,the existing methods could hardly detect surface defects with accuracy.In this study,we propose a differential detection network(YOLO-Differ)based on YOLOv5.Method First,for cases where some defects are similar to background features on the surface of products,according to the studies of biology and psychology,predators use perceptual filters bound to specific features to separate target animals from the back-ground during predation.In other words,they capture camouflaged targets by utilizing frequency domain features.The fre-quency signal strength of the target is lower than that of the background,and this difference helps us find targets similar to the background.Therefore,a novel method is proposed for the first time to integrate frequency cues in the object detection network,thus addressing the issue of inaccurate localization caused by defects that resemble the background,thereby enhancing the distinguishability between defects and the background.Second,a fine-grained classification branch is added after the detection module of the network to address the issue of small differences in defect features among different types.The vision Transformer(ViT)classification network is used as the corrective classifier in this branch to extract subtle distin-guishing features of defects.Specifically,it divides the defective image into N blocks small enough to allow its inherent attention mechanism to capture important regions in the image.At the same time,Transformer performs global relationship modeling on different patches and gives each patch the importance of affecting classification results.This large range of relationship modeling and importance settings enable it to locate subtle differences in features and focus on important fea-tures of defects.Therefore,YOLO-Differ is divided into five parts:RGB feature extraction,frequency feature extraction,feature fusion,detection head,and fine-grained classification.First,RGB feature extraction,which consists of the back-bone network and neck,is responsible for extracting the basic RGB feature information and fusing RGB features of different scales to obtain improved detection results.Next,RGB images are converted to YCbCr image space,and its results are pro-cessed through discrete cosine transform(DCT)and frequency enhancement to obtain their frequency features.The feature fusion module aligns and fuses the RGB features with frequency features.Then,the fused features are fed into the detec-tion head to obtain defect localization information and preliminary classification results.Finally,the defect images are cropped in accordance with the location information and fed into the fine-grained classifier for secondary classification to obtain the final classification results of defects.Result In the experiment,YOLO-Differ models were compared with seven object detection models on three datasets,and YOLO-Differ consistently achieved optimal results.Compared with the current state-of-the-art models,the mean average precision(mAP)improved by 3.6%,2.4%,and 0.4%on each respective dataset.Conclusion Compared with similar models,the YOLO-Differ model exhibits higher detection accuracy and stronger generality.
Data-driven garment animation is a current topic of interest in the computer graphics industry. Existing approaches generally establish the mapping between a single human pose or a temporal pose sequence, and garment deformation, but it is difficult to quickly generate diverse clothed human animations. We address this problem with a method to automatically synthesize dressed human animations with temporal consistency from a specified human motion label. At the heart of our method is a two-stage strategy. Specifically, we first learn a latent space encoding the sequence-level distribution of human motions utilizing a transformer-based conditional variational autoencoder (Transformer-CVAE). Then a garment simulator synthesizes dynamic garment shapes using a transformer encoder–decoder architecture. Since the learned latent space comes from varied human motions, our method can generate a variety of styles of motions given a specific motion label. By means of a novel beginning of sequence (BOS) learning strategy and a self-supervised refinement procedure, our garment simulator is capable of efficiently synthesizing garment deformation sequences corresponding to the generated human motions while maintaining temporal and spatial consistency. We verify our ideas experimentally. This is the first generative model that directly dresses human animation.
In this paper, by replacing the exponential memory kernel function of a tabu learning single-neuron model with the power-law memory kernel function, a novel Caputo’s fractional-order tabu learning single-neuron model and a network of two interacting fractional-order tabu learning neurons are constructed firstly. Different from the integer-order tabu learning model, the order of the fractional-order derivative is used to measure the neuron’s memory decay rate and then the stabilities of the models are evaluated by the eigenvalues of the Jacobian matrix at the equilibrium point of the fractional-order models. By choosing the memory decay rate (or the order of the fractional-order derivative) as the bifurcation parameter, it is proved that Hopf bifurcation occurs in the fractional-order tabu learning single-neuron model where the value of bifurcation point in the fractional-order model is smaller than the integer-order model’s. By numerical simulations, it is shown that the fractional-order network with a lower memory decay rate is capable of producing tangent bifurcation as the learning rate increases from 0 to 0.4. When the learning rate is fixed and the memory decay increases, the fractional-order network enters into frequency synchronization firstly and then enters into amplitude synchronization. During the synchronization process, the oscillation frequency of the fractional-order tabu learning two-neuron network increases with an increase in the memory decay rate. This implies that the higher the memory decay rate of neurons, the higher the learning frequency will be.
This paper presents a self-supervised multi-layer garment animation generation network. The complexity inherent in multi-layer garments, particularly the diverse interactions between layers, poses challenges in generating continuous, stable, physically accurate, and visually realistic garment deformation animations. To tackle these challenges, we present the Self-Supervised Multi-Layer Garment Animation Generation Network (SMLN). The architecture of SMLN is based on graph neural networks, which represents garment models uniformly as graph structures, thereby naturally depicting the hierarchical structure of garments and capturing the relationships between garment layers. Unlike existing multi-layer garment deformation methods, we model interaction forces such as friction and repulsion between garment layers, translating physical laws consistent with dynamics into network constraints. We penalize garment deformation regions that exceed these constraints. Furthermore, instead of the traditional post-processing method of fixed vertex displacement calculation for handling collision interactions, we add an additional repulsion constraint layer within the network to update the corresponding repulsive force acceleration, thereby adaptively managing collisions between garment layers. Our self-supervised modeling approach enables the network to learn without relying on garment sample datasets. Experimental results demonstrate that our method is capable of generating visually plausible multi-layer garment deformation effects, surpassing existing methods in both visual quality and evaluation metrics.
Current detection methods for three dimensional(3D)point cloud data easily identify the local area of low-curvature cylindrical surfaces as planes in a model,but these methods can achieve the fast and accurate identification of only a single element.We propose a fast primitive detection method for point cloud data that can quickly and accurately detect both planar and cylindrical surfaces simultaneously.The proposed method is divided into two stages:coarse recognition and refinement.First,the point cloud is divided into small-grained patches,the patch characteristics are calculated,and the planar and cylindrical patches are roughly identified.Next,according to the filter conditions,the planar patches adjacent to the cylindrical patches are filtered,and then the patches with identical characteristics are combined to obtain the complete planar and cylindrical surfaces.Our experiments show that the proposed method is superior to two popular recognition methods when used to analyze data concerning five mechanical components.Moreover,the proposed method does not exhibit the omission and misidentification errors demonstrated by the other two methods,and the proposed method is more accurate in terms of the surface parameter estimation and segmentation when multiple cylindrical surfaces are connected.
Coloring line art images based on the colors of reference images is an important stage in animation production, which is time-consuming and tedious. In this paper, we propose a deep architecture to automatically color line art videos with the same color style as the given reference images. Our framework consists of a color transform network and a temporal constraint network. The color transform network takes the target line art images as well as the line art and color images of one or more reference images as input, and generates corresponding target color images. To cope with larger differences between the target line art image and reference color images, our architecture utilizes non-local similarity matching to determine the region correspondences between the target image and the reference images, which are used to transform the local color information from the references to the target. To ensure global color style consistency, we further incorporate Adaptive Instance Normalization (AdaIN) with the transformation parameters obtained from a style embedding vector that describes the global color style of the references, extracted by an embedder. The temporal constraint network takes the reference images and the target image together in chronological order, and learns the spatiotemporal features through 3D convolution to ensure the temporal consistency of the target image and the reference image. Our model can achieve even better coloring results by fine-tuning the parameters with only a small amount of samples when dealing with an animation of a new style. To evaluate our method, we build a line art coloring dataset. Experiments show that our method achieves the best performance on line art video coloring compared to the state-of-the-art methods and other baselines.
Phase unwrapping is an important part of fringe projection profilometry(FPP), which greatly affects the efficiency and accuracy of reconstruction. Phase unwrapping methods with deep learning achieve single-frequency phase unwrapping without additional cameras. However, existing methods have low accuracy in the real complex scene, and can not process data whose resolution is greater than the resolution of training data. This paper introduces a neural convolutional network named as VRNet which achieves accurate and single-frequency phase unwrapping without extra cameras. VRNet with encoder-decoder structure gets multi-scale feature maps through feeding the wrapped phase map into the encoder, then fuses the feature maps recursively by using the proposed feature fusion module to accomplish precise prediction. In order to further improve the accuracy of phase unwrapping, this paper presents a phase correction method based on the distribution characteristics of the absolute phase. The method divides the cross-section of the absolute phase map into several curves and identifies a misclassified pixel by comparing its absolute phase value with the value of neighboring curves. In contrast to existing methods, the method is row-independent and does not require segmentation of image. Moreover, this paper accomplishes the prediction of high-resolution data through the phase stitching strategy and fine-tuning the phase correction method. Extensive experiments show that the proposed method is able to achieve high-accuracy and single-frequency phase unwrapping in real scenes which consist of at least one complex object, and is also effective for wrapped phase maps with a resolution larger than the training data.
Coloring line art images based on the colors of reference images is a crucial stage in animation production, which is time-consuming and tedious. This paper proposes a deep architecture to automatically color line art videos with the same color style as the given reference images. Our framework consists of a color transform network and a temporal refinement network based on 3U-net. The color transform network takes the target line art images as well as the line art and color images of the reference images as input and generates corresponding target color images. To cope with the large differences between each target line art image and the reference color images, we propose a distance attention layer that utilizes non-local similarity matching to determine the region correspondences between the target image and the reference images and transforms the local color information from the references to the target. To ensure global color style consistency, we further incorporate Adaptive Instance Normalization (AdaIN) with the transformation parameters obtained from a multiple-layer AdaIN that describes the global color style of the references extracted by an embedder network. The temporal refinement network learns spatiotemporal features through 3D convolutions to ensure the temporal color consistency of the results. Our model can achieve even better coloring results by fine-tuning the parameters with only a small number of samples when dealing with an animation of a new style. To evaluate our method, we build a line art coloring dataset. Experiments show that our method achieves the best performance on line art video coloring compared to the current state-of-the-art methods.
Trajectory prediction plays a crucial role in autonomous driving. Existing mainstream research and continuoual learning-based methods all require training on complete datasets, leading to poor prediction accuracy when sudden changes in scenarios occur and failing to promptly respond and update the model. Whether these methods can make a prediction in real-time and use data instances to update the model immediately(i.e., online learning settings) remains a question. The problem of gradient explosion or vanishing caused by data instance streams also needs to be addressed. Inspired by Hedge Propagation algorithm, we propose Expert Attention Network, a complete online learning framework for trajectory prediction. We introduce expert attention, which adjusts the weights of different depths of network layers, avoiding the model updated slowly due to gradient problem and enabling fast learning of new scenario's knowledge to restore prediction accuracy. Furthermore, we propose a short-term motion trend kernel function which is sensitive to scenario change, allowing the model to respond quickly. To the best of our knowledge, this work is the first attempt to address the online learning problem in trajectory prediction. The experimental results indicate that traditional methods suffer from gradient problems and that our method can quickly reduce prediction errors and reach the state-of-the-art prediction accuracy.