
Parametric garment pattern models are widely used to generate 2D and 3D garment representations in reconstruction and simulation tasks. However, existing models are often difficult to adapt for specific purposes due to limited variations in base pattern shapes and the rigid, hardcoded handling of body and garment measurements. To address these challenges, we propose GarmentoPIA, a system that semi-automatically generates parametric garment pattern models from garment drafting literature using an LLM-based intelligent agent module. With GarmentoPIA, users can select garment drafting references and generate various pattern models that incorporate tailoring parameters specified in the source material. To improve adaptability across references, the system employs a prompt self-refinement mechanism that iteratively updates its instructions during model generation. This generation process is preceded by a manual preprocessing phase that first normalizes raw drafting instructions into a structured format. We furthermore introduce a garment domain-specific language (DSL), composed of explicit function calls corresponding to drafting components, to produce executable models that, once generated, run independently of any LLM—enhancing usability and accessibility. We validate the effectiveness of our approach through quantitative and qualitative evaluations.
Few-shot medical image segmentation (FSMIS) aims to achieve accurate segmentation with extremely limited labeled data. However, medical images are characterized by blurred boundaries, incomplete information, and high noise, leading to high uncertainty in model predictions. Most existing FSMIS methods, including recent descriptor-based state-of-the-art methods, generate segmentation results without explicitly modeling uncertainty at both the descriptor and prediction levels, which limits their reliability in complex regions. To address these problems, we propose DU-Net, a dual-level uncertainty-aware network designed to achieve reliable segmentation. Evidence extraction networks are incorporated to estimate the uncertainty of foreground and background descriptors. To suppress the interference of high-uncertainty descriptors, the descriptor-level uncertainty-aware similarity maps fusion module dynamically modulates the contribution of each descriptor during fusion. To enhance the prediction accuracy, our evidential dual-branch prediction module is designed to enable joint optimization of segmentation and evidential uncertainty estimation in a multi-task learning fashion. We impose consistency constraints on the predictions of the above two branches in low-uncertainty regions for mutual supervision. Experiments on three public medical image datasets demonstrate that our method outperforms current state-of-the-art methods.
This study introduces an extended reality (XR)-based interactive editing framework that allows users to sample colors directly from the physical world, apply them to virtual content, and register the results within the real environment. Existing augmented reality (AR) systems often suffer from scanning delays and low color fidelity, while virtual reality (VR) systems rely on predefined palettes that limit contextual creativity. Recent mixed reality (MR) authoring environments improve spatial continuity but still provide limited support for user-driven acquisition and direct editing of real-world colors or materials. To address these issues, we designed a cyclical interaction loop—sense, edit, display, and register—that integrates real world color sampling, 2D canvas editing, 3D result display, and spatial registration. We have evaluated the framework through three complementary studies: a within-subject comparison of AR, VR, and XR workflows, an XR ablation study, and a technical evaluation of reference-based color reproduction and XR sampling behavior. Findings showed that XR was associated with lower perceived workload and higher creativity support ratings than the scan-based AR workflow. Although XR showed longer task completion time than VR in the descriptive task time results, it received higher creativity support ratings and stronger hedonic-quality responses. Interviews emphasized the novelty and intuitiveness of sampling real-world colors and seeing them immediately reflected in physical space. These results suggest that XR's strengths are associated with the integration of real-world color sampling and contextual alignment in a unified creative experience. The proposed framework informs future XR authoring tools and possible applications in stage design, exhibition prototyping, and creative education.
Unsupervised real-world image super-resolution (SR) faces critical challenges due to the complex, unknown degradation distributions in practical scenarios. Due to a significant domain gap, existing methods struggle to generalize from synthetic low-resolution (LR) and high-resolution (HR) image pairs to real-world data. In this paper, we propose an unsupervised real-world SR method based on rectified flow to capture and model real-world degradation effectively, synthesizing LR-HR training pairs with realistic degradation. Specifically, given unpaired LR and HR images, we propose a novel rectified flow degradation module (RFDM) that introduces degradation-transformed LR (DT-LR) images as intermediaries. By modeling the degradation trajectory continuously and invertibly, RFDM better captures real-world degradation and enhances the realism of generated LR images. Additionally, we propose a Fourier prior guided degradation module (FGDM) that leverages structural information embedded in Fourier phase components to ensure precise modeling of real-world degradation. Finally, the LR images are processed by both FGDM and RFDM, producing final synthetic LR images with real-world degradation. The synthetic LR images are paired with given HR images to train off-the-shelf SR networks. Extensive experiments on real-world datasets demonstrate that our method significantly improves the performance of off-the-shelf approaches in real-world scenarios.
Real-time photorealistic rendering in virtual reality applications is challenging due to the requirements of high-quality visual perception, large display resolution with a wide field of view, and strict frame-rate constraints. Foveated rendering is an effective approach to address these issues by maintaining high quality in the foveal area while reducing quality in the periphery. Unlike traditional rasterization and ray tracing algorithms, the foveated mechanism cannot be directly integrated with a modern radiance-cache-based real-time global illumination pipeline, such as GI 1.0. We propose a real-time global illumination algorithm based on fovea-driven screen space radiance probes to adopt different caching and rendering strategies for different areas. By adaptively locating more screen probes in the foveal area, a higher rendering quality is achieved there. At the same time, only sparse probes need to be placed in the peripheral area, which leads to significant storage and computational savings. We further introduce a mipmap-based temporal probe reuse strategy with strict reuse in the foveal area, but coarse reuse in the peripheral area to improve performance. As a result, our FovGI-1.0 is able to improve quality in the foveal area without perceived quality decrease in the peripheral area, delivering up to 1.31× the frame rates. We also conducted user studies to demonstrate that the perceived quality of our method has a high visual similarity to the ground truth rendered by using fullresolution probe placement.
Neural style transfer (NST) can create impressive artworks by transferring a reference style to a content image. Current image-to-image NST methods lack the fine-grained control often demanded for artistic editing. To mitigate this limitation, we propose a user-oriented interactive style transfer (IST) method, using which a harmonious image like drawing can be interactively created. Our IST method can serve as a brush, dipping style from anywhere, and then painting to any region of the target content image. To control the action scope, we formulate a fluid simulation algorithm, which takes styles as pigments around the position of brush interaction, and uses diffusion in style or content images according to similarity maps. By dipping and painting, even employing a single style image can produce thousands of eye-catching works. Our method expands the creative capabilities of NST.
Image style transfer aims to merge the artistic characteristics of a style image with the spatial structure of a content image, maintaining the content identity and showing the desired style. Previous works based on convolutional neural networks lack the ability to capture global features, leading to the content leakage issue. In this work, we propose a method StyTips, short for style via transformer filter prompts, a novel approach for style transfer that distills essential content and injects style information via learned prompts. To achieve efficient and high quality stylization, we adopt a hierarchical vision transformer encoder to extract multi-scale features of the content and style images, learning disentangled style features of color and strokes. Inspired by the prompt learning technique, we have designed a fine grained style alignment block in the decoder, consisting of a filter-using self-attention module to distill necessary content information and a style-aware cross-attention module to inject style features. StyTips takes full advantage of the attention mechanism to align color and strokes with the purified content features in appropriate proportions. Qualitative and quantitative experiments demonstrate that StyTips can prevent content leakage and generate high-quality stylized images. StyTips provides controllable style transfer with diverse ways for users to manipulate styles in a fine-grained manner and to determine the strength of style, giving it more practical significance to real-world applications than traditional approaches. The code is available at https://github. com/I2-Multimedia-Lab/StyTips.
We present a practical and unbiased sampling technique for volumetric light sources defined by signed distance functions (SDFs). SDFs can compactly represent complex procedural shapes and support Boolean operations for flexible composition. However, their use as emitters remains underexplored in photorealistic rendering due to the absence of efficient sampling strategies. Our key insight is to model the interior of an SDF as a uniform volume and to project volumetric samples onto the unit sphere centered at the shading point, yielding a solid angle distribution that captures the emitter's relative spatial layout. We derive the directional probability densities for analytic primitives and generalize the formulation to arbitrary SDFs via Monte Carlo volume estimation. Leveraging robust sphere tracing, our method enables accurate and efficient sampling of SDF emitters without requiring explicit surface parameterization, voxelization, or precomputed tables. Compared to baseline approaches such as uniform directional sampling or surface approximations, our technique achieves significantly lower variance and runtime in next-event estimation, while broadening the expressive power of light sources in rendering.
3D Gaussian splatting demonstrates significant advantages in high-fidelity rendering for offline reconstruction of objects or scenes. However, its integration with existing simultaneous localization and mapping (SLAM) systems, especially in monocular scenarios, still suffers from limited localization accuracy and poor rendering quality. In this paper, we propose a new monocular 3D Gaussian splatting SLAM, which achieves high fidelity online 3D Gaussian splatting based reconstruction given monocular video input with significance-guided pruning. Our key idea is to maintain structural compactness when optimizing the 3D Gaussians, by adaptively pruning them using a global significance evaluation based on multidimensional cues such as visibility, opacity, and volume coefficient. Specifically, we use a frame-to-model pipeline that jointly optimizes camera poses and 3D Gaussians within a sliding-window framework, ensuring a globally consistent 3D Gaussian representation for high-fidelity rendering with significance-guided pruning. Furthermore, an online monocular depth estimation model is incorporated to extract depth priors from input images, to effectively initialize the 3D Gaussian attributes for better camera tracking. Extensive experiments on Replica and TUM datasets demonstrate that our approach substantially improves both tracking performance and rendering fidelity, and thus provides state-of-the-art results.
Image appearance transfer plays a significant role in interior design by allowing designers to efficiently explore different design styles according to clients' preferences. Given as input an example image and a target scene, the goal is to efficiently produce scene images that exhibit the desired appearance while maintaining the harmonization of the entire scene. For interior designs composed of multiple objects, the key to object-aware appearance transfer is to prevent scrappy appearance features from being distributed all over the generated images. In this paper, we utilize a pre-trained vision�language model (VLM) and a text-to-image generative model to solve the object-aware appearance transfer task. Specifically, we propose a VLM-assisted align-and-complement strategy using scene graph representation to determine object appearance with comprehensive considerations in the target scene. In addition, when conducting multi-object appearance transfer, we propose a multiple contrastive loss to learn object-aware appearance features from single examples and to manipulate the compositional conditions to precisely control the transfer. We have constructed an evaluation dataset and performed comparative experiments to demonstrate the effectiveness of our method of object-aware appearance transfer for interior designs. Both qualitative and quantitative evaluations demonstrate that our method successfully transfers object appearances from an example image to target scenes without feature leakage, achieving superior visual effects to competing solutions.
Medical visual question answering (VQA) is a key medical AI challenge, but scarce data limits progress. Current methods prioritize more pre-training data, overlooking medical data's inherent constraints (ethics, privacy, specialization) which cause slower accumulation of data than public web data. Sole reliance on more data risks rapid performance plateaus. To address these challenges, this paper proposes the cross-modality discriminative pattern identification model (CMDPI), which fundamentally rethinks training methodologies to provide a data-efficient framework for medical VQA. During pre-training, CMDPI identifies inherent biases in conventional techniques under conditions of data scarcity and introduces a co-regularization approach that integrates multiple pretraining techniques for regularization to enhance model generalizability. For fine-tuning, a difference reconstruction mechanism is proposed that effectively preserves unique discriminative features, and head mixup is introduced to further remedy the issue of overfitting. Experimental results demonstrate that CMDPI achieves performance comparable to or surpassing existing methods while requiring substantially less pre-training data. Our work shows the viability of optimizing training paradigms rather than pursuing indiscriminate data scaling for advancing medical VQA systems.
Federated learning has emerged as a promising paradigm for privacy-preserving collaborative learning. Recently, personalized federated learning has attracted growing interest due to its ability to handle statistical heterogeneity across clients, such as hospitals or mobile devices. However, existing approaches often align local models with global representations that are domain-variant, biased toward dominant clients, or unstable during training, thereby compromising fairness and generalization. To address these limitations, we propose FedDTR, a novel personalized federated learning framework that leverages domain-invariant text-derived label embeddings as stable global references. These embeddings serve as global label anchors for cross-client alignment and provide intradomain global priors. Specifically, FedDTR pulls sample embeddings toward their corresponding class anchors while pushing them away from those of other classes, thereby enhancing intra-class compactness and interclass separability. Furthermore, an intra-domain prior module exploits local client data to estimate domain- specific global priors, enabling better modeling of the underlying data distribution. We evaluate FedDTR on benchmark datasets spanning computer vision and natural language processing, and demonstrate its consistent superiority over state-of-the-art methods.
Reconstructing volumetric smoke from 2D images remains challenging in computer graphics. Although physics-based differentiable rendering offers a promising solution, existing methods are computationally expensive and impractical for interactive use. We introduce an efficient physics-based differentiable renderer tailored for smoke reconstruction, based on two insights: (i) for the optically thin-to-moderately dense smoke targeted in this work, appearance is well approximated by single scattering, allowing us to avoid costly multiple scattering simulation, and (ii) direction-independent terms can be precomputed to avoid redundant forward and backward passes. Our technical contributions include a pre-computation strategy that eliminates repetitive calculations, and an efficient ray marching method that computes density derivatives without costly Monte Carlo estimation or automatic differentiation. Comprehensive evaluation shows that our approach achieves real-time performance while maintaining high reconstruction quality. In contrast to current best methods that require minutes per frame, our method enables interactive smoke reconstruction while maintaining competitive visual fidelity. We demonstrate its utility in applications such as augmented reality and fluid motion reconstruction.
Mesh denoising algorithms explore recovering smooth surfaces from 3D meshes that are corrupted with noise. They aim to remove noise while maintaining geometric features. Although existing mesh denoising methods have shown promising results, they typically involve many parameters and settings that can be complex to non-experts. Also, parameter tuning is usually manual and realized through trial and error. To address this, we propose a new fairer benchmark for mesh denoising algorithms. This benchmark consolidates existing algorithms into a single platform, organizes various noisy mesh data, and provides optimal parameter settings for different algorithms. Specifically, we propose a more effective and scientific method for setting parameters by utilizing Bayesian optimization to find optimal parameters. We use the upper confidence bound strategy to balance exploration and exploitation needs. We conduct extensive experiments to compare existing mesh denoising algorithms quantitatively and qualitatively. These experiments show that there does not exist a single method that is optimal for all shapes and noise types. Additionally, we provide all the optimal parameter settings of the state-of-the-art methods that are available to the community.
Image restoration is the inverse problem of recovering high-quality images from knowledge of degraded images; it includes image super-resolution, image denoising, image deblurring, etc. The objective of image restoration methods is to minimize the error defined by the loss function between the network output image and the corresponding ground truth. Distortion-oriented loss functions are fundamental and commonly used in image restoration. However, these functions treat all pixels as equally important without differentiating between sharp and blurred edge areas, which does not match human visual perception. As a result, these methods can produce accurate but blurred images. To address this issue and achieve both accurate and perceptually satisfactory results, we propose a novel perception-enhanced distortion-oriented loss (PE loss) for image restoration, inspired by the Mach band effect. This effect demonstrates that sharp edges are perceived as having better quality than blurred edges by the human visual system. Our approach includes designing a blur factor map that detects blurred pixels and penalizes them by amplifying their error. The PE loss is a simple yet effective plug-and-play method, and we apply it to state-of-the-art networks. Extensive quantitative and qualitative experiments show that our method can restore images with sharp edges and high perceptual quality.
Integrating artificial intelligence (AI) into computer-aided design (CAD) has shown the potential to transform design and manufacturing processes, enabling more efficient, intuitive, and intelligent workflows. In recent years, the application of AI to 3D CAD model generation tasks has gradually emerged. To better enable researchers to understand the current research status of the AI-based CAD generation field and to inspire them to conduct further research, this survey explores the role of AI in 3D CAD model generation tasks that utilize various representations and conditions, ranging from traditional machine learning to LLM-based approaches. Additionally, AI applications in other extended CAD areas are also touched upon in the survey. Finally, we analyze current progress, identify challenges and limitations faced by this field, and propose possible directions for future work.
Urban green spaces such as parks and gardens are indispensable in both virtual and real-world environments. Therefore, planning such spaces is highly valuable. While scene synthesis literature has limited interest in this topic, many existing parametric design and procedural content generation approaches can be adapted to generate urban green spaces. However, these approaches heavily rely on manual work or are prone to producing monotonously repeated objects. This paper presents a framework that can automatically plan urban green spaces. Tailored to urban green space design, our framework comprises three steps: road system generation, region type planning, and model placement. First, it constructs undirected graphs to generate a sound road system for an empty site and divides the space into separate regions. Then it applies a genetic algorithm to plan suitable surface and vegetation for every region. Finally, it places landscape models based on various patterns and adds embellishments to complete an appealing urban green space. Our framework enables the automatic production of urban green spaces. Through extensive experiments, we demonstrate that the generated results are plausible and reasonable.
LiDAR point cloud data is indispensable for autonomous driving systems, enabling critical 3D perception tasks such as 3D object detection and segmentation. However, a shortage of labeled LiDAR data hinders the development of robust deep learning algorithms for these tasks. Data augmentation, as an important and effective technique to increase labeled data, has been applied to LiDAR data in different ways such as geometric transformation, mixup, and inserting synthetic objects. In this paper, we focus on exploring how to utilize synthetic objects to augment LiDAR data in a more effective manner. Given synthetic objects which are in the form of CAD models, we consider three factors that significantly impact the realism and utility of augmented data: insertion pose, point cloud spatial distribution, and point intensity value. Unlike previous insertion methods which only addressed one or two of these factors, our novel framework, LPA Aug, learns to place synthetic objects in and adjust them to existing scenes. A pose prediction module first learns to generate object placement locations and orientations based on scene context. Then, a novel 2D image-based distribution adjustment module adjusts the spatial distribution of points sampled from the inserted synthetic object. Finally, a further intensity prediction module learns in a 2D manner to predict the intensity of each object point. We evaluate LPA-Aug by using the augmented data for 3D object detection on the KITTI dataset. The results demonstrate the superiority of LPA-Aug over prior methods of LiDAR data augmentation.
Low-light image/video analysis is essential for various applications, e.g., night surveillance and photography, high-speed imaging, and autonomous vehicles. Under such conditions, cameras suffer from low signal-to-noise ratio, which degrades image quality severely and poses challenges for downstream tasks such as object detection. Data-driven methods have achieved enormous success for normal-light image/video restoration and high-level vision tasks. However, the lack of a high-quality benchmark dataset with accurate semantic annotations for low-light images and especially videos greatly hinders research progress. In this paper, we contribute the first multi-illuminance, multi-camera, low-light dataset, DarkVision, serving both image/video enhancement and object detection applications. We provide bright and dark pairs with pixel-wise registration, in which the bright counterpart provides a reliable reference for enhancement and annotation. This dataset comprises 13,455 images of 900 static scenes with objects from 15 categories, and 89,411 frames of 32 dynamic scenes with 4 categories of objects. For each scene, images/videos were captured at 5 illuminance levels using three cameras of different quality grades; average photon numbers can be reliably estimated from the calibration curves for quantitative studies. The static images and dynamic videos respectively contain around 7344 and 320,667 object instances in total. With DarkVision, we establish baselines for image/video enhancement and object detection by representative algorithms. To demonstrate an exemplary application of DarkVision, we propose two simple yet effective approaches to improve the performance of video enhancement and object detection respectively by exploiting temporal cues. Furthermore, we study the relationship between image enhancement and object detection. We believe DarkVision can help to advance the state-of-the art in both low-light image/video enhancement and object detection, as well as benefiting cross-task studies.
We propose a multi-view weakly-supervised 3D human pose estimation network. In this network, we fuse multi-view features and generate the final 3D pose, supervising the 3D pose using human body segmentation generated by a human parsing network. Using human body segmentation provides powerful supervision for the network. We further propose the Pose2Seg algorithm to transform 3D pose into simple simulated segmentation for all views, which allows the network to utilize the human body segmentation to supervise the 3D pose generated by the network, without the need for complex rendering processes and estimation of human body shape. We demonstrate the effectiveness of our method on various datasets and compare our method to other state-of-the-art algorithms, showing its advantages in terms of both quality and ability to generalize.