Automatic full-arch tooth segmentation and identification from intraoral scan (IOS) meshes is a fundamental task in digital orthodontics and restorative processes. However, it presents a structural difficulty: inter-tooth boundaries are thin and require precise local reasoning, while tooth identity depends on global arrangement and long-range geometric context along the dental arch. Existing mesh-based methods lack geometrically aligned cross-resolution correspondence. In this paper, we propose a Self-Parameterization Multi-Scale Mesh Segmentation Network (SPMMSegNet), which is an end-to-end mesh-native architecture designed to address this challenge through scale-structured propagation. Feature aggregation is restricted to intrinsic surface topology at fine resolutions to preserve boundary detail, while Euclidean aggregation is introduced at the coarsest scale to capture intra-tooth and arch-level context. The use of bijective surface correspondence enables geometrically consistent cross-scale feature transport. Specifically designed arch-aware positional encoding and an order-consistency regularization further incorporate anatomical structure into learning. By combining surface topology, global context, and anatomical priors within a unified mesh-based framework, SPMMSegNet provides an effective solution for full-arch tooth identification.
Registration of multiview point clouds conventionally relies on extensive pairwise matching to build a pose graph for global synchronization, which is computationally expensive and inherently ill-posed without holistic geometric constraints. This paper proposes FUSER, the first feed-forward multiview registration transformer that jointly processes all scans in a unified, compact latent space to directly predict global poses without any pairwise estimation. To maintain tractability, FUSER encodes each scan into low-resolution superpoint features via a sparse 3D CNN that preserves absolute translation cues, and performs efficient intra- and inter-scan reasoning through a Geometric Alternating Attention module. Particularly, we transfer 2D attention priors from off-the-shelf foundation models to enhance 3D feature interaction and geometric consistency. Building upon FUSER, we further introduce FUSER-DF, an SE(3)^N diffusion refinement framework to correct FUSER's estimates via denoising in the joint SE(3)^N space. FUSER acts as a surrogate multiview registration model to construct the denoiser, and a prior-conditioned SE(3)^N variational lower bound is derived for denoising supervision. Extensive experiments on 3DMatch, ScanNet and ArkitScenes demonstrate that our approach achieves the superior registration accuracy and outstanding computational efficiency.
This paper presents Paver, an interactive tool for designing structured patterns composed of 3D elements on free-form 3D surfaces. To address the inherent complexity of direct pattern creation in 3D space, we introduce a parametric design model that allows users to define patterns on a 2D parameter plane. Each pattern is defined by an underlying tessellation and isometric elements. A bijective mapping then lifts these 2D designs onto the target input surface. To ensure structural coherence and accommodate surface curvature, we model the tessellation as a 2D mass-spring system and propose an optimization method to adjust the particle positions on the 2D parameter plane under well-designed system forces. Once optimized, the tessellation is transferred to the 3D surface, generating the final element-based pattern. Paver supports diverse pattern types and offers real-time feedback, enabling designers to iteratively explore and refine complex surface decorations with ease. Experimental results are presented to demonstrate the effectiveness of our Paver.
Reconstructing precise CAD modeling sequences from point clouds remains a challenging task, especially for objects with complex geometry and topology. In this paper, by formulating the CAD sequence reconstruction as a Markov decision process, we introduce ReACT, a novel Reward-informed Autoregressive decision Cad Transformer architecture for robust CAD sequence prediction. Beyond previous imitation-only approaches, our key innovation is to frame the CAD Transformer under a reinforcement learning paradigm and thereby integrate reward-inspired heuristic learning into our architecture. This allows ReACT to effectively leverage shape-aware long-term reward feedback to guide the inference of (nearly) optimal CAD commands. Specifically, conditioned on past tokens, comprising the historical CAD states, sketch-extrude commands (i.e., actions) and associated geometric rewards, ReACT autoregressively outputs the most promising CAD commands in a causal manner. In particular, we develop a novel scaffold-aware CAD state representation that integrates global point-command features with an incrementally constructed surface point scaffold, enabling fine-grained geometric reasoning for subsequent reconstruction prediction. Moreover, an effective local barrel points-guided dense reward function is designed to jointly evaluate surface fidelity and command efficiency for reliable reward guidance. Extensive evaluations on the DeepCAD and Fusion360 benchmarks demonstrate that ReACT can achieve superior CAD reconstruction quality, even for objects with complex shapes.
This paper proposes GM-R^2, a novel Generative Matching Learning framework for unsupervised geometric descriptor learning and correspondence matching. By reformulating descriptor learning as geometry-conditioned cross-view image generation, GM-R^2 leverages the proxy supervisory signal from structurally aligned view synthesis to implicitly enforce feature consistency across correspondence, enabling robust 3D matching. To instantiate GM-R^2, we introduce Denoising-Agnostic Coupled ControlNet conditioned on depth maps as the required geometry-conditioned cross-view generator. It effectively extends the single-view generation of naive ControlNet to the cross-view via coupled depth-map input design and further remove the latent noise dependency to support geometry-only inference (expected by 3D matching). Moreover, we present Zoomable Equirectangular Projection for intrinsics-free point cloud-to-depth mapping that adaptively zooms into the angular region occupies by the narrow-FOV input for dense range-map acquisition. Extensive experiments on 3DMatch and ScanNet datasets verify the superior precision of our GM-R^2, even surpassing supervised methods.
Learning-based CAD modeling shows great promise in automating parametric design, yet existing approaches often overlook the incremental and state-dependent nature of sketch construction. We present CADSketcher, a query-driven bidirectional framework for completing partial parametric sketches by internalizing the non-linear construction logic of interactive CAD processes. At the core of CADSketcher are two key innovations. First, a bidirectional sketch learner recovers both prior and posterior contexts from arbitrary-span partial sketches via a bidirectional query mechanism, enabling exploration of multiple plausible modeling trajectories. Second, a confidence-guided completion pipeline adaptively determines the expansion direction through a confidence gate and ensures executable instruction generation using a validity compiler, while a progressive context updater preserves sketch consistency throughout the evolving sketch state. In addition, a hybrid positional encoding integrates global modeling progression with local geometric semantics, reinforcing structural coherence during both learning and completion. Extensive experiments demonstrate that CADSketcher achieves superior geometric validity and instruction consistency across diverse sketch completion tasks, offering a robust and interpretable framework toward intelligent CAD automation.
3D Gaussian Splatting (3D-GS) enables real-time, photorealistic view synthesis, but its radiance is tied to the captured lighting and its point-like representation is incompatible with standard digital content creation tools, game engines, and simulation platforms. Existing surface extraction methods typically defer converting splats to polygons until after photometric optimization, resulting in blurred edges, inflated planar regions, and loss of fine geometric details due to the continuous nature of Gaussian densities. This paper proposes a surface-oriented, differentiable optimization framework to overcome these issues. Our approach introduces a triangle-soup optimization pipeline that transforms a pre-trained 3D-GS model into a high-quality manifold mesh, preserving sharp features and fine geometry. Specifically, it starts with sharpening the opacity of input splats to better align with underlying surfaces. Each visible splat is then replaced by a compact triangle pair, followed by an image-driven, geometry-aware optimization that refines the resulting triangle-soup. This yields a compact, feature-preserving, and size-adaptive triangle set that can be sealed into a watertight surface using standard wrapping software tools with minimal post-processing loss. Experiments on synthetic, indoor, and object-level reconstruction benchmarks show that our method improves mesh quality over representative GS-to-mesh baselines and achieves competitive geometric accuracy against recent GS-based surface reconstruction methods.
Detecting Schelling Points—salient 3D mesh landmarks that serve as natural reference points for shape analysis—is a challenging problem in geometry processing. While existing CNN-based methods struggle with limited receptive fields and poor geometric context modeling, this paper proposes {\em SchellingFormer}, a novel Laplacian matrix-guided Geometric Transformer that effectively captures long-range dependencies and discriminative geometric features for robust Schelling point prediction. Our framework consists of two key components: (i) a hybrid geometric feature embedding module that integrates handcrafted descriptors (coordinates, Gaussian curvature, and curvature differences) to encode local geometry, and (ii) a Laplacian-driven vector attention mechanism, where spatial relationships encoded by the Laplacian matrix guide feature aggregation with the Transformer. This approach enables adaptive, geometry-aware message passing and contextual representation learning. Extensive experiments demonstrate that SchellingFormer outperforms state-of-the-art methods across multiple evaluation metrics. Our work bridges the gap between spectral mesh analysis and Transformer-based learning, offering a powerful tool for 3D shape understanding tasks such as shape matching and saliency detection.
This paper addresses the challenge of estimating local surface differential properties, specifically surface normals and curvatures, from raw 3D point clouds. Traditional methods either rely on fitting pre-defined analytic surfaces risking model bias, or directly regress normals and curvatures overlooking their intrinsic geometric correlation. We propose a learning-based approach that locally fits osculating implicit quadrics to recover both normals and curvatures simultaneously. Drawing on classical differential geometry, we exploit the fact that every point on a C² surface admits an osculating quadric in Monge form that exactly reproduces local differential properties. However, the Monge frame itself depends on the very differential quantities being estimated. To bypass this circularity, we reformulate the Monge-form quadric as an implicit representation in a canonical local frame derived solely from point coordinates, enabling supervised learning without requiring Monge frame alignment. This reformulation allows us to construct a ground-truth dataset of such local-frame quadrics and train a neural network to predict per-point weights and offsets for a robust weighted least squares fitting process. The learned offsets account for the deviations of neighboring points from the idealized osculating surface. We further incorporate stable curvature formulations into the training loss alongside normal supervision to enhance estimation fidelity. Extensive experiments on diverse datasets demonstrate that our method outperforms prior approaches in normal and curvature estimation from raw point clouds.
The orientation of surface normals in 3D point cloud is a fundamental problem in computer vision and graphics. Determining a globally consistent orientation solely from the point cloud is however challenging due to the global scope of the problem and the discrete nature of point cloud, particularly in the presence of noise, outliers, holes, thin structures, and complex topologies. This paper presents an efficient, robust, and global algorithm for generating consistent normal orientation of a dense 3D point cloud. The basic idea is to transform the original binary normal orientation problem to finding a relaxed sign field on a Delaunay graph, which can be achieved by solving a sparse linear system. The Delaunay graph is constructed by triangulating a level set of an implicit function defined from the input point cloud. The shape diameter function is estimated to serve as a prior for determining an appropriate level value such that the level set implicitly defines the inner and outer shells enclosing the input point clouds. As such, our algorithm leverages the strengths of the shape diameter function, Delaunay triangulation, and the least-square techniques, making the underlying processes take both geometry and topology into consideration, and thus provides an efficient and robust solution for handling point clouds with complicated geometry and topology. Extensive experiments on various shapes with noise and outliers confirm the effectiveness and robustness of our algorithm.
In this paper, we propose a novel 3D registration paradigm, Generative Point Cloud Registration, which bridges advanced 2D generative models with 3D matching tasks to enhance registration performance. Our key idea is to generate cross-view consistent image pairs that are well-aligned with the source and target point clouds, enabling geometric-color feature fusion to facilitate robust matching. To ensure high-quality matching, the generated image pair should feature both 2D-3D geometric consistency and cross-view texture consistency. To achieve this, we introduce Match-ControlNet, a matching-specific, controllable 2D generative model. Specifically, it leverages the depth-conditioned generation capability of ControlNet to produce images that are geometrically aligned with depth maps derived from point clouds, ensuring 2D-3D geometric consistency. Additionally, by incorporating a coupled conditional denoising scheme and coupled prompt guidance, Match-ControlNet further promotes cross-view feature interaction, guiding texture consistency generation. Our generative 3D registration paradigm is general and could be seamlessly integrated into various registration methods to enhance their performance. Extensive experiments on 3DMatch and ScanNet datasets verify the effectiveness of our approach.
This paper introduces a new class of linkage mechanisms called surface-conformable mechanisms or simply conformable mechanisms. A conformable mechanism conforms to a freeform surface in one of its configurations, in which the mechanism’s joints and links are exactly on the surface. Conformable mechanisms can be stowed compactly when not in use and accomplish complex motion transfer tasks when deployed. This paper aims to model and design conformable mechanisms for 3D path and motion generation. To achieve this goal, we enumerate topologies of conformable mechanisms, and model their geometry in the parameterization space of a freeform surface for surface conformity. To ensure a working and fabricable mechanism, we propose an efficient approach to processing the freeform surface by first removing a portion of the surface that collides with the moving links and joints and then removing disconnected patches and fragile features from the surface. Taking the modeling and processing as a foundation, we propose an optimization-based approach to designing a conformable mechanism for generating a target 3D path/motion, while preserving the mechanism’s appearance in the stowed state. We demonstrate the effectiveness of our approach by designing conformable mechanisms that conform to various freeform surfaces, evaluating their kinematic performance in 3D path and motion generation, validating their functionality with a 3D printed prototype, and showing three applications of these mechanisms.
We propose a new method for computing smooth and integrable cross fields on 2D and 3D surfaces. our approach first computes smooth cross fields by minimizing the Dirichlet energy. Unlike existing optimization-based methods, our technique determines the singularity configuration-i.e., the number, locations, and indices of singularities-by iteratively adjusting them. Singularities can move, merge and split, akin to the behavior of like charges repelling and unlike charges attracting. Once all singularities stop moving, we obtain a cross field with (locally) the lowest Dirichlet energy. In simply connected domains, this cross field is guaranteed to be integrable. However, this property does not hold in multiply connected domains. To make a smooth cross field integrable, we construct a vector field $\bf c$c that characterizes the deviation of the cross field from a curl-free field. We then optimize the locations of singularities by moving them along the field lines of $\bf c$c. Our method is fundamentally different from existing integer programming-based approaches, as it avoids combinatorial optimization. It is fully automatic and includes a parameter to control the number of singularities. Our method is well suited for smooth models where exact boundary alignment and sparse hard directional constraints are desired, and can guide seamless conformal parameterization and T-junction-free quadrangulation.
This paper presents PartCAD, a novel framework for reconstructing CAD modeling sequences directly from point clouds by projection-guided, part-aware geometry reasoning. It consists of (1) an autoregressive approach that decomposes point clouds into part-aware latent representations, serving as interpretable anchors for CAD generation; (2) a projection guidance module that provides explicit cues about underlying design intent via triplane projections; and (3) a non-autoregressive decoder to generate sketch-extrusion parameters in a single forward pass, enabling efficient and structurally coherent CAD instruction synthesis. By bridging geometric signals and semantic understanding, PartCAD tackles the challenge of reconstructing editable CAD models—capturing underlying design processes—from 3D point clouds. Extensive experiments show that PartCAD significantly outperforms existing methods for CAD instruction generation in both accuracy and robustness. The work sheds light on part-driven reconstruction of interpretable CAD models, opening new avenues in reverse engineering and CAD automation.
The 3D-printed mesh screen with dense penetrating pores is a new structure for massive manufacturing of molded pulp package products. However, some of the pores may be clogged by the printing material powder during the printing process. Such defects negatively affect the quality of the pulp packages produced using the mesh screen mold. To pinpoint the defects, we design a model-based robotic visual sensing system, called RoboCam, which uses a robotic arm to carry a high-resolution camera for full inspection of a mold consisting of joined mesh screens. To inspect the entire mold, RoboCam plans the camera poses to capture multiple images of the mold and render synthesized images as references for identifying the clogged pores. In particular, we propose novel designs to rectify the inherent run-time pose errors of the robotic system for ensuring the reference quality and to accelerate the reference rendering for reducing inspection latency. Extensive evaluation shows that RoboCam’s design outperforms various baselines, including three existing computer vision and convolution neural network-based inspection systems. RoboCam achieves a recall rate of 94.95% within 528 seconds latency for inspecting an entire mold with 13,000 designed pores.
The article presents a new method for constructing self-supporting surfaces using arch beams that are designed to convert their thrust into supporting force, thereby eliminating shear stress and bending moments. Our method allows for the placement of the arch beams on the boundary or within a surface and partitions the surface into multiple self-supporting parts. The use of arch beams enhances stability and durability, adds aesthetic appeal, and allows for greater flexibility in the design process. We develop an iterative algorithm for designing self-supporting surfaces with arch beams that enables the user to control the shape of the beams and surface through intuitive parameters and specify the desired location of the arch beams. We verify the physical stability of the structure using finite element analysis. Experimental results show that our method can produce visually pleasing self-supporting surfaces that satisfy the equilibrium equation with high accuracy.
This paper introduces ZeroMatch, a novel zero-shot RGB-D point cloud registration framework, aimed at achieving robust 3D matching on unseen data without any task-specific training. Our core idea is to utilize the powerful zero-shot image representation of Stable Diffusion, achieved through extensive pre-training on large-scale data, to enhance point-cloud geometric descriptors for robust matching. Specifically, we combine the handcrafted geometric descriptor FPFH with Stable-Diffusion features to create point descriptors that are both locally and contextually aware, enabling reliable RGB-D registration with zero-shot capability. This approach is based on our observation that Stable-Diffusion features effectively encode discriminative global contextual cues, naturally alleviating the feature ambiguity that FPFH often encounters in scenes with repetitive patterns or low overlap. To further enhance cross-view consistency of Stable-Diffusion features for improved matching, we propose a coupled-image input mode that concatenates the source and target images into a single input, replacing the original single-image mode. This design achieves both inter-image and prompt-to-image consistency attentions, facilitating robust cross-view feature interaction and alignment. Finally, we leverage feature nearest neighbors to construct putative correspondences for hypothesize-and-verify transformation estimation. Extensive experiments on 3DMatch, ScanNet, and ScanLoNet verify the excellent zero-shot matching ability of our method. [Code]
Recent advances have equipped 3D Gaussian Splatting with texture parameterizations to capture spatially varying attributes, improving the performance of both appearance modeling and downstream tasks. However, the added texture parameters introduce significant memory efficiency challenges. Rather than proposing new texture formulations, we take a step back to examine the characteristics of existing textured Gaussian methods and identify two key limitations in common: (1) Textures are typically defined in canonical space, leading to inefficient sampling that wastes textures' capacity on low-contribution regions; and (2) texture parameterization is uniformly assigned across all Gaussians, regardless of their visual complexity, resulting in over-parameterization. In this work, we address these issues through two simple yet effective strategies: adaptive sampling based on the Gaussian density distribution and error-driven anisotropic parameterization that allocates texture resources according to rendering error. Our proposed ASAP Textured Gaussians, short for Adaptive Sampling and Anisotropic Parameterization, significantly improve the quality efficiency tradeoff, achieving high-fidelity rendering with far fewer texture parameters.