Online monocular 3D reconstruction enables dense scene recovery from streaming video but remains fundamentally limited by the stability-adaptation dilemma: the reconstruction model must rapidly incorporate novel viewpoints while preserving previously accumulated scene structure. Existing streaming approaches rely on uniform or attention-based update mechanisms that often fail to account for abrupt viewpoint transitions, leading to trajectory drift and geometric inconsistencies over long sequences. We introduce PAS3R, a pose-adaptive streaming reconstruction framework that dynamically modulates state updates according to camera motion and scene structure. Our key insight is that frames contributing significant geometric novelty should exert stronger influence on the reconstruction state, while frames with minor viewpoint variation should prioritize preserving historical context. PAS3R operationalizes this principle through a motion-aware update mechanism that jointly leverages inter-frame pose variation and image frequency cues to estimate frame importance. To further stabilize long-horizon reconstruction, we introduce trajectory-consistent training objectives that incorporate relative pose constraints and acceleration regularization. A lightweight online stabilization module further suppresses high-frequency trajectory jitter and geometric artifacts without increasing memory consumption. Extensive experiments across multiple benchmarks demonstrate that PAS3R significantly improves trajectory accuracy, depth estimation, and point cloud reconstruction quality in long video sequences while maintaining competitive performance on shorter sequences.
Document dewarping is crucial for many applications. However, existing learning-based methods rely heavily on supervised regression with annotated data without fully leveraging the inherent geometric properties of physical documents. Our key insight is that a well-dewarped document is defined by its axis-aligned feature lines. This property aligns with the inherent axis-aligned nature of the discrete grid geometry in planar documents. Harnessing this property, we introduce three synergistic contributions: for the training phase, we propose an axis-aligned geometric constraint to enhance document dewarping; for the inference phase, we propose an axis alignment preprocessing strategy to reduce the dewarping difficulty; and for the evaluation phase, we introduce a new metric, Axis-Aligned Distortion (AAD), that not only incorporates geometric meaning and aligns with human visual perception but also demonstrates greater robustness. As a result, our method achieves state-of-the-art performance on multiple existing benchmarks, improving the AAD metric by 18.2% to 34.5%.
This paper introduces an interactive design method for developable surfaces, centered on a data-driven approach to optimize surface patches for developability. Surface patches are the fundamental components of an entire surface, typically represented by triangular meshes. We propose a novel learning-based method that effectively transforms patches with arbitrary boundaries into their closest developable surfaces. Based on this method, our tools enable real-time, drag-and-drop design of developable surfaces and support piecewise developable approximation through interactive inputs. Experimental results demonstrate that this method provides a fast computational foundation for the interactive design of developable surfaces, enhancing design flexibility while exhibiting excellent robustness and generalization. The piecewise developable approximation of the model, guided by human-computer collaborative segmentation, achieved higher overall approximation accuracy, fewer patches, and lifelike papercraft outcomes. This offers greater flexibility to meet the application requirements of complex real-world scenarios and provides a new paradigm for integrating deep learning with interactive geometry design.
Recent advances in diffusion-based generative models have enabled real-time audio-driven avatar generation and unified audio-visual synthesis, providing a promising foundation for interactive avatar systems. However, extending these models to real-time interactive streaming remains challenging, as the generation horizon is unknown in advance and cross-modal identity consistency gradually degrades during long-term generation. To address these challenges, we propose OmniMate, a unified framework for open-ended real-time interactive audio-visual avatar generation. OmniMate jointly synthesizes visual content, speech, and audio effects in real time, enabling natural and immersive multi-turn interactions. To achieve adaptive response progression, we introduce a Generation Progress Controller (GPC) that explicitly models the generation progress of each streaming chunk, allowing the model to complete responses according to the desired progress and achieve seamless transitions between execution and listening states. To preserve long-term cross-modal identity consistency, we propose a Multi-Reference Conditioning Module (MRCM), which leverages multiple reference images and a reference speech segment to provide persistent visual and speaker identity cues throughout long-duration streaming interactions. Extensive experiments on an interaction-oriented adaptation of VerseBench demonstrate that OmniMate achieves high-quality, low-latency streaming generation while maintaining strong long-term audio-visual consistency. The results further show that OmniMate supports realistic, coherent, and responsive interactive avatar experiences over extended multi-turn conversations.
Large Language Models (LLMs) demonstrate impressive capabilities, yet their outputs often suffer from misalignment with human preferences due to the inadequacy of weak supervision and a lack of fine-grained control. Training-time alignment methods like Reinforcement Learning from Human Feedback (RLHF) face prohibitive costs in expert supervision and inherent scalability limitations, offering limited dynamic control during inference. Consequently, there is an urgent need for scalable and adaptable alignment mechanisms. To address this, we propose W2S-AlignTree, a pioneering plug-and-play inference-time alignment framework that synergistically combines Monte Carlo Tree Search (MCTS) with the Weak-to-Strong Generalization paradigm for the first time. W2S-AlignTree formulates LLM alignment as an optimal heuristic search problem within a generative search tree. By leveraging weak model's real-time, step-level signals as alignment proxies and introducing an Entropy-Aware exploration mechanism, W2S-AlignTree enables fine-grained guidance during strong model's generation without modifying its parameters. The approach dynamically balances exploration and exploitation in high-dimensional generation search trees. Experiments across controlled sentiment generation, summarization, and instruction-following show that W2S-AlignTree consistently outperforms strong baselines. Notably, W2S-AlignTree raises the performance of Llama3-8B from 1.89 to 2.19, a relative improvement of 15.9 on the summarization task.
Recent diffusion-based models have enabled realistic audio-driven avatar generation in real-time streaming. However, existing approaches struggle to maintain visual temporal consistency and fail to explicitly perceive user intent in complex interactive streaming scenarios. To address these challenges, we propose InteractiveAvatar, a real-time infinite-streaming video generation framework that supports visually consistent avatar video generation and intent-aware interactions. With autoregressive distillation, InteractiveAvatar achieves real-time str-eaming generation of human avatars over arbitrarily long durations. For visual consistency, we introduce a Long-Short Visual Memory (LSVM) mechanism that flexibly compresses historical visual information into compact tokens, preserving both short-range coherence and long-term consistency. To generate avatars with speeches and actions aligned with user intent, we propose a Reasoning-Reaction Module (RRM), which incorporates a State-Cycling strategy and a Cache-Switching mechanism. Extensive experimental results over diverse scenarios demonstrate that our method achieves state-of-the-art visual consistency in long-duration generation, while enabling complex user-avatar interaction in real time.
Document rectification in real-world scenarios poses significant challenges due to extreme variations in camera perspectives and physical distortions. Driven by the insight that complex transformations can be decomposed and resolved progressively, we introduce a novel multi-stage framework that progressively reverses distinct distortion types in a coarse-to-fine manner. Specifically, our framework first performs a global affine transformation to correct perspective distortions arising from the camera's viewpoint, then rectifies geometric deformations resulting from physical paper curling and folding, and finally employs a content-aware iterative process to eliminate fine-grained content distortions. To address limitations in existing evaluation protocols, we also propose two enhanced metrics: layout-aligned OCR metrics (AED/ACER) for a stable assessment that decouples geometric rectification quality from the layout analysis errors of OCR engines, and masked AD/AAD (AD-M/AAD-M) tailored for accurately evaluating geometric distortions in documents with incomplete boundaries. Extensive experiments show that our method establishes new state-of-the-art performance on multiple challenging benchmarks, yielding a substantial reduction of 14.1%–34.7% in the AAD metric and demonstrating superior efficacy in real-world applications. The code will be publicly available at https://github.com/chaoyunwang/ArbDR.
Many problems in Euclidean geometry, arising in computational design and fabrication, amount to a system of constraints, which is challenging to solve. We suggest a new general approach to the solution, which is to start with analogous problems in isotropic geometry. Isotropic geometry can be viewed as a structure-preserving simplification of Euclidean geometry. The solutions found in the isotropic case give insight and can initialize optimization algorithms to solve the original Euclidean problems. We illustrate this general approach with three examples: quad-mesh mechanisms, composite asymptotic-geodesic gridshells, and asymptotic gridshells with constant node angle.
Topological disclinations, crystallographic defects that break rotation lattice symmetry, have attracted great interest and exhibited wide applications in cavities, waveguides, and lasers. However, topological disclinations have thus far been predominantly restricted to two-dimensional (2D) systems owing to the substantial challenges in constructing such defects in three-dimensional (3D) systems and characterizing their topological features. Here we report the theoretical proposal and experimental demonstration of a 3D topological disclination that exhibits fractional (1/2) charge and zero-dimensional (0D) topological bound states, realized by cutting-and-gluing a 3D acoustic topological crystalline insulator. Using acoustic pump-probe measurements, we directly observe 0D topological disclination states at the disclination core, consistent with the tight-binding model and full-wave simulation results. Our results extend the research frontier of topological disclinations and open a new paradigm for exploring the interplay between momentum-space band topology and the real-space defect topology in 3D and higher dimensions.
Weingarten surfaces are characterized by a functional relation between their principal curvatures. Such a specialty makes them suitable for building surface paneling in architectural applications, as the curvature relation implies approximate local congruence on the surface thus the molds for paneling can be largely reused. In this work, we aim at a novel task of Weingarten surface approximation. Given a surface mesh with arbitrary topology, we optimize its shape to make it as Weingarten as possible. We devise a curvature-based optimization approach based on the fact that the 2D principal curvature plots of a Weingarten surface comprise a group of 1D curves that encode the curvature relations. Our approach alternatively performs two steps. The first step transforms the principal curvature plots from a 2D region to 1D curves in order to explore the curvature relations. The second step deforms the shape such that its curvatures conform to the corresponding transformed curvature plots. We demonstrate the effectiveness of our work on a variety of shapes with different topologies. Hopefully our work would bring inspiration on the study of general Weingarten surfaces with arbitrary topology and curvature relation.
Vision Transformers (ViTs) have achieved state-of-the-art performance on various computer vision tasks. However these models are memory-consuming and computation-intensive, making their deployment and efficient inference on edge devices challenging. Model quantization is a promising approach to reduce model complexity. Prior works have explored tailored quantization algorithms for ViTs but unfortunately retained floating-point (FP) scaling factors, which not only yield non-negligible re-quantization overhead, but also hinder the quantized models to perform efficient integer-only inference. In this paper, we propose H-ViT, a dedicated post-training quantization scheme (e.g., symmetric uniform quantization and layer-wise quantization for both weights and part of activations) to effectively quantize ViTs with fewer Power-of-Two (PoT) scaling factors, thus minimizing the re-quantization overhead and memory consumption. In addition, observing serious inter-channel variation in LayerNorm inputs and outputs, we propose Power-of-Two quantization (PTQ), a systematic method to reducing the performance degradation without hyper-parameters. Extensive experiments are conducted on multiple vision tasks with different model variants, proving that H-ViT offers comparable(or even slightly higher) INT8 quantization performance with PoT scaling factors when compared to the counterpart with floating-point scaling factors. For instance, we reach 78.43 top-1 accuracy with DeiT-S on ImageNet, 51.6 box AP and 44.8 mask AP with Cascade Mask R-CNN (Swin-B) on COCO.
Learned image compression (LIC) has recently benefited from Transformer- and state space models (SSM)- based backbones for modeling long-range dependencies. However, the former typically incurs quadratic complexity, whereas the latter often disrupts neighborhood continuity by flattening 2D features into 1D sequences. To address these issues, we propose a compact Hybrid Convolution and Frequency State Space Network (HCFSSNet) for LIC. HCFSSNet combines convolutional layers for local detail modeling with a Vision Frequency State Space (VFSS) block for complementary long-range contextual aggregation. Specifically, the VFSS block consists of a Vision Omni-directional Neighborhood State Space (VONSS) module, which scans features along horizontal, vertical, and diagonal directions to better preserve 2D neighborhood relations, and an Adaptive Frequency Modulation Module (AFMM), which performs discrete cosine transform-based adaptive reweighting of frequency components. In addition, we introduce a Frequency Swin Transformer Attention Module (FSTAM) in the hyperprior path to enhance frequency-aware side information modeling. Experiments on the benchmark datasets show that the proposed HCFSSNet achieves a competitive rate-distortion performance against recent LIC codecs. The source code and models will be made publicly available.
In this paper, we propose SparseSCIGaussian, a novel method for achieving high-quality novel view synthesis under sparse input conditions. Previous methods often rely heavily on depth or neural priors, which can lead to generalization challenges and significant quality degradation on complex datasets. These limitations arise primarily from the insufficient scene information available in sparse regular images. To overcome these issues, our approach utilizes images captured through Snapshot Compressive Imaging (SCI) as input. SCI-captured images inherently encode richer scene information compared to regular images, thereby substantially improving the quality of novel view synthesis under sparse input conditions. Moreover, SCI images can be conveniently captured using a software-implemented encoder, making them as accessible as traditional images. Experimental results demonstrate that our method improves 2.65 dB (13.04 %) in PSNR compared to previous methods, and further exhibits the inherent advantages of using SCI images for sparse input novel view synthesis.
This paper provides computational tools for the modeling and design of quad mesh mechanisms, which are meshes allowing continuous flexions under the assumption of rigid faces and hinges in the edges. We combine methods and results from different areas, namely differential geometry of surfaces, rigidity and flexibility of bar and joint frameworks, algebraic geometry, and optimization. The basic idea to achieve a time-continuous flexion is timediscretization justified by an algebraic degree argument. We are able to prove computationally feasible bounds on the number of required time instances we need to incorporate in our optimization. For optimization to succeed, an informed initialization is crucial. We present two computational pipelines to achieve that: one based on remeshing isometric surface pairs, another one based on iterative refinement. A third manner of initialization proved very effective: We interactively design meshes which are close to a narrow known class of flexible meshes, but not contained in it. Having enjoyed sufficiently many degrees of freedom during design, we afterwards optimize towards flexibility.
The modern theory of quantized polarization has recently extended from 1D dipole moment to multipole moment, leading to the development from conventional topological insulators (TIs) to higher-order TIs, i.e., from the bulk polarization as primary topological index, to the fractional corner charge as secondary topological index. The authors here extend this development by theoretically discovering a higher-order end TI (HOETI) in a real projective lattice and experimentally verifying the prediction using topolectric circuits. A HOETI realizes a dipole-symmetry-protected phase in a higher-dimensional space (conventionally in one dimension), which manifests as 0D topologically protected end states and a fractional end charge. The discovered bulk-end correspondence reveals that the fractional end charge, which is proportional to the bulk topological invariant, can serve as a generic bulk probe of higher-order topology. The authors identify the HOETI experimentally by the presence of localized end states and a fractional end charge. The results demonstrate the existence of fractional charges in non-Euclidean manifolds and open new avenues for understanding the interplay between topological obstructions in real and momentum space.
In the context of surface representations, we find a natural structural similarity between grid surface and image data. Motivated by this inspiration, we propose a novel approach: encoding grid surfaces as geometric images and using image processing methods to address surface optimization-related problems. As a result, we have created the first dataset for grid surface optimization and devised a learning-based grid surface optimization network specifically tailored to geometric images, addressing the surface optimization problem through a data-driven learning of geometric constraints paradigm. We conduct extensive experiments on developable surface optimization, surface flattening, and surface denoising tasks using the designed network and datasets. The results demonstrate that our proposed method not only addresses the surface optimization problem better than traditional numerical optimization methods, especially for complex surfaces, but also boosts the optimization speed by multiple orders of magnitude. This pioneering study successfully applies deep learning methods to the field of surface optimization and provides a new solution paradigm for similar tasks, which will provide inspiration and guidance for future developments in the field of discrete surface optimization. The code and dataset are available at https://github.com/chaoyunwang/GSO-Net.
D-Form is a special piece-wise developable surface formed by aligning the boundaries of two planar domains. It has been widely used in different design scenarios. In this paper, we study how to computationally and intuitively model D-Forms. We present an optimisation-based framework that can efficiently generate D-Form shapes. Our framework can model D-Forms with two approaches based on two different user inputs, including the forward modelling from two given planar domains and, more importantly, the inverse modelling from a given space curve where the planar domains are no longer needed. Our optimisation is devised based on two critical characteristics of D-Forms. Firstly, the constituent developable surfaces of a D-Form are isometrically deformed from planar domains. Secondly, there is a close relationship between a D-Form and the convex hull of its seam. Through extensive evaluation, we demonstrate that our approach can model plausible D-Forms efficiently from various inputs with different geometric properties.
We address the computational design of architectural structures which are based on a grid of intersecting beams that are aligned with the parameter lines of a quad mesh. While previous work mainly put a planarity constraint onto the faces of the mesh, we focus on the planarity of long-range supporting beams which follow selected polylines in the underlying mesh. In addition to that, we impose further constraints including planarity of faces, right node angles and static equilibrium, and discuss in which way these may be combined. Some of the studied meshes are discrete counterparts of certain well-known surfaces in classical geometry, whose knowledge is helpful for initializing the proposed optimization algorithms.
In this article, we investigate geometric properties and modeling capabilities of quad meshes with planar faces whose mesh polylines enjoy the additional property of being contained in a single plane. This planarity is a major benefit in architectural design and building construction: If a structural element is contained in a plane, it can be manufactured on the ground without scaffolding and put into place as a whole. Further, the plane it is contained in serves as part of a so-called support structure. We discuss design of meshes under the requirement that one half of mesh polylines are planar (“P meshes”), and we also investigate the geometry and design of meshes where all polylines enjoy this property (“PP meshes”). We work in the space of planes and with appropriate transformations of that space. We also incorporate further properties relevant for architectural design, such as near-rectangular panels and repetitive nodes. We provide geometric insights, give explicit constructions, and show an approach to geometric modeling of both P meshes and PP meshes, in particular, the case of nearly rectangular panels.
High-fidelity and efficient audio-driven talking head generation has been a key research topic in computer graphics and computer vision. In this work, we study vector image based audio-driven talking head generation. Compared with directly animating the raster image that most widely used in existing works, vector image enjoys its excellent scalability being used for many applications. There are two main challenges for vector image based talking head generation: the high-quality vector image reconstruction w.r.t. the source portrait image and the vivid animation w.r.t. the audio signal. To address these, we propose a novel scalable vector graphic reconstruction and animation method, dubbed VectorTalker. Specifically, for the highfidelity reconstruction, VectorTalker hierarchically reconstructs the vector image in a coarse-to-fine manner. For the vivid audio-driven facial animation, we propose to use facial landmarks as intermediate motion representation and propose an efficient landmark-driven vector image deformation module. Our approach can handle various styles of portrait images within a unified framework, including Japanese manga, cartoon, and photorealistic images. We conduct extensive quantitative and qualitative evaluations and the experimental results demonstrate the superiority of VectorTalker in both vector graphic reconstruction and audio-driven animation.