Human pose and shape estimation (HPS) has attracted increasing attention in recent years. While most existing studies focus on HPS from 2D images or videos with inherent depth ambiguity, there are surging need to investigate HPS from 3D point clouds as depth sensors have been frequently employed in commercial devices. However, real-world sensory 3D points are usually noisy and incomplete, and also human bodies could have different poses of high diversity. To tackle these challenges, we propose a principled framework, PointHPS, for accurate 3D HPS from point clouds captured in real-world settings, which iteratively refines point features through a cascaded architecture. Specifically, each stage of PointHPS performs a series of downsampling and upsampling operations to extract and collate both local and global cues, which are further enhanced by two novel modules: 1) Cross-stage Feature Fusion (CFF) for multi-scale feature propagation that allows information to flow effectively through the stages, and 2) Intermediate Feature Enhancement (IFE) for body-aware feature aggregation that improves feature quality after each stage. To facilitate a comprehensive study under various scenarios, we conduct our experiments on two large-scale benchmarks, comprising i) a dataset that features diverse subjects and actions captured by real commercial sensors in a laboratory environment, and ii) controlled synthetic data generated with realistic considerations such as clothed humans in crowded outdoor scenes. Extensive experiments demonstrate that PointHPS, with its powerful point feature extraction and processing scheme, outperforms State-of-the-Art methods by significant margins across the board. Homepage: https://caizhongang.github.io/projects/PointHPS/.
Open-World Object Detection (OWOD) requires models to detect known objects in training set while identifying unknown ones. Recently, OWOD has been extended to open-vocabulary detectors, where object semantics are guided by textual vocabularies and unknown objects are distinguished using an object wildcard embedding. However, this object wildcard embedding is learned from open-source datasets with limited category coverage and therefore remains far from capable of recognizing all objects. To avoid the dilemma of enumerating object categories, this paper proposes leveraging geometric depth cues to distinguish objects, rather than relying solely on class-level semantics. Specifically, we first introduce a depth-based objectness assessment algorithm that provides a more reasonable quantification of objectness scores using geometric cues. Then, we propose a depth-guided wildcard embedding learning strategy, which learns objectness knowledge through pseudo-labels during both pre-training and OWOD training stages. Extensive experiments demonstrate the superiority of the proposed Depth-Guided Open-World Object Detection (DG-OWOD) method, yielding 5
Reasoning segmentation is an emerging vision-language task that requires reasoning over intricate text queries to precisely segment objects. However, existing methods typically suffer from overthinking, generating verbose reasoning chains that interfere with object localization in multimodal large language models (MLLMs). To address this issue, we propose DR^2Seg, a self-rewarding framework that improves both reasoning efficiency and segmentation accuracy without requiring extra thinking supervision. DR^2Seg employs a two-stage rollout strategy that decomposes reasoning segmentation into multimodal reasoning and referring segmentation. In the first stage, the model generates a self-contained description that explicitly specifies the target object. In the second stage, this description replaces the original complex query to verify its self-containment. Based on this design, two self-rewards are introduced to mitigate overthinking and the associated attention dispersion. Extensive experiments conducted on 3B and 7B variants of Qwen2.5-VL, as well as on both SAM2 and SAM3, demonstrate that DR^2Seg consistently improves reasoning efficiency and overall segmentation accuracy.
Current object pose estimation research remains predominantly model-centric, focusing on architectural innovations and post-processing refinements. This paper introduces a data-centric optimization by proposing a novel, physically grounded rotation representation through principal axes alignment. Our method aligns the object's coordinate system with its inherent geometric axes, derived from inertial properties, yielding three key advantages: Inherent Stability-leveraging the energy-minimizing property of principal axes provides a robust representation that is less sensitive to noise and occlusions; Symmetry-Aware Canonicalization-explicitly resolving rotational ambiguities for symmetric objects at the data level, which fundamentally eliminates label confusion during network training; and Framework Agnosticism-the optimization is applied purely at the dataset level, ensuring plug-and-play compatibility with existing networks without any architectural modification. We validate the framework across diverse category-level and instance-level models. Extensive experiments demonstrate consistent and significant accuracy improvements, while preserving the integrity of the baseline network. This work establishes a new, geometry-driven direction for enhancing pose estimation, circumventing the need for complex network redesign.
Large pre-trained vision-language models (VLMs) like CLIP have shown great potential for solving the unsupervised domain adaptation (UDA) problem. Existing prompt learning for UDA based on the unsupervised-trained VLMs requires distribution alignment between source and target domains in the common space for both vision and language branches. However, it is difficult for rough cross-domain alignment to maintain the discriminative semantic structure of both domains. Besides, the coarse features with non-informative noises due to ignoring the pseudo-label noises may cause failures to concentrate on precise semantics alignment. In this work, we propose a Prompt-Based Invertible Mapping Alignment (PIMA) method to incorporate discriminative domain knowledge into prompt learning, which is featured with refined cross-domain alignment in two separate space with a well-kept structure. Specifically, we design an invertible neural network-based homeomorphism mapping, and then achieve distribution alignment through such invertible mapping for connecting source and target visual feature space, which can preserve the data semantic structure. For better semantic alignment in vision-language space, we develop cross-modal implicit contrastive learning module to regularize non-informative features, which aims to find the low-rankness of implicit representation space. We conducted extensive experiments on three benchmark datasets to prove the advantages of our proposed PIMA over state-of-the-art methods.
Low-level 3D representations, such as point clouds, meshes, NeRFs and 3D Gaussians, are commonly used for modeling 3D objects and scenes. However, cognitive studies indicate that human perception operates at higher levels and interprets 3D environments by decomposing them into meaningful structural parts, rather than low-level elements like points or voxels. Structured geometric decomposition enhances scene interpretability and facilitates downstream tasks requiring component-level manipulation. In this work, we introduce PartGS, a self-supervised part-aware reconstruction framework that integrates 2D Gaussians and superquadrics to parse objects and scenes into an interpretable decomposition, leveraging multi-view image inputs to uncover 3D structural information. Our method jointly optimizes superquadric meshes and Gaussians by coupling their parameters within a hybrid representation. On one hand, superquadrics enable the representation of a wide range of shape primitives, facilitating flexible and meaningful decompositions. On the other hand, 2D Gaussians capture detailed texture and geometric details, ensuring high-fidelity appearance and geometry reconstruction. Operating in a self-supervised manner, our approach demonstrates superior performance compared to state-of-the-art methods across extensive experiments on the DTU, ShapeNet, and real-world datasets.
Multi-view clustering aims to integrate complementary information from multiple views to improve clustering performance. However, existing ensemble-based methods suffer from information loss due to their reliance on single-granularity labels, limiting the discriminative capability of learned representations. Meanwhile, representation and graph fusion-based approaches face challenges such as explicit view alignment and manual weight tuning, making them less effective for heterogeneous views with varying data distributions. To address these limitations, we propose a novel multi-view clustering framework via Multi-granularity Ensemble (MGE), fully using the multi-granularity information across diverse views for accurate and consistent clustering. Specifically, MGE first modifies the hierarchical clustering and then leverages it on each view (including the fused view) to achieve multi-granularity labels. Moreover, the cross-view and cross-granularity fusion strategy is designed to learn a robust co-association similarity matrix, which effectively preserves the fine-grained and coarse-grained structures of multi-view data and facilitates subsequent clustering. Therefore, MGE can provide a comprehensive representation of local and global patterns within data, eliminating the requirement for view alignment and weight tuning. Experiments demonstrate that MGE consistently outperforms state-of-the-art methods across multiple datasets, validating its effectiveness and superiority in handling heterogeneous views.
Radiance fields, including NeRFs and 3D Gaussians, demonstrate great potential in high-fidelity rendering and scene reconstruction, while they require a substantial number of posed images as input. COLMAP is frequently employed for preprocessing to estimate poses. However, COLMAP necessitates a large number of feature matches to operate effectively, and struggles with scenes characterized by sparse features, large baselines, or few-view images. We aim to tackle few-view NeRF reconstruction using only 3 to 6 unposed scene images, freeing from COLMAP initializations. Inspired by the idea of calibration boards in traditional pose calibration, we propose a novel approach of utilizing everyday objects, commonly found in both images and real life, as “pose probes”. By initializing the probe object as a cube shape, we apply a dual-branch volume rendering optimization (object NeRF and scene NeRF) to constrain the pose optimization and jointly refine the geometry. PnP matching is used to initialize poses between images incrementally, where only a few feature matches are enough. PoseProbe achieves state-of-the-art performance in pose estimation and novel view synthesis across multiple datasets in experiments. We demonstrate its effectiveness, particularly in few-view and large-baseline scenes where COLMAP struggles. In ablations, using different objects in a scene yields comparable performance, showing that PoseProbe is robust to the choice of probe objects. Our project page is available at: https://zhirui-gao.github.io/PoseProbe.github.io/.
The scarcity of labeled data poses a significant challenge for deep learning-based medical image segmentation. To address this, this study introduces the novel Foundation Model-based Few-Shot Segmentation (FM-FSS) paradigm. FM-FSS capitalizes on the knowledge distilled from pre-trained foundation models, such as the Segment Anything Model, to enhance segmentation performance in few-shot scenarios. The paradigm designs a feature coupling module that synergizes SAM’s powerful feature extraction capabilities with nnU-Net’s self-configuration strategy, enabling accurate segmentation with minimal labeled data and optional manual prompt inputs. Extensive experiments on a publicly available cardiac CT dataset demonstrate that FM-FSS outperforms state-of-the-art segmentation models. With only 20 labeled images, our method achieves an average Dice score of 94.33% and an ASD of 1.10 mm. Moreover, FM-FSS maintains its label-efficient performance in a one-shot setup, reducing the annotation requirements by at least fourfold. The code and pre-trained models will be released upon acceptance.
3D object detection methods based on point cloud have made significant progress due to providing rich depth information. However, the disability to obtain the complete shape by point cloud because of occlusion and signal loss leads to unsatisfactory detection performance. In the paper, a new multi-level fusion network based on cross attention (MFNCA) for 3D object detection is proposed to achieve impressive detection accuracy, which extracts not only voxel geometry features at multiple layer-level but also shape occupancy features including the missing parts of objects. Specifically, we introduce a sparse skip connection module to aggregate features from different levels and design a channel-wise pooling layer to enhance the global perspective of the model. Furthermore, RoI (Region of Interest) cross attention module is proposed to generate more accurate 3D bounding boxes by fusing multiple critical features. Extensive experiments on the challenging KITTI 3D dataset show that our method achieves promising performance compared with state-of-the-art methods.
In this work, we present Digital Life Project, a framework utilizing language as the universal medium to build autonomous 3D characters, who are capable of engaging in social interactions and expressing with articulated body motions, thereby simulating life in a digital environment. Our framework comprises two primary components: 1) SocioMind: a meticulously crafted digital brain that models personalities with systematic few-shot exemplars, incorporates a reflection process based on psychology principles, and emulates autonomy by initiating dialogue topics; 2) MoMat-MoGen: a text-driven motion synthesis paradigm for controlling the character's digital body. It integrates motion matching, a proven industry technique to ensure motion quality, with cutting-edge advancements in motion generation for diversity. Extensive experiments demonstrate that each module achieves state-of-the-art performance in its respective domain. Collectively, they enable virtual characters to initiate and sustain dialogues autonomously, while evolving their socio-psychological states. Concurrently, these characters can perform contextually relevant bodily movements. Additionally, an extension of DLP enables a virtual character to recognize and appropriately respond to human players' actions.
The computer-aided diagnosis system for esophageal cancer (EC) holds vital significance in EC diagnosis and treatment making, with a primary focus on accurate segmentation of EC-related organs and classification of EC's T-stage. Above two tasks are closely related and crucial in assisting surgeon segment and diagnose cancer early. Note that this paradigm is still at its infancy and limited by closely related open issues: (1) how to link the complementary relationship between these two tasks and improve the originally poor performance? and (2) how to determine whether the tumor has invaded the surrounding muscle layers from CT images? Aiming at these issues, this study develops nn-TransEC, a 3D transfer learning framework that builds upon nnU-Net and synergizes segmentation and classification. nn-TransEC focuses on prompting fine-grained classification of EC's T-stage with the aid of prior segmentation, which is implemented in two parts: (1) A nnUNet-configured multi-task learning network (nn-MTNet) is designed for complementary segmentation of EC-related organs and classification of EC's T-stages with cross-task attention gates and transfer learning. (2) A knowledge-embedded ROI tokenization method (KRT) is defined to mimic the diagnostic workflow of doctors for classifying EC's T-stage. KRT is implemented by cropping the most concerned regions from entire CT volume based on prior segmentation. Experiments have been conducted on a private dataset collected from 169 patients with confirmed EC through pathological diagnosis. Our proposed nn-TransEC is compared against the state-of-the-art counterparts (e.g., nnU-Net and nnFormer), and results demonstrate that: nn-TransEC excels in all compared methods in multi-organ segmentation and classification of EC's T-stages, with 3D Dice of EC and average AUC of T-stages reaching 0.844 and 0.941, respectively. In contrast, the state-of-the-art method nnFormer achieves 0.814 and 0.927, respectively. Meanwhile, nn-TransEC also outperforms state-of-the-art multi-task learning models in joint segmentation and classification, with Hausdorff Distance of EC and average precision of Tstages reaching 8.497 and 0.845, respectively. In contrast, the state-of-the-art method TransMT-Net achieves 12.206 and 0.730, respectively.
Image- and video-based 3D human recovery ( i.e. , pose and shape estimation) have achieved substantial progress. However, due to the prohibitive cost of motion capture, existing datasets are often limited in scale and diversity. In this work, we obtain massive human sequences by playing the video game with automatically annotated 3D ground truths. Specifically, we contribute GTA-Human, a large-scale 3D human dataset generated with the GTA-V game engine, featuring a highly diverse set of subjects, actions, and scenarios. More importantly, we study the use of game-playing data and obtain five major insights. First , game-playing data is surprisingly effective. A simple frame-based baseline trained on GTA-Human outperforms more sophisticated methods by a large margin. For videobased methods, GTA-Human is even on par with the in-domain training set. Second , we discover that synthetic data provides critical complements to the real data that is typically collected indoor. We highlight that our investigation into domain gap provides explanations for our data mixture strategies that are simple yet useful, which offers new insights to the research community. Third , the scale of the dataset matters. The performance boost is closely related to the additional data available. A systematic study on multiple key factors (such as camera angle and body pose) reveals that the model performance is sensitive to data density. Fourth , the effectiveness of GTA-Human is also attributed to the rich collection of strong supervision labels (SMPL parameters), which are otherwise expensive to acquire in real datasets. Fifth , the benefits of synthetic data extend to larger models such as deeper convolutional neural networks (CNNs) and Transformers, for which a significant impact is also observed. We hope our work could pave the way for scaling up 3D human recovery to the real world. Homepage: https://caizhongang.github.io/projects/GTA-Human/ .
Unsupervised image registration commonly adopts U-Net style networks to predict dense displacement fields in the full-resolution spatial domain. For high-resolution volumetric image data, this process is however resource-intensive and time-consuming. To tackle this problem, we propose the Fourier-Net, replacing the expansive path in a U-Net style network with a parameter-free model-driven decoder. Specifically, instead of our Fourier-Net learning to output a full-resolution displacement field in the spatial domain, we learn its low-dimensional representation in a band-limited Fourier domain. This representation is then decoded by our devised model-driven decoder (consisting of a zero padding layer and an inverse discrete Fourier transform layer) to the dense, full-resolution displacement field in the spatial domain. These changes allow our unsupervised Fourier-Net to contain fewer parameters and computational operations, resulting in faster inference speeds. Fourier-Net is then evaluated on two public 3D brain datasets against various state-of-the-art approaches. For example, when compared to a recent transformer-based method, named TransMorph, our Fourier-Net, which only uses 2.2% of its parameters and 6.66% of the multiply-add operations, achieves a 0.5% higher Dice score and an 11.48 times faster inference speed. Code is available at https://github.com/xi-jia/Fourier-Net.
RNA Polymerase II transcribes mRNA. The largest subunit of this complex contains a C-terminal domain (CTD) that is an intrinsically disordered protein with a repetitive amino acid heptad sequence making the domain very difficult to study. Kinases and Phosphatases regulate CTD through post-translational modifications. Small angle x-ray scattering data shows very little change in CTD compaction with phosphorylation, despite the repulsion between the negatively charged phosphate groups. Data also shows an increase in Pro6 isomerization when Ser5 is phosphorylated.
Estimating 6D object pose from a monocular RGB image remains challenging due to factors such as texture-less and occlusion. Although convolution neural network (CNN)-based methods have made remarkable progress, they are not efficient in capturing global dependencies and often suffer from information loss due to downsampling operations. To extract robust feature representation, we propose a Transformer-based 6D object pose estimation approach (Trans6D). Specifically, we first build two transformer-based strong baselines and compare their performance: pure Transformers following the ViT (Trans6D-pure) and hybrid Transformers integrating CNNs with Transformers (Trans6D-hybrid). Furthermore, two novel modules have been proposed to make the Trans6D-pure more accurate and robust: (i) a patch-aware feature fusion module. It decreases the number of tokens without information loss via shifted windows, cross-attention, and token pooling operations, which is used to predict dense 2D-3D correspondence maps; (ii) a pure Transformer-based pose refinement module (Trans6D+) which refines the estimated poses iteratively. Extensive experiments show that the proposed approach achieves state-of-the-art performances on two datasets.
Expressive human pose and shape estimation (EHPS) unifies body, hands, and face motion capture with numerous applications. Despite encouraging progress, current state-of-the-art methods still depend largely on a confined set of training datasets. In this work, we investigate scaling up EHPS towards the first generalist foundation model (dubbed SMPLer-X), with up to ViT-Huge as the backbone and training with up to 4.5M instances from diverse data sources. With big data and the large model, SMPLer-X exhibits strong performance across diverse test benchmarks and excellent transferability to even unseen environments. 1) For the data scaling, we perform a systematic investigation on 32 EHPS datasets, including a wide range of scenarios that a model trained on any single dataset cannot handle. More importantly, capitalizing on insights obtained from the extensive benchmarking process, we optimize our training scheme and select datasets that lead to a significant leap in EHPS capabilities. 2) For the model scaling, we take advantage of vision transformers to study the scaling law of model sizes in EHPS. Moreover, our finetuning strategy turn SMPLer-X into specialist models, allowing them to achieve further performance boosts. Notably, our foundation model SMPLer-X consistently delivers state-of-the-art results on seven benchmarks such as AGORA (107.2 mm NMVE), UBody (57.4 mm PVE), EgoBody (63.6 mm PVE), and EHF (62.3 mm PVE without finetuning). Homepage: https://caizhongang.github.io/projects/SMPLer-X/
RNA polymerase II (Pol II) transcribes protein-coding genes and coordinates co-transcriptional processes such as mRNA maturation and histone modification. The intrinsically disordered C-terminal domain (CTD) of the largest subunit of Pol II is composed of dozens of repeats with the consensus sequence YSPTSPS and serves as a flexible binding scaffold for co-transcriptional regulatory proteins. During transcription, the CTD undergoes constant post-translational modifications. These changing modifications constitute the CTD code, which specifies the position of Pol II on a gene and recruits specific regulatory proteins. Two major CTD phosphorylation marks, pSer5 and pSer2, are characteristic of the early and the late stage of transcription respectively. To characterize structural properties of CTDs with different phosphorylation patterns, we generated different CTD variants using enzymatic approaches and probed their local and global structures with carbon direct-detect NMR and small-angle X-ray scattering (SAXS). Particularly, carbon direct-detect NMR provides higher spectral resolution for intrinsically disordered proteins and allows direct visualization of prolines, which constitute over 20% of the CTD. Together, our structural characterization of the Pol II CTD in different phosphorylation states provides insights for how phosphorylation regulates the ensemble-function relationship of intrinsically disordered proteins.
In this paper, we propose a novel 3D graph convolution based pipeline for category-level 6D pose and size estimation from monocular RGB-D images. The proposed method leverages an efficient 3D data augmentation and a novel vector-based decoupled rotation representation. Specifically, we first design an orientation-aware autoencoder with 3D graph convolution for latent feature learning. The learned latent feature is insensitive to point shift and size thanks to the shift and scale-invariance properties of the 3D graph convolution. Then, to efficiently decode the rotation information from the latent feature, we design a novel flexible vector-based decomposable rotation representation that employs two decoders to complementarily access the rotation information. The proposed rotation representation has two major advantages: 1) decoupled characteristic that makes the rotation estimation easier; 2) flexible length and rotated angle of the vectors allow us to find a more suitable vector representation for specific pose estimation task. Finally, we propose a 3D deformation mechanism to increase the generalization ability of the pipeline. Extensive experiments show that the proposed pipeline achieves state-of-the-art performance on category-level tasks. Further, the experiments demonstrate that the proposed rotation representation is more suitable for the pose estimation tasks than other rotation representations.
In the task of Fine-Grained Image Recognition (FGIR), the overall difference between different types of images is slight, so locating the representative local region in the image is the key to improving the classification accuracy. This idea of FGIR has been widely used in previous work, and has achieved good results on the benchmark dataset. Recently, the proposal of the Vision Transformer (ViT) method, provides a new method for the field of computer vision. Compared with the previous work based on Convolutional Neural Network (CNN), it has achieved better performance. ViT performs well in general image recognition tasks. However, when applied to FGIR tasks, it only pays attention to the global information and does not pay enough attention to the local features with discrimination. In order to make the model pay more attention to differentiated local regions, we propose an attention-based local region merging method Group Attention Transformer (GA-Trans), which evaluates the importance of each patch by using the self-attention weight inside the Transformer, and then aggregates adjacent high weight attention blocks into groups, then randomly select groups for image crop and drop. Through the weight sharing encoder, the global and local regions of the image are classified after obtaining the features respectively, which is convenient to realize the end-to-end training. Comprehensive experiments show that GA-Trans can achieve state-of-the-art performance on multiple benchmark datasets.