
Automatically identification of an optimal, representative or relevant viewpoint for a given 3D object is an important task with several applications in digital marketing, visualization, 3D data management and shape retrieval. However, an objective function to optimize the viewpoint is difficult to define in purely geometric terms, given the semantic and aesthetic bias a user is likely to introduce when presented with a target object. Therefore, supervised approaches are a natural candidate to address the problem. This, however, implies the availability of large scale dataset describing natural point-of-view preferences for a number of object categories. In this work, we present an interactive online system designed to continuously harvest and extend such a dataset. The system involves tasking end-users with answering general qualitative questions about a displayed 3D object, while silently collecting point-of-view information. Contrary to related work, we transform the viewpoint selection problem into a latent objective, allowing for an unbiased observation of user behavior in performing various tasks that involve object visualization. We publicly release both our dataset of collected viewpoint statistics and the tool to harvest them, which is easily configurable by experimenters to collect more data in the future or integrate the approach into online 3D viewers.
We present PhyDeformer, a new deformation method for high-quality garment mesh registration. It operates in two phases: In the first phase, a garment grading is performed to achieve a coarse 3D alignment between the mesh template and the target mesh, accounting for proportional scaling and fit (e.g. length, size). Then, the graded mesh is refined to align with the fine-grained details of the 3D target through an optimization coupled with the Jacobian-based deformation framework. Both quantitative and qualitative evaluations on synthetic and real garments highlight the effectiveness of our method.
The emerging accessibility of 3D point cloud data has catalyzed the evolution of deep-learning methodologies for analysis and processing of 3D data. However, the efficacy of neural networks in this domain is often inhibited by the necessity for extensively labelled datasets. In this study, we investigate the application of self-distillation techniques based on Siamese networks, BYOL and SIMSIAM, to pre-train encoders designed for 3D point cloud processing. These pre-training regimes enable encoders to generate data representations without label reliance, potentially supporting network performance in downstream tasks. The efficacy of these learned representations was assessed using the established evaluation methodologies for pre-training: linear probing and finetuning. We also incorporate an analysis of self-supervised features in a retrieval scenario. Furthermore, the impact of these representations on subsequent applications was evaluated via transfer learning by employing pre-trained models as a foundation for standard test datasets.
The aim of this SHREC 2024 track is to compare different algorithms for retrieving non-rigid complementary shape pairs, applied in the context of 3D objects being more complex (e.g. with many folds and roughness) such as proteins. The dataset used for this benchmark is based on 52 selected protein-protein complexes for which an experimental structure is publicly available. One of the main difficulties of this challenge is the non-inclusion of the shapes derived from the ground truth conformations in the dataset. Different metrics were used to evaluate the retrieval performance (nearest-neighbor, first-tier, second-tier, and true positives) and to evaluate the quality of the predicted poses (TM-score, lDDT, ICS, IPS and DockQ - those metrics are classically used in the Critical Assessment of PRediction of Interactions challenges). Two teams took part in this challenge and were able to return the expected results. This paper discusses these results and prospects of retrieval methods based only on the protein shape information in the absence of atomic data, in a large context of protein-protein docking.
We present a novel geometric deep learning layer that leverages the varifold gradient (VariGrad) to compute feature vector representations of 3D geometric data. These feature vectors can be used in a variety of downstream learning tasks such as classification, registration, and shape reconstruction. Our model's use of parameterization independent varifold representations of geometric data allows our model to be both trained and tested on data independent of the given sampling or parameterization. We demonstrate the efficiency, generalizability, and robustness to resampling demonstrated by the proposed VariGrad layer.
The phase shift algorithm is an important 3D shape reconstruction method in industrial quality inspection and reverse engineering. To retrieve dense and accurate point clouds, the conventional phase shift methods require at least three fringe projection patterns, limiting its application to statics or semi-statics scenes only. In this paper, we propose a novel and low-cost single-shot phase shift 3D reconstruction framework using convolution neural networks (CNN) trained on 3D synthetic fractals. We first design and optimize a novel projection pattern that compresses the phase period orders and the ambiguous phase information into a single image. Then, we train two different CNNs to predict the ambiguous phase information and the period orders separately. The CNNs were trained on randomly generated 3D shapes whose geometric complexity is modeled by recursive shape generation algorithms which can create an unlimited amount of random 3D shapes on the fly. Initial results demonstrate that our method can produce high-quality point clouds from just a pair of 2D images, thus improving the temporal resolution of a phase-shift 3D scanner to the highest possible. As we also include different real-world lighting and textural conditions in the training data set, experiments also demonstrate that our CNN models which were trained on random synthetic fractals only can perform equally well in the real world.
3D morphable models (3DMMs) simultaneously reconstruct facial morphology, expression and pose from 2D images, and thus could be an invaluable tool for capturing and characterizing the face and facial behavior in early childhood. However, 3DMM fitting on infants is a largely unexplored problem. All publicly available 3DMMs are developed for adults, and it is unclear if and to what extent they can be used on videos of infants. In this paper, we compare five state-of-the-art 3DMM fitting methods on data from naturalistic infant-caregiver interactions. Results suggest that it is possible to produce consistent and subject-specific reconstructions of 3D shape identity from multiple frames, but not from a single frame. Qualitative evaluation highlights that facial regions with high texture variation, such as eyes, brows and mouth, are captured with higher accuracy compared to the rest of the face. Thus, even though a 3DMM developed for adults has significant limitations when reconstructing the morphology of the entire facial region of infants, applications that involve analysis of facial behavior can be feasible. Our encouraging results, combined with the unique ability of 3DMMs to disentangle two major sources of noise for expression analysis (i.e., identity bias and pose variations), motivate future research on using 3DMMs to measure the facial behavior of infants.
The generation of 3-dimensional geometric objects in the most efficient way is a thriving research topic with, for example, the development of geometric deep learning, extending classical machine learning concepts to non euclidean data such as graphs or meshes. In this short paper, we study the effect of a reparameterization on two popular mesh and point cloud neural networks in an auto-encoder mode: PointNet [QSMG16] and SpiralNet [BBP ∗ 19]. Finally, we tested a modified version of PointNet that takes orientation into account (through coordinates of the normals) as a first step towards the construction of a geometric deep learning model built with a more flexible metric regarding the parameterization. The experimental results on standardized face datasets show that SpiralNet is more robust to the reparametrization than PointNet in this specific context with the proposed reparameterization.
The Building Information Modelling (BIM) procedure introduces specifications and data exchange formats widely used by the construction industry to describe functional and geometric elements of building structures in the design, planning, cost estimation and construction phases of large civil engineering projects. In this paper we explain how to apply a modern, low-parameter, neural-network-based classification solution to the automatic geometric BIM element labeling, which is becoming an increasingly important task in software solutions for the construction industry. The network is designed so that it extracts features regarding general shape, scale and aspect ratio of each BIM element and be extremely fast during training and prediction. We evaluate our network architecture on a real BIM dataset and showcase accuracy that is difficult to achieve with a generic 3D shape classification network.
Proteins are essential to nearly all cellular mechanism, and often interact through their surface with other cell molecules, such as proteins and ligands. The evolution generates plenty of different proteins, with unique abilities, but also proteins with related functions hence surface, which is therefore of primary importance for their activity. In the present work, we assess the ability of five methods to retrieve similar protein surfaces, using either their shape only (3D meshes), or their shape and the electrostatic potential at their surface, an important surface property. Five different groups participated in this challenge using the shape only, and one group extended its pre-existing algorithm to handle the electrostatic potential. The results reveal both the ability of the methods to detect related proteins and their difficulties to distinguish between topologically related proteins. CCS Concepts • Applied computing → Computational biology; • General and reference → Evaluation;
Cryo-electron tomography (cryo-ET) is an imaging technique that allows three-dimensional visualization of macro-molecular assemblies under near-native conditions. Cryo-ET comes with a number of challenges, mainly low signal-to-noise and inability to obtain images from all angles. Computational methods are key to analyze cryo-electron tomograms. To promote innovation in computational methods, we generate a novel simulated dataset to benchmark different methods of localization and classification of biological macromolecules in tomograms. Our publicly available dataset contains ten tomographic reconstructions of simulated cell-like volumes. Each volume contains twelve different types of complexes, varying in size, function and structure. In this paper, we have evaluated seven different methods of finding and classifying proteins. Seven research groups present results obtained with learning-based methods and trained on the simulated dataset, as well as a baseline template matching (TM), a traditional method widely used in cryo-ET research. We show that learning-based approaches can achieve notably better localization and classification performance than TM. We also experimentally confirm that there is a negative relationship between particle size and performance for all methods.