Reconstructing thin 3D structures is challenging due to their sparsity, scale variation, and complex geometry. Such structures arise in a wide range of domains, including medical imaging of vascular systems and industrial pipe systems. While recent neural methods perform well on dense surfaces, they often fail to recover fine thin geometries. We propose a reconstruction approach based on local depth projections, which provide an efficient and informative 2D representation of thin structures. Specifically, we traverse the 3D model with a sliding box to generate local orthographic depth projections, which are processed by a neural network to reconstruct missing thin structures in 2D. The local reconstructions are subsequently fused back into the 3D model to produce a coherent and detailed shape. Experiments on pulmonary artery reconstruction from CT volumes and industrial pipeline recovery from synthetic and real scans demonstrate improved preservation of fine structural details over existing methods.
Facial micro-expressions are subtle and short-lived facial movements that provide important cues about genuine human emotions. However, modeling and generating them remains difficult because annotated micro-expression data is limited and the underlying facial motions are extremely weak. Existing micro-expression generation methods therefore often suffer from limited quality, weak robustness, and poor generalization. We propose MagPlus, a transferable micro-expression processing pipeline that connects micro-expression analysis with standard facial animation models. Instead of training a dedicated generator from scratch, MagPlus learns to magnify subtle facial motions into the range of regular facial expressions, transforming micro-expressions into signals that are compatible with existing facial expression processing models. The magnified sequence is then used by a standard facial expression model for tasks such as transfer and synthesis. A complementary DeMagPlus module then restores the generated motion back to realistic micro-expression intensity levels while preserving the synthesized dynamics. We evaluate the framework using four facial animation models: FOMM, FSRT, MetaPortrait, and EmoPortraits. None of these models are trained on micro-expression data. Experiments show that MagPlus-DeMagPlus enables pretrained macro-expression models to generate more realistic micro-expression motion without retraining the backbones.
Stipple patterns, point sets whose local density tracks a target image, are traditionally produced by per-density iterative optimizers, which are slow, non-differentiable, and must be re-run from scratch for each new target. Learned alternatives have so far addressed only unconditional point generation; capacity-constrained, image-conditioned stippling has remained out of reach. We present the first diffusion-based sampler that simultaneously satisfies a learned local point-distribution prior and a continuous, image-defined capacity constraint at inference. The method is a ControlNet branch built on top of an optimal-transport-grid point-set diffusion baseline, conditioned on the target density map and a high-resolution image. Two design choices make the combination tractable: training and inference are restricted to the late-stage denoising regime, initialized from a density-weighted rejection sample, and the standard zero-convolution injection is replaced with a sigmoid-gated 1x1 projection that preserves the base model's blue-noise structure under hard density signals. A single trained checkpoint accepts arbitrary target densities at inference, generalizes to point budgets that were not seen during training, and produces stipples in time nearly independent of the output point count. On the Icons-50 benchmark, our learned sampler reaches parity with per-density-optimized baselines on every reported metric while remaining differentiable end-to-end.
Stochastic porous structures, characterized by randomly distributed voids within solid materials, are prevalent in natural systems such as geological formations, biological tissues, and ecosystems. These structures play crucial roles in processes like nutrient transport and water retention, making them a key focus of interdisciplinary research. Traditional design methods for stochastic porous structures often require detailed modeling of the entire structure, leading to high computational costs. To alleviate this, periodic microstructures are commonly used to fill target regions with repetitive units. However, generating large-scale stochastic porous structures that combine smooth connectivity with global randomness using periodic units remains a significant challenge. This paper presents a novel approach for generating periodic stochastic porous microstructures based on Wang tile rules. The proposed method employs a parameterized generative model with a dual-layer structure, incorporating 27 types of periodic periphery configurations and internal pore-tunnel structures formed from randomly distributed Gaussian kernels. This design balances stochasticity with boundary constraints. Simulations and experiments validate the proposed approach, showing that the resulting stochastic porous microstructures exhibit distinct deformation patterns and superior energy absorption compared to periodic microstructures.
A thin shell model refers to a surface or structure, where the object's thickness is considered negligible. In the context of 3D printing, thin shell models are characterized by having lightweight, hollow structures, and reduced material usage. Their versatility and visual appeal make them popular in various fields, such as cloth simulation, character skinning, and for thin-walled structures like leaves, paper, or metal sheets. Nevertheless, optimization of thin shell models without external support remains a challenge due to their minimal interior operational space. For the same reasons, hollowing methods are also unsuitable for this task. In fact, thin shell modulation methods are required to preserve the visual appearance of a two-sided surface which further constrain the problem space. In this paper, we introduce a new visual disparity metric tailored for shell models, integrating local details and global shape attributes in terms of visual perception. Our method modulates thin shell models using global deformations and local thickening while accounting for visual saliency, stability, and structural integrity. Thereby, thin shell models such as bas-reliefs, hollow shapes, and cloth can be stabilized to stand in arbitrary orientations, making them ideal for 3D printing.
Triply periodic minimal surfaces (TPMS) have been extensively studied in functionally graded structures due to their excellent geometric and mechanical properties. However, existing methods often suffer from deformation and distortion due to inadequate TPMS period calculations, consequently compromising the mechanical performance of TPMS structures. In this paper, we investigate the period and phase of TPMS and establish a relationship between the period and cell size. We introduce an optimization method that effectively reduces mean curvature deviations and geometric distortions while enabling localized control of TPMS. Consequently, we minimize distortion in functionally graded TPMS structures. To evaluate the effectiveness of our proposed method, we conduct a comparative analysis with existing TPMS modification methods, both through simulation and physical experiments. Results demonstrate that our structures exhibit reduced distortions, with mean curvature values closer to the standard TPMS. Physical tests indicate that our approach yields greater stiffness and improved energy absorption compared to existing methods for transitioning TPMS period. This offers the additive manufacturing community a novel solution for producing TPMS with controllable period transitions and higher energy absorption.
Interpretation of predictions made by Convolutional Neural Networks (CNNs) is a rapidly growing field of research. A common approach involves enhancing semantic segmentation predictions through the generation of heatmaps that illustrate the significance of individual pixels in the segmentation. Nevertheless, the selection of beneficial features from these heatmaps remains a challenge. This is because the introduced information often contains interfering factors such as mutual features between different objects, background, and insufficient heat map resolution which often diminish its effectiveness. To overcome these limitations, we introduce Refined Weak Slices (RWS). Our main idea is to identify low attention regions in heat maps i.e. weak slices , in conjunction with segmentation accuracy, and utilize them to select effective features across different DNN layers, to enhance segmentation. We then seamlessly integrate these features back into the CNN, thus refining and enhancing the semantic segmentation result with selected features. Through extensive experiments, we demonstrate that incorporating the RWS module into state-of-the-art methods yields a notable improvement in the average mIoU by 2.84% on benchmark datasets (VOC 2012, COCOStuff, ADE20K, Cityscapes) for both ResNet-101 and ResNet-50 architectures. Furthermore, we achieve a maximum improvement of 5.8% with a single CNN. Overall, the combination of RWS and CNNs exhibits excellent performance in image segmentation tasks.
Image alignment and registration methods typically rely on visual correspondences across common regions and boundaries to guide the alignment process. Without them, the problem becomes significantly more challenging. Nevertheless, in real world, image fragments may be corrupted with no common boundaries and little or no overlap. In this work, we address the problem of learning the alignment of image fragments with gaps (i.e., without common boundaries or overlapping regions). Our setting is unsupervised, having only the fragments at hand with no ground truth to guide the alignment process. This is usually the situation in the restoration of unique archaeological artifacts such as frescoes and mosaics. Hence, we suggest a self-supervised approach utilizing self-examples which we generate from the existing data and then feed into an adversarial neural network. Our idea is that available information inside fragments is often sufficiently rich to guide their alignment with good accuracy. Following this observation, our method splits the initial fragments into sub-fragments yielding a set of aligned pieces. Thus, sub-fragmentation allows exposing new alignment relations and revealing inner structures and feature statistics. In fact, the new sub-fragments construct true and false alignment relations between fragments. We feed this data to a spatial transformer GAN which learns to predict the alignment between fragments gaps. We test our technique on various synthetic datasets as well as large scale frescoes and mosaics. Results demonstrate our method's capability to learn the alignment of deteriorated image fragments in a self-supervised manner, by examining inner image statistics for both synthetic and real data.
Importance Stereotypical motor movements (SMMs) are a form of restricted and repetitive behavior (RRB), which is a core symptom of Autism Spectrum Disorder (ASD). Current quantification of SMM severity is extremely limited, with studies relying on coarse and subjective caregiver reports or laborious manual annotation of short video recordings.Objective To demonstrate the utility of a new open-source AI algorithm that can analyze extensive video recordings of children and automatically identify segments with heterogeneous SMMs, thereby enabling their direct and objective quantification.Design, setting, and participants This retrospective cohort study analyzed video recordings from 319 behavioral assessments of 241 children with ASD, 1.4 to 8 years old, who participated in research at the Azrieli National Centre for Autism and Neurodevelopment Research in Israel. Behavioral assessments included cognitive, language, and autism diagnostic observation schedule, 2nd edition (ADOS-2) assessments.Exposures Each assessment was recorded with 2-4 cameras, yielding 580 hours of video footage. We manually annotated 7,352 video segments containing heterogeneous SMMs performed by different children (21.14 hours of video).Main outcomes and measures We used a pose-estimation algorithm (OpenPose) to extract skeletal representations of all individuals in each video frame and trained an object-detection algorithm (YOLOv5) to identify and track the child in each movie. We then used the skeletal representation of the child to train an SMM recognition algorithm using a PoseConv3D model. We used data from 220 children for training and data from the remaining 21 children for testing.Results The algorithm accurately detected 92.53% of manually annotated SMMs in our test data with 66.82% precision. Overall number and duration of algorithm identified SMMs per child were highly correlated with manually annotated number and duration of SMMs (r=0.8 and r=0.88, p<0.001 respectively).Conclusion and relevance These findings demonstrate the ability of the algorithm to capture a highly diverse range of SMMs and quantify them with high accuracy, enabling objective and direct estimation of SMM severity in individual children with ASD. We openly share the “ASDPose” dataset and “ASDMotion” algorithm for further use by the research community.Question Is it possible to train a deep learning algorithm to accurately identify and quantify stereotypical motor movements (SMMs) in video recordings of children with autism?Findings The ASDMotion algorithm was trained and tested with the largest video dataset of ASD children curated to date, comprised of 319 behavioral assessment recordings from 241 ASD children. The algorithm successfully detected 92.53% of manually identified SMMs with 66.82% precision, achieving highly accurate quantification of SMMs per child that were strongly correlated (r≥0.8) with quantification by manual annotation.Meaning This study demonstrates the utility of ASDMotion for objective and direct quantification of SMM severity in children with ASD, offering a new freely available, open-source algorithm and dataset that enable transformative basic and clinical ASD research.### Competing Interest StatementThe authors have declared no competing interest.
ImportanceStereotypical motor movements (SMMs) are a form of restricted and repetitive behavior, which is a core symptom of autism spectrum disorder (ASD). Current quantification of SMM severity is extremely limited, with studies relying on coarse and subjective caregiver reports or laborious manual annotation of short video recordings.ObjectiveTo assess the utility of a new open-source AI algorithm that can analyze extensive video recordings of children and automatically identify segments with heterogeneous SMMs, thereby enabling their direct and objective quantification.Design, Setting, and ParticipantsThis retrospective cohort study included 241 children (aged 1.4 to 8.0 years) with ASD. Video recordings of 319 behavioral assessments carried out at the Azrieli National Centre for Autism and Neurodevelopment Research in Israel between 2017 and 2021 were extracted. Behavioral assessments included cognitive, language, and autism diagnostic observation schedule, 2nd edition (ADOS-2) assessments. Data were analyzed from October 2020 to May 2024.ExposuresEach assessment was recorded with 2 to 4 cameras, yielding 580 hours of video footage. Within these extensive video recordings, manual annotators identified 7352 video segments containing heterogeneous SMMs performed by different children (21.14 hours of video).Main outcomes and measuresA pose estimation algorithm was used to extract skeletal representations of all individuals in each video frame and was trained an object detection algorithm to identify the child in each video. The skeletal representation of the child was then used to train an SMM recognition algorithm using a 3 dimensional convolutional neural network. Data from 220 children were used for training and data from the remaining 21 children were used for testing.ResultsAmong 319 behavioral assessment recordings from 241 children (172 [78%] male; mean [SD] age, 3.97 [1.30] years), the algorithm accurately detected 92.53% (95% CI, 81.09%-95.10%) of manually annotated SMMs in our test data with 66.82% (95% CI, 55.28%-72.05%) precision. Overall number and duration of algorithm-identified SMMs per child were highly correlated with manually annotated number and duration of SMMs (r = 0.8; 95% CI, 0.67-0.93; P < .001; and r = 0.88; 95% CI, 0.74-0.96; P < .001, respectively).Conclusions and relevanceThis study suggests the ability of an algorithm to identify a highly diverse range of SMMs and quantify them with high accuracy, enabling objective and direct estimation of SMM severity in individual children with ASD.
Stochastic porous structures are ubiquitous in natural phenomena and have gained considerable traction across diverse domains owing to their exceptional physical properties. The recent surge in interest in microstructures can be attributed to their impressive attributes, such as a high strength-to-weight ratio, isotropic elasticity, and bio-inspired design principles. Notwithstanding, extant stochastic structures are predominantly generated via procedural modeling techniques, which present notable difficulties in representing geometric microstructures with periodic boundaries, thereby leading to intricate simulations and computational overhead. In this manuscript, we introduce an innovative method for designing stochastic microstructures that guarantees the periodicity of each microstructure unit to facilitate homogenization. We conceptualize each pore and the interconnecting tunnel between proximate pores as Gaussian kernels and leverage a modified version of the minimum spanning tree technique to assure pore connectivity. We harness the dart-throwing strategy to stochastically produce pore locations, tailoring the distribution law to enforce boundary periodicity. We subsequently employ the level-set technique to extract the stochastic microstructures. Conclusively, we adopt Wang tile rules to amplify the stochasticity at the boundary of the microstructure unit, concurrently preserving periodicity constraints among units. Our methodology offers facile parametric control of the designed stochastic microstructures. Experimental outcomes on 3D models manifest the superior isotropy and energy absorption performance of the stochastic porous microstructures. We further corroborate the efficacy of our modeling strategy through simulations of mechanical properties and empirical experiments.
Humans use simple sketches to convey complex concepts and abstract ideas in a concise way. Just a few abstract pencil strokes can carry a large amount of semantic information that can be used as meaningful representation for many applications. In this work, we explore the power of simple human strokes denoted to capture high-level 2D shape semantics. For this purpose, we introduce OneSketch, a crowd-sourced dataset of abstract one-line sketches depicting high-level 2D object features. To construct the dataset, we formulate a human sketching task with the goal of differentiating between objects with a single minimal stroke. While humans are rather successful at depicting high-level shape semantics and abstraction, we investigate the ability of deep neural networks to convey such traits. We introduce a neural network which learns meaningful shape features from our OneSketch dataset. Essentially, the model learns sketch-to-shape relations and encodes them in an embedding space which reveals distinctive shape features. We show that our network is applicable for differentiating and retrieving 2D objects using very simple one-stroke sketches with good accuracy.
Motions in videos are often governed by physical and biological laws such as gravity, collisions, flocking, etc. Accounting for such natural properties is an appealing way to improve realism in future frame video prediction. Nevertheless, the definition and computation of intricate physical and biological properties in motion videos are challenging. In this work, we introduce PhyLoNet, a PhyDNet extension that learns long-term future frame prediction and manipulation. Similar to PhyDNet, our network consists of a two-branch deep architecture that explicitly disentangles physical dynamics from complementary information. It uses a recurrent physical cell (PhyCell) for performing physically-constrained prediction in latent space. In contrast to PhyDNet, PhyLoNet introduces a modified encoder-decoder architecture together with a novel relative flow loss. This enables a longer-term future frame prediction from a small input sequence with higher accuracy and quality. We have carried out extensive experiments, showing the ability of PhyLoNet to outperform PhyDNet on various challenging natural motion datasets such as ball collisions, flocking, and pool games. Ablation studies highlight the importance of our new components. Finally, we show an application of PhyLoNet for video manipulation and editing by a novel class label modification architecture.
Historical documents and archaeological artifacts are hard to process due to natural degradation, fading, spills, tears, overlaid data,, and so on. In this work, we focus on the task of recovering characters and symbols from images of corrupted archaeological artifacts where data is partially erased, occluded, or overwritten by other data. Such phenomena can be widely observed in image datasets of palimpsests and petroglyphs consisting of erased, overwritten, and in general heavily degraded data. Segmentation and binarization are typically applied to such images to detect and recover characters and symbols from their background. However, these methods mainly focus on the visible data while in our case, due to large corruption, both visible and invisible information should be considered. For example, computing the segmentation mask of an occluded character requires also labeling invisible pixels and missing parts. In this work, we introduce a deep neural network that computes character segmentation in palimpsests and petroglyphs while overcoming occlusions, missing parts, and degradation. Our network has inference abilities, thus, not only segmenting the symbol’s foreground pixels but also inferring and completing missing and corrupted parts. Since palimpsests and petroglyphs have very limited annotated ground-truth data, we also introduce data augmentation tools to properly train our network. We demonstrate both qualitative and quantitative performance of our method also including a user study involving expert evaluation.
The reconstruction of 3D point clouds with high-level geometric primitives is highly desirable due to the compactness and effectiveness of this representation. Nevertheless, correctly fitting 3D point sets with matching primitives is challenging due to the high combinatorial nature of this problem. We introduce a recursive neural network architecture that learns to fit 3D points with geometric primitives in an unsupervised manner. Our core idea is to divide-and-conquer the combinatorial complexity of primitive fitting utilizing recursion layers in our neural network architecture. Through recursion, network layers focus on different shape scales and fit primitives in a coarse-to-fine manner. I.e., early recursion layers solve for coarse regions, fitting large primitives while later layers solve for remaining fine-scale regions, fitting smaller primitives. The network model guarantees a global solution as recursion layers collaborate with each other to minimize an accumulative global loss. We present experiments that validate our approach in terms of accuracy, robustness and performance. We also compare to state-of-the-art to demonstrate our method's advantages.
Learning 3D point sets with rotational invariance is an important and challenging problem in machine learning. Through rotational invariant architectures, 3D point cloud neural networks are relieved from requiring a canonical global pose and from exhaustive data augmentation with all possible rotations. In this work, we introduce a rotational invariant neural network by combining recently introduced vector neurons with self-attention layers to build a point cloud vector neuron transformer network (VNT-Net). Vector neurons are known for their simplicity and versatility in representing SO(3) actions and are thereby incorporated in common neural operations. Similarly, Transformer architectures have gained popularity and recently were shown successful for images by applying directly on sequences of image patches and achieving superior performance and convergence. In order to benefit from both worlds, we combine the two structures by mainly showing how to adapt the multi-headed attention layers to comply with vector neurons operations. Through this adaptation attention layers become SO(3) and the overall network becomes rotational invariant. Experiments demonstrate that our network efficiently handles 3D point cloud objects in arbitrary poses. We also show that our network achieves higher accuracy when compared to related state-of-the-art methods and requires less training due to a smaller number of hyperparameters in common classification and segmentation tasks.
We present a method of designing and fabricating pottery artwork through human posture where the users need neither professional skills nor experience with 3D modeling software. This method creatively solves the problem of mapping between the human skeleton and the pottery shape and has a strong ability to model complex shapes while being user-friendly. Our system represents the deformation of a pottery object with four independent operators and provides real-time visual feedback as the user changes their posture. Unlike traditional pottery throwing, where the product has high symmetry, our system supports modeling asymmetric pottery shapes. After obtaining the model, the model can be fabricated directly via ceramic 3D printing, as the models satisfy printability constraints. A user study showed that the designed mapping relationship supports pottery shape deformation with a high degree-of-freedom, and inexperienced users can easily use the system with minimal instruction.
In this article we introduce a differentiable rendering module which allows neural networks to efficiently process 3D data. The module is composed of continuous piecewise differentiable functions defined as a sensor array of cells embedded in 3D space. Our module is learnable and can be easily integrated into neural networks allowing to optimize data rendering towards specific learning tasks using gradient based methods in an end-to-end fashion. Essentially, the module's sensor cells are allowed to transform independently and locally focus and sense different parts of the 3D data. Thus, through their optimization process, cells learn to focus on important parts of the data, bypassing occlusions, clutter, and noise. Since sensor cells originally lie on a grid, this equals to a highly non-linear rendering of the scene into a 2D image. Our module performs especially well in presence of clutter and occlusions as well as dealing with non-linear deformations to improve classification accuracy through proper rendering of the data. In our experiments, we apply our module in various learning tasks and demonstrate that using our rendering module we accomplish efficient classification, localization, and segmentation tasks on 2D/3D cluttered and non-cluttered data.
Applying textures to 3D models are means for creating realistic looking objects. This is especially important in the 3D manufacturing domain as manufactured models should ideally comprise a natural and realistic appearance. Nevertheless, natural material textures usually consist of dense patterns and fine details. Their embedding onto 3D models is typically cumbersome, requiring large processing time and resulting in large size meshes. This paper presents a novel approach for direct embedding of fine scale geometric textures onto 3D printed models by on-the-fly modification of the 3D printer's head. Our idea is to embed 3D textures by revising the 3D printer's G-code, i.e., incorporating texture details through modification of the printer's path. Direct manipulation of the printer's head movement allows for fine-scale texture mapping and editing on-the-fly in the 3D printing process. Thus, our method avoids the computationally expensive texture mapping, mesh processing and manufacturing preprocessing. This allows embedding detailed geometric textures of unlimited density which can model manual manufacturing artifacts and natural material properties. Results demonstrate that our direct G-code textured models are printed robustly and efficiently in both space and time compared to traditional methods.
The main scientific objective of the project is to introduce computer graphics techniques into solving a specific computer vision problem. It will make use of the latest advancements in capturing technology and in graphics illumination algorithms to produce a novel illumination neutralization technique. Objects shall emerge in appearance that is independent from the lighting conditions of the scene. The process shall be optimized in terms of timing constraints to permit its utilization in real-time systems. The system shall be robust eliminating the difficulties imposed by illumination variations.