Additive manufacturing (AM) enables the production of complex geometries but presents challenges related to defect formation, especially in powder bed fusion technologies where defects like porosities are common. Efficient and accurate defect identification in metallic parts is critical but often time-consuming, error-prone, and costly. X-ray computed tomography (XCT) is widely used for non-destructive defect detection. Still, it suffers from high computational demands, numerical artifacts, and reduced spatial resolution due to the reconstruction process. Deep learning defect inspections require large dataset but only small dataset are available, as of now. To address this challenge, a previously published method introduced a 2D defect extractor operating directly in the Radon space. While effective, this approach is limited by its slice-by-slice processing. To overcome this, we propose the Sinogram Defect Augmentation Methodology (SinoDAM), which refines the previous workflow into a 3D approach. SinoDAM generates volumetric variations commonly observed in industrial conditions by incorporating geometric transformations. Operating directly in the sinogram domain ensures the physical plausibility of synthetic data while avoiding inconsistencies introduced during reconstruction. SinoDAM offers an effective solution for synthesizing diverse and realistic dataset in the Radon space. It is designed to support future deep learning applications in defect detection and classification, offering a more scalable and accurate solution for quality control in AM.
Objective.Accurate reconstruction of localized anatomical details is essential in medical image synthesis, particularly when addressing specific clinical requirements such as the identification or measurement of fine structures. Traditional methods for image translation and synthesis are generally optimized for global image reconstruction but often fall short in providing the finesse required for detailed local analysis. This study represents a step toward addressing this challenge by introducing a novel anatomical feature-prioritized (AFP) loss function into the synthesis process.Approach.The AFP loss integrates features from pre-trained task-specific models, such as anatomical segmentation networks, into the image synthesis pipeline to enforce attention to critical structures. This loss function is evaluated across multiple architectures, including GAN-based and CNN-based models, and applied in two cross-modality contexts: (1) lung MR to CT translation with an emphasis on bronchial structure preservation, using a private thoracic dataset; and (2) pelvis MR to CT synthesis, targeting organ and muscle reconstruction, using the public SynthRAD2023 dataset. Feature embeddings from domain-specific segmentation networks are extracted to guide synthesis toward anatomically meaningful outputs.Main results.The AFP loss demonstrated consistent improvements in downstream segmentation accuracy across both domains. For lung airway reconstruction, the Dice coefficient increased from 0.534 with standard L1 loss to 0.584 using AFP loss. In pelvic imaging, bone reconstruction Dice scores improved from 0.738 using L1 loss to 0.780 with AFP loss. These results confirm that the AFP loss improves the reconstruction of anatomical structures while maintaining comparable intensity-based metrics, indicating that global image quality is not compromised.Significance.The proposed AFP loss provides a modular and generalizable approach for embedding anatomical task-awareness into medical image synthesis. By aligning image translation objectives with clinically relevant features, it offers a pathway toward more precise and useful synthetic images for downstream tasks, supporting broader integration of image synthesis in clinical workflows.
A controversial part of the state-of-the-art in deep learning applied to medical imaging concerns the selection of 2D, 2.5D or 3D representations. This study contributes by contrasting the prevalent 2D and 3D approaches and emphasizing 2.5D methods in the context of image translation for lung MR to CT synthesis. The debate between 2D and 3D convolution techniques, particularly the challenge of preserving 3D information within the network, becomes a focal point of investigation. Our approach involves a comprehensive comparison of these conventional methods with various 2.5D approaches. Notably, we propose and evaluate a novel strategy that combines multi-axial and stacked slices 2.5D techniques, with either 2D or 3D convolutions, aiming to provide insights on the preservation of 3D information through these approaches. This comparative analysis extends beyond visual and quantitative assessments, incorporating structural similarity metrics and task-specific evaluations. Additionally, we review the number of parameters, memory cost, and training times, providing a comprehensive understanding of the practical implications of 2.5D networks in terms of computational resources. The findings contribute valuable insights into the intricate balance between 2D, 3D, and 2.5D representations in the context of MR to CT synthesis in the lungs.
In medical image synthesis, the development of robust and reliable baseline methods is crucial due to the complexity and variability of existing techniques. Despite advances with architectures such as GANs and diffusion models, a clear state-of-the-art has yet to be established. This paper introduces a versatile adaptation of the nnU-Net framework as a robust baseline for both cross-modality synthesis and image inpainting tasks. Known for its superior performance in segmentation challenges, nnU-Net's automatic configuration and parameter optimization capabilities have been adapted for these new applications. We evaluate this method on two use cases: pelvis MR to CT translation using the Synthrad2023 challenge dataset and local synthesis using the BraTs 2023 inpainting challenge dataset. Standard synthesis metrics -MAE, MSE, SSIM and PSNR- demonstrate that our adapted nnU-Net outperforms GAN-based methods like pix2pixHD and ranks among the best methods for both challenges. We recommend this adapted nnU-Net as a new benchmark for medical image translation and inpainting tasks, and provide our implementations for public use on GitHub.
Identifying defects in metallic parts from additive manufacturing is crucial. It can be time-consuming and repeti-tive, prone to errors and high costs. In non-destructive testing, characterising defects in these parts is made possible through X-ray computed tomography. Most of the studies provide analysis in the reconstructed images which is computationally demanding, sensible to artefacts and reduces spatial resolution. In this paper, we introduce our iterative sinusoidal Hough Transform (iSHT) for feature extraction directly in Radon space based on our restrained spatial representation method. We designed a 2D numerical phantom to demonstrate the effectiveness of the methodology up to the augmentation of this dataset.
The use of perceptual loss has emerged as a prominent technique for enhancing image-to-image translation tasks by capturing high-level visual similarities between the ground truth and predicted images. In this study, we investigate fine structure generation in MR to CT synthesis and focus on evaluating the effectiveness of perceptual loss under different parameters. Specifically, we assess the impacts of architecture and training data of the pre-trained network, as well as the effects of data normalization and feature selection. To validate our findings, we generate synthetic lung CT images using GANs from an MR dataset of 120 patients, and evaluate them against the corresponding registered CT scans. Our evaluation metrics include measures of structural and visual image quality as well as a taskspecific metric using the segmentation of airway bronchi in the synthesized CT images. Our work highlights the importance of maintaining a balanced integration of low-level and high-level features to achieve high-quality fine structure generation. Additionally, we compare perceptual loss pre-trained on either natural images or medical images, and our results suggest that the dataset used to pre-train the model does not necessarily need to rely on generic medical data, as it is only used as a fixed feature extractor in the context of perceptual loss. By illuminating the impact of latent space computation and feature selection, our study offers valuable insights into improving fine structure generation in MR to CT synthesis and image-to-image translation in general.
This research embarked on a comparative exploration of the holistic segmentation capabilities of Convolutional Neural Networks (CNNs) in both 2D and 3D formats, focusing on cystic fibrosis (CF) lesions. The study utilized data from two CF reference centers, covering five major CF structural changes. Initially, it compared the 2D and 3D models, highlighting the 3D model's superior capability in capturing complex features like mucus plugs and consolidations. To improve the 2D model's performance, a loss adapted to fine structures segmentation was implemented and evaluated, significantly enhancing its accuracy, though not surpassing the 3D model's performance. The models underwent further validation through external evaluation against pulmonary function tests (PFTs), confirming the robustness of the findings. Moreover, this study went beyond comparing metrics; it also included comprehensive assessments of the models' interpretability and reliability, providing valuable insights for their clinical application.
In clinical practice, the modality of choice for lung diagnosis is usually computed tomography (CT), which exposes patients to ionizing radiations and could potentially affect patients’ health. Conversely, MR scan is considered safe and non-invasive but seems challenging due to the low proton density of the lungs and respiratory artifacts. Recently, ultrashort echo-time (UTE) MRI has been developed for lung assessment and shows promising results. In this work, we propose generating 2D synthetic CT slices from UTE MR slices, to improve the image quality and interpretability. Lung MR and CT volumes of 110 patients acquired on the same day were registered using an accurate edge-based non-rigid registration method. We trained and compared paired state-of-the-art generative models based on adversarial, feature-matching and perceptual losses, and also evaluated the impact of conditional batch normalization, namely SPADE [17], on image synthesis. Quantitative and qualitative evaluations showed that this approach was able to synthesize CT images that closely approximate ground truth CT images, and also enables the use of algorithms originally designed for real CT.
: In medical imaging, MR-to-CT synthesis has been extensively studied. The primary motivation is to benefit from the quality of the CT signal, i.e. excellent spatial resolution, high contrast, and sharpness, while avoiding patient exposure to CT ionizing radiation, by relying on the safe and non-invasive nature of MRI. Recent studies have successfully used deep learning methods for cross-modality synthesis, notably with the use of conditional Generative Adversarial Networks (cGAN), due to their ability to create realistic images in a target domain from an input in a source domain. In this study, we examine in detail the different steps required for cross-modality translation using GANs applied to MR-to-CT lung synthesis, from data representation and pre-processing to the type of method and loss function selection. The different alternatives for each step were evaluated using a quantitative comparison of intensities inside the lungs, as well as bronchial segmentations between synthetic and ground truth CTs. Finally, a general guideline for cross-modality medical synthesis is proposed, bringing together best practices from generation to evaluation.
X-ray absorption imaging is used in the medical field since a long time, but recent advance in phase-contrast imaging made it feasible in a clinical setup. X-ray Phase-contrast imaging technique using a Hartmann sensor allows extracting the absorption and phase information in a single acquisition, allowing to extract a phase-shift information with a minimal exposition and deposited dose. An iterative wavefront reconstruction (IR-WF) algorithm is necessary to extract the phase and absorption values from an acquired image. Our method consists of merging the wavefront reconstruction with a computed tomographic iterative reconstruction (IR-CT) to ensure that all images converge to the same result, improving the final 3D volume.
Maintenance is inevitable, time-consuming, expensive, and risky to production and maintenance operators. Porting maintenance support applications to mixed reality (MR) headsets would ease operations. To function, the application needs to anchor 3D graphics onto real objects, i.e. locate and track real-world objects in three dimensions. This task is known in the computer vision community as Six Degree of Freedom Pose Estimation (6-Dof) and is best solved using Convolutional Neural Networks (CNNs). Training them required numerous examples, but acquiring real labeled images for 6-DoF pose estimation is a challenge on its own. In this article, we propose first a thorough review of existing non-synthetic datasets for 6-DoF pose estimations. This allows identifying several reasons why synthetic training data has been favored over real training data. Nothing can replace real images. We show next that it is possible to overcome the limitations faced by previous datasets by presenting a new methodology for labeled images acquisition. And finally, we present a new dataset named NEMA that allows deep learning methods to be trained without the need for synthetic data.
Hyperspectral imaging allows the classification and localization of materials for diverse applications. The existing datasets are either limited to a single image or created for specific applications. In our work, we need a dataset of urban materials for classification. However, the illumination and acquisition conditions are varying over time. This impacts the raw hyperspectral images and drives the need of making images independent from the acquisition conditions. To deal with it, classical approaches are based on the conversion of raw images to radiance and reflectance. Though many studies have been conducted on reflectance images and on ways to improve their robustness to such changes, the selection of a proper radiance image has yet to be studied. In this paper, we first describe the creation of a dataset, then study two methods for correcting the radiance computation and comment the results of the proposed process.
Augmented Reality is increasingly used for visualizing underground networks. However, standard visual cues for depth perception have never been thoroughly evaluated via user experiments in a context involving physical occlusions (e.g., ground) of virtual objects (e.g., elements of a buried network). We therefore evaluate the benefits and drawbacks of two techniques based on combinations of two well-known depth cues: grid and shadow anchors. More specifically, we explore how each combination contributes to positioning and depth perception. We demonstrate that when using shadow anchors alone or shadow anchors combined with a grid, users generate 2.7 times fewer errors and have a 2.5 times lower perceived workload than when only a grid or no visual cues are used. Our investigation shows that these two techniques are effective for visualizing underground objects. We also recommend the use of one technique or another depending on the situation.
Estimating depth from 2D images has become an active field of study in autonomous driving, scene reconstruction , 3D object recognition, segmentation, and detection. Best performing methods are based on Convolutional Neural Networks, and, as the process of building an appropriate set of data requires a tremendous amount of work, almost all of them rely on the same benchmark to compete between each other : The KITTI benchmark. However, most of them will use the ground truth generated by the LiDAR sensor which generates very sparse depth map with sometimes less than 5% of the image density, ignoring the second image that is given for stereo estimation. Recent approaches have shown that the use of both input images given in most of the depth estimation data set significantly improve the generated results. This paper is in line with this idea, we developed a very simple yet efficient model based on the U-NET architecture that uses both stereo images in the training process. We demonstrate the effectiveness of our approach and show high quality results comparable to state-of-the-art methods on the KITTI benchmark.
Asian hornets are considered a pest because of their dangerousness and their impact on the ecosystem. Detecting nests of this species is a difficult task, as they are found in the trees, hidden in the leaves. Our goal is to carry out this detection from images acquired by a drone. We propose in this work a new method, based on the advantages of visible spectrum and FLIR images. We compare two models of state-of-the-art neural networks (YOLO and Mask-RCNN) for this task. The results are presented from the two separate image sets, then by combining the network responses. To do this, a third dataset (for ensemble model) was built by simulating a FLIR acquisition simultaneous with the acquisition in the visible spectrum. Preliminary results show that the best strategy is to use Mask-RCNN on the ensemble model (detection rate of 93%). A discussion on the relevant information present in the images and on taking into account of this information by the networks is also proposed.
In robotic mapping and navigation, of prime importance today with the trend for autonomous cars, simultaneous localization and mapping (SLAM) algorithms often use stereo vision to extract 3D information of the surrounding world. Whereas the number of creative methods for stereo-based SLAM is continuously increasing, the variety of datasets is relatively poor and the size of their contents relatively small. This size issue is increasingly problematic, with the recent explosion of deep learning based approaches, several methods require an important amount of data. Those multiple techniques contribute to enhance the precision of both localization estimation and mapping estimation to a point where the accuracy of the sensors used to get the ground truth might be questioned. Finally, because today most of these technologies are embedded on on-board systems, the power consumption and real-time constraints turn to be key requirements. Our contribution is twofold: we propose an adaptive SLAM method that reduces the number of processed frame with minimum impact error, and we make available a synthetic flexible stereo dataset with absolute ground truth, which allows to run new benchmarks for visual odometry challenges. This dataset is available online at http://alastor.labri.fr/.
Periosteal reactions are frequently used as a proxy for past populations' health. However, macroscopic evidence of periosteal activity observed on infant's dry bones can be difficult to interpret in terms of normal growth process or pathological changes. This could lead to over-interpretation of a poor health status during childhood in past populations. The aim of our study is to propose new distinctive micro-morphological criteria to differentiate between nonadult dry bones presenting physiological sub-periosteal bone growth and periosteal reaction due to pathological conditions. We sampled 12 perinatal human tibiae from two osteoarcheological collections and proceeded a 3D microstructural analysis of the canal network of cortical bone using micro-CT. 3D canal network organization in cortical bone allowed us to distinguish physiological from pathological periosteal reactions. Moreover, our study revealed different types of organization of the cortical bone microstructure corresponding to various stages of bone remodeling. This exploratory study shows the advantages of a non-destructive microscopic analysis using 3D imaging of the cortical canal network organization in order to distinguish between physiological and pathological bone production during early growth. This will contribute to a better understanding of past populations' epidemiology.
Localizing objects is a key challenge for robotics, augmented reality and mixed reality applications. Images taken in the real world feature many objects with challenging factors such as occlusions, motion blur and changing lights. In manufacturing industry scenes, a large majority of objects are poorly textured or highly reflective. Moreover, they often present symmetries which makes the localization task even more complicated. PoseNet is a deep neural network based on GoogleNet that predicts camera poses in indoor room and outdoor streets. We propose to evaluate this method for the problem of industrial object pose estimation by training the network on the T-LESS dataset. We demonstrate with our experiments that PoseNet is able to predict translation and rotation separately with high accuracy. However, our experiments also prove that it is not able to learn translation and rotation jointly. Indeed, one of the two modalities is either not learned by the network, or forgotten during training when the other is being learned. This justifies the fact that future works will require other formulation of the loss as well as other architectures in order to solve the pose estimation general problem.