Holography captures the full three-dimensional optical field of a sample in a single intensity measurement. This enables numerical refocusing to arbitrary axial planes and, in principle, automatic multifocusing in digital holographic microscopy (DHM). Existing work on holographic autofocusing, however, largely targets the simpler single-plane setting and typically relies either on computationally expensive depth sweeps with hand-crafted sharpness metrics or on deep-learning models trained for a single focal distance per hologram. We address the more practical and challenging problem of automatic multifocusing for in-line flow-through DHM, where heterogeneous microscopic objects are distributed across multiple axial depths within a single hologram. We propose a two-stage pipeline for real-world DHM applications. First, a coarse, low-resolution 3D search detects object locations and provides approximate lateral and axial positions using a robust focus metric on downscaled holograms. Second, a deep-learning-based refinement stage processes the resulting object ROIs and predicts the residual defocus to obtain precise axial positions. Within this framework, we compare regression and classification formulations, introduce classification models with non-uniform depth bins that concentrate capacity near focus, and examine recursive refinement. Evaluated on a practical and diverse dataset of living micro-organisms, the proposed system achieves accurate, data-efficient, robust, and computationally efficient multifocusing, making it suitable for quasi-real-time DHM applications.
This paper introduces a noise augmentation technique designed to enhance the robustness of state-of-the-art (SOTA) deep learning models against degraded image quality, a common challenge in long-term recording systems. Our method, demonstrated through the classification of digital holographic images, utilizes a novel approach to synthesize and apply random colored noise, addressing the typically encountered correlated noise patterns in such images. Empirical results show that our technique not only maintains classification accuracy in high-quality images but also significantly improves it when given noisy inputs without increasing the training time. This advancement demonstrates the potential of our approach for augmenting data for deep learning models to perform effectively in production under varied and suboptimal conditions.
A modified version of the well-known Gerchberg-Saxton algorithm is introduced that in the case of sparse samples provides an extremely fast and accurate phase reconstruction utilizing the whole bandwidth of an off-axis hologram.
A hologram, measured by using appropriate coherent illumination, records all substantial volumetric information of the measured sample. It is encoded in its interference patterns and, from these, the image of the sample objects can be reconstructed in different depths by using standard techniques of digital holography. We claim that a 2D convolutional network (CNN) cannot be efficient in decoding this volumetric information spread across the whole image as it inherently operates on local spatial features. Therefore, we propose a method, where we extract the volumetric information of the hologram by mapping it to a volume—using a standard wavefield propagation algorithm—and then feed it to a 3D-CNN-based architecture. We apply this method to a challenging real-life classification problem and compare its performance with an equivalent 2D-CNN counterpart. Furthermore, we inspect the robustness of the methods to slightly defocused inputs and find that the 3D method is inherently more robust in such cases. Additionally, we introduce a hologram-specific augmentation technique, called hologram defocus augmentation, that improves the performance of both methods for slightly defocused inputs. The proposed 3D-model outperforms the standard 2D method in classification accuracy both for in-focus and defocused input samples. Our results confirm and support our fundamental hypothesis that a 2D-CNN-based architecture is limited in the extraction of volumetric information globally encoded in the reconstructed hologram image.
We adopted an unpaired neural network training technique, namely CycleGAN, to generate bright-field microscope-like images from hologram reconstructions. The motivation for unpaired training in microscope applications is that the construction of paired/parallel datasets is cumbersome or sometimes not even feasible, for example, lensless or flow-through holographic measuring setups. Our results show that the proposed method is applicable in these cases and provides comparable results to the paired training. Furthermore, it has some favorable properties even though its metric scores are lower. The CycleGAN training results in sharper and-from this point of view-more realistic object reconstructions compared to the baseline paired setting. Finally, we show that a lower metric score of the unpaired training does not necessarily imply a worse image generation but a correct object synthesis, yet with a different focal representation.
Non-contact visual monitoring of vital signs in neonatology has been demonstrated by several recent studies in ideal scenarios where the baby is calm and there is no medical or parental intervention. Similar to contact monitoring methods (e.g., ECG, pulse oximeter) the camera-based solutions suffer from motion artifacts. Therefore, during care and the infants’ active periods, calculated values typically differ largely from the real ones. In this way, our main contribution to existing remote camera-based techniques is to detect and classify such situations with a high level of confidence. Our algorithms can not only evaluate quiet periods, but can also provide continuous monitoring. Altogether, our proposed algorithms can measure pulse rate, breathing rate, and to recognize situations such as medical intervention or very active subjects using only a single camera, while the system does not exceed the computational capabilities of average CPU-GPU-based hardware. The performance of the algorithms was evaluated on our database collected at the Ist Dept. of Neonatology of Pediatrics, Dept of Obstetrics and Gynecology, Semmelweis University, Budapest, Hungary.
Remote photoplethysmography (RPPG) is a camera-based optical technique for detecting volumetric changes of organs. This technique enables the non-contact measurement of respiration and pulse. Monitoring newborn infants is a challenging task, due to the weak pulse signals, the rather irregular respiratory pattern, and the frequent movements. Therefore, heavy optimization of the sensing and evaluation process is a must in a resource-limited embedded vision system. This is a two-faceted study, with a special focus on low computational complexity. In the field of respiration monitoring, the paper introduces an optimized convolutional neural network (CNN) and a novel, light Long Short-term Memory (LSTM) motion classifier with a narrow CNN layer. From heart rate measurement point of view, a skin segmentation based algorithm is presented. The performance of each algorithm is evaluated on a database collected at the Ist Dept. of Neonatology of Pediatrics, Dept of Obstetrics and Gynecology, Semmelweis University, Budapest, Hungary.
Experimental setup for remote Photopethysmographic measurement was built and Pulse Volume Vector based signal evaluation method was implemented to measure pulse and blood oxygenation. As opposed to the typically used RGB or three monochromatic camera, an RGB-NIR camera was applied in our setup to obtain space and time registered data from the visual and NIR regions. The setup was calibrated; the pulse and blood oxygenation curves were compared to reference signals. The real-time implementation of a state of the art method was also carried out and tested on premature infants.
Experimental setup for remote Photopethysmographic measurement was built and Pulse Volume Vector based signal evaluation method was implemented to measure pulse and blood oxygenation. As opposed to the typically used RGB or three monochromatic camera, an RGB-NIR camera was applied in our setup to obtain space and time registered data from the visual and NIR regions. The setup was calibrated; the pulse and blood oxygenation curves were compared to reference signals.