Fourier phase retrieval aims to reconstruct a real-valued signal from the magnitude of the Fourier transform of the signal. The solution is typically non-unique and forms a union of tori, each of which is generally called a magnitude torus. The non-convex geometry of the magnitude torus, together with additional constraints, further complicates theoretical analysis of phase retrieval algorithms such as projected gradient descent. In this paper, we propose two algorithms for manifold optimization on magnitude torus and analyze their performance directly via ordinary first- and second-order analyses. We further evaluate their performance on binary signals and, along the way, obtain an explicit description of the magnitude torus and its tangent and normal spaces. Our results facilitate further analysis of related optimization algorithms on magnitude torus. Detailed proofs and code can be found in the supplementary material.
Reconstructing complete 3D geometry from monocular clinical endoscopic videos is challenging due to weak texture, repetitive tissue patterns, and severe illumination artifacts. Although emerging dense matching methods exhibit improved resilience to textureless regions, they often produce abundant spurious correspondences across non-overlapping views, which corrupts the correspondence graph and causes structure-from-motion (SfM) pipelines to fail. In this work, we propose a framework for the construction of dense correspondence graph that leverages explicit temporal locality, parallax-driven geometric constraints, and loop-closure revisiting to enable reliable SfM for monocular endoscopic videos. Instead of exhaustively connecting all frame pairs, the approach effectively suppresses invalid inter-frame correspondences while preserving essential long-range geometric relations critical for stable reconstruction. Combined with illumination-aware masking and SfM initialization adapted to endoscopy, the proposed framework achieves substantial improvements in registration robustness and reconstruction completeness for both phantom and real clinical datasets.
Most near-eye augmented reality (AR) displays with a single fixed focal plane struggle to provide accurate depth cues, causing vergence-accommodation conflict and visual discomfort to viewers. Light field display offers a promising solution for reconstructing virtual scenes with full-parallax information and matching focus cues. However, due to non-sequential ray propagation in the diffractive waveguide, ghost artifact arises as the structure of light field is disrupted by the replicated light rays. This disruptive phenomenon compromises depth fidelity and overall image quality and remains as a significant challenge for near-eye light field AR displays. Since light rays emanated from the same virtual object point across different subviews encode the spatial and angular information of the scene and converge at the focal plane of the eye corresponding to the depth of the virtual object, maintaining the correct spatial sequence of subviews is crucial for accurate rendering of virtual image at that depth. In this paper, we propose a polarization-based approach to mitigate the ghost artifact. By differentiating polarization states of adjacent subviews in the light field and incorporating polarizers, repetitive out-coupling of light rays is effectively suppressed. The results show a significant reduction of ghost artifact in the perceived image while preserving the structure of light field and image quality. Our study provides insight into preservation of the structure of light field and offers an effective solution to improve the viewers' visual experience, contributing to the development of more comfortable, perceptually accurate near-eye AR displays.
Phase detection autofocus (PDAF) technology is essential to digital cameras in many applications including professional photography and autonomous navigation. A deep-learning-based autofocus method is advantageous over conventional AF methods in both speed and accuracy; however, the quality of training data remains a major bottleneck. In this paper, we propose a hybrid labeling strategy that leverages the complementary strengths of focus profiles derived from phase and RGB data. The latter has finer spatial resolution but is noisier than the former. By including both types of data for model training, better accuracy and reliability for PDAF can be achieved. Experiments on various scenes demonstrate that the proposed hybrid labeling strategy achieves higher accuracy than monotype labeling strategies, leading to a practical single-camera alternative to multi-camera-based or depth-based labeling solutions that are often clumsy and computationally expensive.
Light field AR glasses can provide better visual comfort than conventional AR glasses; however, studies on user performance comparison between them are notably scarce. In this article, we present a systematic method employing a serial visual search task without confounding factors to quantify and compare the user performance and experience between these two types of AR glasses at two different viewing distances, 30 cm and 60 cm, and in two modes, purely virtual VR mode and virtual-real integration AR mode. The results show that the light field AR glasses led to a significantly faster reaction speed and higher accuracy than the conventional AR glasses at 30 cm in the AR mode. The participant feedback also shows that the former led to better virtual-real integration. User performance and experience of the light field AR glasses remained consistent across different viewing distances. Although the conventional AR glasses had a better search efficiency than the light field AR glasses at 60 cm in both AR and VR modes, it had more negative feedback from the participants. Overall, the design of this experiment successfully allows us to quantify the effect of VAC and underscores the strength of the evaluation method.
Near-eye light field displays offer natural 3-D visual experiences for AR/VR users by projecting light rays onto retina as if the light rays were emanated from a real object. Such displays normally take four-dimensional light field data as input. Given that sizeable existing 3-D contents are in the form of stereo images, we propose a practical approach that generates light field data from such contents at minimal computational cost while maintaining a reasonable image quality. The perceptual quality of light field is ensured by making the baseline of light field subviews consistent with that of the micro-projectors of the light field display and by compensating for the optical artifact of the light field display through digital rectification. The effectiveness and efficiency of the proposed approach is verified through both quantitative and qualitative experiments. The results demonstrate that our light field converter works for real-world light field displays.
Light field displays deliver realistic 3D content by representing scenes through a collection of light rays, offering an immersive and lifelike viewing experience. Among various implementations, spatial-multiplexing light field displays have gained increasing attention for their ability to provide full parallax without requiring micro-displays with high frame rates or complex light modulation. However, light leakage between adjacent subviews leads to stray light and ghost images. Placing a grid-like blocker on the micro-display can separate subviews but requires high manufacturing precision and is best applicable for emissive displays only. In this paper, we consider reflective displays because of their frame rate and brightness advantages and propose a polarization-alternating beam splitter to direct light and address the crosstalk issue for reflective displays. Experimental results show that the proposed approach reduces crosstalk by 95% while preserving the refocusing effect of light field displays.
Light field technology revolutionizes AR displays from traditional "showing image to each eye" to "projecting light field to each retina" and resolves the vergence-accommodation conflict (VAC) to provide continuous focus and seamless integration of virtual objects with real environments. In this paper, we provide an overview of the basic principles of our light field display and demonstrate its advantages over traditional AR displays for close-range applications such as endoscopic ultrasound probe localization, AR-guided gallbladder drainage, and medical training upper endoscopy. In addition, we describe an optical approach to address the inherent spatial-angular tradeoff of light field displays.
Texture generation is crucial for endoscopic 3D reconstruction, as it provides essential visual information for computer-assisted medical diagnosis and surgical procedures. Previous mapping-based methods for texturing 3D stomach depend on high-quality mesh structures for accurate camera view selection. However, it is challenging to obtain high-quality mesh structures in endoscopic 3D reconstruction. To eliminate this dependency, we propose an alternative texture generation method that extracts texture directly from a neural radiance field, removing the need for camera view selection. Furthermore, since endoscopic images often suffer from uneven lighting including local low light and overexposure, we develop a weight mechanism to guide our model in prioritizing the learning of pixels that clearly depict the stomach wall. Experimental results demonstrate that our method is more robust than previous approaches in texturing 3D stomach models and effectively mitigates lighting artifacts, thereby producing high-fidelity textures that are crucial for downstream tasks.
Near-eye light field displays are superior to conventional AR displays because they offer continuous focus and viewing experiences free from visual accommodation conflict (VAC). However, given a fixed number of pixels for representation of the spatio-angular information of a light field, the inherent tradeoff between angular and spatial resolutions presents a great challenge to the widespread adoption of light field technology. To address the challenge, we propose a hybrid super-resolution framework consisting of a digital neural network and an optical neural network and allowing an end-to-end optimization of the downsampling operation for fitting the light field data into a fixed-resolution display panel and the upsampling operation for enhancing the light field quality, all in the frequency domain. Experimental results show that the proposed hybrid framework is a promising approach to quality enhancement of near-eye light field displays.
Semantic segmentation of basal cell carcinoma (BCC) from full-field optical coherence tomography (FF-OCT) images of human skin has received considerable attention in medical imaging. However, it is challenging for dermatopathologists to annotate the training data due to OCT's lack of color specificity. Very often, they are uncertain about the correctness of the annotations they made. In practice, annotations fraught with uncertainty profoundly impact the effectiveness of model training and hence the performance of BCC segmentation. To address this issue, we propose an approach to model training with uncertain annotations. The proposed approach includes a data selection strategy to mitigate the uncertainty of training data, a class expansion to consider sebaceous gland and hair follicle as additional classes to enhance the performance of BCC segmentation, and a self-supervised pre-training procedure to improve the initial weights of the segmentation model parameters. Furthermore, we develop three post-processing techniques to reduce the impact of speckle noise and image discontinuities on BCC segmentation. The mean Dice score of BCC of our model reaches 0.503±0.003, which, to the best of our knowledge, is the best performance to date for semantic segmentation of BCC from FF-OCT images.
Hand tracking algorithms relying on a single camera as the sensing device can only provide relative depth information, resulting in limited practicality. This limitation underscores the necessity for effective and accurate estimation of the absolute distances between hand joints and the camera in the real world. We respond to this pressing need by introducing a methodology that exploits the autofocus functionality of a camera for hand tracking. It takes advantage of the unutilized potential of a camera and removes the need for additional power-demanding and costly depth sensors to accurately estimate the absolute distances of hand joints. Our methodology undergoes rigorous experimental validation and consistently outperforms traditional methods across different lens positions.
Wearing facial masks has become a must in our daily life due to the global COVID-19 pandemic. However, the performance of a face recognition system is severely degraded due to the fact that the face images in the gallery are unmasked faces while the probe face images captured by the camera are masked faces, making the probe face images different from gallery face images in the activated region and the distribution domain. In this paper, we propose a novel face recognition system to address the issue. The system is integrated with a domain adaptation layer and a feature refinement layer. The feature refinement layer is based on the structure of the self-attention mechanism to align activated regions of unmasked faces with those of masked faces. The domain adaptation layer works by adapting the system from the unmasked face domain to the synthetically masked face domain and the real- world masked face domain. The system is tested on real-world data through face verification and face identification. The face verification accuracy is improved by 6.83% for the RMFD_FV dataset and 4.2% for the MFR2 dataset, and the face identification accuracy is improved by 15.43% for the MFRFI dataset.
Playlist continuation involves adding new songs to a playlist based on information of the songs and their relation to the playlist. However, its performance suffers when the candidate songs or the existing songs of the playlist are cold-start songs. In this paper, we address the information scarcity issue by proposing a hybrid model that integrates collaborative filtering and content-based filtering in a regression framework. Specifically, we consider the music tags as a 2D image and exploit the information embedded in or carried by the music tags as contextual feature vectors, which in turn are used as the basis of song recommendation. Then we apply a switching mechanism to determine which recommendation method to use for the songs. The sequential order of songs in a playlist is preserved to create smooth music listening experience. We evaluate the performance of the proposed model on both editor- and user-defined playlists. The results indicate that the proposed model outperforms competing models for cold-start songs.
Most near-eye displays with one fixed focal plane suffer from the vergence–accommodation conflict and cause visual discomfort to users. In contrast, light field displays can provide natural and comfortable 3D visual sensation to users without the conflict. This paper presents a near-eye light field display consisting of a geometric lightguide and a light field generator, along with a collimator to ensure the light rays propagating in the lightguide are collimated. Unlike most lightguides, which reduce thickness by employing total internal reflection that can easily generate stray light, our lightguide directly propagates light rays without total internal reflection. The partially reflective mirrors of the lightguide expand the exit pupil to achieve an eyebox of 13mm(horizontal)×6.5mm(vertical) with an eye relief of 18 mm. The collimator and the light field generator, both having effective focal lengths different in the horizontal and vertical directions, are designed to provide a 40-deg diagonal field of view. The working range of the light field generator, which is 30 cm to infinity, is verified qualitatively and quantitatively by experiments. We optimize the illuminance uniformity and analyze the illuminance variation across the eyebox. Further, we minimize the ghost artifact (referring to the split-up of light fields replicated by the partially reflective mirrors) by orienting the partially reflective mirrors at slightly different angles to enhance the image quality for short-range applications such as medical surgery.
Most near-eye displays with one fixed focal plane suffer from the vergence-accommodation conflict (VAC) and cause visual discomfort to users. In contrast, a light field display with continuous focal planes offers the most natural and comfortable AR/VR visual experiences without VAC and holds the promise to be the ultimate near-eye 3-D display. It projects light rays onto human retina as if the light rays were emanated from a real object. This paper considers a near-eye light field display comprising a light field generator, a collimator, and a geometric waveguide as the three main components. It takes 4-D light field data in the form of an array of 2-D subview images as input and generates a light field as output. The light field generator is the device responsible for converting the light emitted from the display panel to the light representing the light field of a virtual scene. The geometric waveguide along with a collimator ensures that the light rays propagating in the waveguide are collimated. The partially reflective mirrors of the waveguide replicate the optical path to achieve exit pupil expansion (EPE) and a large eyebox. However, existing waveguide eyepieces for near-eye AR/VR displays are not designed for, and hence may not fit light field displays. In this work, we look into a geometric waveguide for light field display and find that the light fields replicated by the partially reflective mirrors cannot perfectly overlap on the user's retina, resulting in the appearance of multiple repetitive images—a phenomenon we call "ghost artifact". This paper delves into the cause of this artifact and develops a solution for applications that require short-range interaction with virtual objects, such as surgical procedures. We define a working range devoid of noticeable ghost artifact based on the angular resolution characteristics of human eye and optimize the orientation of an array of partially reflective mirrors of the waveguide to meet the image quality requirement for short-range interaction. With the optimized waveguide, the ghost artifact is significantly reduced. More results of the optimized waveguide will be shown at the conference.
Medical image-to-image translation is often difficult and of limited effectiveness due to the differences in image acquisition mechanisms and the diverse structure of biological tissues. This work presents an unpaired image translation model between in-vivo optical coherence tomography (OCT) and ex-vivo Hematoxylin and eosin (H&E) stained images without the need for image stacking, registration, post-processing, and annotation. The model can generate high-quality and highly accurate virtual medical images, and is robust and bidirectional. Our framework introduces random noise to (1) blur redundant features, (2) defend against self-adversarial attacks, (3) stabilize inverse conversion, and (4) mitigate the impact of OCT speckles. We also demonstrate that our model can be pre-trained and then fine-tuned using images from different OCT systems in just a few epochs. Qualitative and quantitative comparisons with traditional image-to-image translation models show the robustness of our proposed signal-to-noise ratio (SNR) cycle-consistency method.
Histopathology for tumor margin assessment is time-consuming and expensive. High-resolution full-field optical coherence tomography (FF-OCT) images fresh tissues rapidly at cellular resolution and potentially facilitates evaluation. Here, we define FF-OCT features of normal and neoplastic skin lesions in fresh ex vivo tissues and assess its diagnostic accuracy for malignancies. For this, normal and neoplastic tissues were obtained from Mohs surgery, imaged using FF-OCT, and their features were described. Two expert OCT readers conducted a blinded analysis to evaluate their diagnostic accuracies, using histopathology as the ground truth. A convolutional neural network was built to distinguish and outline normal structures and tumors. Of the 113 tissues imaged, 95 (84%) had a tumor (75 basal cell carcinomas [BCCs] and 17 squamous cell carcinomas [SCCs]). The average reader diagnostic accuracy was 88.1%, with a sensitivity of 93.7%, and a specificity of 58.3%. The artificial intelligence (AI) model achieved a diagnostic accuracy of 87.6 ± 5.9%, sensitivity of 93.2 ± 2.1%, and specificity of 81.2 ± 9.2%. A mean intersection-over-union of 60.3 ± 10.1% was achieved when delineating the nodular BCC from normal structures. Limitation of the study was the small sample size for all tumors, especially SCCs. However, based on our preliminary results, we envision FF-OCT to rapidly image fresh tissues, facilitating surgical margin assessment. AI algorithms can aid in automated tumor detection, enabling widespread adoption of this technique.
When the light field of a scene is generated with a finite number of subviews, the defocused regions would appear to be split-up if a camera is used to capture the light field. Yet, the split-up effect is unnoticeable when the light field is viewed directly through a human eye. In this paper, we attribute the unobservability of the split-up effect to the decrease in visual acuity as a function of retinal eccentricity and to the low-pass filtering property of visual attention. Theoretical and experimental results are provided to support our claim. Furthermore, we set an observability criterion for the split-up effect and discuss design strategies for performance improvement of light field displays.
Chia-Kai Liang (梁家愷)合作论文数Google Research22