Multispectral filter arrays (MSFAs) are optical components that incorporate pixelated colored filters to selectively filter light at different wavelengths. MSFAs are used in a range of multispectral imaging (MSI) systems, such as mosaic color image sensing and compressive spectral imaging (CSI). The former being cameras with generalized Bayer color filter masks, and the latter where the color filter array is used in concert with a dispersive element to compressively capture spectral images with a 2D focal plane array. Several approaches to optimize the color patterns in MSFAs have been proposed; notably, the placement of the optical elements in most of the proposed approaches is guided by generalizing the constraints defined by Bayer half a century ago. While Bayer dither matrices place spatial constraints in two dimensions, MSFAs generalize these constraints across three dimensions: spatial (2D) and spectral. This paper proposes a new, simple, yet surprisingly effective optimization strategy inspired by the globally recognized combinatorial puzzle Sudoku. The Sudoku-based approach constrains the spatial distribution of colors in the MSFA to form homogeneous patterns rich in high-frequency components across both spatial and spectral dimensions. These properties are desirable for improving image reconstruction. The Sudoku guidelines naturally provide colored patterns that follow blue-noise distributions characterized by the absence of low-frequency components. Notably, if the Sudoku "single digit" constraints along cells, rows, and columns are further generalized to gradually have inter-cell constraints, the resulting colored patterns become increasingly periodic. The proposed framework thus offers mechanisms to generate a broad spectrum of colored patterns, ranging from aperiodic, random blue-noise patterns to periodic Bayer-like patterns. Simulations demonstrate that mosaic and CSI measurements using Sudoku color codes lead to data reconstructions that are comparable to, if not better than, those produced by state-of-the-art MSFA design strategies. (c) 2025 Optica Publishing Group under the terms of the Optica Open Access Publishing Agreement
In graph signal processing, learning weighted connections between nodes from signals is a fundamental task when the underlying relationships are unknown. With the extension of graphs to hypergraphs, where edges can connect more than two nodes, graph learning methods have similarly been generalized to hypergraphs. However, the absence of a unified framework for calculating total variation has led to divergent definitions of smoothness and, consequently, differing approaches to hyperedge recovery. This challenge is confronted in this work through generalization of several previously proposed hypergraph total variations, allowing ease of substitution into a vector-based optimization. To this end, a novel hypergraph learning method is proposed that recovers a hypergraph topology from time-series signals using convex optimization based on a smoothness prior. This approach, designated Hypergraph Structure Learning with Smoothness (HSLS), addresses key limitations in prior works such as hyperedge selection and convergence issues. Additionally, a process is introduced that limits the span of the hyperedge search and maintains a valid hyperedge selection set, creating a scalable model. Experimental results demonstrate improved performance over state-of-the-art hypergraph inference methods. The method is empirically shown to be robust to total variation terms, biased towards global smoothness, and scalable to larger hypergraphs.
In structured light illumination, lens distortions in both the camera and the projector compromise the accuracy of 3D reconstruction. Typically, existing methods separately compensate for camera and projector lens distortion. In this paper, we report a novel joint distortion model that analytically relates distorted 3D coordinate to its undistorted counterpart, thereby directly recovering distortion-free 3D coordinate from distorted one. First, we conduct a typical 3D scanning to have the distorted 3D coordinate. Second, we derive a set of linear equations of undistorted coordinate, whose coefficient matrix is represented by the distorted 3D coordinate and calibration parameters. Finally, we straightforwardly compute the corrected 3D coordinate using the least square method. Extensive experiments show that, compared with the distorted point cloud, our method effectively reduces the lens distortion of the system by a factor of 5 in root mean squared error, outperforming the existing methods in terms of accuracy.
Blue-noise stacked error diffusion (BSED) is a high-quality multi-level halftoning (multitoning) algorithm based on the blue-noise dithering model. This algorithm transforms a multitoning task into multiple error diffusion halftoning tasks, each producing a sub-halftone with inter-layer dependencies. These sub-halftones are stacked together to form the final visually pleasing multitone with blue-noise characteristics. In this research, the significant potential for parallelizing this multitoning algorithm inspires the design of a novel parallel raster image processor (NPRIP) for highly efficient execution, capable of producing an output pixel every six clock cycles at a clock frequency of 200 MHz. A hardware prototype based on the Xilinx Zynq SoC implements our raster image processor design on the FPGA within the SoC. The prototype is also capable of a conventional software implementation of the BSED algorithm running on the ARM CPU within the SoC. Output images from the prototype are successfully validated in both the spatial and frequency domains. Benchmark results demonstrate a consistent speedup exceeding 30x with FPGA co-processing compared to conventional sequential execution on the ARM CPU, reaching a maximum processing speed of 3.963 megapixels per second.
Objectives: There is a growing interest among practitioners in employing artificial intelligence (AI) to enhance the precision and efficiency of diagnostic methods. The objective of this study is to assess the precision of an AI-based method for facial asymmetry assessment using 3D facial images. Methods: The study included 130 patients (84 female, 46 male), analyzing 3D facial images from the Vectra® M3 imaging system using both manual and AI-based methods. Seven bilateral facial landmarks were identified for manual analysis, calculating the asymmetry index for facial symmetry assessment. An AI-based program was developed to automate the identification of the same landmarks and calculate the asymmetry index. The reliability of the manual measurements was assessed using intraclass correlation coefficients (ICC) with 95% confidence intervals (CI). Precision of automated landmark identification was compared to the manual method. Results: The ICCs for the manual measurements demonstrated moderate to excellent reliability, both within raters (ICC = 0.62–0.99) and between raters (ICC = 0.72–0.96) each calculated with 95% CI. Agreement was observed between the manual and automated methods in calculating the asymmetry index for five landmarks. There was a statistically significant difference between the two methods in determining the asymmetry index for alare (median: 2.05 mm manual vs. 1.54 mm automated, p = 0.0056) and cheilion (median: 2.77 mm manual vs. 2.30 mm automated, p = 0.0081). Conclusions: The AI-based method provides efficient and comparable precision of facial asymmetry analysis using 3D images. The disagreement observed between the two methods can be addressed through further improvement and training of the automated software. This innovative approach opens doors to significant advancements in both research and clinical orthodontics.
Structured light is a method of 3D imaging that involves projecting a sequence of striped or structured patterns onto a target scene and then, based on the warping of stripes across a target surface, a camera can reconstruct the target. For practical reasons relating to changing lens parameters for the projector, binocular structured light improves over monocular structured by performing triangulation between two cameras, using the striped patterns to solve the correspondence matching problem; however, the speed at which these correspondences are found is still an active area of research. And in this paper, we propose a fast-matching algorithm that enhances phase labeling and introduces the concept of triangle plane interpolation based on linear interpolation. The experimental results demonstrate that our method achieves a noticeable speed improvement in computational efficiency and a stable improvement in accuracy as compared to the phase matching along epipolar line pointwisely based on look-up tables.
Structured light illumination is an active 3D scanning technique based on projecting and capturing a set of striped patterns and measuring the warping of the patterns as they reflect off a target object's surface. As designed, each pixel in the camera sees exactly one pixel from the projector; however, there are multi-path situations where a camera pixel sees light from multiple projector positions. In the case of bimodal multi-path, the camera pixel receives light from exactly two positions, which occurs along a step edge where the edge slices through a pixel which, therefore, sees both a foreground and background surface. In this paper, we present a general mathematical model to address this bimodal multi-path issue in a phase-shifting or so-called phase-measuring-profilometry scanner to measure the constructive and destructive interference between the two light paths, and by taking advantage of this interference, separate the paths and make two decoupled depth measurements. We validate our algorithm with both simulations and a number of challenging real-world scenarios, significantly outperforming the state-of-the-art methods.
Balancing speed and accuracy has always been a challenge in 3D reconstruction. One-shot structured light illuminations are of perfect performance on real-time scanning, while the related 3D point clouds are typically of relatively poor quality, especially in regions with rapid height changes. To solve this problem, we propose a one-shot reconstruction scheme based on shearlet transform, which combines spatial and frequency domain information to enhance reconstruction accuracy. First, we apply the shearlet transform to the deformed fringe pattern to obtain the transform coefficients. Second, pixel-wise select the indices associated with the N largest coefficients in magnitude to obtain a new filter. Finally, we refocus globally to extract phase using these filters and generate a reliable quality map based on coefficient magnitudes to guide phase unwrapping. Simultaneously, we utilize the maximum coefficient value to generate a quality map for guiding the phase unwrapping process. Experimental results show that the proposed method is robust in discontinuous regions, resulting in more accurate 3D point clouds.
Representation learning considering high-order relationships in data has recently shown to be advantageous in many applications. The construction of a meaningful hypergraph plays a crucial role in the success of hypergraph-based representation learning methods, which is particularly useful in hypergraph neural networks and hypergraph signal processing. However, a meaningful hypergraph may only be available in specific cases. This paper addresses the challenge of learning the underlying hypergraph topology from the data itself. As in graph signal processing applications, we consider the case in which the data possesses certain regularity or smoothness on the hypergraph. To this end, our method builds on the novel tensor-based hypergraph signal processing framework (t-HGSP) that has recently emerged as a powerful tool for preserving the intrinsic high-order structure of data on hypergraphs. Given the hypergraph spectrum and frequency coefficient definitions within the t-HGSP framework, we propose a method to learn the hypergraph Laplacian from data by minimizing the total variation on the hypergraph (TVL-HGSP). Additionally, we introduce an alternative approach (PDL-HGSP) that improves the connectivity of the learned hypergraph without compromising sparsity and use primal-dual-based algorithms to reduce the computational complexity. Finally, we combine the proposed learning algorithms with novel tensor-based hypergraph convolutional neural networks to propose hypergraph learning-convolutional neural networks (t-HyperGLNN).
In this paper, we propose an adaptive method for phase measuring profilometry to more accurately reconstruct 3D point clouds of high-reflective surfaces. First, we project high-frequency sinusoidal fringe patterns to robustly detect saturated pixels and then obtain their corresponding optimal projection intensity. Second, based on the coordinate mapping relationship between a camera-projector pair, we generate a set of fringe patterns with adaptive intensity. Finally, we project the adaptive fringes patterns to achieve more accurate 3D reconstruction. The experimental results show that our method can more accurately detect saturated pixels, and thus improve the quality of 3D reconstruction.
Graph signal processing (GSP) extends classical signal processing methods to analyzing signals supported over irregular grids represented by graphs. Within the scope of GSP, sampling and reconstruction represent fundamental tools that have received considerable attention. For very large graphs, however, many of the current methods struggle with the computational and memory requirements. Vertex domain and randomized sampling strategies somewhat ameliorate the computational requirements, but these algorithms perform poorly at preserving signal fidelity for graphs with hundreds of thousands of vertices. To address these shortcomings, this article introduces a new and scalable approach that can be easily parallelized. This new approach uses existing graph partitioning algorithms in concert with vertex-domain blue-noise sampling and reconstruction, performed independently across partitions. In the reconstruction, some degree of overlapping is added to the partitions to induce trans-partition smoothness in the recovered signal. We also propose two sampling schemes based on the spatial characteristic of the graph that minimizes the recovery error. The first combines graph partitioning with the Void-and-Cluster algorithm, while the second approach uses Error Diffusion. We conclude this article with experiments on synthetic and real data that show the effectiveness of these new approaches on very large graphs.
Graph signal processing (GSP) techniques are powerful tools that model complex relationships within large datasets, being now used in a myriad of applications in different areas including data science, communication networks, epidemiology, and sociology. Simple graphs can only model pairwise relationships among data which prevents their application in modeling networks with higher-order relationships. For this reason, some efforts have been made to generalize well-known graph signal processing techniques to more complex graphs such as hypergraphs, which allow capturing higher-order relationships among data. In this paper, we provide a new hypergraph signal processing framework (t-HGSP) based on a novel tensor-tensor product algebra that has emerged as a powerful tool for preserving the intrinsic structures of tensors. The proposed framework allows the generalization of traditional GSP techniques while keeping the dimensionality characteristic of the complex systems represented by hypergraphs. To this end, the core elements of the t-HGSP framework are introduced, including the shifting operators and the hypergraph signal. The hypergraph Fourier space is also defined, followed by the concept of bandlimited signals and sampling. In our experiments, we demonstrate the benefits of our approach in applications such as clustering and denoising.
In fringe projection profilometry, applying pre-distortion to fringe patterns reduces the errors caused by projector lens distortion. However, it is important to note that discontinuous fringe patterns, such as binary fringe patterns, introduce additional errors when using pre-distortion methods. While post-undistortion methods are applicable for discontinuous fringe patterns, the computation is typically time-consuming. We propose a linear-grid model for correcting lens distortion. First, we select multiple equidistant points within the grid to calculate the linear parameters and store them as look-up tables (LUTs). Second, by rounding down the captured distorted point to the nearest integer point, we obtain the index value for LUTs. Finally, we achieve real-time compensation for distortion error through linear expressions. The experimental results show that the proposed effectively mitigates the distortion by a factor of 6x in terms of root mean squared error. Additionally, it exhibits a computational speed of 409.50 fps, which is an improvement compared to the traditional iterative model at 39.48 fps and the scale-offset model at 264.48 fps.
A novel reconstruction method for compressive spectral imaging is designed by assuming that the spectral image of interest is sufficiently smooth on a collection of graphs. Since the graphs are not known in advance, we propose to infer them from a panchromatic image using a state-of-the-art graph learning method. Our approach leads to solutions with closed-form that can be found efficiently by solving multiple sparse systems of linear equations in parallel. Extensive simulations and an experimental demonstration show the merits of our method in comparison with traditional methods based on sparsity and total variation and more recent methods based on low-rank minimization and deep-based plug-and-play priors. Our approach may be instrumental in designing efficient methods based on deep neural networks and covariance estimation.
Quick Response (QR) codes usage in e-commerce is on the rise due to their versatility and ability to connect offline and online content, taking over almost every aspect of a business from posters to payments. Thus, many efforts have aimed at improving the visual quality of QR codes to be easily included in publicity designs in billboards and magazines. The most successful approaches, however, are slow since optimization algorithms are required for the generation of each beautified QR code, hindering its online customization. The aim of this paper is the fast generation of visually pleasant and robust QR codes. The proposed framework leverages state-of-the-art deep-learning algorithms to embed a color image into a baseline QR code in seconds while keeping a maximum probability of error during the decoding procedure. Halftoning techniques that exploit the human visual system (HVS) are used to smooth the embedding of the QR code structure in the final QR code image while reinforcing the decoding robustness. Compared to optimization-based methods, our framework provides similar qualitative results but is 3 orders of magnitude faster.
In fringe projection profilometry, inevitable distortion of optical lenses decreases phase accuracy and decreases the quality of 3D point clouds. For camera lens distortion, existing compensation methods include real time look-up tables derived from the related parameters of camera calibration. However, for projector lens distortion, so far, post-undistortion methods iteratively correcting lens distortion are relatively time-consuming while, despite avoiding iteration, pre-distortion methods are not suitable for binary fringe patterns. In this paper, we aim to achieve real-time phase correction for the projector by means of a scale-offset model that characterizes projector distortion by four correction parameters within a small-enough area, and thus we can speed up the post-undistortion by looking up tables. Experiments show that the proposed method can suppress the distortion error by a factor of 20 ×, i.e., the error of root mean square is less than 45 µm/0.7‰, while also proposed improving the computation speed by a factor of 50× over traditional iterative post-undistortion.
Structured light illumination is a 3D scanning technique based on projecting a series of striped patterns and reconstructing depth based on the observed warping of the pattern across the target surface. Extensively studied, real-time performance is of paramount importance in many applications. In this Letter, we build on prior researches with epipolar geometry to further simplify the computations of 3D point clouds, with the symmetrical epipolar features derived by extending epipolar geometry over normalized calibration matrices of the camera and projector. Experiments show that the proposed two new processes are of the same accuracy over the normalized calibration matrices with substantially fewer calculations.
Abstract Linking root traits to plant functions can enable crop improvement for yield and ecosystem functions. However, plant breeding efforts targeting belowground traits are limited by appropriate phenotyping methods for large root systems. While advances have been made allowing for imaging large in situ root systems, many of these methods are inaccessible due to expensive technology requirements. The aim of this work was to develop a plant phenotyping platform and analysis method suitable for assessing root traits of large, intact root systems. With the use of a purpose‐built imaging table and automated photo capture system, machine learning‐based image segmentation, and off‐the‐shelf trait analysis software, the developed method yielded results of comparable accuracy to commercial root scanning platforms without requiring access to prohibitively expensive equipment. This methodology enables root studies to move beyond the size limitations of scanner‐based methods, integrate whole‐system traits like root depth distribution, and save time on root image capture.
Quick Response (QR) codes are widely used to connect offline and online content, and thus many efforts have aimed at improving the visual quality of QR codes to be easily included in publicity designs in billboards and magazines. The most successful approaches, however, are slow since optimization algorithms are required for the generation of each beautified QR code, hindering its online customization. The aim of this paper is the fast generation of visually pleasant and robust QR codes. The proposed framework leverages state-of-the-art deep-learning algorithms to embed a color image into a baseline QR code in seconds while keeping a maximum probability of error during the decoding procedure. Halftoning techniques that exploit the human visual system (HVS) are used to smooth the embedding of the QR code structure in the final QR code image while reinforcing the decoding robustness. Compared to optimization-based methods, our framework provides similar qualitative results but is significantly faster.
Efficient sampling of graph signals is essential to graph signal processing. Recently, blue-noise was introduced as a sampling method that maximizes the separation between sampling nodes leading to high-frequency dominance patterns, and thus, to high-quality patterns. Despite the simple inter-pretation of the method, blue-noise sampling is restricted to approximately regular graphs. This study presents an extension of blue-noise sampling that allows the application of the method to irregular graphs. Before sampling with a blue-noise algorithm, the approach regularizes the weights of the edges such that the graph represents a regular structure. Then, the resulting pattern adapts the node's distribution to the local density of the nodes. This work also uses an approach that minimizes the strength of the high-frequency components to recover approximately bandlimited signals. The experimental results show that the proposed methods have superior performance compared to the state-of-the-art techniques.