Image quality in modern imaging systems emerges from the coupled effects of the sensor, optics, and computational reconstruction. Ultra-thin metalenses offer a path toward substantial miniaturization of optical modules, but practical designs often exhibit pronounced chromatic and field-dependent aberrations that necessitate computational reconstruction. In current metalens pipelines, reconstruction models are commonly trained and selected using distortion-based fidelity objectives, such as PSNR, yet these proxies can be weakly correlated with human preference and downstream utility, reflecting the well-known perception–distortion trade-off. We introduce MetaRanker, a human-in-the-loop active ranking framework that formalizes metalens image quality in terms of semantic interpretability, defined as the degree to which humans can reliably recognize objects and structures in the presence of optical artifacts. MetaRanker combines a probabilistic preference model with uncertainty-aware query selection, and leverages vision–language models to provide lightweight semantic priors. Importantly, these priors are used only to guide the sampling of informative comparisons; human judgments remain the primary supervision signal throughout. Across real-world and synthetic metalens datasets with distinct degradation profiles, MetaRanker produces rankings that align most closely with human assessments, while reducing the number of pairwise annotations required by approximately 80
Confidence calibration is crucial for accurate and reliable ordinal classification, yet it remains largely overlooked, with existing calibration studies rarely addressing the unique challenges posed by ordered class labels. We introduce Margin-based Ordinal Classification with Dynamic Regularization for Calibration and Unimodality (MORCU). It combines dynamic log-barrier regularization to enforce structured probability distributions with our Target-Preserving Margin Penalty (TPMP), a newly introduced approach that refines adjacent non-target logits to promote calibration and unimodality. By adaptively balancing structural constraints and confidence estimation, MORCU mitigates both overconfidence and underconfidence, producing well-calibrated probability distributions aligned with ordinal relationships. Experimental results across diverse benchmark datasets demonstrate consistent calibration gains and competitive ordinal classification performance, making it well-suited for applications requiring both predictive accuracy and trustworthy confidence estimation. The code is publicly available at https://github.com/labhai/MORCU.
Adjoint optimization is a cornerstone of broadband nanophotonic inverse design, but conventional time-domain implementations face a severe memory bottleneck because they retain forward-field histories at every finite-difference time-domain (FDTD) time step. Here, we show that this full time-step storage is unnecessary for band-limited design objectives. By storing forward fields only at Nyquist-compliant temporal intervals and using the resulting sparse field history during the reverse-time adjoint pass, the proposed method enables on-the-fly gradient accumulation without retaining full forward- or adjoint-field histories. This Nyquist-sampled adjoint FDTD framework preserves the two-simulation scaling of time-domain adjoint optimization while substantially reducing the dominant field-storage cost. Gradient verification confirms that Nyquist-compliant sampling reproduces conventional full-storage adjoint gradients with negligible error, whereas undersampling beyond the Nyquist limit produces aliasing-induced gradient degradation. Across four two-dimensional broadband nanophotonic benchmarks and a fully three-dimensional metalens, the method maintains gradient fidelity and optimized device performance while reducing dominant field-storage memory by up to 107x. These results suggest that the principal memory barrier in broadband time-domain adjoint FDTD is not an intrinsic requirement of gradient evaluation, but a consequence of redundant temporal field storage, opening a practical route to large-scale three-dimensional nanophotonic inverse design.
Conventional machine-vision pipelines typically rely on high-quality optics that produce clean, human-interpretable images, and optical design has therefore been driven by image-level criteria such as resolution, aberration correction, and pixel fidelity. However, such optics are often impractical for size-, cost-, or form-factor-constrained applications, where compact meta-optics offer an attractive alternative but operate under strict physical efficiency limits. We propose CODA, a co-design framework that optimizes a continuous-density meta-optic front-end for frozen-model recognition using differentiable image formation and adjoint-gradient updates of Maxwell-based simulations. CODA directly optimizes the cross-entropy loss of a fixed zero-shot CLIP classifier without learned reconstruction, image signal processing, or image-fidelity auxiliary objectives. In a two-dimensional simulated imaging benchmark on ImageNet-100, CODA improves CLIP ViT-L/14 zero-shot accuracy from 53.75 ± 3.57% with a focal-concentration baseline to 65.41 ± 3.99%. The optimized optics further transfer without re-optimization across CLIP, SigLIP, and DINOv2 on ImageNet-100, CIFAR-100, and Food-101. These results demonstrate that, under constrained meta-optic imaging, downstream recognition can be improved by aligning optical design with frozen vision-model objectives rather than conventional image-formation criteria.
Optical skyrmions are structured vector fields with nontrivial polarization topology and subwavelength-scale features. One common approach to generating optical skyrmions is the superposition of a zeroth-order Bessel beam and a higher-order Bessel beam carrying orbital angular momentum, with each beam possessing an orthogonal circular polarization state. However, creating such complex beams typically requires bulky free-space optical setups; therefore, recent efforts have focused on compact optical skyrmion generators based on metasurfaces. Nevertheless, achieving the degrees of freedom required for simultaneous phase and polarization control remains challenging because of the limited design flexibility of conventional meta-atoms. Here, we address this challenge by employing an inverse-design approach and demonstrate a single-layer metasurface that generates high-fidelity optical skyrmions. We employ an adjoint-based topology-optimization method to design a silicon metasurface that converts an incident beam into an optical skyrmion without the need for additional optical components. The optimized metasurface generates an optical skyrmion with skyrmion number (N_sk) = 0.970. This work demonstrates that inverse design can be a promising route to compact skyrmion generators, and our approach provides a basis for near-field particle manipulation and the generation of independent topological bits in dense photonic integration.
Optical vortex beams are of interest for a variety of photonic applications. One approach to generating optical vortices exploits spin-orbit coupling within light propagation in anisotropic media, where the polarization state of light is converted into orbital angular momentum. However, due to low birefringence of conventional anisotropic media, it is required to use bulky crystals to obtain high efficiency. In this context, van der Waals (vdW) crystals emerge as a promising candidate for this mechanism because of their large birefringence and low optical loss. In this work, we investigate optical vortex generation using 8- mu m-thick-hexagonal boron nitride and 23- mu m-thick-molybdenum disulfide (MoS2) crystals in a transmission configuration. Additionally, the efficiency of spin-orbit coupling is quantitatively analysed by varying the beam waist. For the MoS2 crystal, an experimental conversion efficiency reaches up to 0.46, close to the theoretical limit. These findings show the optical vortex generation based on micrometer-thick vdW crystals, providing an effective platform using simple method and without the necessity of nanofabrication.
Adjoint-based topology optimization enables gradient computation for electromagnetic design from only two simulations, independent of problem size. Conventional frequency-domain adjoint methods suit single-frequency objectives but incur computational costs scaling linearly with spectral resolution for broadband design. Time-domain adjoint methods efficiently capture broadband responses, however, their native gradients integrate over the entire excitation bandwidth, preventing independent control of multiple spectral bands. Consequently, multi-band optimization requires separate forward-adjoint simulation pairs per band, eliminating the computational efficiency advantage.We present a multi-objective time-domain adjoint method that computes band-selective gradients via temporal convolution of stored electromagnetic fields. By the convolution theorem, post-processing with band-pass filters isolates designated spectral content without additional simulations. We validate the method on three nanophotonic systems. First, a wavelength-division demultiplexer achieves 8.74s per iteration, outperforming conventional time-domain (14.32s) and frequency-domain (120s) methods. Second, a metalens demonstrates independently prescribed numerical apertures across four spectral bands with less than 5% focal length error. Third, a spectral router achieves 70%–75% routing efficiency, approximately 2× the theoretical limit of absorption-based designs. These results establish that convolution-enabled, band-selective gradients restore the computational advantage of time-domain adjoint optimization for multi-band electromagnetic design.
Medical tabular data are ubiquitous in clinical research, but deep learning for tables remains underexplored because reliable labels often require costly expert adjudication, even though structured clinical variables are routinely available in tabular form. Self-supervised learning can leverage these unlabeled tables, and recent binning-based pretexts offer a promising inductive bias, but existing objectives fix a single global quantile discretization and apply feature-agnostic supervision. We propose Adaptive Binning, a training-adaptive discretization pretext for tabular SSL that couples discretization to learning through a feature-wise coarse-to-fine curriculum. Motivated by the spectral bias of neural networks and the principles of curriculum learning, our method progressively refines discretization per feature upon plateau detection and selects representation-aware splits to jointly improve value-space concentration and representation-space coherence. A heterogeneity-aware objective unifies categorical reconstruction with ordinal supervision for numerical features, and experiments on public medical tabular datasets under unified evaluation protocols show consistent gains for linear probing and fine-tuning without dataset-specific discretization tuning. We further introduce a medical tabular SSL benchmark with standardized protocols to support reproducible progress in this underexplored domain. Our code is available at https://github.com/labhai/Adaptive-Binning.
Pairwise comparison labeling is emerging as it yields higher inter-rater reliability than conventional classification labeling, but exhaustive comparisons require quadratic cost. We propose Dodgersort, which leverages CLIP-based hierarchical pre-ordering, a neural ranking head and probabilistic ensemble (Elo, BTL, GP), epistemic–aleatoric uncertainty decomposition, and information-theoretic pair selection. It reduces human comparisons while improving the reliability of the rankings. In visual ranking tasks in medical imaging, historical dating, and aesthetics, Dodgersort achieves a 11–16 × more ranking information per comparison than baselines, yielding Pareto-optimal accuracy–efficiency trade-offs.
Optical logic gates offer a promising route toward high-speed, energy-efficient photonic computing. Still, conventional intensity-encoded architectures often require a reference or bias light to implement nontrivial Boolean functions such as NAND and NOR. Here, we propose an inverse-designed optical logic gate that encodes binary information in guided spatial modes rather than optical intensity. In the proposed scheme, the logical states “0” and “1” are represented by the TE$_{00}$ and TE$_{10}$ modes of a multimode silicon waveguide, enabling both logic states to carry optical power and eliminating the need for an external reference beam. Using full three-dimensional adjoint-based topology optimization, we design six two-input, single-output Boolean gates on a silicon-on-insulator platform, including AND, OR, XOR, NAND, NOR, and NXOR, within $4 \times 4~\mu \rm {m}^{2}$ footprint. The optimized devices realize the target truth-table behavior by selectively transmitting the desired optical modes while suppressing unwanted mode crosstalk, ensuring that each input combination yields the correct logical output (0,1) for the corresponding optical modes (TE$_{00}$, TE$_{10}$). The XOR gate maintains a signal-to-noise ratio(SNR) above 5 dB across the 1.5–1.6 $\mu$m wavelength band despite being optimized only at 1.55 $\mu$m, owing to distributed multimode interference rather than narrowband resonance. We also show that the same physical structure can switch between XOR and NXOR operations by tuning only the relative phase between the two input signals. These results demonstrate compact, broadband, and reconfigurable mode-encoded Boolean logic elements while identifying the power and phase requirements that must be addressed for their extension toward multistage photonic circuits.
Designing free-form photonic devices is fundamentally challenging due to the vast number of possible geometries and the complex requirements of fabrication constraints. Traditional inverse-design approaches-whether driven by human intuition, global optimization, or adjoint-based gradient methods-often involve intricate binarization and filtering steps, while recent deep-learning strategies demand prohibitively large numbers of simulations (105-106). To overcome these limitations, we present AdjointDiffusion, a physics-guided framework that integrates adjoint sensitivity gradients into the sampling process of diffusion models. AdjointDiffusion begins by training a diffusion network on a synthetic, fabrication-aware dataset of binary masks. During inference, we compute the adjoint gradient of a candidate structure and inject this physics-based guidance at each denoising step, steering the generative process toward high-Figure of Merit (FoM) solutions without requiring meticulous binarization or filtering. We show that our method achieves approximately 15% higher FoM at equal simulation cost compared to state-of-the-art nonlinear optimizers (e.g., Method of Moving Asymptotes (MMA), Sequential Least-Squares Quadratic Programming (SLSQP)), or requires about 3x fewer simulations to reach the same FoM, all while ensuring fabrication-aware manufacturability. Compared to pure deep-learning approaches, our method requires similar to 103x fewer simulations. By eliminating complex binarization schedules and minimizing simulation overhead, AdjointDiffusion offers a simulation-efficient and fabrication-aware inverse-design algorithm with the nonconvex optimization capabilities of deep learning. Our open-source implementation is available at https://github.com/dongjin-seo2020/AdjointDiffusion.
Post-hoc calibration corrects reported confidence, yet a multiclass calibrator can also change the associated top-1 prediction. Accuracy captures only the net effect of these changes on correctness, not how often predictions change; the Top-1 Prediction Change Rate (TPCR) instead measures this frequency. We propose Calibrator-Output Repair for Top-1 Decision Preservation (CORD), the first post-fit adapter to impose exact prediction preservation by repairing the full calibrated probability vector. From the original and calibrated outputs alone, CORD determines the mass assigned to the original top-1. The calibrated conditional distribution allocates the remaining mass over the other classes, yielding a repaired vector whose own argmax recovers the original prediction. On the calibration split, CORD coordinates the repaired masses to retain the calibrated outputs' mean mass on original predictions whenever attainable. The adapter alters neither the fitted calibrator nor its direct output, fits no additional supervised map, and requires no user- or validation-tuned hyperparameter. Across CIFAR-10/100 and ImageNet-1K, CORD attains zero TPCR by construction and lowers mean ECE, NLL, and Brier relative to the corresponding direct outputs in every dataset; paired gains persist under distribution shift and across calibration-set sizes. CORD thus removes the preservation constraint from calibrator fitting and assigns exact recovery of the original decision to subsequent output repair. Our code is available at https://github.com/labhai/ORCU.
Optical spin-orbit coupling provides a promising, fabrication-free route for developing ultra-compact optical vortex generators. However, the conversion efficiency has been theoretically limited to 0.5. Here, we demonstrate enhanced vortex generation efficiency by employing a Bessel beam as the input and propagating it through van der Waals (vdW) crystals. The large birefringence of vdW crystals and the single transverse wave vector of a Bessel beam allow a near unity spin-orbit conversion efficiency and a topological charge transition of $\ell \rightarrow \ell + 2$. Through combined analytical and experimental investigations, we demonstrate a conversion efficiency of up to 0.82 in hexagonal boron nitride (hBN) crystals with a thickness of $27.4\,μ\mathrm{m}$. The higher efficiency of Bessel input beams over Gaussian beams is attributed to their distinct transverse wave vector distribution of constituent plane wave components. Furthermore, we demonstrate the dependence of conversion efficiency on the numerical aperture (NA) of the objective lens, which is in good alignment with theoretical predictions. These demonstrations provide a fabrication-free route to highly efficient optical vortex generation via microscale vdW materials platforms.
ABSTRACT Nanophotonic color routers promise higher photon utilization than absorptive color filters for complementary metal‐oxide‐semiconductor (CMOS) image sensors, yet most reported designs remain optimized for normal incidence, with efficiency dropping to about half of its peak by . Here, we present an angle‐specific inverse‐design framework in which the sensor plane is partitioned into local angular zones, and each zone is assigned a dedicated color‐router unit cell. To enable practical assembly of these heterogeneous unit cells, we introduce an optical structural similarity (OSS) constraint that enforces permittivity‐level continuity between neighboring cells during topology optimization without additional full‐wave simulations. In a three‐dimensional freeform library, the optimized unit cells achieve average optical efficiencies of 91.1%, 80.9%, and 83.9% for the red, green, and blue channels, respectively, with optical cross talk below 2.0%. OSS reduces the average stitching error from 24.6% to 5.99%, suggesting that it may provide an effective means of suppressing stitching errors in metasurface designs based on the locally periodic approximation. The framework further extends to a five‐layer architecture with a comparable stitching error of 5.65%, establishing a practical route to oblique‐incidence‐robust, library‐based color routing and, more broadly, to large‐area metasurfaces requiring spatially varying functionality with robust inter‐cell compatibility.
With continued pixel miniaturization, conventional micro‐lenses and absorptive color filters in CMOS image sensors suffer from significant performance degradation and low photon influx as the pixel pitch approaches the diffraction limit. While numerous studies have explored nanophotonic color routers as a solution, most lack a systematic analysis of optical efficiency under broad incident angles, which is critical for practical CIS applications. In this work, an inverse design framework based on automatic differentiation is developed to achieve robustness against incident angle variations. The design maintains a high optical efficiency of 78% on average within a ±12° incident angle range for unpolarized light.The trade‐off relation between the acceptance angle range and the achievable optical efficiency of the device is also identified. The proposed optimized method is then demonstrated and can be further extended to design broadband broad‐angle color routers. These findings pave the way for more efficient and versatile CMOS image sensors, offering substantial improvements over existing methods by ensuring efficient light utilization under realistic illumination conditions.
Multi-view imaging, such as mammography and chest radiography, is a standard component of clinical practice. However, medical images are often unregistered and contain view-specific artifacts or irrelevant background cues that can obscure diagnostically relevant findings. Many existing methods directly fuse per-view representations, allowing such irrelevant content to contaminate the fused embedding and reducing robustness under varying view configurations. We propose OTCHA, a confidence-aware latent hub token alignment module based on optimal transport (OT) that refines patch tokens before fusion for multi-view classification. OTCHA introduces a set of learnable latent hub tokens shared across views. For each view, we compute an OT plan between patch tokens and hub tokens that jointly considers feature similarity and geometry, and augment the OT formulation with token-conditional dustbins to enable partial matching and discard irrelevant tokens. The resulting transport plan provides token-wise matching confidence, which gates hub-mediated message passing and weights a novel optimal-transport-based representation alignment loss to stabilize refinement. Experiments on three multi-view medical image datasets demonstrate consistent improvements over competing baselines across diverse anatomies and view configurations. Our code is available at https://github.com/labhai/OTCHA.
Detecting terrestrial exoplanets in the habitable zones of nearby stars remains a critical challenge. Such planets can be 10^8 to 10^10 times fainter than their host stars and lie at diffraction-limited angular separations, where starlight strongly obscures the companion signal. Here we present an adaptive quantum measurement method for estimating the number, positions, and brightnesses of mutually incoherent point sources in the sub-Rayleigh, ultra-high-contrast regime, operating at contrasts down to 10^-8 – five orders of magnitude beyond previous quantum imaging approaches to exoplanet detection. The method adopts a spatial-mode basis that is updated to maximize the quantum Fisher information per detected photon. Estimation is performed by maximum likelihood in log-brightness coordinates, and the source count is determined by Bayesian-information-criterion (BIC) model selection directly from photon-count statistics, without a tunable detection threshold. For point sources within sub-Rayleigh separations and with brightness ratios spanning eight orders of magnitude, the method reconstructs complete scenes with a mean success rate of 72.5%. Furthermore, it is robust to misalignment, maintaining a 71.3% success rate under offsets of up to six pixels. These results demonstrate that terrestrial exoplanets can be detected below the Rayleigh limit, a regime previously inaccessible to direct imaging.
Inverse design of thin-film lithium niobate (TFLN) photonic devices is computationally demanding because optical birefringence and fabrication-induced slanted sidewalls generally require three-dimensional electromagnetic models. We introduce a birefringent effective-index (BEI) method to reduce this problem to two dimensions while retaining polarization-dependent slab confinement and a representative cross section of the etched geometry. The method is integrated with adjoint topology optimization and fabrication constraints to design a 30 x 10 um demultiplexer that routes 1550 and 775 nm light to separate output ports. Quantitative comparisons with three-dimensional finite-difference time-domain simulations establish the accuracy and etch-depth dependence of the reduced model. The fabricated device provides mean signal-to-crosstalk ratios of 13.9 dB across 1540-1560 nm and 13.3 dB across 770-780 nm. A two-stage cascaded configuration increases the output extinction ratio to 26.8 dB in the telecom band and 20.3 dB in the near-visible band. This fabrication-aware reduced-dimensional approach enables optimizations of the multifunctional photonic devices for nonlinear optical and quantum applications on the TFLN platform.
Multi-view learning often struggles to effectively leverage images captured from diverse angles and locations. Learning methods for unstructured multi-view images remain largely underexplored. We propose a novel Hierarchical Mutual Distillation for Multi-View Fusion (HMDMV) method, which can handle both structured and unstructured multi-view scenarios. It makes predictions utilizing all possible view combinations: single view, partial multi-view, and full multi-view. The method generates predictions for each view combination and then applies hierarchical mutual distillation to enhance inter-view consistency. An uncertainty-based weighting mechanism further refines the fusion process by adjusting the influence of each view combination according to its prediction confidence, reducing the impact of low-confidence views. Extensive experiments on large-scale structured and unstructured datasets demonstrate that HMDMV consistently achieves state-of-the-art classification accuracy. Another unique advantage of HMDMV is that it provides improved flexibility in inference, allowing for more or fewer view counts in inference than those used in training without additional processing. We also provide a light version with reduced training cost by designing an efficient strategy that randomly samples subsets of view combinations during each training iteration. These results highlight HMDMV's robustness in real-world settings where view availability is variable or incomplete. The code is available at https://github.com/labhai/HMDMV.
Accurate segmentation of thin, tortuous anatomical structures, such as retinal vessels, cerebral vasculature, and facial wrinkles, remains challenging due to low contrast, frequent discontinuities, and severe class imbalance. Although recent convolutional and Transformer-based models have improved performance, they often yield fragmented predictions and fail to recover fine branches. We propose CSWinUNETR, a general-purpose backbone for 2D and 3D thin-structure segmentation. It employs cross-shaped stripe self-attention to model long-range principal-axis context and incorporates cyclic shifts to enhance information exchange across stripes. To better preserve fine-grained details, we further introduce a detail-enhanced multi-scale self-attention module that aggregates contextual features from multi-resolution representations. In addition, we propose sparse-control dynamic snake convolution, which reconstructs reliable dense curvilinear kernels from sparsely predicted control points to better follow tortuous geometry. Extensive experiments on four benchmarks across ophthalmology, neurovascular imaging, and dermatology demonstrate that CSWinUNETR consistently outperforms state-of-the-art methods without task-specific post-processing or topology-aware losses. The code is available at https://github.com/labhai/CSWinUNETR.