In most real-world image-to-image (I2I) scenarios, existing evaluations primarily focus on instruction following and the perceptual quality or aesthetics of the generated images. However, they largely fail to assess whether the output image preserves the semantic correspondence and spatial structure of the input image. To address this limitation, we propose StableI2I, a unified and dynamic evaluation framework that explicitly measures content fidelity and pre--post consistency across a wide range of I2I tasks without requiring reference images, including image editing and image restoration. In addition, we construct StableI2I-Bench, a benchmark designed to systematically evaluate the accuracy of MLLMs on such fidelity and consistency assessment tasks. Extensive experimental results demonstrate that StableI2I provides accurate, fine-grained, and interpretable evaluations of content fidelity and consistency, with strong correlations to human subjective judgments. Our framework serves as a practical and reliable evaluation tool for diagnosing content consistency and benchmarking model performance in real-world I2I systems.
Tropical cyclones (TCs) three-dimensional (3D) structures are crucial for understanding their intensification processes and assessing associated risks. However, observational data remain sparse, especially for complete dynamic and thermodynamic variables. Dropsondes are considered one of the best available in situ observations providing high-quality vertical profiles, but they are extremely sparse in spatial distribution. This study constructs a physics-guided generative AI framework capable of reconstructing complete 3D TC fields, including wind, temperature, and humidity, from sparse dropsonde observations. Our method utilizes a score-based diffusion model, pretrained on global climate simulation data and fine-tuned on high-resolution operational analysis fields, to learn priors of TC structures. By integrating this generative prior with a score-based posterior sampling algorithm and imposing physical constraints of divergence, vorticity, and thermodynamic consistency, we obtain physically consistent TC 3D reconstructions. Results from systematic Observing System Simulation Experiments (OSSEs) and extensive real-world operational cases indicate that our method can reconstruct the complete TC 3D dynamic and thermodynamic structures, providing a data-driven pathway for reconstructing high-dimensional atmospheric states from sparse in situ data.
Multimodal large language models (MLLMs) have achieved remarkable progress in visual understanding tasks such as visual grounding, segmentation, and captioning. However, their ability to perceive perceptual-level image features remains limited. In this work, we present UniPercept-Bench, a unified framework for perceptual-level image understanding across three key domains: Aesthetics, Quality, Structure and Texture. We establish a hierarchical definition system and construct large-scale datasets to evaluate perceptual-level image understanding. Based on this foundation, we develop a strong baseline UniPercept trained via Domain-Adaptive Pre-Training and Task-Aligned RL, enabling robust generalization across both Visual Rating (VR) and Visual Question Answering (VQA) tasks. UniPercept outperforms existing MLLMs on perceptual-level image understanding and can serve as a plug-and-play reward model for text-to-image generation. This work defines Perceptual-Level Image Understanding in the era of MLLMs and, through the introduction of a comprehensive benchmark together with a strong baseline, provides a solid foundation for advancing perceptual-level multimodal image understanding.
Tropical cyclone (TC) inner-core surface wind vectors underpin intensity forecasting and storm-surge prediction, yet direct observations remain scarce: routine aircraft reconnaissance is confined to the North Atlantic and Eastern Pacific and, even there, samples each storm only episodically. CYGNSS is the only satellite that penetrates heavy precipitation to measure inner-core surface winds, but delivers directionless scalar wind speeds and is assimilated by no operational analysis system. Here we show that the full 10 m vector wind field inside the TC inner core can be reconstructed globally at 1.5 km resolution from sparse CYGNSS scalar observations alone, by generalising score-based diffusion assimilation to a nonlinear observation operator and injecting three TC boundary-layer constraints; we further propose a CYGNSS-intrinsic Observation Coverage Sufficiency (OCS) criterion that flags reliable reconstructions without external references. Applied to 4,955 snapshots of 249 TCs across all six active basins (2020-2022), the reconstructions reduce systematic Vmax bias against IBTrACS best-track by ~79% and ~75% relative to ERA5 and CCMP. Independent Tail Doppler Radar validation (47 storms) yields a wind speed RMSE of 6.9 m/s on the 23 coverage-sufficient cases (7.5 m/s overall); ablation across the full sample shows that the physical constraints cut wind-direction RMSE by 60% without degrading speed accuracy. The framework further supports joint assimilation of heterogeneous observations: adding only 11 dropsonde vectors to CYGNSS for TC FIONA (2022) reduces the cross-eye profile RMSE by 42%, outlining a practical pathway for fusing CYGNSS with SFMR, SAR and scatterometer data. The result is a globally consistent, observation-anchored kilometre-scale description of TC inner-core vector winds across all six active basins, including those without routine aircraft reconnaissance.
Performing unsupervised anomaly detection in retinal optical coherence tomography (OCT) images involves training a model solely on anomaly-free samples and detecting anomalies during inference, which reduces the cost of collecting large-scale annotated anomalous data. However, retinal OCT images exhibit significant variations in shape, thickness, and orientation, and lesions often have similar reflectance signals as normal tissues, making anomaly localization highly challenging. Existing methods address these challenges by flattening retinal layers, normalizing thickness, or leveraging reflectance priors, but their reliance on complex pre- and post-processing introduces uncertainties and limits end-to-end clinical applicability. To overcome these issues, we propose a novel semantic augmentation variational autoencoder (SeAugVAE) for unsupervised anomaly detection in retinal OCT images. Specifically, to capture the anatomical variability of normal retinas and thereby enhance anomaly sensitivity, we introduce a self-supervised semantic data augmentation strategy that enforces dual distribution consistency in both image and feature spaces during VAE training. For precise anomaly localization, we develop structural-semantic anomaly attention maps in the inference phase to detect anomalies from both local and global perspectives, and combine them to calculate anomaly score maps as the metric for localizing anomalous regions in images. Extensive experiments on multiple publicly and privately collected Cirrus and Spectralis OCT datasets demonstrate the effectiveness of SeAugVAE in pixel-wise unsupervised anomaly detection across multiple retinal diseases. Our codes are available at https://github.com/xyzhou1121/SeAugVAE.
In this paper, we propose a robust subspace-constrained quadratic model (SCQM) for learning low-dimensional structure from high-dimensional data. Building upon the subspace-constrained quadratic matrix factorization (SQMF) framework, the proposed model accommodates a broad class of noise distributions, including generalized Gaussian and radial Laplace models. This generalization enables reliable performance under both heavy-tailed and light-tailed noise, thereby substantially enhancing robustness across diverse data regimes. To efficiently address the resulting nonconvex optimization problem, we develop a gradient-based algorithm equipped with a backtracking line-search strategy that ensures stable and efficient convergence. In addition, we present a sensitivity analysis of the ℓ_p^p and ℓ_2 loss functions, elucidating their distinct behaviors under varying noise characteristics. Extensive numerical experiments corroborate the theoretical analysis and demonstrate that the proposed approach consistently outperforms existing methods in terms of robustness and reconstruction accuracy.
To fully exploit the advantages of GNSS-R small satellites (large quantity, short revisit cycle) and address key challenges of multisource GNSS-R data-track-wise distribution, uneven spatial coverage, and observation gaps, this study proposes a FusionGAN model integrating spatio-temporal attention mechanisms and physical constraints. This model enables end-to-end reconstruction of full-coverage gridded sea surface wind fields directly from sparse GNSS-R observations in the China Seas region. A two-stage training strategy is adopted: pretraining using 31+ years of cross-calibrated multiplatform wind dataset to lay a foundation for learning how to fill data gaps based on input orbit data, followed by fine-tuning with real GNSS-R observations from Cyclone Global Navigation Satellite System, FY-3 series satellites, and Tianmu-1 to enhance adaptability to practical data. Experimental results, validated through comparisons across low-to-medium wind, high wind, and typhoon scenarios, demonstrate the effectiveness of the proposed model: Finetuned-FusionGAN achieves an overall Bias of -0.03 m/s, a root-mean-square errors (RMSE) of 1.38 m/s (22% reduction), and a correlation coefficient R of 0.90. Specifically, by leveraging physics-constrained losses and spatio-temporal attention, the model reduces input GNSS-R observational data RMSE by 28.5% to 1.18 m/s and improves R by 8.2% to 0.92. In nonorbit regions, it generates physically consistent wind fields with an RMSE of 1.45 m/s, a correlation coefficient R of 0.89, comparable to input GNSS-R observational data accuracy. Moreover, benefiting from its physics-constrained design, the proposed model effectively mitigates observational discrepancies in multisource GNSS-R observations, fills spatial gaps, providing an effective solution for multisource GNSS-R sea surface wind field fusion.
Abstract. Tropical cyclone (TC) inner-core vector wind fields are essential for intensity forecasting, storm surge prediction, and structural climatology. The Cyclone Global Navigation Satellite System (CYGNSS) constellation provides dense temporal sampling and L-band precipitation-penetrating capability over the tropical belt, but its sparse, scalar wind speed retrievals have not been fully assimilated to produce kilometre-scale TC inner-core vector wind fields with global multi-basin coverage. This paper presents the QiFeng-CYGNSS dataset, which combines CYGNSS observations with a physics-guided score-based diffusion assimilation framework to reconstruct spatially complete 10 m vector wind fields from sparse scalar wind speed observations. The dataset covers 249 TCs across six active global basins during January 2020–September 2022, providing 1.5 km resolution vector wind fields at every IBTrACS reporting time with available CYGNSS coverage (4955 snapshots in total), accompanied by observation metadata and pixel-level ensemble uncertainty estimates for 138 major-hurricane snapshots. Independent validation against spaceborne C-band synthetic aperture radar, airborne Tail Doppler Radar, and GPS dropsondes indicates that the reconstructions represent TC inner-core structures at kilometre scales, reducing the absolute Vmax bias relative to ERA5 and CCMP by ~79% and ~75%, respectively, on the full sample. The dataset is freely available from Zenodo at https://doi.org/10.5281/zenodo.20046109 (Han et al., 2026b).
Storm surge forecasting is critical for coastal disaster mitigation, yet existing machine learning approaches remain predominantly regional—each model. Here we present a globally adaptive framework (ST-GNNFormer SurgeCast) that achieves tropical cyclone-induced storm surge forecasting across all ocean basins with a single trained model. Trained on 12,128 samples spanning 1993--2018, the model achieves an event-averaged RMSE of 4.65 cm (R=0.87) on the global test set, outperforming spatial-only and temporal-only baselines by 11.8% and 6.2%. Notably, robust generalisation is confirmed even in data-sparse basins: the North Indian (66 typhoons) and South Pacific (33 typhoons) basins achieve R=0.83 and R=0.83, respectively, demonstrating effective cross-basin knowledge transfer. Autoregressive rolling forecasts extend skillful prediction to 48 hours on a common subset of 34 typhoon tracks with continuous 48-hour data coverage, with RMSE increasing monotonically from 2.63 cm to 4.77 cm while retaining R>0.76. Independent validation against 199 GESLA3.0 in-situ tide gauges achieves an RMSE of 15.07 cm, confirming the model's predictive skill on real observational data beyond the training reanalysis.
Wind energy is an important part of sustainable energy. Potential assessment of wind energy, based on an appropriate wind speed dataset, is crucial for wind energy development. The assessment of wind energy potential is constrained by factors such as temporal scales, land cover types, and wind grades. It is essential to identify the optimal wind speed product under these constraints for accurately characterizing wind speed and its volatility. However, existing studies often lack a comprehensive evaluation of wind speed products considering these limitations, and wind speed volatility is rarely considered in the optimal wind speed product selection. To address these issues, this paper aims to select the optimal wind speed product under multi-factor constraints, including temporal scales, land cover types, and wind grades, for wind resource assessment from CFSR, CN05.1, ERA5-Land, GLDAS, JRA55, and MERRA2. Major findings are summarized as follows: (1) CN05.1 is the most suitable product for characterizing wind speed across mainland China under multi-factor constraints, followed by ERA5-land. (2) The wind speed volatility of all six datasets is underestimated compared to actual observations. Among them, JRA55 demonstrates the best capability to depict wind speed volatility in China, followed by MERRA2 and ERA5-land. (3) ERA5-land is the optimal product for wind energy resource assessments across mainland China, offering relatively accurate characterizations of both wind speed and wind speed volatility.
Diffusion models have demonstrated exceptional capabilities in image restoration, yet their application to video super-resolution (VSR) faces significant challenges in balancing fidelity with temporal consistency. Our evaluation reveals a critical gap: existing approaches consistently fail on severely degraded videos--precisely where diffusion models' generative capabilities are most needed. We identify that existing diffusion-based VSR methods struggle primarily because they face an overwhelming learning burden: simultaneously modeling complex degradation distributions, content representations, and temporal relationships with limited high-quality training data. To address this fundamental challenge, we present DiffVSR, featuring a Progressive Learning Strategy (PLS) that systematically decomposes this learning burden through staged training, enabling superior performance on complex degradations. Our framework additionally incorporates an Interweaved Latent Transition (ILT) technique that maintains competitive temporal consistency without additional training overhead. Experiments demonstrate that our approach excels in scenarios where competing methods struggle, particularly on severely degraded videos. Our work reveals that addressing the learning strategy, rather than focusing solely on architectural complexity, is the critical path toward robust real-world video super-resolution with diffusion models.
This paper presents an advanced graph convolutional network model, enhanced with Wasserstein distance-based adversarial learning (WD-ACGN), addressing the limitations of existing single-station and less explored multi-station water level forecasting approaches. The model features a novel coupled module for effectively capturing short and long-term dependencies, and a hybrid distance-based adaptive graph learning approach for spatial dependencies. This spatial dependency analysis enables rapid deployment in different coastal settings. Adversarial learning with gradient penalty further refines the model’s performance. Our model, applied to datasets from China’s Zhejiang coast and Daya Bay, outperforms baselines with a notable 12-h average root mean square error of 6.77 cm at 16 Zhejiang stations, proving its efficacy in varied maritime environments. Ablation studies validate the contribution of each model component, highlighting their collective impact on overall efficacy. Notably, the model showcases robustness in tropical cyclone scenarios and reliable results when tested with real-world observational data, underlining its potential for versatile applications in ocean engineering.
This paper explores the periodic scheduling problem in robot manufacturing units, considering constraints like multiple handling robots and "no waiting" time window characteristics. The research strives to address a global optimization challenge that integrates production line scheduling design and robot handling scheduling, aiming to minimize processing cycle time. To attain these objectives, a load-carrying initialization encoding, grounded in a splitting strategy, is introduced. To enhance computational efficiency, we incorporate the inter-zone handling processes derived from the splitting strategy into a mixed integer linear programming (MILP) model, solving it with CPLEX. The algorithm framework utilizes an adaptive iterative neighborhood search algorithm (S-AILS) rooted in a splitting strategy, proposing various transformation search strategies both within and across regions. We introduce interval optimal records to store the currently found optimal solution, using this saved solution as the starting point for iterative optimization, thus eliminating redundant searches. The algorithm's efficiency and effectiveness are validated through benchmark problems and dataset experiments.
This study evaluates and calibrates wind products derived from Global Navigation Satellite System Reflectometry (GNSS-R) using data from the FengYun-3E (FY-3E) global navigation satellite system occultation sounder II (GNOS-II) and TIANMU missions. The research highlights the significance of remote sensing for the accurate measurement of sea surface wind speeds in nearshore areas, which are crucial for environmental monitoring and climate studies. Initial comparisons with National Data Buoy Center (NDBC) measurements revealed root - mean - square errors (RMSE) of 2.49 m/s for FY-3E GNOS-II Beidou navigation satellite system (BDS) signals and 2.13 m/s for global positioning system (GPS) signals. For the TIANMU mission, the RMSE values were 3.21 m/s for BDS, 3.13 m/s for GPS, 2.91 m/s for GLONASS (GLO), and 2.91 m/s for Galileo (GAL) signals. To improve accuracy, especially in the complex nearshore environments, a deep learning calibration model incorporating residual blocks was employed. This model significantly enhanced the performance compared to a basic neural network. An ablation study confirmed that including residual blocks reduced RMSE by over 20% across all signal types. The calibrated model achieved substantial accuracy improvements in the test set, reducing RMSE to 1.03 m/s for FY BDS (improvement of 60%), 0.99 m/s for FY GPS (improvement of 54%), 1.57 m/s (improvement of 51%), 1.36 m/s for TIANMU GPS (57% improvement), 1.26 m/s for TIANMU GLO (improvement of 56%), and 1.50 m/s for TIANMU GAL (improvement 47%).
To address the issue of small target features in drone-captured images being easily affected by noise, a small object detection method with an attention mechanism under dual coordinate systems is proposed. Firstly, a dual coordinate systems-based attention mechanism module is constructed, which can simultaneously capture positional correlation features and directional correlation features between pixels. Then, an adaptive weighted Efficient Inter-section over Union (EIoU) loss function is designed, which adaptively adjusts weights based on the positional relationship between the detection box and the ground truth box. Experimental results show that the mAP0.5:0.95 (mean Average Precision) and the AP50 (Average Precision) can reach 40.33 % and 62.00 % respectively based on the proposed method, and the detection performance has been significantly improved compared to the baseline models.
Logistics distribution requires competitive planning of distribution routes to reduce costs and improve customer satisfaction. As a result, it has garnered significant attention and research. This paper addresses the challenges faced by organizations with alternative fuel-powered vehicle fleets, such as limited vehicle driving range and refueling infrastructure, by proposing a more realistic Green Vehicle Routing Problem (G-VRP). The objective of this study is to optimize the total length of the final planning route and reduce distribution costs. To achieve this, a Hybrid Two-stage Heuristic Algorithm (HTSHA) is introduced to solve the G-VRP with capacitated service stations. Additionally, this paper accounts for the possibility of service stations temporarily failing to provide services and presents a rescheduling strategy for the remaining distribution tasks. The results demonstrate the practical significance of this research in promoting the application of new energy vehicles in urban logistics and distribution.
This study evaluates and calibrates wind products derived from Global Navigation Satellite System Reflectometry(GNSS-R)using data from the FengYun-3E(FY-3E)global navigation satellite system occultation sounder Ⅱ(GNOS-Ⅱ)and Tianmu-1 missions.The research highlights the significance of remote sensing for the accurate measurement of sea surface wind speeds in nearshore areas,which are crucial for environmental monitoring and climate studies.Initial comparisons with National Data Buoy Center(NDBC)measurements revealed root-mean-square errors(RMSE)of 2.49 m/s for FY-3E GNOS-Ⅱ Beidou navigation satellite system(BDS)signals and 2.13 m/s for global positioning system(GPS)signals.For the Tianmu-1 mission,the RMSE values were 3.21 m/s for BDS,3.13 m/s for GPS,2.91 m/s for GLONASS(GLO),and 2.91 m/s for Galileo(GAL)signals.To improve accuracy,especially in the complex nearshore environments,a deep learning calibration model incorporating residual blocks was employed.This model significantly enhanced the performance compared to a basic neural network.An ablation study confirmed that including residual blocks reduced RMSE by over 20%across all signal types.The calibrated model achieved substantial accuracy improve-ments in the test set,reducing RMSE to 1.03 m/s for FY BDS(improvement of 60%),0.99 m/s for FY GPS(improvement of 54%),1.57 m/s(improvement of 51%),1.36 m/s for Tianmu-1 GPS(57%improvement),1.26 m/s for Tianmu-1 GLO(improvement of 56%),and 1.50 m/s for Tianmu-1 GAL(improvement 47%).
Transformer-based models like ViViT and TimeSformer have advanced video understanding by effectively modeling spatiotemporal dependencies. Recent video generation models, such as Sora and Vidu, further highlight the power of transformers in long-range feature extraction and holistic spatiotemporal modeling. However, directly applying these models to real-world video super-resolution (VSR) is challenging, as VSR demands pixel-level precision, which can be compromised by tokenization and sequential attention mechanisms. While recent transformer-based VSR models attempt to address these issues using smaller patches and local attention, they still face limitations such as restricted receptive fields and dependence on optical flow-based alignment, which can introduce inaccuracies in real-world settings. To overcome these issues, we propose Dual Axial Spatial×Temporal Transformer for Real-World Video Super-Resolution (DualX-VSR), which introduces a novel dual axial spatial×temporal attention mechanism that integrates spatial and temporal information along orthogonal directions. DualX-VSR eliminates the need for motion compensation, offering a simplified structure that provides a cohesive representation of spatiotemporal information. As a result, DualX-VSR achieves high fidelity and superior performance in real-world VSR task.