
Reliable prediction of the Indian Summer Monsoon (ISM) requires accurate representation of large-scale circulation features such as the Low-Level Jet (LLJ) and Tropical Easterly Jet (TEJ), yet sparse observational coverage over the Indian Ocean region continues to limit model validation and improvement. This study validated horizontal line-of-sight (HLOS) winds from the ESA Aeolus satellite, carrying the first spaceborne Doppler wind lidar (ALADIN), against high-resolution radiosonde observations at Gadanki (13.5°N, 79.2°E) during 2019–2021, with spatial intercomparisons against reanalysis datasets extended through 2022. Observation days were classified into clear-sky and cloudy-sky conditions using infrared brightness temperature to assess Aeolus retrievals under varying cloud regimes. Rayleigh-clear retrievals showed strong agreement with radiosondes (correlation coefficient = 0.95, bias = 0.11 m s−1), while Mie-cloudy retrievals performed notably weaker (correlation coefficient = 0.35, bias = 3.40 m s−1). Agreement improved with altitude, with the highest correlation (0.97) observed in the upper troposphere–lower stratosphere (UTLS). HLOS wind differences between Aeolus and radiosondes generally remained within ±2 m s−1. The vertical structure and intensity of the LLJ and TEJ derived from Aeolus agreed most closely with radiosonde observations, followed by ERA5, MERRA-2, and NCEP-2 reanalyses, in order of increasing deviation. Spatial deviations were small relative to ERA5 and MERRA-2 but substantially larger relative to NCEP-2. These findings demonstrate that Aeolus provides reliable HLOS wind measurements for characterizing the vertical structure and seasonal evolution of Indian Summer Monsoon circulation, while supporting the evaluation of atmospheric reanalysis datasets over observationally sparse regions.
An azimuth wide-swath spotlight synthetic aperture radar (SAR) imaging method utilizing multifrequency subpulse (MFSP) diversity is proposed. Initially, the MFSP scheme is formulated by dividing the transmit pulse into multiple subpulses, where each is assigned a distinct carrier frequency offset with bandwidth-level spacing. Different subpulses are then steered toward multiple desired subscenes to achieve wide azimuth beam coverage. After receiving the aliased signals from different subscenes, the unambiguous subscene echoes are extracted through a set of bandpass filters. Subsequently, interpolation along the line-of-sight is employed to accurately reconstruct the original sub-images based on the squint geometry of each subscene. Ultimately, following both geometric and radiometric corrections, the complete azimuth-extended spotlight image is generated by merging subscene imagery according to their geometric relationships. Specifically, the key imaging characteristics and parameter design strategies are discussed. The effectiveness of the proposed method is validated through comparisons with traditional methods, demonstrating its capability to expand azimuth width without degrading azimuth resolution or introducing severe wavefront-curvature errors.
Infrared target detection supports the continuous observation of traffic participants and low-altitude targets across aerial and fixed-view imaging settings, particularly under weak or changing illumination. However, infrared targets are often small, weakly textured, and easily confused with thermal noise and background clutter. Large pretrained vision models offer strong representation and generalization capabilities, but their parameter counts and computational costs make direct deployment on edge platforms with limited resources impractical. To transfer these capabilities to lightweight models, this paper proposes knowledge distillation at the label level based on filtered teacher detections. A large teacher adapted to the infrared domain first generates candidate boxes. Candidate boxes are selected using confidence thresholds, class reliability, and spatial relationships with ground truth annotations. The retained teacher boxes and the original annotations jointly form the student training targets, while the student retains its standard detection loss and original inference structure. On an independent sequence-level test set, the YOLOv5n baseline obtains an mAP@0.5 of 0.2277 and an mAP@0.5:0.95 of 0.1070, whereas filtered label distillation obtains 0.2547 and 0.1153, respectively, over three matched seeds. The filtering rules and thresholds are fixed before this evaluation, and the fixed Epoch-30 checkpoint is used for every run. Both student architectures are deployed on RK3588.
Estuarine wetlands are highly dynamic ecosystems, and the vegetation serves as a critical indicator of ecological health. Accurate mapping of different vegetation types remains challenging due to spectral similarities and the high dimensionality of time-series data. This research introduces a Google Earth Engine (GEE)-based hierarchical framework for mapping eight typical vegetation types in the Liaohe Estuary using Sentinel-2 imagery from 2023 to 2025. To address data redundancy, a novel feature selection algorithm based on Mahalanobis distance and class separability (FSMD-CS) was developed, reducing 225 dimensions to seven optimal variables. Integrated with a hierarchical decision tree calibrated using the SEaTH approach, the framework achieved an overall accuracy of 86.39% (kappa = 0.832), surpassing single-temporal spectral imagery classification and unoptimized MPS feature classification, which achieved OAs of 72.31% and 82.68%, respectively. The majority of selected features originate from the early green-up and late senescence stages, indicating that seasonal phenological metrics offer superior discrimination of vegetation types compared to peak-summer spectral data. Overall, the proposed framework provides an efficient and interpretable solution for fine-scale estuarine vegetation mapping.
The identification and delineation of subsurface resistive anomalies beneath sedimentary cover remain challenging for the magnetotelluric (MT) method. In this study, we compared the attenuation behavior of electromagnetic (EM) fields and their response characteristics under surface and/or subsurface observation configurations using both synthetic and field data. The synthetic results show that the conductive cover layer substantially suppresses the EM responses of the subsurface high-resistivity body, resulting in only weak relative response perturbations at surface stations. In contrast, the relative response amplitudes recorded at underground stations are enhanced, indicating that underground EM observations offer advantages in resolving subsurface resistive anomalies. The results further reveal distinct attenuation behaviors of the electric and magnetic fields, with the electric field being more sensitive to variations in the bulk conductivity of the sedimentary layer. Furthermore, one-dimensional Bayesian probabilistic inversions of both synthetic and field datasets indicate that underground observations provide more robust and reliable estimates of the resistivity structure. These findings suggest that further exploration of deep underground observations has considerable potential for detecting and characterizing weak EM responses from resistive targets in sedimentary environments, while also improving the reliability of inversion results and reducing their uncertainty.
Landslides are frequent and destructive geological disasters. Accurate landslide identification is essential for post-disaster reconstruction and preventing secondary disasters. Deep learning has shown considerable potential for recognizing landslide objects from remote-sensing images; however, existing models still suffer from insufficient detection accuracy in scenarios with complex backgrounds, blurred boundary localization, sample-class imbalance, and difficulty in balancing detection speed and segmentation accuracy. To address these issues, this study investigates landslide identification in the Great Bend of the Yarlung Zangbo River region using improved deep learning models and heterogeneous optical remote-sensing imagery. (1) By introducing the convolutional block attention module (CBAM) into YOLOv8, 82.79% precision was achieved, and the recall improved by 9.56% compared to the original model, reaching a mean average precision of 75.27% while maintaining computational efficiency, outperforming YOLOv5 and the original YOLOv8. (2) Replacing the cross-entropy loss with Focal Loss in DeepLabV3+ improved the landslide edge segmentation by dynamically adjusting the weights of difficult and easy samples. Compared with the original DeepLabV3+, the precision and recall of the DeepLabV3+-FL semantic-segmentation model were improved by 0.26% and 1.93%, respectively, with the mean pixel accuracy and mean intersection over union reaching 81.76% and 62.59%, respectively. Overall, the two improved models enhanced the accuracy of landslide identification and resistance to interference, demonstrating potential for landslide monitoring and emergency response.
The development of temporally consistent long-term normalized difference vegetation index (NDVI)climate data records is essential for global change and ecosystem research. The Medium Resolution Spectral Imager-II (MERSI-II) sensor aboard China’s Fengyun-3D (FY-3D) satellite shares similar spectral characteristics with Moderate Resolution Imaging Spectroradiometer (MODIS), offering potential for synergistic applications. However, systematic biases arising from differences in sensor design, radiometric calibration, and atmospheric correction hinder their direct combination. This study established a full-chain framework that integrated cross-calibration of surface reflectance using quasi-synchronous FY-3D/MODIS observations and a MERSI-II-specific atmospheric correction scheme based on the 6S radiative transfer model. After correction, the FY-3D NDVI shows substantially improved consistency with MODIS, achieving a reduction in root mean square error of over 25.9%, an increase in correlation coefficient of approximately 5%, and a decrease in mean absolute error of about 40%. Spatial biases are within ±0.1 over most global land areas, with robust performance across vegetation types and climate zones. Based on this technical framework, a fused FY-3D and MODIS NDVI climate data record was established, which has been operationalized at the Beijing Climate Center for global vegetation monitoring. This work provides a transferable framework for integrating Chinese Fengyun satellite data with international datasets like MODIS/VIIRS.
Polarimetric synthetic aperture radar (PolSAR) calibration quality assessment is essential for verifying the reliability of polarimetric calibration results and ensuring the accuracy of subsequent quantitative applications. The corner-reflector-based assessment is accurate but depends on field deployment and maintenance, whereas the distributed-target-based methods is easier to automate but is sensitive to mixed scattering within image patches and to non-unique histogram peaks. In addition, the parameter distribution may contain multiple peaks, affecting the uniqueness and stability of assessment. To address these problems, this paper proposes a robust distributed-target-based PolSAR calibration quality assessment (RD-PCQA) method without corner reflectors (CRs) to address these problems. The proposed method first uses hypothesis testing of confidence interval method for PolSAR calibration (PCHTCI) to extract high-quality distributed targets and combines the polarimetric correlation coefficient RHHVV to select volume-scattering-dominant targets, thereby improving the physical consistency of samples used for channel imbalance amplitude (CIA) estimation. Second, a high-proportion distributed-target constraint is used to refine the samples for channel imbalance phase (CIP) and polarimetric crosstalk estimation, reducing the influence of nonideal scatterers on parameter estimation. Finally, a unique peak searching strategy based on progressively enlarged statistical scales is proposed to suppress the effect of multi-peak distributions on assessment result. Experiments were conducted using three GF-3 PolSAR images acquired over the SAR calibration site in Etuoke Banner, Ordos, Inner Mongolia, China, with CR results used as references, and considering finite sample uncertainty. The experimental results show that, compared with the conventional distributed-target-based method, the proposed method is closer to the corresponding CR mean in all comparisons, with the mean absolute deviations for CIA, CIP and crosstalk scenarios reduced to 0.024 dB, 2.043° and 3.069 dB, respectively. Therefore, it demonstrates the effectiveness and practical potential of the proposed method for PolSAR calibration quality assessment without CRs.
Ionospheric scintillation and radio frequency interference (RFI) affect Global Navigation Satellite System (GNSS) signals, making their distinction important for ionospheric monitoring and interference detection. This study identifies GNSS signal anomalies associated with scintillation and RFI using COSMIC-2 1 Hz precise orbit determination (POD) observations. Using 60 s per-satellite windows, dual-frequency time-series inputs and 16-dimensional statistical features were extracted, and weak labels for Normal, Scintillation, and RFI were generated from the amplitude scintillation index S4 and the RFI index. An InceptionTimeLite-FiLM-DeepSets joint multi-satellite model, where FiLM denotes feature-wise linear modulation, was trained and tested on 2024 data and directly applied to the full-year 2025 dataset. On the 2024 test set, recall for Normal, Scintillation, and RFI was 98.9%, 76.1%, and 58.5%, respectively, with a Macro-F1 of 0.8227 and a Matthews correlation coefficient of 0.7120. Predicted Scintillation occurrence rates were higher at low magnetic latitudes and during the postsunset premidnight period, whereas predicted RFI occurrence rates were concentrated over North Africa, the Middle East, South Asia, and Southeast Asia. These results show that, within the weak-label framework, COSMIC-2 1 Hz POD observations can support GNSS signal anomaly classification and spatiotemporal distribution analysis, providing a complementary approach for long-term anomaly analysis.
Target recognition methods based on feature extraction of bistatic radar polarimetric characteristics face the challenge of feature variation caused by independent rotation of transmit and receive polarization bases. Based on Huynen decomposition and Cameron decomposition theories, this paper rigorously derives ten types of bistatic polarization roll-invariants covering scattering intensity, polarization modulation capability, scattering mechanism, symmetry and reciprocity. Simulations are systematically carried out on dihedral, sphere and trihedral structures, followed by anechoic chamber experimental verification using a self-built bistatic full-polarimetric electromagnetic scattering measurement system. Simulation and experimental results demonstrate that all proposed invariants strictly satisfy roll-invariance under independent rotation of transmit and receive polarization bases. Polarization roll-invariants of different structures exhibit distinct variation laws with bistatic angles, which can provide robust and reliable feature support for bistatic polarimetric radar target recognition.
Hyperspectral image (HSI) classification benefits from rich spectral information; however, high dimensionality of HSI data increases computational cost, noise sensitivity, and the risk of overfitting when labeled samples are limited. Most pretrained computer vision networks are designed for three-channel inputs, making direct application to hyperspectral cubes difficult. Conventional principal component analysis (PCA) ranks components by total variance without distinguishing useful signal variance from noise-related variance, which can reduce the reliability of the resulting representation when only a few components are retained. This paper proposes a data-augmented Noise-Adjusted Principal Component Analysis (DA-NAPCA) framework for deep learning-based HSI classification. By accounting for estimated noise covariance, NAPCA orders the transformed components by signal-to-noise ratio rather than total variance, while data augmentation mitigates the overfitting risk when labeled samples are limited. Unlike typical NAPCA/MNF applications, which select the number of retained components empirically, DA-NAPCA deliberately retains three noise-adjusted components to form a compact three-channel representation, enabling pretrained models designed for three-channel inputs to be fine-tuned without modifying their input layers. The framework is evaluated using a 3D convolutional neural network (3D-CNN) for spatial–spectral feature learning and a pretrained EfficientNet-B0 model for lightweight transfer learning. Although this paper uses 3D-CNN and EfficientNet-B0 as illustrative examples, the proposed DA-NAPCA framework is a representation-level preprocessing approach and does not require architecture-specific modification. Experiments conducted on the Indian Pines, University of Pavia, and Salinas datasets compare DA-NAPCA with RGB, band selection, PCA-based dimensionality reduction, and ablation variants. Across the three datasets, DA-NAPCA achieved mean overall accuracies of 93.11–94.71% with 3D-CNN and 95.93–97.44% with EfficientNet-B0. Compared with the second-best baseline method, DA-NAPCA improved overall accuracy by 2.75–7.58 percentage points with 3D-CNN and 1.28–2.12 percentage points with EfficientNet-B0. These results demonstrate that combining a compact noise-adjusted representation with spatial augmentation provides an effective input representation for deep learning-based HSI classification.
Ground-based GNSS tomography reconstructs three-dimensional tropospheric water-vapor fields from slant observations. However, unconstrained solutions rely heavily on station and satellite geometry. This study introduces a GNSS-only approach that uses precise SP3 orbit products to create pseudo-slant observations in satellite directions not tracked by individual receivers. These directions combine with existing IWV and horizontal-gradient estimates. This completes the ray distribution without adding external atmospheric constraints or independent water-vapor information. The method is tested in Hong Kong and Iceland with GPS-only, GLONASS-only, observed multi-constellation, and SP3-completed setups. GPS-derived gradients were generally in line with the full multi-GNSS solution, but GLONASS-only gradients showed larger differences. Overall, the mapped directions improved or maintained the inversion’s effective rank and numerical stability. They did not systematically degrade the retrieved water-vapor profiles, although the advantages decreased when the additional rays were geometrically redundant.
Landscape character change in mountain traditional villages is difficult to assess from a single perspective because roof-material replacement, material–color deviation, and visual exposure vary under complex terrain and settlement configurations. This study develops a multi-source UAV framework for reproducible, spatially explicit assessment of courtyard-scale visual–material integrity (CI), a spatially observable component of landscape character integrity. Using 749 courtyards in nine nationally designated traditional villages in Shangluo, China, the framework integrates UAV orthophotos, 3D mesh models, and point-cloud data to derive material penetration rate (PR), material–color conflict (CC), and standardized visual exposure (VC). PR and CC represent baseline material–color loss, whereas VC is incorporated as an exposure-amplification condition. Formula-structure sensitivity analysis against blinded ratings of 50 sampled courtyards showed that the proposed formulation had the highest rank consistency with expert judgments (Spearman’s ρ = 0.949). The frozen framework also showed a strong association with blinded professional ratings in geographically independent Longnan villages (ρ = 0.912, p < 0.001). XGBoost–SHAP identified courtyard impervious-surface ratio, primary material, and roof form as the leading model-associated predictors of CI variation. The framework supports courtyard-scale CI assessment, priority screening, and repeat-survey reassessment under comparable acquisition conditions.
Semantic segmentation of remote sensing images is challenging because multi-scale irregular objects in complex scenes often exhibit large intra-class variability, high inter-class similarity, and sparse spatial distributions. These factors hinder accurate boundary delineation and reliable contextual modeling among spatially distant but semantically related regions. Considering the capability of graph neural networks in modeling irregular relationships, we propose DGCR-Net, a dynamic graph contextual reasoning network for semantic segmentation of remote sensing imagery. Specifically, DGCR-Net integrates a ResNet18 encoder with a multi-stage decoder composed of cascaded dynamic graph reasoning blocks (DGRBs), which adaptively infer complex contextual dependencies among irregular objects and progressively refine multi-scale semantic representations. A semantic graph adapter (SGA) is incorporated at each skip connection to enhance encoder features and project them into graph-compatible representations, ensuring robust contextual reasoning. Extensive experiments on the Vaihingen, Potsdam, LoveDA, and UAVid datasets demonstrate that DGCR-Net achieves competitive performance, with mIoU scores of 83.4%, 86.5%, 53.9%, and 69.4%, respectively.
In contrast to the well-established vicarious methods for satellite radiometric calibration, studies on polarimetric calibration remain relatively limited. This study presented a sun glint based in-flight polarization calibration method in which measurements of the degree of linear polarization (DOLP) at the top of atmosphere (TOA) are directly compared with theoretical simulations generated by an atmosphere-ocean coupled vector radiative transfer model. A well calibrated reference band is introduced to retrieve the instantaneous sea surface wind speed (WS), thereby improving the accuracy of simulated TOA DOLP. The dependencies of WS retrieval error on reference band wavelength, actual WS, and solar-viewing geometry are analyzed. The results show that the WS retrieval error generally decreases with increasing reference band wavelength and increases with WS and sun glint angle. A reference band with a wavelength longer than 670 nm and the highest radiometric calibration accuracy is recommended for WS retrieval. Under typical conditions, a WS retrieval error of approximately ±0.33 m/s is predicted, which is substantially lower than the ±2 m/s uncertainty commonly associated with meteorological reanalysis WS data. Theoretical TOA DOLP simulation errors are then quantified by considering typical uncertainties in atmospheric and oceanic input parameters across spectral bands from blue to SWIR under various solar-viewing geometries. Among the investigated factors, aerosol optical depth and aerosol model are the dominant sources to DOLP calibration uncertainty, accounting for more than 80% of the total error budget in most spectral bands. Chlorophyll concentration mainly affects the short visible bands. The WS effect can be reduced by using reference band retrievals instead of meteorological reanalysis data. The wind direction effect is negligible near the center of glint spots and under small solar zenith angle conditions, but it increases considerably with the sun glint angle under oblique illumination geometry condition. The contributions from ozone, and water vapor are negligible. Overall, the total TOA DOLP error increases with solar zenith angle, viewing zenith angle (i.e., longer atmospheric optical paths), and sun glint angle (i.e., weaker sun glint brightness). The errors exhibit weak wavelength dependence, with relative lower values in the short visible and SWIR bands and slightly higher values in the NIR bands. The typical total DOLP calibration error ranges from 0.0091 to 0.0125 across blue to SWIR spectral bands (according to the center of sun glint region with solar zenith angle of 30°). This study presents theoretical guidance for satellite in-flight polarization calibration using sun glint across blue to SWIR bands. As budgeted, this method can be effective for the validation of polarimeters with moderate DOLP accuracy (e.g., 0.01–0.02). Nevertheless, it may not be ideally suited to act as an absolute reference for high-accuracy polarimeters (e.g., 0.002).
Rapid and non-destructive monitoring of soil total nitrogen (STN) is important for precision nutrient management in ecologically fragile agricultural systems. This study established a controlled spectral resampling experiment to simulate the multispectral responses of Sentinel-2, WorldView-3, GF-6, and Landsat-9 from laboratory ASD hyperspectral measurements of Hemerocallis citrina fields. Two-dimensional (DI, RI, and NDI) and three-dimensional (TBI1–TBI5) spectral indices were constructed and evaluated using eight machine learning algorithms across five phenological stages. The results demonstrate that: (1) Three-dimensional spectral indices exhibited substantially higher sensitivity to STN than conventional two-dimensional indices, with TBI3 showing the strongest overall correlation; among the simulated sensor configurations, GF-6 delivered the best mean performance due to its dual red-edge bands. (2) Genetic algorithm-optimized backpropagation neural network (GA-BPNN) effectively addressed the local-minima limitation of standard backpropagation neural networks (BPNNs) and displayed robust generalization under multi-sensor and multi-phenological scenarios. (3) Phenology-specific modeling reduced spectral heterogeneity caused by pooling across growth stages, increasing R2 by 20.28% and decreasing RMSE by 18.35%, with leaf expansion and bolting identified as optimal estimation windows. These findings elucidate the spectral response potential explanations of STN across phenological stages and provide a reference for applying multi-source simulated remote sensing data to nutrient monitoring in specialty agricultural systems.
This study is devoted to few-shot oriented object detection in aerial images, aiming to enhance detection performance for novel object classes, using only limited supervised samples. Currently, most few-shot object detection models adopt the two-stage fine-tuning approach (TFA), which consists of a base training stage and a few-shot fine-tuning stage. However, the region proposal network (RPN) suffers from foreground–background classification confusion and the rotation angle ambiguity of the square-like bounding boxes. These issues significantly degrade the detection performance for novel categories. To this end, we propose a confusion-resistant learning (CRL) for few-shot aerial oriented object detection. CRL contains a classification reweighting scheme (CRS) and an Edge-Vectors Cosine Similarity (EVCS) Loss. First, the CRS utilizes credible bounding box regression outputs from the base training stage to assist foreground–background classification learning. This process suppresses classification confusion and improves accuracy for novel categories. Second, we propose an EVCS Loss, which builds upon the Kalman filtering IoU (KFIoU) loss. The EVCS Loss alleviates rotation angle confusion for square-like boxes by maximizing the cosine similarity between the edges of the ground-truth and predicted bounding boxes. In addition, CRL can be plugged into the existing two-stage oriented object detectors. Extensive experiments on DOTA and DIOR-R oriented object detection benchmarks show that, compared with the ReDet-KFIoU baseline, our CRL achieves up to 2.4% overall AP50 improvement and 1.8% novel-class AP50 improvement on DOTA, and yields up to 2.0% overall AP gain and 2.3% novel-class AP gain on DIOR-R, providing direct quantitative evidence for the effectiveness of our method.
In real-world wildfire monitoring, haze, glare, refraction artifacts, and related visual effects can cause false alarms and missed detections. This study explores two implementations: ContextFireVLM and ContextFireAgent. ContextFireVLM fine-tunes Llama-3.2 Vision with LoRA to improve the base model’s ability to identify and analyze wildfires; ContextFireAgent equips an agent with YOLO as a tool to improve the same capabilities. The first and second assessors independently inspect the full image, and the second assessor invokes YOLO only when it reports uncertainty; when their judgments agree, the final reviewer preserves the agreed decision. Because the two implementations use different models, inputs, and training pipelines, these results do not establish a causal comparison or general superiority between LoRA adaptation and the agent design.
Rock glaciers are important geomorphic indicators of mountain permafrost conditions, yet the mechanisms controlling spatially heterogeneous deformation within individual rock glaciers remain poorly understood. This study investigates the Baishuigou Rock Glacier (BRG) in the Qilian Mountains by integrating multi-year InSAR observations, borehole stratigraphic data, and ground-temperature measurements to characterize its deformation and explore its relationship with permafrost degradation. The results reveal pronounced spatial heterogeneity in surface displacement. The main body of the rock glacier exhibits persistent displacement, with mean annual LOS displacement rates ranging from −6.93 to 13.19 mm yr−1, corresponding to relatively well-preserved ice-rich frozen ground. The frontal zone shows stronger and more variable displacement under more degraded subsurface conditions, whereas the relict rock glacier area exhibits relatively weak displacement dominated by seasonal thermal responses. Comparisons among the borehole sites further indicate that spatial differences in surface displacement correspond to variations in active-layer thickness, ground-ice conditions, and thermal state. The integrated observations suggest a possible spatial transition from relatively well-preserved ice-rich frozen ground in the upper part of the rock glacier toward more degraded conditions downstream. This study demonstrates that integrating surface deformation with subsurface structural and thermal observations provides stronger physical constraints for understanding spatially heterogeneous rock glacier deformation and its relationship with frozen-ground conditions.
Reliable crop yield estimation is fundamental to food security and efficient agricultural management. However, current deep learning models still face limitations in selecting and integrating multi-source features, and their high predictive accuracy is often accompanied by limited interpretability. This study introduces a Bayesian Optimization–Temporal Convolutional Network–Bidirectional Long Short-Term Memory–Dual Attention (BO-TCBDA) deep learning framework for winter wheat yield estimation. Using Henan Province, China, as the study area, county-level winter wheat yield from 2013 to 2022 was estimated using the Enhanced Vegetation Index (EVI), Leaf Area Index (LAI), Solar-Induced Chlorophyll Fluorescence (SIF), and climate data. The proposed model was compared with five commonly used machine learning and deep learning models. BO-TCBDA achieved the best performance, with an R2 of 0.823 and an RMSE of 561.26 kg/ha. SIF improved the predictive performance of all models, with statistically significant gains observed in the deep learning models. The dual-attention mechanism provided interpretable insights by revealing relatively balanced contributions among the input features and highlighting the grain-filling stage through temporal attention. Furthermore, SHAP-based cross-validation analysis identified T12, corresponding to the latter part of the jointing stage, as the period with the highest contribution to yield prediction. The model also achieved an R2 of approximately 0.80 about 25 days before harvest. Overall, BO-TCBDA provides an accurate and interpretable approach for county-level winter wheat yield estimation and supports regional food security assessments and precision agriculture.