Few-shot remote sensing scene classification aims to recognize land-use or land-cover categories from only a few labeled samples per class. Vision-language models (VLMs) such as CLIP provide strong zero-shot capability via semantic priors; however, integrating few-shot supervision into CLIP-based inference remains challenging due to the misalignment of semantic priors and the unreliability of task-adapted visual evidence under limited samples. In this letter, we propose reliability-calibrated CLIP (RC-CLIP), a training-free reliability-calibrated inference framework that operates at the decision level without fine-tuning CLIP encoders. RC-CLIP estimates the reliability of visual evidence through prototype-semantic discrepancy, prototype dispersion, and supervision sufficiency and adaptively balances few-shot predictions with zero-shot priors. Extensive experiments on multiple remote sensing benchmarks demonstrate that RC-CLIP consistently outperforms representative inference-time adaptation methods under the full-class few-shot protocol, especially in the challenging one-shot regime. Moreover, RC-CLIP improves discrimination among semantically ambiguous classes, yielding more robust and interpretable predictions.
更多
查看译文
关键词
CLIP,few-shot learning,reliability-aware inference,remote sensing scene classification,vision-language models (VLMs),CLIP,few-shot learning,reliability-aware inference,remote sensing scene classification,vision-language models (VLMs)