Vision-Language Models for Remote Sensing: Chain-of-Thought Reasoning and Reinforcement Learning Approaches Uzaktan Algilama için Görüntü-Dil Modelleri: Düşünce Zinciri Muhakemesi ve Pekiştirmeli Ö?grenme Yaklaşimlari


KÖKSAL A., ALATAN A. A.

34th Signal Processing and Communications Applications Conference, SIU 2026, İstanbul, Türkiye, 7 - 10 Temmuz 2026, (Tam Metin Bildiri)

  • Yayın Türü: Bildiri / Tam Metin Bildiri
  • Doi Numarası: 10.1109/siu71813.2026.11636937
  • Basıldığı Şehir: İstanbul
  • Basıldığı Ülke: Türkiye
  • Anahtar Kelimeler: chain-of-thought reasoning, GRPO, reinforcement learning, remote sensing, vision-language models
  • Orta Doğu Teknik Üniversitesi Adresli: Evet

Özet

In recent years, vision-language models (VLMs) have achieved significant advances in natural language processing and visual understanding. However, their effective use in specialized domains such as remote sensing has been constrained by high computational costs and domain-specific data requirements. This paper summarizes three approaches developed for remote sensing image analysis. First, TinyRS, a compact 2B-parameter VLM, and its reasoning-augmented variant TinyRS-R1 with chain-of-thought (CoT) reasoning and Group Relative Policy Optimization (GRPO) alignment are presented. Second, the SAMChat model, focusing on military installation detection, outperforms larger general-purpose models through domain-specific fine-tuning and reinforcement learning. Finally, a few-shot reinforcement learning with verifiable rewards (RLVR) framework demonstrates that meaningful performance gains can be achieved with as few as a single training example. Experimental results show that 2B-parameter models achieve competitive or superior performance compared to 7B-scale models.