An Enhanced Deep Learning Framework For Land Target Detection And Captioning In Remote Sensing Imagery
DOI:
https://doi.org/10.51483/IJAIML.6.6s.2026.1093-1103Keywords:
Image Captioning, Multimodal learning, Geospatial analysis, Land-use classification, Scene interpretation, Visual semantics.Abstract
Remote sensing image analysis supports ecological observing, town development, and tragedy organization. Though existing methods separately handle target detection and caption generation, limiting contextual understanding and producing less informative descriptions of complex aerial scenes. To address these challenges, this research proposes a Scene-Aware Swin Transformer Captioning Network (SASTCN) for integrated land target detection and remote sensing image captioning. The proposed model utilizes a ST encoder to acquire hierarchical visual features for the effective representation of land targets. Moreover, the scene-aware embedding module acquires high-level contextual information in relation to various land use categories. Additionally, a semantic attribute learning module detects discriminative attributes of the detected land targets. This way, these complementary visual, scene, and semantic attributes are merged into a multimodal feature fusion approach to form a comprehensive semantic representation. Finally, the captions are generated using a transformer-based decoder, which includes positional encoding and beam search decoding to ensure contextual and semantic coherency in the produced sentences. Moreover, comprehensive text preprocessing, vocabulary formation, and sequence encoding are included to improve the language modeling and caption generation process. The proposed model is programmed in Python programming language and tested on the benchmark dataset called Remote Sensing Image Caption Dataset (RSICD). Experiments prove that the proposed SASTCN recognizes various geographical targets and produces correct captions and reaches a Semantic Propositional Image Caption Evaluation (SPICE) score of 0.6031 and increases the contextual and semantic coherency of the produced sentences.





