Bibliographic Details
| Title: |
Learning to enhance areal video captioning with visual question answering. |
| Authors: |
Al Mehmadi, Shima M.1 (AUTHOR), Bazi, Yakoub1 (AUTHOR), Al Rahhal, Mohamad M.2 (AUTHOR) mmalrahhal@ksu.edu.sa, Zuair, Mansour1 (AUTHOR) |
| Source: |
International Journal of Remote Sensing. Sep2024, Vol. 45 Issue 18, p6395-6407. 13p. |
| Subjects: |
Drone aircraft, Remote sensing, Videos |
| Abstract: |
The utilization of Unmanned Aerial Vehicles (UAV) in remote sensing (RS) has witnessed a significant surge, offering valuable insights into Earth dynamics and human activities. However, this has led to a substantial increase in the volume of video data, rendering manual screening and analysis impractical. Consequently, there is a pressing need for the development of automated interpretation models for these aerial videos. In this paper, we propose a novel approach that leverages visual dialogue to enhance aerial video captioning. Our model adopts an encoder-decoder architecture, integrating a Visual Question Answering (VQA) task before the captioning task. The VQA task aims to enrich the captioning process by soliciting additional information about the image content. Specifically, our video encoder utilizes ViT-L/16, while the decoder employs Generative Pre-trained Transformer-2 (Distill-GPT-2). To validate our model, we introduce a novel benchmark dataset named CapERA-VQA, comprising videos accompanied by sets of questions, answers, and captions. Through experimental validation, we demonstrate the effectiveness of our proposed approach in enhancing the automated captioning of aerial videos. [ABSTRACT FROM AUTHOR] |
|
Copyright of International Journal of Remote Sensing is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) |
| Database: |
Engineering Source |