Learning to enhance areal video captioning with visual question answering.
Saved in:
| Title: | Learning to enhance areal video captioning with visual question answering. |
|---|---|
| Authors: | Al Mehmadi, Shima M.1 (AUTHOR), Bazi, Yakoub1 (AUTHOR), Al Rahhal, Mohamad M.2 (AUTHOR) mmalrahhal@ksu.edu.sa, Zuair, Mansour1 (AUTHOR) |
| Source: | International Journal of Remote Sensing. Sep2024, Vol. 45 Issue 18, p6395-6407. 13p. |
| Subjects: | Drone aircraft, Remote sensing, Videos |
| Abstract: | The utilization of Unmanned Aerial Vehicles (UAV) in remote sensing (RS) has witnessed a significant surge, offering valuable insights into Earth dynamics and human activities. However, this has led to a substantial increase in the volume of video data, rendering manual screening and analysis impractical. Consequently, there is a pressing need for the development of automated interpretation models for these aerial videos. In this paper, we propose a novel approach that leverages visual dialogue to enhance aerial video captioning. Our model adopts an encoder-decoder architecture, integrating a Visual Question Answering (VQA) task before the captioning task. The VQA task aims to enrich the captioning process by soliciting additional information about the image content. Specifically, our video encoder utilizes ViT-L/16, while the decoder employs Generative Pre-trained Transformer-2 (Distill-GPT-2). To validate our model, we introduce a novel benchmark dataset named CapERA-VQA, comprising videos accompanied by sets of questions, answers, and captions. Through experimental validation, we demonstrate the effectiveness of our proposed approach in enhancing the automated captioning of aerial videos. [ABSTRACT FROM AUTHOR] |
| Copyright of International Journal of Remote Sensing is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 179637986 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Learning to enhance areal video captioning with visual question answering. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Al+Mehmadi%2C+Shima+M%2E%22">Al Mehmadi, Shima M.</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Bazi%2C+Yakoub%22">Bazi, Yakoub</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Al+Rahhal%2C+Mohamad+M%2E%22">Al Rahhal, Mohamad M.</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> mmalrahhal@ksu.edu.sa</i><br /><searchLink fieldCode="AR" term="%22Zuair%2C+Mansour%22">Zuair, Mansour</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22International+Journal+of+Remote+Sensing%22">International Journal of Remote Sensing</searchLink>. Sep2024, Vol. 45 Issue 18, p6395-6407. 13p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Drone+aircraft%22">Drone aircraft</searchLink><br /><searchLink fieldCode="DE" term="%22Remote+sensing%22">Remote sensing</searchLink><br /><searchLink fieldCode="DE" term="%22Videos%22">Videos</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: The utilization of Unmanned Aerial Vehicles (UAV) in remote sensing (RS) has witnessed a significant surge, offering valuable insights into Earth dynamics and human activities. However, this has led to a substantial increase in the volume of video data, rendering manual screening and analysis impractical. Consequently, there is a pressing need for the development of automated interpretation models for these aerial videos. In this paper, we propose a novel approach that leverages visual dialogue to enhance aerial video captioning. Our model adopts an encoder-decoder architecture, integrating a Visual Question Answering (VQA) task before the captioning task. The VQA task aims to enrich the captioning process by soliciting additional information about the image content. Specifically, our video encoder utilizes ViT-L/16, while the decoder employs Generative Pre-trained Transformer-2 (Distill-GPT-2). To validate our model, we introduce a novel benchmark dataset named CapERA-VQA, comprising videos accompanied by sets of questions, answers, and captions. Through experimental validation, we demonstrate the effectiveness of our proposed approach in enhancing the automated captioning of aerial videos. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of International Journal of Remote Sensing is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=179637986 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/01431161.2024.2388875 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 13 StartPage: 6395 Subjects: – SubjectFull: Drone aircraft Type: general – SubjectFull: Remote sensing Type: general – SubjectFull: Videos Type: general Titles: – TitleFull: Learning to enhance areal video captioning with visual question answering. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Al Mehmadi, Shima M. – PersonEntity: Name: NameFull: Bazi, Yakoub – PersonEntity: Name: NameFull: Al Rahhal, Mohamad M. – PersonEntity: Name: NameFull: Zuair, Mansour IsPartOfRelationships: – BibEntity: Dates: – D: 15 M: 09 Text: Sep2024 Type: published Y: 2024 Identifiers: – Type: issn-print Value: 01431161 Numbering: – Type: volume Value: 45 – Type: issue Value: 18 Titles: – TitleFull: International Journal of Remote Sensing Type: main |
| ResultId | 1 |