Learning to enhance areal video captioning with visual question answering.

Saved in:
Bibliographic Details
Title: Learning to enhance areal video captioning with visual question answering.
Authors: Al Mehmadi, Shima M.1 (AUTHOR), Bazi, Yakoub1 (AUTHOR), Al Rahhal, Mohamad M.2 (AUTHOR) mmalrahhal@ksu.edu.sa, Zuair, Mansour1 (AUTHOR)
Source: International Journal of Remote Sensing. Sep2024, Vol. 45 Issue 18, p6395-6407. 13p.
Subjects: Drone aircraft, Remote sensing, Videos
Abstract: The utilization of Unmanned Aerial Vehicles (UAV) in remote sensing (RS) has witnessed a significant surge, offering valuable insights into Earth dynamics and human activities. However, this has led to a substantial increase in the volume of video data, rendering manual screening and analysis impractical. Consequently, there is a pressing need for the development of automated interpretation models for these aerial videos. In this paper, we propose a novel approach that leverages visual dialogue to enhance aerial video captioning. Our model adopts an encoder-decoder architecture, integrating a Visual Question Answering (VQA) task before the captioning task. The VQA task aims to enrich the captioning process by soliciting additional information about the image content. Specifically, our video encoder utilizes ViT-L/16, while the decoder employs Generative Pre-trained Transformer-2 (Distill-GPT-2). To validate our model, we introduce a novel benchmark dataset named CapERA-VQA, comprising videos accompanied by sets of questions, answers, and captions. Through experimental validation, we demonstrate the effectiveness of our proposed approach in enhancing the automated captioning of aerial videos. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Remote Sensing is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 179637986
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Learning to enhance areal video captioning with visual question answering.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Al+Mehmadi%2C+Shima+M%2E%22">Al Mehmadi, Shima M.</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Bazi%2C+Yakoub%22">Bazi, Yakoub</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Al+Rahhal%2C+Mohamad+M%2E%22">Al Rahhal, Mohamad M.</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> mmalrahhal@ksu.edu.sa</i><br /><searchLink fieldCode="AR" term="%22Zuair%2C+Mansour%22">Zuair, Mansour</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22International+Journal+of+Remote+Sensing%22">International Journal of Remote Sensing</searchLink>. Sep2024, Vol. 45 Issue 18, p6395-6407. 13p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Drone+aircraft%22">Drone aircraft</searchLink><br /><searchLink fieldCode="DE" term="%22Remote+sensing%22">Remote sensing</searchLink><br /><searchLink fieldCode="DE" term="%22Videos%22">Videos</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The utilization of Unmanned Aerial Vehicles (UAV) in remote sensing (RS) has witnessed a significant surge, offering valuable insights into Earth dynamics and human activities. However, this has led to a substantial increase in the volume of video data, rendering manual screening and analysis impractical. Consequently, there is a pressing need for the development of automated interpretation models for these aerial videos. In this paper, we propose a novel approach that leverages visual dialogue to enhance aerial video captioning. Our model adopts an encoder-decoder architecture, integrating a Visual Question Answering (VQA) task before the captioning task. The VQA task aims to enrich the captioning process by soliciting additional information about the image content. Specifically, our video encoder utilizes ViT-L/16, while the decoder employs Generative Pre-trained Transformer-2 (Distill-GPT-2). To validate our model, we introduce a novel benchmark dataset named CapERA-VQA, comprising videos accompanied by sets of questions, answers, and captions. Through experimental validation, we demonstrate the effectiveness of our proposed approach in enhancing the automated captioning of aerial videos. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of International Journal of Remote Sensing is the property of Taylor & Francis Ltd and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=179637986
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/01431161.2024.2388875
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 6395
    Subjects:
      – SubjectFull: Drone aircraft
        Type: general
      – SubjectFull: Remote sensing
        Type: general
      – SubjectFull: Videos
        Type: general
    Titles:
      – TitleFull: Learning to enhance areal video captioning with visual question answering.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Al Mehmadi, Shima M.
      – PersonEntity:
          Name:
            NameFull: Bazi, Yakoub
      – PersonEntity:
          Name:
            NameFull: Al Rahhal, Mohamad M.
      – PersonEntity:
          Name:
            NameFull: Zuair, Mansour
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 15
              M: 09
              Text: Sep2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 01431161
          Numbering:
            – Type: volume
              Value: 45
            – Type: issue
              Value: 18
          Titles:
            – TitleFull: International Journal of Remote Sensing
              Type: main
ResultId 1