Cross-modal spatio-temporal fusion weakly supervised video anomaly detection based on large-scale vision-language models.

Saved in:
Bibliographic Details
Title: Cross-modal spatio-temporal fusion weakly supervised video anomaly detection based on large-scale vision-language models.
Authors: Pan, Lihu1 (AUTHOR) panlh@tyust.edu.cn, Peng, Shouxin1 (AUTHOR) 2692446165@qq.com, Zhang, Rui1 (AUTHOR) zhangrui@tyust.edu.cn, Xu, Sendren Sheng-Dong2 (AUTHOR) sdxu@pme.nthu.edu.tw, Xie, Binhong1 (AUTHOR) binhongxie@tyust.edu.cn
Source: Multimedia Systems. Apr2026, Vol. 32 Issue 2, p1-16. 16p.
Subjects: Large scale systems, Spatial analysis (Statistics), Supervised learning, Metadata, Signal detection
Abstract: Video anomaly detection plays a vital role in the development of intelligent industry and smart city by effectively identifying and responding to abnormal events, which enhances production efficiency and the safety of urban operations. Existing large-scale visual-language models still suffer from insufficient extraction of spatio-temporal features and semantic disconnection in the fusion of multimodal data. To address this, this paper proposes a weakly supervised video anomaly detection method based on large-scale visual-language models with cross-modal spatio-temporal fusion (WSDL-CAM). In this task, we designed a Spatio-Temporal Context Adapter Module (STCA) that effectively uses the Mamba network's ability to capture long-range dependencies in the spatio-temporal context information extracted by the spatio-temporal graph convolutional network, achieving sufficient extraction of spatio-temporal features from the video. In addition, we designed an Image-Text Semantic Perception Fusion Module (ITSPF) that associates fine-grained visual features with text features to judge semantic consistency, further enhancing the semantic relevance between images and text. Finally, we implemented hyperparameter self-optimization in training and evaluation through Bayesian optimization algorithms. We conducted extensive experiments on two large-scale datasets with real-world surveillance scenarios (XD-Violence and UCF-Crime). Experimental results show that our model has achieved state-of-the-art performance in video anomaly detection. [ABSTRACT FROM AUTHOR]
Copyright of Multimedia Systems is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 191288860
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Cross-modal spatio-temporal fusion weakly supervised video anomaly detection based on large-scale vision-language models.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Pan%2C+Lihu%22">Pan, Lihu</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> panlh@tyust.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Peng%2C+Shouxin%22">Peng, Shouxin</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> 2692446165@qq.com</i><br /><searchLink fieldCode="AR" term="%22Zhang%2C+Rui%22">Zhang, Rui</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> zhangrui@tyust.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Xu%2C+Sendren+Sheng-Dong%22">Xu, Sendren Sheng-Dong</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> sdxu@pme.nthu.edu.tw</i><br /><searchLink fieldCode="AR" term="%22Xie%2C+Binhong%22">Xie, Binhong</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> binhongxie@tyust.edu.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Multimedia+Systems%22">Multimedia Systems</searchLink>. Apr2026, Vol. 32 Issue 2, p1-16. 16p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Large+scale+systems%22">Large scale systems</searchLink><br /><searchLink fieldCode="DE" term="%22Spatial+analysis+%28Statistics%29%22">Spatial analysis (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Supervised+learning%22">Supervised learning</searchLink><br /><searchLink fieldCode="DE" term="%22Metadata%22">Metadata</searchLink><br /><searchLink fieldCode="DE" term="%22Signal+detection%22">Signal detection</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Video anomaly detection plays a vital role in the development of intelligent industry and smart city by effectively identifying and responding to abnormal events, which enhances production efficiency and the safety of urban operations. Existing large-scale visual-language models still suffer from insufficient extraction of spatio-temporal features and semantic disconnection in the fusion of multimodal data. To address this, this paper proposes a weakly supervised video anomaly detection method based on large-scale visual-language models with cross-modal spatio-temporal fusion (WSDL-CAM). In this task, we designed a Spatio-Temporal Context Adapter Module (STCA) that effectively uses the Mamba network's ability to capture long-range dependencies in the spatio-temporal context information extracted by the spatio-temporal graph convolutional network, achieving sufficient extraction of spatio-temporal features from the video. In addition, we designed an Image-Text Semantic Perception Fusion Module (ITSPF) that associates fine-grained visual features with text features to judge semantic consistency, further enhancing the semantic relevance between images and text. Finally, we implemented hyperparameter self-optimization in training and evaluation through Bayesian optimization algorithms. We conducted extensive experiments on two large-scale datasets with real-world surveillance scenarios (XD-Violence and UCF-Crime). Experimental results show that our model has achieved state-of-the-art performance in video anomaly detection. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Multimedia Systems is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=191288860
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s00530-025-02158-w
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 16
        StartPage: 1
    Subjects:
      – SubjectFull: Large scale systems
        Type: general
      – SubjectFull: Spatial analysis (Statistics)
        Type: general
      – SubjectFull: Supervised learning
        Type: general
      – SubjectFull: Metadata
        Type: general
      – SubjectFull: Signal detection
        Type: general
    Titles:
      – TitleFull: Cross-modal spatio-temporal fusion weakly supervised video anomaly detection based on large-scale vision-language models.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Pan, Lihu
      – PersonEntity:
          Name:
            NameFull: Peng, Shouxin
      – PersonEntity:
          Name:
            NameFull: Zhang, Rui
      – PersonEntity:
          Name:
            NameFull: Xu, Sendren Sheng-Dong
      – PersonEntity:
          Name:
            NameFull: Xie, Binhong
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 04
              Text: Apr2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 09424962
          Numbering:
            – Type: volume
              Value: 32
            – Type: issue
              Value: 2
          Titles:
            – TitleFull: Multimedia Systems
              Type: main
ResultId 1