Cross-modal progressive modeling for neuro-visual representation learning.

Saved in:
Bibliographic Details
Title: Cross-modal progressive modeling for neuro-visual representation learning.
Authors: Sun, Yueming1 (AUTHOR), Pu, Jiyao1 (AUTHOR), Sun, Kaili1 (AUTHOR), Fu, Zeyu2 (AUTHOR), Duan, Haoran3 (AUTHOR), Long, Yang1 (AUTHOR) yang.long@durham.ac.uk
Source: Neurocomputing. Jun2026, Vol. 683, pN.PAG-N.PAG. 1p.
Subjects: Electroencephalography, Neurophysiology
Abstract: Neural decoding from scalp signals requires models that respect spatial, temporal, and spectral structure while leveraging strong visual priors. In this paper, we introduce CFT-NET for disentangled neural visual representation together with a progressive visual–semantic adaptation (PVSA) framework that aligns EEG embeddings to pretrained visual backbones under a contrastive objective followed by pairwise matching. CFT-NET integrates Frequency-Separated Weights (FSW), Spatial-Context Aggregation (SCA), and Adaptive Temporal Filtering (ATF) to explicitly extract spectral, spatial, and temporal factors. PVSA consists of an instance-guided visual encoder and a visual-guided semantic decoder linked by cross attention, enabling fine-grained neuro–image interaction. On THINGS-EEG and THINGS-MEG datasets, the approach consistently outperforms state-of-the-art baselines in both subject-dependent and subject-independent zero-shot classification. By aligning model architecture with visual cognition principles and coupling it to strong visual priors, our method narrows the gap between neural activity and visual cognition, provides novel cross-modal neural decoding method that achieves competitive performance against recent state-of-the-art baselines. [ABSTRACT FROM AUTHOR]
Copyright of Neurocomputing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 193147630
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Cross-modal progressive modeling for neuro-visual representation learning.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Sun%2C+Yueming%22">Sun, Yueming</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Pu%2C+Jiyao%22">Pu, Jiyao</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Sun%2C+Kaili%22">Sun, Kaili</searchLink><relatesTo>1</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Fu%2C+Zeyu%22">Fu, Zeyu</searchLink><relatesTo>2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Duan%2C+Haoran%22">Duan, Haoran</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Long%2C+Yang%22">Long, Yang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> yang.long@durham.ac.uk</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Neurocomputing%22">Neurocomputing</searchLink>. Jun2026, Vol. 683, pN.PAG-N.PAG. 1p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Electroencephalography%22">Electroencephalography</searchLink><br /><searchLink fieldCode="DE" term="%22Neurophysiology%22">Neurophysiology</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Neural decoding from scalp signals requires models that respect spatial, temporal, and spectral structure while leveraging strong visual priors. In this paper, we introduce CFT-NET for disentangled neural visual representation together with a progressive visual–semantic adaptation (PVSA) framework that aligns EEG embeddings to pretrained visual backbones under a contrastive objective followed by pairwise matching. CFT-NET integrates Frequency-Separated Weights (FSW), Spatial-Context Aggregation (SCA), and Adaptive Temporal Filtering (ATF) to explicitly extract spectral, spatial, and temporal factors. PVSA consists of an instance-guided visual encoder and a visual-guided semantic decoder linked by cross attention, enabling fine-grained neuro–image interaction. On THINGS-EEG and THINGS-MEG datasets, the approach consistently outperforms state-of-the-art baselines in both subject-dependent and subject-independent zero-shot classification. By aligning model architecture with visual cognition principles and coupling it to strong visual priors, our method narrows the gap between neural activity and visual cognition, provides novel cross-modal neural decoding method that achieves competitive performance against recent state-of-the-art baselines. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Neurocomputing is the property of Elsevier B.V. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193147630
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1016/j.neucom.2026.133450
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 1
        StartPage: N.PAG
    Subjects:
      – SubjectFull: Electroencephalography
        Type: general
      – SubjectFull: Neurophysiology
        Type: general
    Titles:
      – TitleFull: Cross-modal progressive modeling for neuro-visual representation learning.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Sun, Yueming
      – PersonEntity:
          Name:
            NameFull: Pu, Jiyao
      – PersonEntity:
          Name:
            NameFull: Sun, Kaili
      – PersonEntity:
          Name:
            NameFull: Fu, Zeyu
      – PersonEntity:
          Name:
            NameFull: Duan, Haoran
      – PersonEntity:
          Name:
            NameFull: Long, Yang
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 28
              M: 06
              Text: Jun2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 09252312
          Numbering:
            – Type: volume
              Value: 683
          Titles:
            – TitleFull: Neurocomputing
              Type: main
ResultId 1