WAGRank: A word ranking model based on word attention graph for keyphrase extraction.

Saved in:
Bibliographic Details
Title: WAGRank: A word ranking model based on word attention graph for keyphrase extraction.
Authors: Bian, Rong1,2 (AUTHOR), Cheng, Bing1,3 (AUTHOR) bc2@amss.ac.cn
Source: Intelligent Data Analysis. Jul2025, Vol. 29 Issue 4, p866-888. 23p.
Subjects: Language models, Word frequency, Granger causality test, Statistical significance, Terms & phrases
Abstract: Keyphrase extraction is an essential task of identifying representative words or phrases in document processing. Main traditional models rely on each word frequency feature in a document and its associated corpus. There are two major limitations of the word frequency method: first, it fails to fully exploit semantic information in the document, that is, it is a bag-of-word method; second, it tends to be influenced by local word frequency in the short current text when the linked corpus is not available or incomplete. This paper proposes WAGRank, a novel unsupervised ranking model on a word attention graph, where nodes are words and edges are semantic relations between words. To assign edge weights, two interpretable statistical methods of assessing correlation strength between words are designed using attention mechanism. WAGRank depends on word semantics rather than frequency only in the current text, using external knowledge stored in a pre-trained language model. WAGRank was evaluated on two publicly available datasets against twelve baselines, presenting its effectiveness and robustness. Besides, the Granger causality test illustrated that word attention has a statistically significant predictive effect on word frequency, providing a more reasonable explanation for word frequency analysis. [ABSTRACT FROM AUTHOR]
Copyright of Intelligent Data Analysis is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 186498100
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: WAGRank: A word ranking model based on word attention graph for keyphrase extraction.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Bian%2C+Rong%22">Bian, Rong</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Cheng%2C+Bing%22">Cheng, Bing</searchLink><relatesTo>1,3</relatesTo> (AUTHOR)<i> bc2@amss.ac.cn</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Intelligent+Data+Analysis%22">Intelligent Data Analysis</searchLink>. Jul2025, Vol. 29 Issue 4, p866-888. 23p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Word+frequency%22">Word frequency</searchLink><br /><searchLink fieldCode="DE" term="%22Granger+causality+test%22">Granger causality test</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+significance%22">Statistical significance</searchLink><br /><searchLink fieldCode="DE" term="%22Terms+%26+phrases%22">Terms & phrases</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Keyphrase extraction is an essential task of identifying representative words or phrases in document processing. Main traditional models rely on each word frequency feature in a document and its associated corpus. There are two major limitations of the word frequency method: first, it fails to fully exploit semantic information in the document, that is, it is a bag-of-word method; second, it tends to be influenced by local word frequency in the short current text when the linked corpus is not available or incomplete. This paper proposes WAGRank, a novel unsupervised ranking model on a word attention graph, where nodes are words and edges are semantic relations between words. To assign edge weights, two interpretable statistical methods of assessing correlation strength between words are designed using attention mechanism. WAGRank depends on word semantics rather than frequency only in the current text, using external knowledge stored in a pre-trained language model. WAGRank was evaluated on two publicly available datasets against twelve baselines, presenting its effectiveness and robustness. Besides, the Granger causality test illustrated that word attention has a statistically significant predictive effect on word frequency, providing a more reasonable explanation for word frequency analysis. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Intelligent Data Analysis is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=186498100
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/1088467X241296257
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 23
        StartPage: 866
    Subjects:
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Word frequency
        Type: general
      – SubjectFull: Granger causality test
        Type: general
      – SubjectFull: Statistical significance
        Type: general
      – SubjectFull: Terms & phrases
        Type: general
    Titles:
      – TitleFull: WAGRank: A word ranking model based on word attention graph for keyphrase extraction.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Bian, Rong
      – PersonEntity:
          Name:
            NameFull: Cheng, Bing
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 07
              Text: Jul2025
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1088467X
          Numbering:
            – Type: volume
              Value: 29
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Intelligent Data Analysis
              Type: main
ResultId 1