WAGRank: A word ranking model based on word attention graph for keyphrase extraction.
Saved in:
| Title: | WAGRank: A word ranking model based on word attention graph for keyphrase extraction. |
|---|---|
| Authors: | Bian, Rong1,2 (AUTHOR), Cheng, Bing1,3 (AUTHOR) bc2@amss.ac.cn |
| Source: | Intelligent Data Analysis. Jul2025, Vol. 29 Issue 4, p866-888. 23p. |
| Subjects: | Language models, Word frequency, Granger causality test, Statistical significance, Terms & phrases |
| Abstract: | Keyphrase extraction is an essential task of identifying representative words or phrases in document processing. Main traditional models rely on each word frequency feature in a document and its associated corpus. There are two major limitations of the word frequency method: first, it fails to fully exploit semantic information in the document, that is, it is a bag-of-word method; second, it tends to be influenced by local word frequency in the short current text when the linked corpus is not available or incomplete. This paper proposes WAGRank, a novel unsupervised ranking model on a word attention graph, where nodes are words and edges are semantic relations between words. To assign edge weights, two interpretable statistical methods of assessing correlation strength between words are designed using attention mechanism. WAGRank depends on word semantics rather than frequency only in the current text, using external knowledge stored in a pre-trained language model. WAGRank was evaluated on two publicly available datasets against twelve baselines, presenting its effectiveness and robustness. Besides, the Granger causality test illustrated that word attention has a statistically significant predictive effect on word frequency, providing a more reasonable explanation for word frequency analysis. [ABSTRACT FROM AUTHOR] |
| Copyright of Intelligent Data Analysis is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Text: Availability: 1 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 186498100 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: WAGRank: A word ranking model based on word attention graph for keyphrase extraction. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Bian%2C+Rong%22">Bian, Rong</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Cheng%2C+Bing%22">Cheng, Bing</searchLink><relatesTo>1,3</relatesTo> (AUTHOR)<i> bc2@amss.ac.cn</i> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Intelligent+Data+Analysis%22">Intelligent Data Analysis</searchLink>. Jul2025, Vol. 29 Issue 4, p866-888. 23p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Word+frequency%22">Word frequency</searchLink><br /><searchLink fieldCode="DE" term="%22Granger+causality+test%22">Granger causality test</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+significance%22">Statistical significance</searchLink><br /><searchLink fieldCode="DE" term="%22Terms+%26+phrases%22">Terms & phrases</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Keyphrase extraction is an essential task of identifying representative words or phrases in document processing. Main traditional models rely on each word frequency feature in a document and its associated corpus. There are two major limitations of the word frequency method: first, it fails to fully exploit semantic information in the document, that is, it is a bag-of-word method; second, it tends to be influenced by local word frequency in the short current text when the linked corpus is not available or incomplete. This paper proposes WAGRank, a novel unsupervised ranking model on a word attention graph, where nodes are words and edges are semantic relations between words. To assign edge weights, two interpretable statistical methods of assessing correlation strength between words are designed using attention mechanism. WAGRank depends on word semantics rather than frequency only in the current text, using external knowledge stored in a pre-trained language model. WAGRank was evaluated on two publicly available datasets against twelve baselines, presenting its effectiveness and robustness. Besides, the Granger causality test illustrated that word attention has a statistically significant predictive effect on word frequency, providing a more reasonable explanation for word frequency analysis. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Intelligent Data Analysis is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=186498100 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1177/1088467X241296257 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 23 StartPage: 866 Subjects: – SubjectFull: Language models Type: general – SubjectFull: Word frequency Type: general – SubjectFull: Granger causality test Type: general – SubjectFull: Statistical significance Type: general – SubjectFull: Terms & phrases Type: general Titles: – TitleFull: WAGRank: A word ranking model based on word attention graph for keyphrase extraction. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Bian, Rong – PersonEntity: Name: NameFull: Cheng, Bing IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 07 Text: Jul2025 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1088467X Numbering: – Type: volume Value: 29 – Type: issue Value: 4 Titles: – TitleFull: Intelligent Data Analysis Type: main |
| ResultId | 1 |