A Method of Extractive Text Summarization Using Document Semantic Graph With Node Ranking.

Saved in:
Bibliographic Details
Title: A Method of Extractive Text Summarization Using Document Semantic Graph With Node Ranking.
Authors: Li, Zhenhao1 (AUTHOR), Liu, Miao1 (AUTHOR) liumiao@gzhu.edu.cn, Chen, Wenbin1 (AUTHOR), Zheng, Ligang1 (AUTHOR), Murray, Richard1 (AUTHOR) rmurray@wiley.com
Source: International Journal of Intelligent Systems. 11/21/2025, Vol. 2025, p1-19. 19p.
Subjects: Automatic summarization, Text summarization, Artificial neural networks, Conceptual structures
Abstract: With the rise of neural networks and pre‐trained models such as BERT, abstractive text summarization techniques have received widespread attention. Nevertheless, traditional extractive text summarization methods still hold substantial research value due to their low computational cost, interpretability, and robustness. In algorithms like TextRank and its variants, graph nodes are typically constructed based on surface‐level lexical features. These graphs often fail to incorporate many contextual relationships, such as coreference relationships among nodes, resulting in fragmented representations of key concepts. For edge construction, a sliding window of size T is commonly used to connect word nodes within the window. However, these methods often fall short in modeling the rich contextual dependencies embedded in the document. Several recent studies have demonstrated that semantic graphs can effectively improve the accuracy of text summarization. In this paper, we construct a more interpretable semantic graph from syntax trees and propose a novel unsupervised algorithm based on the personalized PageRank algorithm for summary extraction. We utilize tree transformation methods to enrich word‐level information for graph construction, define node‐merging rules to reduce graph complexity, use coreference chains to merge coreferring entities across sentences for enriching contextual links, and introduce the concept of Meta Node sets to capture thematic relationships that are not fully represented by syntactic dependencies or coreference chains alone. By clustering semantically related words, Meta Nodes enhance the graph's ability to reflect deeper contextual coherence across the document. Compared with previous TextRank‐based methods, our improvement yields significant ROUGE score boosts on the CNN‐DM dataset. While the method was developed and evaluated using English‐language datasets, its underlying design is language agnostic and can be adapted to other languages with suitable linguistic tools. [ABSTRACT FROM AUTHOR]
Copyright of International Journal of Intelligent Systems is the property of Wiley-Blackwell and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
Description
Abstract:With the rise of neural networks and pre‐trained models such as BERT, abstractive text summarization techniques have received widespread attention. Nevertheless, traditional extractive text summarization methods still hold substantial research value due to their low computational cost, interpretability, and robustness. In algorithms like TextRank and its variants, graph nodes are typically constructed based on surface‐level lexical features. These graphs often fail to incorporate many contextual relationships, such as coreference relationships among nodes, resulting in fragmented representations of key concepts. For edge construction, a sliding window of size T is commonly used to connect word nodes within the window. However, these methods often fall short in modeling the rich contextual dependencies embedded in the document. Several recent studies have demonstrated that semantic graphs can effectively improve the accuracy of text summarization. In this paper, we construct a more interpretable semantic graph from syntax trees and propose a novel unsupervised algorithm based on the personalized PageRank algorithm for summary extraction. We utilize tree transformation methods to enrich word‐level information for graph construction, define node‐merging rules to reduce graph complexity, use coreference chains to merge coreferring entities across sentences for enriching contextual links, and introduce the concept of Meta Node sets to capture thematic relationships that are not fully represented by syntactic dependencies or coreference chains alone. By clustering semantically related words, Meta Nodes enhance the graph's ability to reflect deeper contextual coherence across the document. Compared with previous TextRank‐based methods, our improvement yields significant ROUGE score boosts on the CNN‐DM dataset. While the method was developed and evaluated using English‐language datasets, its underlying design is language agnostic and can be adapted to other languages with suitable linguistic tools. [ABSTRACT FROM AUTHOR]
ISSN:08848173
DOI:10.1155/int/5530784