A two-stage text summarization method based on an improved PEGASUS model and adaptive error correction mechanism.

Saved in:
Bibliographic Details
Title: A two-stage text summarization method based on an improved PEGASUS model and adaptive error correction mechanism.
Authors: ZHANG, Hang1, WU, Jun1 wujun@yzu.edu.cn
Source: Computer Engineering & Science / Jisuanji Gongcheng yu Kexue. Feb2026, Vol. 48 Issue 2, p309-318. 10p.
Subjects: Text summarization, Automatic summarization, Hierarchical clustering (Cluster analysis), Language models, Recurrent neural networks, Machine learning
Abstract: To address the issues of word redundancy and poor readability in extractive summarization, as well as semantic confusion, logical inconsistency, and exposure bias in abstractive summarization, this paper proposes a two-stage text summarization method based on an improved PEGASUS model and an adaptive error correction mechanism, employing a hybrid summarization technique. In the extraction stage, text vectors are obtained using the BERT model, combined with a Bi-GRU and a graph structure. An improved MMR algorithm is utilized to effectively reduce redundancy in candidate summaries, enhancing summary precision. In the generation stage, the extracted sentences are processed by the PEGASUS model, incorporating hierarchical clustering technology and introducing an adaptive error correction mechanism to solve the out-of-vocabulary (OOV) problem. Additionally, a contrastive learning framework is adopted to significantly mitigate exposure bias. Experimental results demonstrate that the model established by our method achieves significant improvements in ROUGE scores on the NLPCC dataset, with average increases of 2.66 percentage points, 0.84 percentage points, and 1.81 percentage points across various metrics compared to models established by existing hybrid methods. This method not only improves summary quality but also exhibits superior performance in resolving OOV problem and exposure bias. [ABSTRACT FROM AUTHOR]
Copyright of Computer Engineering & Science / Jisuanji Gongcheng yu Kexue is the property of Computer Engineering & Science and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Description
Abstract:To address the issues of word redundancy and poor readability in extractive summarization, as well as semantic confusion, logical inconsistency, and exposure bias in abstractive summarization, this paper proposes a two-stage text summarization method based on an improved PEGASUS model and an adaptive error correction mechanism, employing a hybrid summarization technique. In the extraction stage, text vectors are obtained using the BERT model, combined with a Bi-GRU and a graph structure. An improved MMR algorithm is utilized to effectively reduce redundancy in candidate summaries, enhancing summary precision. In the generation stage, the extracted sentences are processed by the PEGASUS model, incorporating hierarchical clustering technology and introducing an adaptive error correction mechanism to solve the out-of-vocabulary (OOV) problem. Additionally, a contrastive learning framework is adopted to significantly mitigate exposure bias. Experimental results demonstrate that the model established by our method achieves significant improvements in ROUGE scores on the NLPCC dataset, with average increases of 2.66 percentage points, 0.84 percentage points, and 1.81 percentage points across various metrics compared to models established by existing hybrid methods. This method not only improves summary quality but also exhibits superior performance in resolving OOV problem and exposure bias. [ABSTRACT FROM AUTHOR]
ISSN:1007130X
DOI:10.3969/j.issn.1007-130X.2026.02.012