Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence.
Saved in:
| Title: | Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence. |
|---|---|
| Authors: | SURABHI, ANURADHA1 surabhiaim12023@gmail.com, MARTHA, SHESHIKALA1 |
| Source: | Journal of Information Science & Engineering. May2026, Vol. 42 Issue 3, p793-804. 12p. |
| Subjects: | Text summarization, Recall (Information retrieval), Natural language processing, Statistical accuracy, Language models, Benchmark problems (Computer science), Cohesion (Linguistics) |
| Abstract: | This study investigates the effectiveness of abstractive text summarization in the context of scientific documents using 40 diverse Large Language Models (LLMs). Unlike traditional extractive approaches that often produoe fragmented and loss coherent summaries, our work focuses on enhancing semantic fidelity, coherence, and comprehensive content coverage. Through a recall-oriented evaluation supported by BERT and METEOR metrics, our experimental results show that models such as Claude v2.1, Qwen-14B, Zephyr-7B, and Phi-3 emerged as top performers, achieving outstanding Fl scores above 0.93 and METEOR scores as high as 1.00. These models demonstrated a strong ability to retain critical information while producing fluent, human-like summaries. Our findings provide valuable benchmarks for selecting high-performing LLMs in summarization tasks and offer a foundation for future advancements, including domain adaptation, fact-checking integration, and multimodal summarization approaches in real-world Natural Language Processing applications. [ABSTRACT FROM AUTHOR] |
| Copyright of Journal of Information Science & Engineering is the property of Institute of Information Science, Academia Sinica and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Links: – Type: pdflink Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 193961070 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22SURABHI%2C+ANURADHA%22">SURABHI, ANURADHA</searchLink><relatesTo>1</relatesTo><i> surabhiaim12023@gmail.com</i><br /><searchLink fieldCode="AR" term="%22MARTHA%2C+SHESHIKALA%22">MARTHA, SHESHIKALA</searchLink><relatesTo>1</relatesTo> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Journal+of+Information+Science+%26+Engineering%22">Journal of Information Science & Engineering</searchLink>. May2026, Vol. 42 Issue 3, p793-804. 12p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Text+summarization%22">Text summarization</searchLink><br /><searchLink fieldCode="DE" term="%22Recall+%28Information+retrieval%29%22">Recall (Information retrieval)</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+accuracy%22">Statistical accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Benchmark+problems+%28Computer+science%29%22">Benchmark problems (Computer science)</searchLink><br /><searchLink fieldCode="DE" term="%22Cohesion+%28Linguistics%29%22">Cohesion (Linguistics)</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: This study investigates the effectiveness of abstractive text summarization in the context of scientific documents using 40 diverse Large Language Models (LLMs). Unlike traditional extractive approaches that often produoe fragmented and loss coherent summaries, our work focuses on enhancing semantic fidelity, coherence, and comprehensive content coverage. Through a recall-oriented evaluation supported by BERT and METEOR metrics, our experimental results show that models such as Claude v2.1, Qwen-14B, Zephyr-7B, and Phi-3 emerged as top performers, achieving outstanding Fl scores above 0.93 and METEOR scores as high as 1.00. These models demonstrated a strong ability to retain critical information while producing fluent, human-like summaries. Our findings provide valuable benchmarks for selecting high-performing LLMs in summarization tasks and offer a foundation for future advancements, including domain adaptation, fact-checking integration, and multimodal summarization approaches in real-world Natural Language Processing applications. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Journal of Information Science & Engineering is the property of Institute of Information Science, Academia Sinica and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193961070 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.6688/JISE.202605_42(3).0018 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 12 StartPage: 793 Subjects: – SubjectFull: Text summarization Type: general – SubjectFull: Recall (Information retrieval) Type: general – SubjectFull: Natural language processing Type: general – SubjectFull: Statistical accuracy Type: general – SubjectFull: Language models Type: general – SubjectFull: Benchmark problems (Computer science) Type: general – SubjectFull: Cohesion (Linguistics) Type: general Titles: – TitleFull: Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: SURABHI, ANURADHA – PersonEntity: Name: NameFull: MARTHA, SHESHIKALA IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 05 Text: May2026 Type: published Y: 2026 Identifiers: – Type: issn-print Value: 10162364 Numbering: – Type: volume Value: 42 – Type: issue Value: 3 Titles: – TitleFull: Journal of Information Science & Engineering Type: main |
| ResultId | 1 |