Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence.

Saved in:
Bibliographic Details
Title: Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence.
Authors: SURABHI, ANURADHA1 surabhiaim12023@gmail.com, MARTHA, SHESHIKALA1
Source: Journal of Information Science & Engineering. May2026, Vol. 42 Issue 3, p793-804. 12p.
Subjects: Text summarization, Recall (Information retrieval), Natural language processing, Statistical accuracy, Language models, Benchmark problems (Computer science), Cohesion (Linguistics)
Abstract: This study investigates the effectiveness of abstractive text summarization in the context of scientific documents using 40 diverse Large Language Models (LLMs). Unlike traditional extractive approaches that often produoe fragmented and loss coherent summaries, our work focuses on enhancing semantic fidelity, coherence, and comprehensive content coverage. Through a recall-oriented evaluation supported by BERT and METEOR metrics, our experimental results show that models such as Claude v2.1, Qwen-14B, Zephyr-7B, and Phi-3 emerged as top performers, achieving outstanding Fl scores above 0.93 and METEOR scores as high as 1.00. These models demonstrated a strong ability to retain critical information while producing fluent, human-like summaries. Our findings provide valuable benchmarks for selecting high-performing LLMs in summarization tasks and offer a foundation for future advancements, including domain adaptation, fact-checking integration, and multimodal summarization approaches in real-world Natural Language Processing applications. [ABSTRACT FROM AUTHOR]
Copyright of Journal of Information Science & Engineering is the property of Institute of Information Science, Academia Sinica and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Links:
  – Type: pdflink
Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 193961070
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22SURABHI%2C+ANURADHA%22">SURABHI, ANURADHA</searchLink><relatesTo>1</relatesTo><i> surabhiaim12023@gmail.com</i><br /><searchLink fieldCode="AR" term="%22MARTHA%2C+SHESHIKALA%22">MARTHA, SHESHIKALA</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Journal+of+Information+Science+%26+Engineering%22">Journal of Information Science & Engineering</searchLink>. May2026, Vol. 42 Issue 3, p793-804. 12p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Text+summarization%22">Text summarization</searchLink><br /><searchLink fieldCode="DE" term="%22Recall+%28Information+retrieval%29%22">Recall (Information retrieval)</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+language+processing%22">Natural language processing</searchLink><br /><searchLink fieldCode="DE" term="%22Statistical+accuracy%22">Statistical accuracy</searchLink><br /><searchLink fieldCode="DE" term="%22Language+models%22">Language models</searchLink><br /><searchLink fieldCode="DE" term="%22Benchmark+problems+%28Computer+science%29%22">Benchmark problems (Computer science)</searchLink><br /><searchLink fieldCode="DE" term="%22Cohesion+%28Linguistics%29%22">Cohesion (Linguistics)</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study investigates the effectiveness of abstractive text summarization in the context of scientific documents using 40 diverse Large Language Models (LLMs). Unlike traditional extractive approaches that often produoe fragmented and loss coherent summaries, our work focuses on enhancing semantic fidelity, coherence, and comprehensive content coverage. Through a recall-oriented evaluation supported by BERT and METEOR metrics, our experimental results show that models such as Claude v2.1, Qwen-14B, Zephyr-7B, and Phi-3 emerged as top performers, achieving outstanding Fl scores above 0.93 and METEOR scores as high as 1.00. These models demonstrated a strong ability to retain critical information while producing fluent, human-like summaries. Our findings provide valuable benchmarks for selecting high-performing LLMs in summarization tasks and offer a foundation for future advancements, including domain adaptation, fact-checking integration, and multimodal summarization approaches in real-world Natural Language Processing applications. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Journal of Information Science & Engineering is the property of Institute of Information Science, Academia Sinica and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=193961070
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.6688/JISE.202605_42(3).0018
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 793
    Subjects:
      – SubjectFull: Text summarization
        Type: general
      – SubjectFull: Recall (Information retrieval)
        Type: general
      – SubjectFull: Natural language processing
        Type: general
      – SubjectFull: Statistical accuracy
        Type: general
      – SubjectFull: Language models
        Type: general
      – SubjectFull: Benchmark problems (Computer science)
        Type: general
      – SubjectFull: Cohesion (Linguistics)
        Type: general
    Titles:
      – TitleFull: Evaluating Large Language Models for Abstractive Summarization: A Benchmark Study on Recall, Fidelity, and Content Coherence.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: SURABHI, ANURADHA
      – PersonEntity:
          Name:
            NameFull: MARTHA, SHESHIKALA
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 05
              Text: May2026
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 10162364
          Numbering:
            – Type: volume
              Value: 42
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Journal of Information Science & Engineering
              Type: main
ResultId 1