Document vector embedding based extractive text summarization system for Hindi and English text.

Saved in:
Bibliographic Details
Title: Document vector embedding based extractive text summarization system for Hindi and English text.
Authors: Rani, Ruby1 (AUTHOR) ruby73_scs@jnu.ac.in, Lobiyal, D. K.1 (AUTHOR)
Source: Applied Intelligence. Jun2022, Vol. 52 Issue 8, p9353-9372. 20p.
Subjects: Text summarization, Latent semantic analysis, English fiction, Hindi language, Standard language
Abstract: Nowadays, several automatic text summarization (ATS) methods have been proposed for resource-rich languages, such as English, Chinese. However, resource-limited languages like Hindi realized very little attention from researchers. The lack of resources still makes the ATS task for the Hindi language a challenging and open problem. Capturing semantic features and hidden relationships among the text units are the two main characteristics of an informative summary. In the current work, we propose an ATS model based on the document vector method to explore the semantic relations existing in the document. Moreover, we suggest two algorithms: sentence ranking and summary generation based on three main characteristics including, redundancy, diversity, and compression rate to create a clear and coherent summary. The proposed model is language-independent with some language-specific preprocessing. Further, we evaluate our model on two different language datasets as literary novels in Hindi and DUC 2007 news articles in English. We apply the ROUGE metric to measure the performance of the generated summaries. Besides, we also compare the proposed model against four baseline methods: TextRank, Lexrank, Latent Semantic Analysis (LSA), and Mudasir et al. models. The overall macro-Average F-Score (18.5% for Hindi, 26% for English) for very short length summaries of sizes 5% and 15% compression rates produced by our model is higher than the baseline approaches. In case of very lengthy summaries of size 50% compression rate, our model has the highest Macro-Average values, 18% for the Hindi novels and 25% for the English news articles against all the comparison methods. From the result analysis, we perceive that the proposed model beats all the baselines from the experimental outcomes and leads to diverse, least-redundant, semantic-rich, and compressed text summary generation. [ABSTRACT FROM AUTHOR]
Copyright of Applied Intelligence is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 157151811
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Document vector embedding based extractive text summarization system for Hindi and English text.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Rani%2C+Ruby%22">Rani, Ruby</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> ruby73_scs@jnu.ac.in</i><br /><searchLink fieldCode="AR" term="%22Lobiyal%2C+D%2E+K%2E%22">Lobiyal, D. K.</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Applied+Intelligence%22">Applied Intelligence</searchLink>. Jun2022, Vol. 52 Issue 8, p9353-9372. 20p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Text+summarization%22">Text summarization</searchLink><br /><searchLink fieldCode="DE" term="%22Latent+semantic+analysis%22">Latent semantic analysis</searchLink><br /><searchLink fieldCode="DE" term="%22English+fiction%22">English fiction</searchLink><br /><searchLink fieldCode="DE" term="%22Hindi+language%22">Hindi language</searchLink><br /><searchLink fieldCode="DE" term="%22Standard+language%22">Standard language</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Nowadays, several automatic text summarization (ATS) methods have been proposed for resource-rich languages, such as English, Chinese. However, resource-limited languages like Hindi realized very little attention from researchers. The lack of resources still makes the ATS task for the Hindi language a challenging and open problem. Capturing semantic features and hidden relationships among the text units are the two main characteristics of an informative summary. In the current work, we propose an ATS model based on the document vector method to explore the semantic relations existing in the document. Moreover, we suggest two algorithms: sentence ranking and summary generation based on three main characteristics including, redundancy, diversity, and compression rate to create a clear and coherent summary. The proposed model is language-independent with some language-specific preprocessing. Further, we evaluate our model on two different language datasets as literary novels in Hindi and DUC 2007 news articles in English. We apply the ROUGE metric to measure the performance of the generated summaries. Besides, we also compare the proposed model against four baseline methods: TextRank, Lexrank, Latent Semantic Analysis (LSA), and Mudasir et al. models. The overall macro-Average F-Score (18.5% for Hindi, 26% for English) for very short length summaries of sizes 5% and 15% compression rates produced by our model is higher than the baseline approaches. In case of very lengthy summaries of size 50% compression rate, our model has the highest Macro-Average values, 18% for the Hindi novels and 25% for the English news articles against all the comparison methods. From the result analysis, we perceive that the proposed model beats all the baselines from the experimental outcomes and leads to diverse, least-redundant, semantic-rich, and compressed text summary generation. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Applied Intelligence is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=157151811
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10489-021-02871-9
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 9353
    Subjects:
      – SubjectFull: Text summarization
        Type: general
      – SubjectFull: Latent semantic analysis
        Type: general
      – SubjectFull: English fiction
        Type: general
      – SubjectFull: Hindi language
        Type: general
      – SubjectFull: Standard language
        Type: general
    Titles:
      – TitleFull: Document vector embedding based extractive text summarization system for Hindi and English text.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Rani, Ruby
      – PersonEntity:
          Name:
            NameFull: Lobiyal, D. K.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 06
              Text: Jun2022
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 0924669X
          Numbering:
            – Type: volume
              Value: 52
            – Type: issue
              Value: 8
          Titles:
            – TitleFull: Applied Intelligence
              Type: main
ResultId 1