Document vector embedding based extractive text summarization system for Hindi and English text.
Saved in:
| Title: | Document vector embedding based extractive text summarization system for Hindi and English text. |
|---|---|
| Authors: | Rani, Ruby1 (AUTHOR) ruby73_scs@jnu.ac.in, Lobiyal, D. K.1 (AUTHOR) |
| Source: | Applied Intelligence. Jun2022, Vol. 52 Issue 8, p9353-9372. 20p. |
| Subjects: | Text summarization, Latent semantic analysis, English fiction, Hindi language, Standard language |
| Abstract: | Nowadays, several automatic text summarization (ATS) methods have been proposed for resource-rich languages, such as English, Chinese. However, resource-limited languages like Hindi realized very little attention from researchers. The lack of resources still makes the ATS task for the Hindi language a challenging and open problem. Capturing semantic features and hidden relationships among the text units are the two main characteristics of an informative summary. In the current work, we propose an ATS model based on the document vector method to explore the semantic relations existing in the document. Moreover, we suggest two algorithms: sentence ranking and summary generation based on three main characteristics including, redundancy, diversity, and compression rate to create a clear and coherent summary. The proposed model is language-independent with some language-specific preprocessing. Further, we evaluate our model on two different language datasets as literary novels in Hindi and DUC 2007 news articles in English. We apply the ROUGE metric to measure the performance of the generated summaries. Besides, we also compare the proposed model against four baseline methods: TextRank, Lexrank, Latent Semantic Analysis (LSA), and Mudasir et al. models. The overall macro-Average F-Score (18.5% for Hindi, 26% for English) for very short length summaries of sizes 5% and 15% compression rates produced by our model is higher than the baseline approaches. In case of very lengthy summaries of size 50% compression rate, our model has the highest Macro-Average values, 18% for the Hindi novels and 25% for the English news articles against all the comparison methods. From the result analysis, we perceive that the proposed model beats all the baselines from the experimental outcomes and leads to diverse, least-redundant, semantic-rich, and compressed text summary generation. [ABSTRACT FROM AUTHOR] |
| Copyright of Applied Intelligence is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.) | |
| Database: | Engineering Source |
| FullText | Text: Availability: 0 |
|---|---|
| Header | DbId: egs DbLabel: Engineering Source An: 157151811 AccessLevel: 6 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Document vector embedding based extractive text summarization system for Hindi and English text. – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Rani%2C+Ruby%22">Rani, Ruby</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> ruby73_scs@jnu.ac.in</i><br /><searchLink fieldCode="AR" term="%22Lobiyal%2C+D%2E+K%2E%22">Lobiyal, D. K.</searchLink><relatesTo>1</relatesTo> (AUTHOR) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="JN" term="%22Applied+Intelligence%22">Applied Intelligence</searchLink>. Jun2022, Vol. 52 Issue 8, p9353-9372. 20p. – Name: Subject Label: Subjects Group: Su Data: <searchLink fieldCode="DE" term="%22Text+summarization%22">Text summarization</searchLink><br /><searchLink fieldCode="DE" term="%22Latent+semantic+analysis%22">Latent semantic analysis</searchLink><br /><searchLink fieldCode="DE" term="%22English+fiction%22">English fiction</searchLink><br /><searchLink fieldCode="DE" term="%22Hindi+language%22">Hindi language</searchLink><br /><searchLink fieldCode="DE" term="%22Standard+language%22">Standard language</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: Nowadays, several automatic text summarization (ATS) methods have been proposed for resource-rich languages, such as English, Chinese. However, resource-limited languages like Hindi realized very little attention from researchers. The lack of resources still makes the ATS task for the Hindi language a challenging and open problem. Capturing semantic features and hidden relationships among the text units are the two main characteristics of an informative summary. In the current work, we propose an ATS model based on the document vector method to explore the semantic relations existing in the document. Moreover, we suggest two algorithms: sentence ranking and summary generation based on three main characteristics including, redundancy, diversity, and compression rate to create a clear and coherent summary. The proposed model is language-independent with some language-specific preprocessing. Further, we evaluate our model on two different language datasets as literary novels in Hindi and DUC 2007 news articles in English. We apply the ROUGE metric to measure the performance of the generated summaries. Besides, we also compare the proposed model against four baseline methods: TextRank, Lexrank, Latent Semantic Analysis (LSA), and Mudasir et al. models. The overall macro-Average F-Score (18.5% for Hindi, 26% for English) for very short length summaries of sizes 5% and 15% compression rates produced by our model is higher than the baseline approaches. In case of very lengthy summaries of size 50% compression rate, our model has the highest Macro-Average values, 18% for the Hindi novels and 25% for the English news articles against all the comparison methods. From the result analysis, we perceive that the proposed model beats all the baselines from the experimental outcomes and leads to diverse, least-redundant, semantic-rich, and compressed text summary generation. [ABSTRACT FROM AUTHOR] – Name: AbstractSuppliedCopyright Label: Group: Ab Data: <i>Copyright of Applied Intelligence is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.) |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=157151811 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10489-021-02871-9 Languages: – Code: eng Text: English PhysicalDescription: Pagination: PageCount: 20 StartPage: 9353 Subjects: – SubjectFull: Text summarization Type: general – SubjectFull: Latent semantic analysis Type: general – SubjectFull: English fiction Type: general – SubjectFull: Hindi language Type: general – SubjectFull: Standard language Type: general Titles: – TitleFull: Document vector embedding based extractive text summarization system for Hindi and English text. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Rani, Ruby – PersonEntity: Name: NameFull: Lobiyal, D. K. IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Text: Jun2022 Type: published Y: 2022 Identifiers: – Type: issn-print Value: 0924669X Numbering: – Type: volume Value: 52 – Type: issue Value: 8 Titles: – TitleFull: Applied Intelligence Type: main |
| ResultId | 1 |