Likelihood corpus distribution: an efficient topic modelling scheme for Bengali document class identification.

Saved in:
Bibliographic Details
Title: Likelihood corpus distribution: an efficient topic modelling scheme for Bengali document class identification.
Authors: Das Dawn, Debapratim1 (AUTHOR) debapratimdd@gmail.com, Khan, Abhinandan1,2 (AUTHOR), Shaikh, Soharab Hossain3 (AUTHOR), Pal, Rajat Kumar1 (AUTHOR)
Source: Sādhanā: Academy Proceedings in Engineering Sciences. Sep2024, Vol. 49 Issue 3, p1-19. 19p.
Subjects: Identification documents, Latent semantic analysis, Artificial intelligence, Corpora, Library science, Document clustering, Sports sciences
Abstract: The learning quality of humans depends on the sense of contemplation. Textual documents are a huge part of the literature on contemplation which effortlessly creates perception. Automatic document class identification or organisation is a machine learning function to understand the psychological and emotional content of the text in a concise way. The problem of identification of documents falls in the field of library science, information science and artificial intelligence. The research progress of class identification of documents has been made in various most spoken languages. Numerous research works have been published in European and Asian languages. However, there is a gap in the literature when it comes to any less resource language, especially Bengali. Consequently, this work portrays an efficient topic modelling approach for Bengali document class identification. It proposes a Dirichlet-polynomial clustering model likelihood corpus distribution (LCD), which is based on a Bayesian numerical prototype. Experiments are done to prove the efficiency of LCD over various topic modelling algorithms, such as latent Dirichlet allocation (LDA), LDA with bag-of-words (LDA-BOW), latent semantic indexing (LSI), and hierarchical Dirichlet process (HDP). For performance evaluation, we considered five real-world datasets of Bengali corpora, such as science, sports, computer, season, and epic in this work. The coherence score of different modelling algorithms is compared to find the best model for each dataset separately. [ABSTRACT FROM AUTHOR]
Copyright of Sādhanā: Academy Proceedings in Engineering Sciences is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 178527531
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Likelihood corpus distribution: an efficient topic modelling scheme for Bengali document class identification.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Das+Dawn%2C+Debapratim%22">Das Dawn, Debapratim</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> debapratimdd@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Khan%2C+Abhinandan%22">Khan, Abhinandan</searchLink><relatesTo>1,2</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Shaikh%2C+Soharab+Hossain%22">Shaikh, Soharab Hossain</searchLink><relatesTo>3</relatesTo> (AUTHOR)<br /><searchLink fieldCode="AR" term="%22Pal%2C+Rajat+Kumar%22">Pal, Rajat Kumar</searchLink><relatesTo>1</relatesTo> (AUTHOR)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Sādhanā%3A+Academy+Proceedings+in+Engineering+Sciences%22">Sādhanā: Academy Proceedings in Engineering Sciences</searchLink>. Sep2024, Vol. 49 Issue 3, p1-19. 19p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Identification+documents%22">Identification documents</searchLink><br /><searchLink fieldCode="DE" term="%22Latent+semantic+analysis%22">Latent semantic analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Corpora%22">Corpora</searchLink><br /><searchLink fieldCode="DE" term="%22Library+science%22">Library science</searchLink><br /><searchLink fieldCode="DE" term="%22Document+clustering%22">Document clustering</searchLink><br /><searchLink fieldCode="DE" term="%22Sports+sciences%22">Sports sciences</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The learning quality of humans depends on the sense of contemplation. Textual documents are a huge part of the literature on contemplation which effortlessly creates perception. Automatic document class identification or organisation is a machine learning function to understand the psychological and emotional content of the text in a concise way. The problem of identification of documents falls in the field of library science, information science and artificial intelligence. The research progress of class identification of documents has been made in various most spoken languages. Numerous research works have been published in European and Asian languages. However, there is a gap in the literature when it comes to any less resource language, especially Bengali. Consequently, this work portrays an efficient topic modelling approach for Bengali document class identification. It proposes a Dirichlet-polynomial clustering model likelihood corpus distribution (LCD), which is based on a Bayesian numerical prototype. Experiments are done to prove the efficiency of LCD over various topic modelling algorithms, such as latent Dirichlet allocation (LDA), LDA with bag-of-words (LDA-BOW), latent semantic indexing (LSI), and hierarchical Dirichlet process (HDP). For performance evaluation, we considered five real-world datasets of Bengali corpora, such as science, sports, computer, season, and epic in this work. The coherence score of different modelling algorithms is compared to find the best model for each dataset separately. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Sādhanā: Academy Proceedings in Engineering Sciences is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=178527531
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s12046-024-02470-7
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 1
    Subjects:
      – SubjectFull: Identification documents
        Type: general
      – SubjectFull: Latent semantic analysis
        Type: general
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Corpora
        Type: general
      – SubjectFull: Library science
        Type: general
      – SubjectFull: Document clustering
        Type: general
      – SubjectFull: Sports sciences
        Type: general
    Titles:
      – TitleFull: Likelihood corpus distribution: an efficient topic modelling scheme for Bengali document class identification.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Das Dawn, Debapratim
      – PersonEntity:
          Name:
            NameFull: Khan, Abhinandan
      – PersonEntity:
          Name:
            NameFull: Shaikh, Soharab Hossain
      – PersonEntity:
          Name:
            NameFull: Pal, Rajat Kumar
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 02562499
          Numbering:
            – Type: volume
              Value: 49
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Sādhanā: Academy Proceedings in Engineering Sciences
              Type: main
ResultId 1