Optimization of frequent item set mining parallelization algorithm based on spark platform.

Saved in:
Bibliographic Details
Title: Optimization of frequent item set mining parallelization algorithm based on spark platform.
Authors: Fan, Deng1 (AUTHOR) wzideng@qq.com, Jiabin, Wang1 (AUTHOR) fatwang@hqu.edu.cn, Sheng, Lv1 (AUTHOR) 2651113234@qq.com
Source: Information Retrieval Journal. Dec2024, Vol. 27 Issue 1, p1-19. 19p.
Subjects: Boolean matrices, Data warehousing, Artificial intelligence, Image processing, Information theory
Abstract: In this paper, we propose a new method that combines the parallelism of the Spark-based platform with fast frequent mining, called STB_Apriori. Previous research has shown that traditional frequent itemset mining algorithms have high overhead when faced with large datasets and high-dimensional data computation, and generate a large number of candidate itemsets; at the same time, when faced with diverse user requirements, they often generate very sparse and diverse data. In order to solve the problem of fast mining of massive data, our idea originates from the capability of Spark distributed computing and the common optimisation ideas in Apriori mining, by using the efficient operator BitSet to achieve transaction compression, bit storage and data manipulation by Boolean matrices, and at the same time by parallelising the processing and optimising the algorithmic logic to achieve fast and frequent mining. In experiments on real-world datasets, our model consistently outperforms five widely used methods by a significant margin on very large data and maintains its excellence in the remaining cases, proving its effectiveness on real-world tasks, while further analysis shows that increasing the number of distributed nodes also incrementally and continuously improves performance. [ABSTRACT FROM AUTHOR]
Copyright of Information Retrieval Journal is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
Text:
  Availability: 1
Header DbId: egs
DbLabel: Engineering Source
An: 182471776
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Optimization of frequent item set mining parallelization algorithm based on spark platform.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Fan%2C+Deng%22">Fan, Deng</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> wzideng@qq.com</i><br /><searchLink fieldCode="AR" term="%22Jiabin%2C+Wang%22">Jiabin, Wang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> fatwang@hqu.edu.cn</i><br /><searchLink fieldCode="AR" term="%22Sheng%2C+Lv%22">Sheng, Lv</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> 2651113234@qq.com</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Information+Retrieval+Journal%22">Information Retrieval Journal</searchLink>. Dec2024, Vol. 27 Issue 1, p1-19. 19p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Boolean+matrices%22">Boolean matrices</searchLink><br /><searchLink fieldCode="DE" term="%22Data+warehousing%22">Data warehousing</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+intelligence%22">Artificial intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Image+processing%22">Image processing</searchLink><br /><searchLink fieldCode="DE" term="%22Information+theory%22">Information theory</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: In this paper, we propose a new method that combines the parallelism of the Spark-based platform with fast frequent mining, called STB_Apriori. Previous research has shown that traditional frequent itemset mining algorithms have high overhead when faced with large datasets and high-dimensional data computation, and generate a large number of candidate itemsets; at the same time, when faced with diverse user requirements, they often generate very sparse and diverse data. In order to solve the problem of fast mining of massive data, our idea originates from the capability of Spark distributed computing and the common optimisation ideas in Apriori mining, by using the efficient operator BitSet to achieve transaction compression, bit storage and data manipulation by Boolean matrices, and at the same time by parallelising the processing and optimising the algorithmic logic to achieve fast and frequent mining. In experiments on real-world datasets, our model consistently outperforms five widely used methods by a significant margin on very large data and maintains its excellence in the remaining cases, proving its effectiveness on real-world tasks, while further analysis shows that increasing the number of distributed nodes also incrementally and continuously improves performance. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Information Retrieval Journal is the property of Springer Nature and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=182471776
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/s10791-024-09470-5
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 19
        StartPage: 1
    Subjects:
      – SubjectFull: Boolean matrices
        Type: general
      – SubjectFull: Data warehousing
        Type: general
      – SubjectFull: Artificial intelligence
        Type: general
      – SubjectFull: Image processing
        Type: general
      – SubjectFull: Information theory
        Type: general
    Titles:
      – TitleFull: Optimization of frequent item set mining parallelization algorithm based on spark platform.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Fan, Deng
      – PersonEntity:
          Name:
            NameFull: Jiabin, Wang
      – PersonEntity:
          Name:
            NameFull: Sheng, Lv
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 12
              Text: Dec2024
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 13864564
          Numbering:
            – Type: volume
              Value: 27
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Information Retrieval Journal
              Type: main
ResultId 1