A Parallel Random Forest Algorithm for Big Data in a Spark Cloud Computing Environment.

Saved in:
Bibliographic Details
Title: A Parallel Random Forest Algorithm for Big Data in a Spark Cloud Computing Environment.
Authors: Chen, Jianguo1, Li, Kenli1, Tang, Zhuo1, Bilal, Kashif2, Yu, Shui3, Weng, Chuliang4, Li, Keqin1
Source: IEEE Transactions on Parallel & Distributed Systems. Apr2017, Vol. 28 Issue 4, p919-933. 15p.
Subjects: SPARK (Computer program language), Big data, Cloud computing, Parallel programs (Computer programs), Random forest algorithms
Abstract: With the emergence of the big data age, the issue of how to obtain valuable knowledge from a dataset efficiently and accurately has attracted increasingly attention from both academia and industry. This paper presents a Parallel Random Forest (PRF) algorithm for big data on the Apache Spark platform. The PRF algorithm is optimized based on a hybrid approach combining data-parallel and task-parallel optimization. From the perspective of data-parallel optimization, a vertical data-partitioning method is performed to reduce the data communication cost effectively, and a data-multiplexing method is performed is performed to allow the training dataset to be reused and diminish the volume of data. From the perspective of task-parallel optimization, a dual parallel approach is carried out in the training process of RF, and a task Directed Acyclic Graph (DAG) is created according to the parallel training process of PRF and the dependence of the Resilient Distributed Datasets (RDD) objects. Then, different task schedulers are invoked for the tasks in the DAG. Moreover, to improve the algorithm's accuracy for large, high-dimensional, and noisy data, we perform a dimension-reduction approach in the training process and a weighted voting approach in the prediction process prior to parallelization. Extensive experimental results indicate the superiority and notable advantages of the PRF algorithm over the relevant algorithms implemented by Spark MLlib and other studies in terms of the classification accuracy, performance, and scalability. With the expansion of the scale of the random forest model and the Spark cluster, the advantage of the PRF algorithm is more obvious. [ABSTRACT FROM PUBLISHER]
Copyright of IEEE Transactions on Parallel & Distributed Systems is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Engineering Source
FullText Text:
  Availability: 0
Header DbId: egs
DbLabel: Engineering Source
An: 121854137
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Parallel Random Forest Algorithm for Big Data in a Spark Cloud Computing Environment.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Chen%2C+Jianguo%22">Chen, Jianguo</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Li%2C+Kenli%22">Li, Kenli</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Tang%2C+Zhuo%22">Tang, Zhuo</searchLink><relatesTo>1</relatesTo><br /><searchLink fieldCode="AR" term="%22Bilal%2C+Kashif%22">Bilal, Kashif</searchLink><relatesTo>2</relatesTo><br /><searchLink fieldCode="AR" term="%22Yu%2C+Shui%22">Yu, Shui</searchLink><relatesTo>3</relatesTo><br /><searchLink fieldCode="AR" term="%22Weng%2C+Chuliang%22">Weng, Chuliang</searchLink><relatesTo>4</relatesTo><br /><searchLink fieldCode="AR" term="%22Li%2C+Keqin%22">Li, Keqin</searchLink><relatesTo>1</relatesTo>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22IEEE+Transactions+on+Parallel+%26+Distributed+Systems%22">IEEE Transactions on Parallel & Distributed Systems</searchLink>. Apr2017, Vol. 28 Issue 4, p919-933. 15p.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22SPARK+%28Computer+program+language%29%22">SPARK (Computer program language)</searchLink><br /><searchLink fieldCode="DE" term="%22Big+data%22">Big data</searchLink><br /><searchLink fieldCode="DE" term="%22Cloud+computing%22">Cloud computing</searchLink><br /><searchLink fieldCode="DE" term="%22Parallel+programs+%28Computer+programs%29%22">Parallel programs (Computer programs)</searchLink><br /><searchLink fieldCode="DE" term="%22Random+forest+algorithms%22">Random forest algorithms</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: With the emergence of the big data age, the issue of how to obtain valuable knowledge from a dataset efficiently and accurately has attracted increasingly attention from both academia and industry. This paper presents a Parallel Random Forest (PRF) algorithm for big data on the Apache Spark platform. The PRF algorithm is optimized based on a hybrid approach combining data-parallel and task-parallel optimization. From the perspective of data-parallel optimization, a vertical data-partitioning method is performed to reduce the data communication cost effectively, and a data-multiplexing method is performed is performed to allow the training dataset to be reused and diminish the volume of data. From the perspective of task-parallel optimization, a dual parallel approach is carried out in the training process of RF, and a task Directed Acyclic Graph (DAG) is created according to the parallel training process of PRF and the dependence of the Resilient Distributed Datasets (RDD) objects. Then, different task schedulers are invoked for the tasks in the DAG. Moreover, to improve the algorithm's accuracy for large, high-dimensional, and noisy data, we perform a dimension-reduction approach in the training process and a weighted voting approach in the prediction process prior to parallelization. Extensive experimental results indicate the superiority and notable advantages of the PRF algorithm over the relevant algorithms implemented by Spark MLlib and other studies in terms of the classification accuracy, performance, and scalability. With the expansion of the scale of the random forest model and the Spark cluster, the advantage of the PRF algorithm is more obvious. [ABSTRACT FROM PUBLISHER]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of IEEE Transactions on Parallel & Distributed Systems is the property of IEEE and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=egs&AN=121854137
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1109/TPDS.2016.2603511
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 919
    Subjects:
      – SubjectFull: SPARK (Computer program language)
        Type: general
      – SubjectFull: Big data
        Type: general
      – SubjectFull: Cloud computing
        Type: general
      – SubjectFull: Parallel programs (Computer programs)
        Type: general
      – SubjectFull: Random forest algorithms
        Type: general
    Titles:
      – TitleFull: A Parallel Random Forest Algorithm for Big Data in a Spark Cloud Computing Environment.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Chen, Jianguo
      – PersonEntity:
          Name:
            NameFull: Li, Kenli
      – PersonEntity:
          Name:
            NameFull: Tang, Zhuo
      – PersonEntity:
          Name:
            NameFull: Bilal, Kashif
      – PersonEntity:
          Name:
            NameFull: Yu, Shui
      – PersonEntity:
          Name:
            NameFull: Weng, Chuliang
      – PersonEntity:
          Name:
            NameFull: Li, Keqin
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 04
              Text: Apr2017
              Type: published
              Y: 2017
          Identifiers:
            – Type: issn-print
              Value: 10459219
          Numbering:
            – Type: volume
              Value: 28
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: IEEE Transactions on Parallel & Distributed Systems
              Type: main
ResultId 1