Building a training dataset for classification under a cost limitation.

Saved in:
Bibliographic Details
Title: Building a training dataset for classification under a cost limitation.
Authors: Chen, Yen-Liang1 (AUTHOR) ylchen@mgt.ncu.edu.tw, Cheng, Li-Chen2 (AUTHOR) lijen.cheng@gmail.com, Zhang, Yi-Jun1 (AUTHOR) valorelove@gmail.com
Source: Electronic Library. 2021, Vol. 39 Issue 1, p77-96. 20p.
Subject Terms: *Cataloging, *Machine learning, *Algorithms, Conceptual structures, Research funding, Data analytics, Data mining
Abstract: Purpose: A necessary preprocessing of document classification is to label some documents so that a classifier can be built based on which the remaining documents can be classified. Because each document differs in length and complexity, the cost of labeling each document is different. The purpose of this paper is to consider how to select a subset of documents for labeling with a limited budget so that the total cost of the spending does not exceed the budget limit, while at the same time building a classifier with the best classification results. Design/methodology/approach: In this paper, a framework is proposed to select the instances for labeling that integrate two clustering algorithms and two centroid selection methods. From the selected and labeled instances, five different classifiers were constructed with good classification accuracy to prove the superiority of the selected instances. Findings: Experimental results show that this method can establish a training data set containing the most suitable data under the premise of considering the cost constraints. The data set considers both "data representativeness" and "data selection cost," so that the training data labeled by experts can effectively establish a classifier with high accuracy. Originality/value: No previous research has considered how to establish a training set with a cost limit when each document has a distinct labeling cost. This paper is the first attempt to resolve this issue. [ABSTRACT FROM AUTHOR]
Copyright of Electronic Library is the property of Emerald Publishing Limited and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Education Research Complete
FullText Text:
  Availability: 0
Header DbId: ehh
DbLabel: Education Research Complete
An: 150406194
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Building a training dataset for classification under a cost limitation.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Chen%2C+Yen-Liang%22">Chen, Yen-Liang</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> ylchen@mgt.ncu.edu.tw</i><br /><searchLink fieldCode="AR" term="%22Cheng%2C+Li-Chen%22">Cheng, Li-Chen</searchLink><relatesTo>2</relatesTo> (AUTHOR)<i> lijen.cheng@gmail.com</i><br /><searchLink fieldCode="AR" term="%22Zhang%2C+Yi-Jun%22">Zhang, Yi-Jun</searchLink><relatesTo>1</relatesTo> (AUTHOR)<i> valorelove@gmail.com</i>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Electronic+Library%22">Electronic Library</searchLink>. 2021, Vol. 39 Issue 1, p77-96. 20p.
– Name: Subject
  Label: Subject Terms
  Group: Su
  Data: *<searchLink fieldCode="DE" term="%22Cataloging%22">Cataloging</searchLink><br />*<searchLink fieldCode="DE" term="%22Machine+learning%22">Machine learning</searchLink><br />*<searchLink fieldCode="DE" term="%22Algorithms%22">Algorithms</searchLink><br /><searchLink fieldCode="DE" term="%22Conceptual+structures%22">Conceptual structures</searchLink><br /><searchLink fieldCode="DE" term="%22Research+funding%22">Research funding</searchLink><br /><searchLink fieldCode="DE" term="%22Data+analytics%22">Data analytics</searchLink><br /><searchLink fieldCode="DE" term="%22Data+mining%22">Data mining</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Purpose: A necessary preprocessing of document classification is to label some documents so that a classifier can be built based on which the remaining documents can be classified. Because each document differs in length and complexity, the cost of labeling each document is different. The purpose of this paper is to consider how to select a subset of documents for labeling with a limited budget so that the total cost of the spending does not exceed the budget limit, while at the same time building a classifier with the best classification results. Design/methodology/approach: In this paper, a framework is proposed to select the instances for labeling that integrate two clustering algorithms and two centroid selection methods. From the selected and labeled instances, five different classifiers were constructed with good classification accuracy to prove the superiority of the selected instances. Findings: Experimental results show that this method can establish a training data set containing the most suitable data under the premise of considering the cost constraints. The data set considers both "data representativeness" and "data selection cost," so that the training data labeled by experts can effectively establish a classifier with high accuracy. Originality/value: No previous research has considered how to establish a training set with a cost limit when each document has a distinct labeling cost. This paper is the first attempt to resolve this issue. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Electronic Library is the property of Emerald Publishing Limited and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=ehh&AN=150406194
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1108/EL-07-2020-0209
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 20
        StartPage: 77
    Subjects:
      – SubjectFull: Cataloging
        Type: general
      – SubjectFull: Machine learning
        Type: general
      – SubjectFull: Algorithms
        Type: general
      – SubjectFull: Conceptual structures
        Type: general
      – SubjectFull: Research funding
        Type: general
      – SubjectFull: Data analytics
        Type: general
      – SubjectFull: Data mining
        Type: general
    Titles:
      – TitleFull: Building a training dataset for classification under a cost limitation.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Chen, Yen-Liang
      – PersonEntity:
          Name:
            NameFull: Cheng, Li-Chen
      – PersonEntity:
          Name:
            NameFull: Zhang, Yi-Jun
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Text: 2021
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-print
              Value: 02640473
          Numbering:
            – Type: volume
              Value: 39
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Electronic Library
              Type: main
ResultId 1