Cost-Benefit Arbitration Between Multiple Reinforcement-Learning Systems.

Saved in:
Bibliographic Details
Title: Cost-Benefit Arbitration Between Multiple Reinforcement-Learning Systems.
Authors: Kool, Wouter, Gershman, Samuel J., Cushman, Fiery A.
Source: Psychological Science (0956-7976). Sep2017, Vol. 28 Issue 9, p1321-1333. 13p. 2 Diagrams, 2 Charts, 3 Graphs.
Subjects: Arbitration & award, Reinforcement learning, Human behavior
Abstract: Human behavior is sometimes determined by habit and other times by goal-directed planning. Modern reinforcementlearning theories formalize this distinction as a competition between a computationally cheap but inaccurate modelfree system that gives rise to habits and a computationally expensive but accurate model-based system that implements planning. It is unclear, however, how people choose to allocate control between these systems. Here, we propose that arbitration occurs by comparing each system's task-specific costs and benefits. To investigate this proposal, we conducted two experiments showing that people increase model-based control when it achieves greater accuracy than model-free control, and especially when the rewards of accurate performance are amplified. In contrast, they are insensitive to reward amplification when model-based and model-free control yield equivalent accuracy. This suggests that humans adaptively balance habitual and planned action through on-line cost-benefit analysis. [ABSTRACT FROM AUTHOR]
Copyright of Psychological Science (0956-7976) is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract. (Copyright applies to all Abstracts.)
Database: Psychology and Behavioral Sciences Collection
FullText Text:
  Availability: 0
Header DbId: pbh
DbLabel: Psychology and Behavioral Sciences Collection
An: 125138737
AccessLevel: 6
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Cost-Benefit Arbitration Between Multiple Reinforcement-Learning Systems.
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kool%2C+Wouter%22">Kool, Wouter</searchLink><br /><searchLink fieldCode="AR" term="%22Gershman%2C+Samuel+J%2E%22">Gershman, Samuel J.</searchLink><br /><searchLink fieldCode="AR" term="%22Cushman%2C+Fiery+A%2E%22">Cushman, Fiery A.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="JN" term="%22Psychological+Science+%280956-7976%29%22">Psychological Science (0956-7976)</searchLink>. Sep2017, Vol. 28 Issue 9, p1321-1333. 13p. 2 Diagrams, 2 Charts, 3 Graphs.
– Name: Subject
  Label: Subjects
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Arbitration+%26+award%22">Arbitration & award</searchLink><br /><searchLink fieldCode="DE" term="%22Reinforcement+learning%22">Reinforcement learning</searchLink><br /><searchLink fieldCode="DE" term="%22Human+behavior%22">Human behavior</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Human behavior is sometimes determined by habit and other times by goal-directed planning. Modern reinforcementlearning theories formalize this distinction as a competition between a computationally cheap but inaccurate modelfree system that gives rise to habits and a computationally expensive but accurate model-based system that implements planning. It is unclear, however, how people choose to allocate control between these systems. Here, we propose that arbitration occurs by comparing each system's task-specific costs and benefits. To investigate this proposal, we conducted two experiments showing that people increase model-based control when it achieves greater accuracy than model-free control, and especially when the rewards of accurate performance are amplified. In contrast, they are insensitive to reward amplification when model-based and model-free control yield equivalent accuracy. This suggests that humans adaptively balance habitual and planned action through on-line cost-benefit analysis. [ABSTRACT FROM AUTHOR]
– Name: AbstractSuppliedCopyright
  Label:
  Group: Ab
  Data: <i>Copyright of Psychological Science (0956-7976) is the property of Sage Publications Inc. and its content may not be copied or emailed to multiple sites without the copyright holder's express written permission. Additionally, content may not be used with any artificial intelligence tools or machine learning technologies. However, users may print, download, or email articles for individual use. This abstract may be abridged. No warranty is given about the accuracy of the copy. Users should refer to the original published version of the material for the full abstract.</i> (Copyright applies to all Abstracts.)
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=pbh&AN=125138737
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1177/0956797617708288
    Languages:
      – Code: eng
        Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 13
        StartPage: 1321
    Subjects:
      – SubjectFull: Arbitration & award
        Type: general
      – SubjectFull: Reinforcement learning
        Type: general
      – SubjectFull: Human behavior
        Type: general
    Titles:
      – TitleFull: Cost-Benefit Arbitration Between Multiple Reinforcement-Learning Systems.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kool, Wouter
      – PersonEntity:
          Name:
            NameFull: Gershman, Samuel J.
      – PersonEntity:
          Name:
            NameFull: Cushman, Fiery A.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Text: Sep2017
              Type: published
              Y: 2017
          Identifiers:
            – Type: issn-print
              Value: 09567976
          Numbering:
            – Type: volume
              Value: 28
            – Type: issue
              Value: 9
          Titles:
            – TitleFull: Psychological Science (0956-7976)
              Type: main
ResultId 1