Why Higher Working Memory Capacity May Help You Learn: Sampling, Search, and Degrees of Approximation
Saved in:
| Title: | Why Higher Working Memory Capacity May Help You Learn: Sampling, Search, and Degrees of Approximation |
|---|---|
| Language: | English |
| Authors: | Lloyd, Kevin, Sanborn, Adam, Leslie, David, Lewandowsky, Stephan |
| Source: | Cognitive Science. Dec 2019 43(12). |
| Availability: | Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA |
| Peer Reviewed: | Y |
| Page Count: | 43 |
| Publication Date: | 2019 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Short Term Memory, Bayesian Statistics, Cognitive Ability, Individual Differences, Inferences, Correlation, Classification, Learning Strategies, Models, Learning Processes |
| DOI: | 10.1111/cogs.12805 |
| ISSN: | 1551-6709 |
| Abstract: | Algorithms for approximate Bayesian inference, such as those based on sampling (i.e., Monte Carlo methods), provide a natural source of models of how people may deal with uncertainty with limited cognitive resources. Here, we consider the idea that individual differences in working memory capacity (WMC) may be usefully modeled in terms of the number of samples, or "particles," available to perform inference. To test this idea, we focus on two recent experiments that report positive associations between WMC and two distinct aspects of categorization performance: the ability to learn novel categories, and the ability to switch between different categorization strategies ("knowledge restructuring"). In favor of the idea of modeling WMC as a number of particles, we show that a single model can reproduce both experimental results by varying the number of particles--increasing the number of particles leads to both faster category learning and improved strategy-switching. Furthermore, when we fit the model to individual participants, we found a positive association between WMC and best-fit number of particles for strategy switching. However, no association between WMC and best-fit number of particles was found for category learning. These results are discussed in the context of the general challenge of disentangling the contributions of different potential sources of behavioral variability. |
| Abstractor: | As Provided |
| Entry Date: | 2019 |
| Accession Number: | EJ1237804 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFROgfbBMTBefBvj9N10vXfAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNlO8XNZcFOSSErkdAIBEICBm7JdRnQAym7DzKSw38nuTcIA6DxGOBE4iRfCERjZ4dC9Ny3sDtJDZPir-o2FnSh9dLyik6FUnRmHyhU_amVbUfN5uk41-jUaUT_EihZKaUv4zqA68oz1gNFFoXuawBsoWlB7V0g7U9A8tdLEn6H5gYWT0HBImcTiM0phU38_NXO6WUdKt3Wr9Lexae-cD14jbtDEf6muMc4h-fxX Text: Availability: 1 Value: <anid>AN0140459439;cgn01dec.19;2019Dec23.03:49;v2.2.500</anid> <title id="AN0140459439-1">Why Higher Working Memory Capacity May Help You Learn: Sampling, Search, and Degrees of Approximation </title> <p>Algorithms for approximate Bayesian inference, such as those based on sampling (i.e., Monte Carlo methods), provide a natural source of models of how people may deal with uncertainty with limited cognitive resources. Here, we consider the idea that individual differences in working memory capacity (WMC) may be usefully modeled in terms of the number of samples, or "particles," available to perform inference. To test this idea, we focus on two recent experiments that report positive associations between WMC and two distinct aspects of categorization performance: the ability to learn novel categories, and the ability to switch between different categorization strategies ("knowledge restructuring"). In favor of the idea of modeling WMC as a number of particles, we show that a single model can reproduce both experimental results by varying the number of particles—increasing the number of particles leads to both faster category learning and improved strategy‐switching. Furthermore, when we fit the model to individual participants, we found a positive association between WMC and best‐fit number of particles for strategy switching. However, no association between WMC and best‐fit number of particles was found for category learning. These results are discussed in the context of the general challenge of disentangling the contributions of different potential sources of behavioral variability.</p> <p>Keywords: Working memory; Category learning; Knowledge partitioning; Strategy switching; Approximate Bayesian inference; Particle filtering</p> <hd id="AN0140459439-2">Introduction</hd> <p>How to deal with uncertainty arising from noisy and incomplete information is a ubiquitous challenge for natural and artificial agents alike. Bayesian statistics provides a rigorous system for representing and reasoning about such uncertainty, yielding a principled method for updating beliefs in the light of new evidence (Bernardo &amp; Smith, [<reflink idref="bib12" id="ref1">12</reflink>]). Human behavior is often well described in terms of Bayesian inference, from "low level" sensorimotor (Körding &amp; Wolpert, [<reflink idref="bib51" id="ref2">51</reflink>]) and perceptual (Yuille &amp; Kersten, [<reflink idref="bib96" id="ref3">96</reflink>]) phenomena, to "high level" competencies, such as causal reasoning (Griffiths &amp; Tenenbaum, [<reflink idref="bib44" id="ref4">44</reflink>]), category learning (Sanborn, Navarro, &amp; Griffiths, [<reflink idref="bib79" id="ref5">79</reflink>]), and predictions about future everyday events (Griffiths &amp; Tenenbaum, [<reflink idref="bib45" id="ref6">45</reflink>]; reviews include Chater &amp; Oaksford, [<reflink idref="bib18" id="ref7">18</reflink>]; Sanborn &amp; Chater, [<reflink idref="bib77" id="ref8">77</reflink>]; Tenenbaum, Kemp, Griffiths, &amp; Goodman, [<reflink idref="bib89" id="ref9">89</reflink>]).</p> <p>How humans frequently—though by no means always (e.g., Tversky &amp; Kahnman, [<reflink idref="bib91" id="ref10">91</reflink>])—achieve this consistency with Bayesian principles is less clear. Though simple in principle, exact Bayesian calculations are frequently intractable in real‐world settings, leading to a need for approximations. In statistics and computer science, this challenge has been met through the development of powerful, general purpose techniques for approximate Bayesian inference, such as Monte Carlo methods (Gelfand &amp; Smith, [<reflink idref="bib35" id="ref11">35</reflink>]; Robert &amp; Casella, [<reflink idref="bib75" id="ref12">75</reflink>]), which allow for the practical application of Bayesian methods in complex domains.</p> <p>The practical success of these techniques has naturally led to an interest in whether they also tell us something about how people reason under uncertainty. That is, they provide one source of hypotheses about the nature of the psychological and neural mechanisms that underlie how people process probabilistic information (Chater &amp; Oaksford, [<reflink idref="bib18" id="ref13">18</reflink>]; Doya, Ishii, Pouget, &amp; Rao, [<reflink idref="bib31" id="ref14">31</reflink>]). Since the aim of these algorithms is to approximate the normative solution to a computational problem—that is, to approximate Bayesian inference—they have been called <emph>rational process models</emph> when considered as candidate psychological mechanisms (Griffiths, Vul, &amp; Sanborn, [<reflink idref="bib46" id="ref15">46</reflink>]; Sanborn et al., [<reflink idref="bib79" id="ref16">79</reflink>]). This distinguishes them from traditional process models in cognitive psychology, which are typically rich in postulated psychological mechanisms but often poor in terms of normative foundations (cf. Anderson, [<reflink idref="bib2" id="ref17">2</reflink>]).</p> <p>Importantly, Monte Carlo methods can in principle approximate probabilistic inference arbitrarily well when sufficient time and memory are available, thereby providing a benchmark for ideal performance. At the same time, these methods display systematic deviations from the normative solution when resources are limited. Such "qualitative fingerprints" associated with different species of approximation may then be particularly illuminating when considering human cognition, where it is generally assumed that information processing capacity is limited (Daw, Courville, &amp; Dayan, [<reflink idref="bib28" id="ref18">28</reflink>]; Gigerenzer &amp; Goldstein, [<reflink idref="bib38" id="ref19">38</reflink>]; Kahneman, [<reflink idref="bib49" id="ref20">49</reflink>]; Simon, [<reflink idref="bib84" id="ref21">84</reflink>]).</p> <p>One such limitation has long been associated with working memory (Cowan, [<reflink idref="bib24" id="ref22">24</reflink>]; Miller, [<reflink idref="bib63" id="ref23">63</reflink>]), defined in cognitive psychology as the memory system responsible for temporary storage and manipulation of task‐relevant information (Baddeley, [<reflink idref="bib8" id="ref24">8</reflink>]; Baddeley &amp; Hitch, [<reflink idref="bib9" id="ref25">9</reflink>]). Individual differences in working memory capacity (WMC), such as measured in the complex span paradigm (Daneman &amp; Carpenter, [<reflink idref="bib26" id="ref26">26</reflink>]), have been found to predict performance on a variety of cognitive tasks, including conventional intelligence tests (Conway, Jarrold, Kane, Miyake, &amp; Towse, [<reflink idref="bib21" id="ref27">21</reflink>]). Indeed, WMC may account for up to one half of the variance in general intelligence (Conway, Kane, &amp; Engle, [<reflink idref="bib22" id="ref28">22</reflink>]).</p> <p>However, the exact nature of the WMC limitation that underpins such individual differences remains the subject of debate, with proposals variously emphasizing decay of representations (e.g., Baddeley, Thompson, &amp; Buchanan, [<reflink idref="bib10" id="ref29">10</reflink>]), resource constraints (e.g., Just &amp; Carpenter, [<reflink idref="bib48" id="ref30">48</reflink>]), or interference (e.g., Oberauer &amp; Kliegl, [<reflink idref="bib70" id="ref31">70</reflink>]; see Oberauer, Farrell, Jarrold, &amp; Lewandowsky, [<reflink idref="bib69" id="ref32">69</reflink>] for a recent discussion). Indeed, opinions continue to differ as to whether working memory is best conceptualized as discrete, for example, comprising a limited number of "slots," or as a more continuous "resource" that can be flexibly distributed across representations in memory (Ma, Husain, &amp; Bays, [<reflink idref="bib60" id="ref33">60</reflink>]; Suchow, Fougnie, Brady, &amp; Alvarez, [<reflink idref="bib88" id="ref34">88</reflink>]).</p> <p>Our approach in the current work is to consider WMC limitations within the broader context of probabilistic inference, asking whether WMC may be usefully modeled as a constraint on the amount of <emph>inferential</emph> resources available. The implication is that at least in tasks involving uncertainty, enhanced performance in individuals with higher WMC may be attributable to an ability to better approximate "ideal" Bayesian solutions.</p> <p>To begin to explore this idea, we focus on recent experiments showing positive associations between WMC and performance on category learning tasks (Lewandowsky, [<reflink idref="bib56" id="ref35">56</reflink>]; Lewandowsky, Yang, Newell, &amp; Kalish, [<reflink idref="bib58" id="ref36">58</reflink>]; Sewell &amp; Lewandowsky, [<reflink idref="bib81" id="ref37">81</reflink>], [<reflink idref="bib82" id="ref38">82</reflink>]). This focus is motivated by two considerations. First, category learning tasks are well characterized as probabilistic inference problems, requiring participants to reason about possible underlying category structures. Even when the mapping between stimuli and category labels is deterministic, participants face epistemic uncertainty regarding the nature of this mapping. Normative solutions to such problems, as well as how these solutions may be practically approximated—notably via Monte Carlo methods—have received substantial attention (Anderson, [<reflink idref="bib2" id="ref39">2</reflink>]; Goodman, Tenenbaum, Feldman, &amp; Griffiths, [<reflink idref="bib42" id="ref40">42</reflink>]; Sanborn et al., [<reflink idref="bib79" id="ref41">79</reflink>]). We build on this previous work here. Secondly, WMC appears to be positively associated with two distinct aspects of categorization: the ability to acquire novel categories (i.e., category learning; Lewandowsky, [<reflink idref="bib56" id="ref42">56</reflink>]), and the ability to flexibly switch between different categorization strategies (sometimes referred to as "knowledge restructuring"; Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref43">82</reflink>]). Previous work has explored how such positive associations may arise in formal category learning models (Lewandowsky, [<reflink idref="bib56" id="ref44">56</reflink>]; Sewell &amp; Lewandowsky, [<reflink idref="bib81" id="ref45">81</reflink>], [<reflink idref="bib82" id="ref46">82</reflink>]) but has treated these aspects of categorization separately, and via different models and mechanisms; the possibility that WMC may influence both category learning and knowledge restructuring via a single mechanism has not been explored, and we seek such a common mechanism in the present article.</p> <p>The key assumptions of the current work are that individuals approximate Bayesian solutions to category learning problems by sampling from probability distributions (i.e., via Monte Carlo inference) and, more important, that an individual's WMC directly translates into how many samples, or hypotheses, he or she is able to represent at one time. We show that this simple equating of WMC with the number of active hypotheses allows us to reproduce the positive associations between WMC and both aspects of categorization performance—category learning and knowledge restructuring—with a single mechanism. Before describing the modeling approach and results in detail, we briefly summarize the basic ideas behind Monte Carlo methods and the target experimental results.</p> <hd id="AN0140459439-3">Monte Carlo as a psychological mechanism</hd> <p>In the Bayesian paradigm, background knowledge gives rise to a constrained set of candidate hypotheses <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="script"&gt;H&lt;/mi&gt;&lt;/math&gt; </ephtml> for the true state of nature, and to associated degrees of belief <emph>P</emph>(<emph>h</emph>) in each candidate in the set <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mi mathvariant="script"&gt;H&lt;/mi&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . The sum of all beliefs about the true state of nature is fixed to 1. Such "prior" beliefs are updated in the light of observed data <emph>d</emph> to yield "posterior" beliefs <emph>P</emph>(<emph>h</emph>\<emph>d</emph>) via Bayes' theorem,</p> <p> <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0003" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;munder&gt;&lt;mo movablelimits="false"&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;msup&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mi mathvariant="script"&gt;H&lt;/mi&gt;&lt;/mrow&gt;&lt;/munder&gt;&lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;mrow&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;/mrow&gt;&lt;msup&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;/mfenced&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;msup&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;&amp;#8242;&lt;/mo&gt;&lt;/msup&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>where the likelihood <emph>P</emph>(<emph>d</emph>\vert<emph>h</emph>) quantifies how expected the data are under each candidate hypothesis.</p> <p>As we will describe in detail below, for our purposes the state of nature is the true category structure that participants are required to learn; the set of candidate hypotheses is the space of all possible category structures that a participant is assumed to be able to generate; and the observed data are the particular category instances presented to participants that they must categorize and for which they subsequently receive feedback about the correct category label.</p> <p>While Bayes' theorem is simple to write down, it leads to complex practical issues such as the source of the prior distribution, the choice of likelihood function, and how to compute and summarize the posterior distribution if the hypothesis space <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="script"&gt;H&lt;/mi&gt;&lt;/math&gt; </ephtml> is very large—such as when <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="script"&gt;H&lt;/mi&gt;&lt;/math&gt; </ephtml> is the space of all possible categories.</p> <p>In Monte Carlo methods, the basic idea is to approximate the target distribution <emph>P</emph>(<emph>h</emph>\<emph>d</emph>) by drawing samples from it. In other words, one represents <emph>P</emph>(<emph>h</emph>\<emph>d</emph>) with a set of samples {<emph>h</emph><sups>(</sups><emph><sups>i</sups></emph><sups>)</sups>} ∼ <emph>P</emph>(<emph>h</emph>\<emph>d</emph>) from that distribution, each randomly selected with a frequency proportional to its probability in the full distribution.</p> <p>In the case where beliefs are updated sequentially as new information arrives—as in the experiments we consider below, where participants receive feedback trial by trial—one attempts to approximate a <emph>sequence</emph> of target distributions, and so we are more specifically interested in the idea of <emph>sequential Monte Carlo</emph>, or "particle filtering" (Doucet, de Freitas, &amp; Gordon, [<reflink idref="bib29" id="ref47">29</reflink>]). As we will describe in more detail, one way of promoting a good approximation to posterior distributions in this instance is to propose local changes to a current hypothesis <emph>h</emph>, and to accept or reject the proposed variant <emph>h</emph><sups>′</sups> as a function of its posterior probability. This latter process can be thought of in terms of continuous exploration, or <emph>search</emph>, of the hypothesis space for regions of high probability.</p> <p>These two characteristics of Monte Carlo inference—representation by a limited number of hypotheses, and inference as involving an active process of exploration, or search, of the posterior—draw parallels with working memory, which is typically characterized not only as limited in capacity but also as <emph>active</emph> memory (Baddeley, [<reflink idref="bib8" id="ref48">8</reflink>]). In other words, if WMC is the number of hypotheses that one can actively maintain and manipulate at a given time, and if these latter processes can be cast in terms of probabilistic inference, then a possible analogy between working memory processes and Monte Carlo inference presents itself.</p> <p>Of course, the idea that <emph>sampling</emph> plays a role in psychological mechanisms has a long tradition in psychology (Busemeyer, [<reflink idref="bib17" id="ref49">17</reflink>]; Estes, [<reflink idref="bib33" id="ref50">33</reflink>]; Restle, [<reflink idref="bib74" id="ref51">74</reflink>]; Stewart, Chater, &amp; Brown, [<reflink idref="bib86" id="ref52">86</reflink>]), though not typically in the context of approximating Bayesian inference. More recent work has explicitly considered sample‐based inference as a possible psychological mechanism (recent reviews include Griffiths et al., [<reflink idref="bib46" id="ref53">46</reflink>]; Suchow, Bourgin, &amp; Griffiths, [<reflink idref="bib87" id="ref54">87</reflink>]). For example, Vul and Pashler ([<reflink idref="bib93" id="ref55">93</reflink>]) argued that the "wisdom of crowds" effect, where the error of a judgment averaged over individuals is substantially smaller than the average error of individual judgments, is consistent with individuals using only a limited number of samples to form estimates (cf. Lewandowsky, Griffiths, &amp; Kalish, [<reflink idref="bib57" id="ref56">57</reflink>]). Other work has focused on apparent suboptimalities displayed in people's sensitivity to the ordering of information when they must update their beliefs over time. Such order effects have been successfully captured by models employing sequential inference with limited samples in a variety of domains, including change detection (Brown &amp; Steyvers, [<reflink idref="bib15" id="ref57">15</reflink>]), garden path effects in sentence processing (Levy, Reali, &amp; Griffiths, [<reflink idref="bib55" id="ref58">55</reflink>]), and category learning (Sanborn et al., [<reflink idref="bib79" id="ref59">79</reflink>]).</p> <hd id="AN0140459439-4">Working memory capacity and category learning</hd> <p>Despite the central importance of both working memory and categorization in cognition, until recently the relationship between these abilities received scant attention. The nature of this relationship is of interest not only to provide further constraints on adequate theories of these faculties, but also in light of recent arguments for the existence of multiple categorization systems that rely to differing degrees on distinct memory systems. One salient hypothesis is that category learning tasks that can be solved with relatively simple, verbalizable rules ("rule‐based" tasks) rely especially on working memory, while tasks with solutions that generally defy description in terms of simple rules ("information‐integration" tasks) do not (Ashby &amp; Maddox, [<reflink idref="bib5" id="ref60">5</reflink>], [<reflink idref="bib6" id="ref61">6</reflink>]; Ashby &amp; O'Brien, [<reflink idref="bib7" id="ref62">7</reflink>]).</p> <p>In contrast to this proposal, recent studies have found a positive association between WMC and category learning performance, regardless of whether the categorization task is rule‐based (Lewandowsky, [<reflink idref="bib56" id="ref63">56</reflink>]) or based on information‐integration (Lewandowsky et al., [<reflink idref="bib58" id="ref64">58</reflink>]). Interestingly, WMC has also been found to be positively associated with a somewhat distinct aspect of categorization, namely the ability to flexibly switch between different categorization strategies (Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref65">82</reflink>])—a capacity that the authors refer to as "knowledge restructuring." These apparently disparate findings, which we describe next, form the target of the current work.</p> <hd id="AN0140459439-5">A positive association between WMC and category learning</hd> <p>Lewandowsky ([<reflink idref="bib56" id="ref66">56</reflink>]) used a battery of four working memory tasks (memory updating, operation span, sentence span, and spatial short‐term memory tasks—refer to the original paper for further detail and references) to measure the WMC of participants before testing their category learning performance on the six classical problem types of Shepard, Hovland, and Jenkins ([<reflink idref="bib83" id="ref67">83</reflink>]) (henceforth "SHJ"). Each problem type involves learning to assign each of a set of eight stimuli to category <emph>A</emph> or <emph>B</emph> based on their values on three binary dimensions (Fig. A); half of the stimuli are assigned to category <emph>A</emph>, and the other half to category <emph>B</emph>. There are 72 possible assignments that satisfy these conditions, but these reduce to six "types" assuming interchangeability of dimensions and labels (Fig. B). The problem types vary with respect to the number of stimulus dimensions that are relevant for classification. For example, in a Type I problem, only a single dimension is relevant; in a Type VI problem, by contrast, all three dimensions are relevant.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0001.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0001.jpg" title="The six category learning problem types of Shepard et al. ([83]). (A) Each one of eight stimuli is defined by its unique combination of values on three dimensions (e.g., color, size, and shape) that correspond to the edges of the cube. (B) In each problem type, four stimuli are assigned to category A (filled circles), and the remaining four stimuli are assigned to category B (open circles). (C) Learning curves for each problem type, averaged over all participants, measured by Lewandowsky (data replotted from Lewandowsky, [56]). (D) Overall proportion of errors for high‐ and low‐WMC participants (median split by WMC score) for each problem type. Error bars represent +1 SE." /> </p> <p></p> <p>Consistent with the classical results, Lewandowsky found that the average trend of participants was to learn a Type I problem fastest, a Type VI problem the slowest, with Types II–V clustered in between (Fig. C). Crucially, structural equation modeling of WMC and category learning measures also revealed that WMC was positively related to category learning performance in each problem type (see Lewandowsky, [<reflink idref="bib56" id="ref68">56</reflink>] for details). In Fig. D, we replot the data to show the overall proportion of errors for each problem type given the median split of participants into high‐ and low‐WMC groups based on their WMC scores. There is a clear trend for high‐WMC participants to make fewer errors on each type of problem. Entering errors into a 2 (WMC: low, high) × 6 (Problem: I, II, III, IV, V, VI) × 12 (Block: 1–12) repeated measures anova confirmed that high‐WMC participants were more accurate than low‐WMC participants (<emph>F</emph>(<reflink idref="bib1" id="ref69">1</reflink>, 111) = 13.63, <emph>p</emph> &lt; .01), with no significant interactions between WMC and the other factors. Low‐WMC participants made significantly more errors on each problem type, with the exception of Type IV.</p> <hd id="AN0140459439-7">A positive association between WMC and knowledge restructuring</hd> <p>Sewell and Lewandowsky ([<reflink idref="bib82" id="ref70">82</reflink>]) found that higher WMC (where WMC was assessed using the same battery of measures as in Lewandowsky, [<reflink idref="bib56" id="ref71">56</reflink>]) was associated not only with better category learning performance, consistent with the findings of Lewandowsky ([<reflink idref="bib56" id="ref72">56</reflink>]), but also with an improved ability to switch between categorization strategies when instructed to do so—an ability assumed to reflect knowledge restructuring (Sewell &amp; Lewandowsky, [<reflink idref="bib81" id="ref73">81</reflink>]).</p> <p>Like the SHJ problems, the basic task in the studies by Sewell and Lewandowsky ([<reflink idref="bib82" id="ref74">82</reflink>]) was to learn to assign stimuli to category <emph>A</emph> or <emph>B</emph>. Here, stimuli were rectangles that varied with respect to three features (height, the position a vertical bar located along their base, and color). Stimuli were assigned to category <emph>A</emph> or <emph>B</emph> depending on their position in stimulus space (Fig. A). Height and bar offset were continuous dimensions, whereas color could take only one of two values (e.g., blue or red). Training stimuli (filled circles, Fig. A) were clustered into two separate regions of category space, with categories arranged so that partial category boundaries (solid lines, Fig. A) could not be integrated in a coherent manner—that is, neither partial boundary could be extended in a way that allowed accurate classification of training stimuli in the other cluster, thereby encouraging co‐ordination of multiple partial rules (for fuller discussion, see Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref75">82</reflink>]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0002.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0002.jpg" title="Knowledge restructuring task of Sewell and Lewandowsky ([82]). (A) Experimental stimuli. These were rectangles (two examples shown at top) that varied with respect to their height, position of a vertically oriented bar along their base, and color (e.g., blue or red). Stimuli were assigned to category A or B depending on their position in stimulus space. Filled circles denote training stimuli, open squares denote test stimuli, and solid lines indicate the partial rule boundaries. (B) Ideal response profiles associated with the context‐insensitive (CI; top row) and knowledge‐partitioning (KP; bottom row) categorization strategies. Shading indicates the probability with which a test stimulus should be classified as belonging to category A (darker color indicates a higher probability). Ideal performance in the different contexts (i.e., test stimulus presented in blue or red) is shown in the left and right columns of panels, respectively. (C) Context sensitivity across all transfer tests for knowledge‐partitioning (KP)‐first and context‐insensitive (CI)‐first conditions. Error bars indicate ±1 SEM. (D) Mean absolute change in context sensitivity (CS) for participants with WMC scores in the top and bottom quartiles (&quot;High&quot; and &quot;Low&quot; WMC, respectively) for Session 1 (i.e., between transfer tests 1 and 2) and Session 2 (i.e., between transfer tests 3 and 4). Error bars indicate +1 SE. Figures A–C after Sewell and Lewandowsky ([82])." /> </p> <p></p> <p>Importantly, equally good categorization performance in this task could be obtained by learning any one of a number of different strategies. For example, a participant could use the color of the rectangle to decide whether height (for blue rectangles) or bar position (for red rectangles) predicted category <emph>A</emph> or <emph>B</emph>—this was named a <emph>knowledge‐partitioning</emph> (KP) strategy. Alternatively, a participant could attend to whether bar position was to the left or right of center in order to then diagnose category membership based on either height or, again, bar position—thereby ignoring the color dimension entirely. This latter was named a <emph>context‐insensitive</emph> (CI) strategy.</p> <p>The crucial experimental manipulation was to encourage a participant, using verbal instruction, to first learn one of these two strategies—by hinting that the problem could be solved using bar position (for a participant assigned to the "CI‐first" experimental group) or color (for a participant assigned to the "KP‐first" experimental group)—before giving the participant an unexpected instruction to switch to using the alternative strategy. The degree to which participants' predictions conformed to a CI or KP strategy could be assessed via their generalization performance on a set of test stimuli (open squares, Fig. A), since generalization performance should be either insensitive (CI strategy) or sensitive (KP strategy) to the color of the presented stimuli (Fig. B). On the basis of their generalization pattern, participants were assigned a "context sensitivity" score, summarizing the degree to which their performance best conformed to a CI (context sensitivity close to 0) or KP (context sensitivity close to 1) strategy.</p> <p>Regardless of whether participants were encouraged to use a CI or KP strategy in the first instance, they were able to shift between strategies without any training on the novel strategy (Fig. C), an ability assumed to reflect knowledge restructuring (Sewell &amp; Lewandowsky, [<reflink idref="bib81" id="ref76">81</reflink>]). More important for our purposes, however, was the finding of a significant positive correlation between WMC and the extent of knowledge restructuring, the latter being measured in terms of the absolute change in context sensitivity in each test session (see Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref77">82</reflink>], for full details of the structural equation modeling approach and results). Fig. D shows the average change in context sensitivity for participants with WMC scores in the top and bottom quartiles, for Session 1 (i.e., changes between transfer tests 1 and 2) and Session 2 (i.e., changes between transfer tests 3 and 4). Entering these change scores into a 2 (WMC: low, high) × 2 (Condition: CI‐first, KP‐first) × 2 (Session: 1, 2) repeated measures anova confirmed a main effect of WMC on change in context sensitivity (<emph>F</emph>(<reflink idref="bib1" id="ref78">1</reflink>, 47) = 4.42, <emph>p</emph> &lt; .05). High‐WMC participants had significantly higher changes in context sensitivity in Session 1 (<emph>t</emph>(<reflink idref="bib48" id="ref79">48</reflink>) = 2.81, <emph>p</emph> &lt; .01), though not in Session 2 (<emph>t</emph>(<reflink idref="bib48" id="ref80">48</reflink>) = 1.17, ns); we defer discussion of this, and further subtleties of the experimental results, until later (see Section 4).</p> <p>The results of Sewell and Lewandowsky ([<reflink idref="bib82" id="ref81">82</reflink>]) thus suggest that WMC supports not just standard category learning but also the flexible application of different categorization strategies.</p> <hd id="AN0140459439-9">Modeling approach</hd> <p>The hypothesis of the current study was that by equating working memory capacity (WMC) with the number of samples available for inference in a Bayesian category learning model, positive associations between WMC, category learning, and knowledge restructuring would naturally arise, consistent with the experimental findings.</p> <p>Our model can be described as comprising three parts: (a) a model of how participants are assumed to <emph>represent</emph> categories, specified in terms of an explicit process whereby categories can be constructed (i.e., a "generative model"); (b) a procedure by which participants are assumed to <emph>infer</emph> categories in light of their prior assumptions and the experimental stimuli; and (c) a means for translating participants' beliefs about categories into <emph>choice</emph>, that is, a prediction of the category label associated with a stimulus before receiving feedback about the true label.</p> <hd id="AN0140459439-10">Category representation</hd> <p>Many representational formats for categories have been discussed in the literature, including rules (Bruner, Goodnow, &amp; Austin, [<reflink idref="bib16" id="ref82">16</reflink>]; Goodman et al., [<reflink idref="bib42" id="ref83">42</reflink>]; Nosofsky, Palmeri, &amp; McKinley, [<reflink idref="bib68" id="ref84">68</reflink>]), prototypes (Posner &amp; Keele, [<reflink idref="bib72" id="ref85">72</reflink>]; Rosch, [<reflink idref="bib76" id="ref86">76</reflink>]), exemplars (Kruschke, [<reflink idref="bib52" id="ref87">52</reflink>]; Medin &amp; Schaffer, [<reflink idref="bib62" id="ref88">62</reflink>]; Nosofsky, [<reflink idref="bib66" id="ref89">66</reflink>]), or some mixture of these (Anderson, [<reflink idref="bib3" id="ref90">3</reflink>]; Ashby, Alfonso‐Reese, Turken, &amp; Waldron, [<reflink idref="bib4" id="ref91">4</reflink>]; Love, Medin, &amp; Gureckis, [<reflink idref="bib59" id="ref92">59</reflink>]). In the current work, we chose to work within the framework of <emph>classification and regression tree</emph> (CART) models (Breiman, Friedman, Olshen, &amp; Stone, [<reflink idref="bib14" id="ref93">14</reflink>]), which can be considered a type of rule‐based representation. This choice was largely pragmatic. First, CART models offer an intuitive format for the categories used in the experimental tasks of interest, which are readily described in terms of simple, verbalizable rules (i.e., "rule‐based," in the terms of Ashby &amp; Maddox, [<reflink idref="bib5" id="ref94">5</reflink>]) and that also suggest an ordering on rules (particularly the task of Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref95">82</reflink>]; see below). Second, as we will describe, these models are amenable to a Bayesian formulation (Chipman, George, &amp; McCulloch, [<reflink idref="bib19" id="ref96">19</reflink>]), which is obviously crucial for our purposes.</p> <p>Most broadly, CART models (Breiman et al., [<reflink idref="bib14" id="ref97">14</reflink>]) provide a flexible method for specifying the conditional distribution of a response variable (e.g., a category label) given a collection of input predictors (e.g., stimulus features). In the experiments we consider, category labels are always binary, <emph>y</emph> ∊ {<emph>A</emph>, <emph>B</emph>}, and each stimulus to be categorized is represented by a <emph>p</emph>‐dimensional feature vector <bold>x</bold> = (<emph>x</emph><subs>1</subs>, <emph>x</emph><subs>2</subs>, ..., <emph>x<subs>p</subs></emph>).[<reflink idref="bib1" id="ref98">1</reflink>] The models work by recursively partitioning the input space into axis‐aligned cuboids—imagine making a series of axis‐aligned "slices" through the input space—and applying a simple conditional model to each region; the sequence of partitions on the input space can be represented as a binary tree (Fig. A).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0003.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0003.jpg" title="Representing categories with a classification tree. (A) Consider the stimulus space of Sewell and Lewandowsky ([82]), which comprises three stimulus dimensions (color, height, and bar position) and can be represented as a cube (left). A single partition of this space into two subspaces can be achieved by selecting one of the stimulus dimensions (here, bar position) and splitting the space on that dimension at a particular location. This partitioning can be represented by a simple binary tree (right). The root node η (which is also an &quot;internal&quot; node) is associated with the full stimulus space B(η). In this example, node η is split on the dimension corresponding to bar position (κ(η) = bar position) at a location τ(η). This partitions the input space into two blocks, B(ηL) and B(ηR), associated with the &quot;leaf&quot; nodes ηL and ηR. (B) Tree corresponding to a knowledge‐partitioning (KP) strategy; the initial split is on the color dimension. (C) Tree corresponding to a context‐insensitive (CI) strategy; the initial split is on the bar position dimension. (D) In the model, proposed modifications to trees may be of three types, each involving the initial random selection of a node (shaded red): grow selects a leaf node for expansion (i.e., splitting); prune selects an internal node and renders it a leaf node by deleting all nodes below it; and change selects an internal node and assigns it a new rule (i.e., a splitting dimension and location)." /> </p> <p></p> <p>Formally, a binary tree structure <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/math&gt; </ephtml> consists of a hierarchy of nodes <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . Nodes with children, or leaves, are referred to as <emph>internal</emph> nodes, while nodes without children are referred to as <emph>leaf</emph> nodes (Fig. A, right). The set of internal nodes for <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/math&gt; </ephtml> is denoted <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msub&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/msub&gt;&lt;/math&gt; </ephtml> , and the set of leaves is denoted <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msub&gt;&lt;mi&gt;L&lt;/mi&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/msub&gt;&lt;/math&gt; </ephtml> . Each internal node <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;I&lt;/mi&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> has exactly two children, called the left child η<emph><subs>L</subs></emph> and right child η<emph><subs>R</subs></emph>. Each node is associated with a block <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8838;&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> of the input space as follows (cf. Fig. A, left): The root node is associated with the entire input space, while each further internal node splits its block into two parts by selecting a single dimension κ(η) = {1, ..., <emph>p</emph>} and location τ(η) so that</p> <p> <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0013" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;B(&amp;#951;L)=B(&amp;#951;)&amp;#8745;{x:x&amp;#954;(&amp;#951;)&amp;#8804;&amp;#964;(&amp;#951;)}andB(&amp;#951;R)=B(&amp;#951;)&amp;#8745;{x:x&amp;#954;(&amp;#951;)&amp;#62;&amp;#964;(&amp;#951;)}.&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>The block of input space associated with a node η is determined by the ranges on each dimension <emph>j</emph> that it covers, and we denote the corresponding range <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced close="]" open="[" separators=""&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . We call the tuple <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#954;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#964;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> the <emph>decision tree</emph>.</p> <p>In addition to a decision tree <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;/math&gt; </ephtml> with <emph>K</emph> leaf nodes, a CART model has a parameter Θ = (θ<subs>1</subs>, θ<subs>2</subs>, ..., θ<emph><subs>K</subs></emph>), which associates parameter value θ<emph><subs>k</subs></emph> with the <emph>k</emph>th leaf node. If a stimulus <bold>x</bold> lies in the region of the <emph>k</emph>th leaf node, then <emph>y</emph>\vert<bold>x</bold> has distribution <emph>f</emph>(<emph>y</emph>\vertθ<emph><subs>k</subs></emph>) for some parametric family <emph>f</emph>. It is typically assumed that, conditional on <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#920;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> , <emph>y</emph> values within a leaf node are i.i.d., and furthermore, that <emph>y</emph> values across leaf nodes are independent. Thus, letting <emph>n<subs>k</subs></emph> denote the number of observations assigned to the <emph>k</emph>th leaf node and letting <emph>y<subs>k</subs></emph><subs>,</subs><emph><subs>i</subs></emph> denote the <emph>i</emph>th observation of <emph>y</emph> assigned to leaf <emph>k</emph>,</p> <p>1 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0018" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#920;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8719;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;/munderover&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8719;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/munderover&gt;&lt;mi&gt;f&lt;/mi&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>where <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;/msubsup&gt;&lt;msub&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the total number of observations. As we will make more precise below, for us, the parameter θ<emph><subs>k</subs></emph> is the probability that a stimulus within the <emph>k</emph>th leaf node has category label <emph>A</emph>.</p> <p>This provides a general framework for representing categories, but we require a more detailed specification for the experiments of interest. We now do this for the categorization task used by Sewell and Lewandowsky ([<reflink idref="bib82" id="ref99">82</reflink>]), described above. The SHJ tasks employed in Lewandowsky ([<reflink idref="bib56" id="ref100">56</reflink>]) are simpler and are straightforwardly modeled with only minor modifications.</p> <p>In the Sewell–Lewandowsky task, the stimulus on each trial <emph>t</emph> comprised a three‐dimensional input <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mtext&gt;bar position&lt;/mtext&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> , <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mtext&gt;height&lt;/mtext&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> , <emph>x<subs>t</subs></emph><subs>,3</subs> = color<emph><subs>t</subs></emph> ∊ {blue = 0, red = 1}).[<reflink idref="bib2" id="ref101">2</reflink>] On training trials, participants made a category prediction before observing the binary category label <emph>y<subs>t</subs></emph> ∊ {<emph>A</emph>, <emph>B</emph>}. The "ideal" knowledge‐partitioning (KP) and context‐insensitive (CI) strategies which participants were encouraged to learn and deploy can be naturally represented in tree form (Fig. B,C).</p> <p>In the Bayesian framework, we need to specify some prior beliefs about the state of nature. In the current case, the relevant prior beliefs concern category structure which, by modeling assumption, can be formalized as a prior distribution on decision trees. Such a prior can be imposed implicitly by specifying a stochastic process for generating such trees. Following Chipman et al. ([<reflink idref="bib19" id="ref102">19</reflink>]), we set the prior probability of a node η in tree structure <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/math&gt; </ephtml> being split into children nodes to be</p> <p>2 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0023" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mtext&gt;SPLIT&lt;/mtext&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mi mathvariant="normal"&gt;&amp;#945;&lt;/mi&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;d&lt;/mi&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#946;&lt;/mi&gt;&lt;/msup&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>where <emph>d</emph><subs>η</subs> denotes the depth of the node (the depth of the root node is zero), and α &lt; 1 and β ≥ 0 are parameters controlling expected tree size. Under this specification, the probability <emph>p</emph><subs>SPLIT</subs> is a decreasing function of node depth, and it decreases more steeply for large β (cf. fig. 3 of Chipman et al., [<reflink idref="bib19" id="ref103">19</reflink>]). In all simulations, we fix α = 0.95 and β = 1, which gives a prior mean on the number of terminal nodes ≈ 3.7 (Chipman et al., [<reflink idref="bib19" id="ref104">19</reflink>]), but results are essentially identical for other reasonable parameterizations.</p> <p>In addition to a prior on tree structure <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0024" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi&gt;&amp;#8868;&lt;/mi&gt;&lt;/math&gt; </ephtml> achieved through a prior on a node's probability of splitting, we need to specify the prior probability of a node η splitting on each stimulus dimension κ(η) = {1, ..., <emph>p</emph>} and location τ(η). We generally assume that the probability of splitting on each dimension is equal, that is,</p> <p>3 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0025" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#954;&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="1em" /&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;...&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Conditional on the choice of dimension, a split location is assumed to be drawn uniformly from the node's range on the relevant dimension:</p> <p>4 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0026" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#964;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#954;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mo&gt;&amp;#8764;&lt;/mo&gt;&lt;mi mathvariant="script"&gt;U&lt;/mi&gt;&lt;/mrow&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;R&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#951;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>However, consideration of the information given to participants at the outset of Sewell and Lewandowsky's experiment leads us to a slightly different prior for the root node η<subs>0</subs>. In particular, in the experiment, participants were initially told that stimulus color (KP‐first condition) or bar position (CI‐first condition) reliably indicated whether height or bar position was diagnostic of stimulus category. We assume that this information is reflected in the prior probability of splitting the root node η<subs>0</subs> on a particular dimension. Thus, we introduce a "bias" parameter <emph>b</emph> to indicate that splits of the root node η<subs>0</subs> on one dimension should be regarded as much more likely than on the others. Letting <emph>j</emph><sups>∗</sups> indicate the dimension highlighted by instruction, we can write this prior probability as</p> <p>5 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0027" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#954;&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#951;&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfenced close="" open="{" separators=""&gt;&lt;mrow&gt;bif&amp;#954;(&amp;#951;0)=j&amp;#8727;,1-b2otherwise.&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Setting <emph>b</emph> &lt; 1, which would give nonzero probability to alternative splits at the root, might reflect incomplete confidence in the experimenter's instructions, for example.</p> <p>In addition, participants were not only guided to a particular initial dimension—bar position or color—but effectively also to an initial split location. Thus, in the KP‐first condition, attention was drawn to the color of the stimulus, while in the CI‐first condition, participants were explicitly told that the relevant feature was whether the bar was to the left or right of center. We therefore assume that split locations for the highlighted dimension at the root node are known. Note that the question of split location is actually irrelevant in the case of the (binary) color dimension since all split locations on (0, 1) are equivalent in terms of the resulting partition. However, this dimension can be treated as continuous for ease of presentation and without consequence for modeling outcomes.</p> <p>The preceding specifies a simple prior distribution on decision trees <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> that can be summarized as a process of deciding whether to split each node and, if so, selecting a splitting dimension and location. To complete the model specification, we also require a likelihood model <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> that gives the conditional probabilities of stimulus labels given the tree structure. In this case, we simply assume that the <emph>k</emph>th leaf node has an associated probability θ<emph><subs>k</subs></emph> of generating label <emph>A</emph>,</p> <p>6 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0030" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msubsup&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/msubsup&gt;&lt;msup&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>and that this probability is an i.i.d. draw from a Beta distribution,</p> <p>7 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0031" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mover&gt;&lt;mo&gt;&amp;#8764;&lt;/mo&gt;&lt;mi mathvariant="italic"&gt;iid&lt;/mi&gt;&lt;/mover&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>Standard analytical simplification for this beta‐binomial model yields the marginal likelihood</p> <p>8 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0032" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msup&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#915;&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#915;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#915;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mfenced&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;/msup&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8719;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;K&lt;/mi&gt;&lt;/munderover&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#915;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;kA&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#915;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;kA&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#915;&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>where <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;kA&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> are respectively the number of instances of category <emph>A</emph> and the total number of data points in the partition of leaf <emph>k</emph> up to trial <emph>t</emph>. Note that for a given tree, this likelihood is higher for leaves assigned observations with homogeneous labels (i.e., with labels that are either mostly <emph>A</emph> or mostly <emph>B</emph>). These are exactly the partitions that constitute "good" solutions to the categorization problem.</p> <hd id="AN0140459439-12">Inference</hd> <p>Given the model specified above, we assume that participants seek to represent the sequence of posterior distributions over possible trees <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> as they successively predict and receive information about stimulus labels over trials. Generally, a brute force procedure of enumerating all possible trees, a space which dramatically increases in size with <emph>t</emph>, is not a plausible model of how participants perform inference. Instead, we assume that people's beliefs are represented by a relatively small number of samples from these posterior distributions which can be updated over time. In other words, we model participants as performing <emph>particle filtering</emph> (Daw &amp; Courville, [<reflink idref="bib27" id="ref105">27</reflink>]; Doucet et al., [<reflink idref="bib29" id="ref106">29</reflink>]; Sanborn, Griffiths, &amp; Navarro, [<reflink idref="bib78" id="ref107">78</reflink>]).</p> <p>As mentioned above, two aspects of the inference process which we now describe draw parallels with working memory. First, similar to the idea that there is a limit on the number of items that can be held in working memory (Cowan, [<reflink idref="bib24" id="ref108">24</reflink>]), we assume there is a bounded number of hypotheses about category structure—in this case, the samples/particles which correspond to particular tree structures—that can be entertained at a given time. Second, similar to the notion that working memory is <emph>active</emph> (Baddeley, [<reflink idref="bib8" id="ref109">8</reflink>]), involving the manipulation rather than merely passive storage of items, we assume that inference involves a continuing process whereby local transformations to current hypotheses are proposed, and which may be accepted or rejected. The latter process promotes diversity in the hypothesis set and continuous exploration of the hypothesis space.</p> <p>In detail, we assume that on a given trial <emph>t</emph>, a participant's beliefs are represented by a small set of <emph>L</emph> possible trees <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;L&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> with associated weights <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0037" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;L&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> proportional to their posterior probability. This set of trees constitutes the limited set of hypotheses putatively maintained in a working memory of capacity <emph>L</emph>. With the observation of the stimulus and category label on the next trial <emph>t</emph> + 1, a proper reweighting of the <emph>l</emph>th tree is given by the following update (Chopin, [<reflink idref="bib20" id="ref110">20</reflink>]):</p> <p>9 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0038" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;wt+1(l)&amp;#8733;wt(l)p(T(l)|x1:t+1,y1:t+1)p(T(l)|x1:t,y1:t)&amp;#8733;wt(l)p(y1:t+1|T(l),x1:t+1)p(y1:t|T(l),x1:t)=wt(l)p(yt+1|T(l),xt+1,y1:t).&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>As standard within‐particle filtering methods (Doucet et al., [<reflink idref="bib29" id="ref111">29</reflink>]), this reweighting process can be alternated with a <emph>resampling</emph> stage in which very unlikely trees, that is, those with very low weights, are discarded to be replaced by replicates of more probable trees. A simple way of doing this is to sample <emph>L</emph> times with replacement from the set {<emph>T</emph><sups>(</sups><emph><sups>l</sups></emph><sups>)</sups>} with probabilities proportional to the updated weights <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;L&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> (Gordon, Salmond, &amp; Smith, [<reflink idref="bib43" id="ref112">43</reflink>]).</p> <p>Additionally, this resampled particle set can then be "rejuvenated" (Chopin, [<reflink idref="bib20" id="ref113">20</reflink>]; Gilks &amp; Berzuini, [<reflink idref="bib40" id="ref114">40</reflink>]), reintroducing diversity and allowing continuous exploration of alternative solutions. This is the "active" step which, we suggest, recalls conceptions of working memory as involving active manipulation of currently stored items. Specifically, we may, without altering the targeted posterior distribution of interest, propose transformations of trees from a Markov chain transition kernel <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;q&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> and accept or reject these proposals such that we retain the appropriate stationary distribution <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . Closely following the transition kernel suggested by Chipman et al. ([<reflink idref="bib19" id="ref115">19</reflink>]), we consider the scheme where for each tree <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> , a new tree <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/math&gt; </ephtml> is proposed by randomly choosing among three possible transformations (Fig. D):</p> <p></p> <ulist> <item> GROW: Randomly select a leaf node, then draw a splitting dimension and location from the prior (Eqs. 3 and 4). Not permitted if the split leads to an empty node (i.e., a partition with no assigned data points).</item> <p></p> <item> PRUNE: Randomly select an internal node, then turn it into a leaf node by deleting all nodes below it. Not permitted if the tree comprises only the root node.</item> <p></p> <item> CHANGE: Randomly select an internal node, then randomly reassign it a splitting dimension and location by a draw from the prior. Not permitted if the reassigned split is inconsistent with splits of nodes below the selected node.</item> </ulist> <p>This proposed tree <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/math&gt; </ephtml> is then accepted with probability</p> <p>10 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0045" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;&amp;#945;&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mo movablelimits="true"&gt;min&lt;/mo&gt;&lt;mfenced close="}" open="{" separators=""&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;q&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;q&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;&amp;#8727;&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>as per the standard Metropolis‐Hastings algorithm (Gelman, Carlin, Stern, &amp; Rubin, [<reflink idref="bib37" id="ref116">37</reflink>]). This simple "resample‐move" algorithm (Chopin, [<reflink idref="bib20" id="ref117">20</reflink>]; Gilks &amp; Berzuini, [<reflink idref="bib40" id="ref118">40</reflink>]) is summarized in Algorithm 1.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0010.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0010.jpg" title="." /> </p> <p></p> <p>Why might the number of samples/particles be expected to influence category learning? The basic intuition comes from viewing the category learning process as one of <emph>search</emph> (Fig.). In particular, "good" category structures are those that partition stimuli into regions with homogeneous labels (<emph>A</emph> or <emph>B</emph>), and these are the category structures that have high posterior probability. In the sample‐based inference procedure we consider, the population of particles will seek out regions of high posterior probability, and the rate at which these regions are found may plausibly depend on the number of particles.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0004.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0004.jpg" title="Category learning as search. In the formulation here, category learning is conceptualized as a process of search for category structures h∈H that have a high posterior probability, p(h\d), given both the prior distribution on category structures and the observed data, D. In the sample‐based inference procedure considered, this search is enacted by a particle set (black circles) whose positions may be changed through the acceptance of proposed local changes to the corresponding category structure. Proposals that result in a category structure with higher posterior probability (arrows) will be accepted more often. With a larger number of particles (right), this search may be more efficient, in that high probability structures will be discovered more quickly." /> </p> <p></p> <p>So far, we have suggested a particle filtering scheme for representing a sequence of posterior distributions over category structures, where that structure is assumed to be specified by a classification tree. However, we have not yet addressed the issue of <emph>strategy switching</emph>. Thus, in the Sewell–Lewandowsky experiment, participants were able to immediately switch between different categorization strategies when instructed to do so, and in the absence of further training.</p> <p>We model such switches as a simple <emph>reweighting</emph> operation on the set of trees. Take the specific example where a participant has initially been encouraged to use the CI strategy and after <emph>t</emph> training sessions has in mind the set of weighted trees <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mrow&gt;&lt;mo&gt;{&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;}&lt;/mo&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;L&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> approximating the target distribution under the prior appropriate to the CI strategy. We denote this target distribution <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0048" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;CI&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . The experimenter then instructs the participant to change to using the KP strategy. Assuming that the set of trees remains fixed, the associated tree weights now need to be changed to reflect the new target distribution <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;KP&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> . This can be achieved by an <emph>importance weighting</emph> step, treating <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;CI&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> as the importance distribution. In particular, denoting a particle's weight before and after the instruction to switch as <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> , respectively, the relevant reweighting is as follows:</p> <p>11 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0053" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;&amp;#8733;&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;KP&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;CI&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>which, under the specified model, becomes particularly simple:</p> <p>12 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0054" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;/mrow&gt;&lt;/msubsup&gt;&lt;mo&gt;&amp;#8733;&lt;/mo&gt;&lt;mfenced close="" open="{" separators=""&gt;&lt;mrow&gt;wt(l)-&amp;#215;1-b2/bif&amp;#954;(&amp;#951;0)=bar position,wt(l)-if&amp;#954;(&amp;#951;0)=height,wt(l)-&amp;#215;b/1-b2if&amp;#954;(&amp;#951;0)=color.&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>To switch in the reverse direction—from the KP to CI strategy—the appropriate reweighting involves the ratio <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;CI&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;KP&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> , with the appropriate alterations made to Eq. 12.</p> <p>Again, why might a greater number of particles improve ability to switch between strategies? Consider the cartoon example in Fig. A, depicting the posterior probability <emph>P</emph>(<emph>h</emph>\<emph>d</emph>) of different possible category structures <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0056" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;h&lt;/mi&gt;&lt;mo&gt;&amp;#8712;&lt;/mo&gt;&lt;mi&gt;H&lt;/mi&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> given a stimulus set <emph>D</emph>. In this example, two particular category structures, <emph>h</emph><subs>1</subs> and <emph>h</emph><subs>2</subs>, are most probable, and equally so, and we can think of these as being two equally valid categorization strategies, as in the Sewell–Lewandowsky task. Again, this probability distribution will be represented by a set of particles with locations (i.e., particular category structures) drawn from this distribution, along with corresponding weights that are proportional to the posterior probabilities of those locations.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0005.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0005.jpg" title="Particle diversity and flexibility of behavior. Cartoon of how different numbers of particles affect the model's ability to switch between different categorization strategies. (A) Given the observed data D, comprising a set of stimuli and their category labels, there is a posterior distribution P(h\d) over the set of possible category structures h∈H. Here, two particular category structures h1 and h2 are equally probable, and can be considered as two equally valid categorization strategies. The distribution can be approximated by a set of particles, where each particle has a particular location (circles), corresponding to a category structure h, and a weight, which is proportional to the posterior probability (vertical, dashed lines). (B) The instruction to use a particular strategy is conceptualized as biasing the posterior distribution so that particular category structures are more probable, in this case category structures in the region of h1. Whether regions of lower probability are represented in the approximation depends on the number of particles: If there are many particles, some are likely to be located in regions of lower probability, such as around h2 (upper); if there are fewer particles, there may be no particles in this region (lower). (C) The instruction to switch strategy is conceptualized as leading to a change in the posterior distribution, and a corresponding change in the particle weights (upper); however, in the case of fewer particles, there may be no particles immediately available to represent the change in distribution (lower)." /> </p> <p></p> <p>Now assume that the effect of an instruction to use a particular strategy is to increase the posterior probability of category structures that accord with that strategy, in this case those in the region of <emph>h</emph><subs>1</subs> (Fig. B). Such a change in posterior distribution, driven by the different priors underlying the distinct strategies, is exactly what we assumed when suggesting that strategy‐switching is mediated by a reweighting of particles (see above). Depending on the number of particles available, how well this collection of particles represents the true posterior distribution—especially in regions of lower probability—may differ. With a sufficiently large number of particles, at least some particles should be allocated to regions of lower probability, such as around <emph>h</emph><subs>2</subs> (Fig. B, upper). However, with a decreasing number of particles, representation of the posterior distribution may become impoverished to the extent that such regions of low probability may not contain any particles at all (Fig. B, lower). In other words, the shift in "mental set" associated with a switch in categorization strategy is here implemented by a change in posterior distribution; the participant's immediate ability to represent this change is assumed to depend in some sense on the diversity of the current hypothesis set.</p> <p>The possible relevance to knowledge restructuring is what these different degrees of approximation to the true posterior may entail when instructed to switch categorization strategy. Intuitively, if fewer resources have been devoted to representing alternative strategies in the first place, however unlikely, then it may be more difficult to entertain these alternatives when instructed to do so. In our particular formulation of the switching process, we considered a simple formulation in which the immediate effect of an instruction to switch strategy is that the locations of the particles remain the same, but the relative weightings of particles are updated according to the new posterior distribution (Fig. C). In particular, if there are particles located in the region of <emph>h</emph><subs>2</subs>, these will immediately be updated (Fig. C, upper), and the new categorization strategy can be immediately deployed. By contrast, if there are no particles located in the region of <emph>h</emph><subs>2</subs>, no up‐weighting can occur and the alternative strategy is initially unavailable (Fig. C, lower).</p> <hd id="AN0140459439-16">Choice</hd> <p>We have so far described a process for performing inference (i.e., particle filtering) under an assumed generative model for the structure of categories (i.e., CART). What is still missing is a model of how participants finally generate a guess about a stimulus's category label before they receive feedback in the form of the true label. We consider two possible choice rules: one in which a participant chooses the category label with the highest probability ("maximum‐probability rule"), and another in which a participant chooses a category label stochastically in accord with their probabilities ("probability‐matching rule"). Since there is no explore‐exploit dilemma in the categorization tasks we consider—full information about the correct label is always received, regardless of choice—participants should always select the label they think is most likely (i.e., maximum‐probability rule). On the other hand, given that probability‐matching behavior has sometimes been observed in this domain (e.g., Estes, Campbell, Hatsopoulos, &amp; Hurwitz, [<reflink idref="bib34" id="ref119">34</reflink>]; Gluck &amp; Bower, [<reflink idref="bib41" id="ref120">41</reflink>]), we considered it possible that participants also used this strategy, despite it being suboptimal in the tasks considered.</p> <p>From the above, a sample‐based approximation to the predictive probability that a stimulus <bold>x</bold><emph><subs>t</subs></emph><subs>+1</subs> has label <emph>y<subs>t</subs></emph><subs>+1</subs> = <emph>A</emph> is given by</p> <p>13 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0058" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;p(yt+1=A|x1:t+1,y1:t)=&amp;#8721;Tp(yt+1=A|x1:t+1,y1:t,T)p(T|x1:t,y1:t)&amp;#8776;1L&amp;#8721;l=1Lp(yt+1=A|x1:t+1,y1:t,T(l))=1L&amp;#8721;l=1LE&amp;#952;k|x1:t+1,y1:t,T(l)[&amp;#952;k],&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>noting that</p> <p> <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0059" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;p(yt+1=A|x1:t+1,y1:t,T(l))=&amp;#8747;p(yt+1=A|x1:t+1,y1:t,&amp;#952;k,T(l))p(&amp;#952;k|x1:t+1,y1:t,T(l))d&amp;#952;k=&amp;#8747;&amp;#952;kp(&amp;#952;k|x1:t+1,y1:t,T(l))d&amp;#952;k=E&amp;#952;k|x1:t+1,y1:t,T(l)[&amp;#952;k].&lt;/mrow&gt;&lt;/math&gt; </ephtml> </p> <p>Equation 13 simply says that an approximation to the predictive probability in this case is given by an unweighted average of posterior means for θ<emph><subs>k</subs></emph>, where <emph>k</emph> for the <emph>l</emph>th particle is the index of the leaf node relevant to the input <bold>x</bold><emph><subs>t</subs></emph><subs>+1</subs> in <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0060" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/math&gt; </ephtml> . For the leaf model used in the current case, the posterior mean is given by</p> <p>14 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0061" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="double-struck"&gt;E&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msup&gt;&lt;mi mathvariant="script"&gt;T&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;msub&gt;&lt;mi mathvariant="normal"&gt;&amp;#952;&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;kA&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>where, again, <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0062" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mi mathvariant="italic"&gt;kA&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> and <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0063" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msubsup&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msubsup&gt;&lt;/math&gt; </ephtml> are, respectively, the number of instances of category <emph>A</emph> and the total number of data points in the partition of leaf <emph>k</emph> up to trial <emph>t</emph>. The deterministic maximum‐probability rule would choose the category label with the highest predictive probability, but more generally we consider the ϵ‐greedy form</p> <p>15 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0064" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#1013;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mn mathvariant="double-struck"&gt;1&lt;/mn&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;~&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;mover accent="true"&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;~&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;B&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;0.5&lt;/mn&gt;&lt;mi mathvariant="normal"&gt;&amp;#1013;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>where <emph>P<subs>t</subs></emph><subs>+1</subs>(<emph>A</emph>) is the probability of guessing category <emph>A</emph> on trial (<emph>t</emph> + 1), <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0065" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;~&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is shorthand for the sample‐based approximation given in Eq. 13, ϵ is the probability of guessing a category label according to the flip of a fair coin, and <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0066" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msub&gt;&lt;mn mathvariant="double-struck"&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#183;&lt;/mo&gt;&lt;/msub&gt;&lt;/math&gt; </ephtml> is the indicator function. In other words: choose the most probable label with probability (1 − ϵ), or with probability ϵ simply flip a coin. When ϵ = 0, we recover the deterministic case.</p> <p>The probability‐matching rule takes the slightly different form</p> <p>16 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0067" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;&amp;#1013;&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;~&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mi&gt;A&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;0.5&lt;/mn&gt;&lt;mi mathvariant="normal"&gt;&amp;#1013;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>so that the probability of guessing a category label is a linear combination of its predictive probability <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0068" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mover accent="true"&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mo&gt;~&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> (again, using shorthand for the probability given in Eq. 13) and the guessing rate ϵ; strict probability‐matching is obtained when ϵ = 0.</p> <p>Given that sample‐based inference will itself tend to introduce stochasticity, we should comment on the addition of a guessing rate ϵ, which, for ϵ &gt; 0, will provide an additional source of variability. Briefly, our motivation was simply the (common) observation that model fit was improved by including this parameter; the behavior of participants tended to exhibit levels of variability beyond what our model would generate with ϵ = 0, even with a single particle. As such, ϵ captures our ignorance about such variability, which may arise from sources distinct from sample‐based inference (e.g., lapses in attention, lack of motivation, etc.). Of course, the price to be paid for this improvement in fit, as we will see below, is that apportioning responsibility for behavioral variability to different components of the model—inference versus choice—becomes all the more difficult.</p> <hd id="AN0140459439-17">Model‐fitting and analysis</hd> <p>Models of varying degrees of complexity were fit to the data by finding the combination of the parameters of our category‐learning model (described above) that maximized the likelihood of the observed sequence of category predictions. Models varied in the number of parameters to be fit, lying on a spectrum from the simplest case, which required that all participants be fit by a single set of parameters, to the most complex case, in which each participant was fit with a separate set of parameters. Formally, denoting an observed sequence of predictions over <emph>T</emph> trials by <emph>c</emph><subs>1:</subs><emph><subs>T</subs></emph> and the full set of parameters by Φ = {<emph>L</emph>, <emph>b</emph>, α, β, <emph>a</emph><subs>0</subs>, <emph>b</emph><subs>0</subs>, ϵ} (see Table), the general aim was to find the (free) parameters Φ that maximized the probability</p> <p>17 <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0069" display="block" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;&amp;#934;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;munderover&gt;&lt;mo movablelimits="false"&gt;&amp;#8719;&lt;/mo&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;/munderover&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;c&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/msub&gt;&lt;mrow&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi mathvariant="bold"&gt;&amp;#934;&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mi mathvariant="bold"&gt;x&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;y&lt;/mi&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml></p> <p>with the trial‐by‐trial probabilities extracted from Eq. 15 or Eq. 16, as appropriate.</p> <p>Model parameters</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="top"&gt;&lt;tr&gt;&lt;th align="left"&gt;Parameters&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left"&gt;Fixed&lt;/th&gt;&lt;th align="left"&gt;Free&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#945;&amp;#160;=&amp;#160;0.95&lt;/td&gt;&lt;td align="left"&gt;L: number of particles&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#946;&amp;#160;=&amp;#160;1&lt;/td&gt;&lt;td align="left"&gt;b: bias&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;b&lt;sub&gt;0&lt;/sub&gt;&amp;#160;=&amp;#160;a&lt;sub&gt;0&lt;/sub&gt;&lt;/td&gt;&lt;td align="left"&gt;a&lt;sub&gt;0&lt;/sub&gt;: Beta shape parameter&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&amp;#160;&lt;/td&gt;&lt;td align="left"&gt;&amp;#1013;: random guessing rate&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Best‐fit parameters for a given model were defined as those maximizing the average likelihood in a grid search. The grid was defined as follows: number of particles <emph>L</emph> logarithmically spaced on the interval [<reflink idref="bib1" id="ref121">1</reflink>, 100], yielding 34 values; guessing rate uniformly spaced ϵ ∊ (0, 0.02, 0.04, ..., 0.2); and shape <emph>a</emph><subs>0</subs> ∊ (0.01, 0.1, 0.5, 1). In the knowledge restructuring case, we also included three possible values of bias, <emph>b</emph> ∊ {0.5, 0.75, 0.9}. The grid values were chosen to reflect our a priori assumptions about plausible parameter values. That is, we expected participants to be more plausibly modeled as instantiating relatively few particles (hence the logarithmic scale), and as expressing noise levels in the lower range (hence the upper limit of 0.2 on the guessing rate ϵ). The choice of comparatively finely spaced ϵ values was motivated by the expectation that <emph>L</emph> and ϵ would at least partly trade off with each other, so effort was made to make the resolution of these parameters comparable in order to minimize the possibility of bias. In addition, we included the case where the number of particles was set to a much larger number (<emph>L</emph> = 10,000); this was to provide a comparison model that approximated the full posterior distribution much more closely than when the number of particles was more restricted.</p> <p>Since the estimate of the likelihood was generally less reliable with fewer particles (due to greater variability in the algorithm's behavior), the number of simulation runs was chosen so that an "effective" number of particles would be constant, thereby facilitating a fair comparison between the fits of different numbers of particles. We set the effective number of particles to 1,000, so that the number of simulation runs was determined by rounding to the nearest integer the result of 1,000/<emph>L</emph> (i.e., the 1‐particle case was run 1,000 times, the 100‐particle case was run 10 times, etc.).</p> <p>As mentioned in Section 2.3, we additionally compared two different choice models. Modulo the effect of the guessing rate ϵ, either a stimulus was deterministically assigned to the most likely category (maximum‐probability choice rule), or it was probabilistically assigned to a category in proportion to that category's predictive probability (probability‐matching choice rule).</p> <p>In evaluating the fit of different models, we used the Bayesian information criterion (BIC) to select the best‐fitting model (Schwarz, [<reflink idref="bib80" id="ref122">80</reflink>]). That is, we chose the model <emph>M</emph> for which the quantity <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0070" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mtext&gt;BIC&lt;/mtext&gt;&lt;mo&gt;&amp;#8801;&lt;/mo&gt;&lt;mo&gt;-&lt;/mo&gt;&lt;mi&gt;log&lt;/mi&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mfenced close=")" open="(" separators=""&gt;&lt;mrow&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;/msub&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mfrac&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mfrac&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;g&lt;/mi&gt;&lt;mrow&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> was minimized, where <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0071" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;mi&gt;P&lt;/mi&gt;&lt;mo&gt;(&lt;/mo&gt;&lt;mi&gt;D&lt;/mi&gt;&lt;mo&gt;|&lt;/mo&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;/msub&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/math&gt; </ephtml> is the value of the likelihood function (see above) given the maximum likelihood estimate <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0072" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi mathvariant="normal"&gt;&amp;#934;&lt;/mi&gt;&lt;mo&gt;^&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;/msub&gt;&lt;/math&gt; </ephtml> of the model parameters, <emph>k</emph> is the number of estimated parameters in the model, and <emph>n</emph> is the number of data points (i.e., the number of trials). The BIC provides a simple yet principled (via its approximation of the Bayes factor) approach to model comparison (for more details, see, for example, Kass &amp; Raftery, [<reflink idref="bib50" id="ref123">50</reflink>]).</p> <p>To assess relationships between best‐fitting model parameters and participants' WMC scores, we used two methods. The first was simply to measure the correlation between parameters and WMC scores, and determine whether the correlation was significantly different from zero. While this method has the advantage of being straightforward, the strength of the correlation can be reduced both by imprecision in the estimates of the best‐fitting model parameters and tradeoffs between parameters in fitting the data. While these issues cannot be entirely avoided, we developed a second measure to mitigate them that involved estimating a function that mapped WMC scores to a particular parameter of interest as part of the fitting procedure. To do so, we again used BIC scores to compare slope‐intercept models (in which the parameter of interest was a linear function of the individual WMC scores) against intercept‐only models (in which the parameter was fixed across participants and thus did not depend on WMC scores). In cases in which there is a relationship between a parameter and WMC score, the slope‐intercept model should perform better, since the slope helps to capture that relationship. Our second measure helps address imprecision in estimating parameters because the parameters fit in the slope‐intercept model are the best‐fitting values that are consistent with a relationship with WMC, so if the individual parameters are somewhat imprecise but still consistent with a relationship to WMC, then the slope‐intercept model would still perform best. Additionally, because of the concern about parameter tradeoffs in fitting the data, we allowed the other parameters in both the slope‐intercept and intercept‐only models to freely vary, so that these other parameters could trade off against the linear relationship between the parameter of interest and WMC in the way that allowed the best fit to the data. When comparing details of model fit with a participant's WMC score, we always used for the latter the average of that participant's scores over the battery of working memory tasks used in Lewandowsky ([<reflink idref="bib56" id="ref124">56</reflink>]) and Sewell and Lewandowsky ([<reflink idref="bib82" id="ref125">82</reflink>]).</p> <hd id="AN0140459439-18">Results</hd> <p></p> <hd id="AN0140459439-19">Category learning</hd> <p></p> <hd id="AN0140459439-20">Simulations</hd> <p>Both Lewandowsky ([<reflink idref="bib56" id="ref126">56</reflink>]) and Sewell and Lewandowsky ([<reflink idref="bib82" id="ref127">82</reflink>]) found that working memory capacity (WMC) was positively correlated with category learning performance, such that participants with higher WMC tended to make fewer categorization errors. We hypothesized that a greater number of particles would have a similar effect because, on average, one might expect the search for a "good" (i.e., more probable) category structure to progress faster, and with less chance of getting stuck at local maxima, with a higher number of particles (Fig.). Here, we focus on simulating the classical SHJ tasks used by Lewandowsky ([<reflink idref="bib56" id="ref128">56</reflink>]). Since we always found that the probability‐matching choice rule yielded better fits to the data than the maximum‐probability rule (see Table below), the simulation results always reflect use of the former.</p> <p>Model comparison, SHJ tasks</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="top"&gt;&lt;tr&gt;&lt;th align="left"&gt;Model&lt;/th&gt;&lt;th align="left"&gt;No. Free Parameters&lt;/th&gt;&lt;th align="left"&gt;NLL&lt;/th&gt;&lt;th align="left"&gt;BIC&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;3&lt;/td&gt;&lt;td align="left"&gt;40,801 (45,411)&lt;/td&gt;&lt;td align="left"&gt;40,819 (45,429)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;2&lt;/td&gt;&lt;td align="left"&gt;18&lt;/td&gt;&lt;td align="left"&gt;39,669 (44,131)&lt;/td&gt;&lt;td align="left"&gt;39,775 (44,237)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;3&lt;/td&gt;&lt;td align="left"&gt;115&lt;/td&gt;&lt;td align="left"&gt;40,664 (45,251)&lt;/td&gt;&lt;td align="left"&gt;41,342 (45,928)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;4&lt;/td&gt;&lt;td align="left"&gt;115&lt;/td&gt;&lt;td align="left"&gt;38,569 (41,119)&lt;/td&gt;&lt;td align="left"&gt;39,246 (41,796)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;5&lt;/td&gt;&lt;td align="left"&gt;115&lt;/td&gt;&lt;td align="left"&gt;39,336 (45,171)&lt;/td&gt;&lt;td align="left"&gt;40,013 (45,848)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;6&lt;/td&gt;&lt;td align="left"&gt;339&lt;/td&gt;&lt;td align="left"&gt;37,649 (40,827)&lt;/td&gt;&lt;td align="left"&gt;39,645 (42,824)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;7&lt;/td&gt;&lt;td align="left"&gt;2,034&lt;/td&gt;&lt;td align="left"&gt;34,284 (34,915)&lt;/td&gt;&lt;td align="left"&gt;46,261 (46,892)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>3 Notes.</item> <item>4 We compared model fit under different constraints of the number of parameters. Model 1: single set of parameters {<emph>L</emph>, <emph>a</emph><subs>0</subs>, ϵ} fixed across all participants and problem types. Model 2: single set of parameters per problem type, fixed across participants. Model 3: different number of particles <emph>L</emph> per participant, fixed across problems, with {<emph>a</emph><subs>0</subs>, ϵ} fixed across participants. Model 4: different guessing rate ϵ per participant, fixed across problems, with {<emph>L</emph>, <emph>a</emph><subs>0</subs>} fixed across participants. Model 5: different shape <emph>a</emph><subs>0</subs> per participant, fixed across problems, with {<emph>L</emph>, ϵ} fixed across participants. Model 6: single set of parameters per participant, fixed across problem types. Model 7: single set of parameters per participant‐problem type. Values for the maximum‐probability choice rule are shown in parentheses. BIC, Bayesian information criterion; NLL, negative log likelihood.</item> </ulist> <p>Fig. A shows the overall average error rate for simulations as the number of particles is increased from 1 to 20 while keeping other parameter values fixed (<emph>a</emph><subs>0</subs> = 1, ϵ = 0); each data point represents 113 simulation runs, where each simulation run uses a stimulus sequence of 192 trials observed by one of the 113 participants in Lewandowsky ([<reflink idref="bib56" id="ref129">56</reflink>]). For each problem type, increasing the number of particles does indeed lead to a decrease in the average proportion of errors, though the size of this effect is rather modest and quickly asymptotes (note that the <emph>x</emph>‐axis here indicates the number of particles—not block number, as in Fig. C).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0006.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0006.jpg" title="Effect of model parameters on category learning in the SHJ problems. Overall proportion of errors (i.e., averaged across blocks) for each problem type as each parameter is varied. (A; B) Number of particles L; other parameters fixed a0 = 1, ϵ = 0. (C; D) Noise ϵ; other parameters fixed a0 = 1, L = 1. (E; F) Shape a0; other parameters fixed L = 1, ϵ = 0. Numbers in (E) for Problem VI indicate the average number of nodes in the final classification tree. Each data point represents an average of 113 simulation runs; error bars in lower panels indicate +1 SD. Note the reversed x‐axes for C and E." /> </p> <p></p> <p>Note that even without attempting to fit the parameters of the model, the ordering of error rates produced by the model for the different problem types conforms to the basic SHJ pattern of results—Type I easiest and Type VI hardest, with Types II–V clustered in between. Briefly, this is because of the so‐called automatic Occam's razor, which refers to a preference for simpler, or more parsimonious, hypotheses, and which arises naturally within the Bayesian framework (Goodman et al., [<reflink idref="bib42" id="ref130">42</reflink>]; MacKay, [<reflink idref="bib61" id="ref131">61</reflink>]).</p> <p>It is also interesting to note that the difference in the simulated error rates between the Type II problem and, for example, Type IV increases—up to a point—as the number of particles grows. An advantage in learning Type II relative to Type IV problems has been reported in the experimental literature (e.g., Nosofsky, Gluck, Palmeri, McKinley, &amp; Glauthier, [<reflink idref="bib67" id="ref132">67</reflink>]; Shepard et al., [<reflink idref="bib83" id="ref133">83</reflink>]), though this has not always been found, as in Lewandowsky ([<reflink idref="bib56" id="ref134">56</reflink>]) (cf. Kurtz, Levering, Stanton, Romero, &amp; Morris, [<reflink idref="bib54" id="ref135">54</reflink>]). Given our basic hypothesis that WMC reflects number of particles, this simulation result prompts the question of whether Type II advantage depends on WMC.</p> <p>To investigate this further, we revisited the data of Lewandowsky ([<reflink idref="bib56" id="ref136">56</reflink>]), splitting participants into low‐ and high‐WMC groups according to a median split of WMC scores and entering blockwise error rates into a 2 (WMC: low, high) × 2 (Problem: II, IV) × 12 (Block: 1–12) repeated measures anova. In addition to a main effect of WMC (<emph>F</emph>(<reflink idref="bib1" id="ref137">1</reflink>, 111) = 7.65, <emph>p</emph> &lt; .01), we found a significant three‐way interaction between WMC*Problem*Block (<emph>F</emph>(<reflink idref="bib11" id="ref138">11</reflink>, 1221) = 2.19, <emph>p</emph> = .01). High‐WMC participants performed significantly better in terms of proportion correct on Type II (<emph>M</emph> = 0.92, <emph>SD</emph> = 0.10) than on Type IV (<emph>M</emph> = 0.88, <emph>SD</emph> = 0.11; <emph>t</emph>(<reflink idref="bib56" id="ref139">56</reflink>) = 2.19, <emph>p</emph> = .02), while low‐WMC participants did not perform significantly differently on these two problem types (Type II: <emph>M</emph> = 0.85, <emph>SD</emph> = 0.14; Type IV: <emph>M</emph> = 0.85, <emph>SD</emph> = 0.14, <emph>t</emph>(<reflink idref="bib55" id="ref140">55</reflink>) = 0.06, n.s.). Learning curves are shown in the Appendix (Fig. S1). This result is consistent with our basic hypothesis, as we expect a Type II advantage to appear, or become stronger, with more particles (i.e., higher WMC).</p> <p>The effect of a larger number of particles across problem types is further illustrated in Fig. B, where we compare the overall error proportions for the extreme case of 1 particle vs. 100 particles. A larger number of particles reduces the error rate for each problem type, and in a manner that qualitatively resembles that observed in the experimental data when participants are grouped according to WMC score (cf. Fig. D). Indeed, a rank‐ordering of problem types by the extent to which performance is better for higher WMC/particles revealed a significant positive correlation (Spearman's rank‐order correlation <emph>r<subs>s</subs></emph>(<reflink idref="bib4" id="ref141">4</reflink>) = .94, <emph>p</emph> &lt; .05). In other words, the problem types that show greatest difference between high‐ and low‐WMC participants tend also to be those where an increased number of particles also makes the most difference (from greatest to smallest advantage, the experimental pattern follows the order VI, II, III, I, V, IV; our simulations follow the order VI, II, III, V, I, IV).</p> <p>We also examined the effect on performance of varying the other free parameters (i.e., guessing rate ϵ and shape <emph>a</emph><subs>0</subs>). Fig. C shows, unsurprisingly, that the proportion of errors decreases linearly as ϵ decreases. Since this rate of decrease is essentially uniform across problem types, the amount of improvement in each problem type is roughly the same (Fig. D). A simple inverse association between WMC and guessing rate therefore fails to capture differential effects of WMC on performance of the problem types (Spearman's rank‐order correlation <emph>r<subs>s</subs></emph>(<reflink idref="bib4" id="ref142">4</reflink>) = −.37, <emph>p</emph> = .50, n.s.).</p> <p>Decreasing <emph>a</emph><subs>0</subs> generally leads to a lower error rate—recall that a higher <emph>a</emph><subs>0</subs> entails a higher tolerance for "mixed" categories (i.e., instances of both <emph>A</emph> and <emph>B</emph>; cf. Section 2.1)—with Problem Type VI proving a notable exception (Fig. E,F). Briefly, what happens in the latter case is that the model becomes increasingly intolerant of the intermediate tree manipulations necessary to reach a more satisfactory solution; this can be observed, for example, in the decreasing average number of nodes in the final tree as <emph>a</emph><subs>0</subs> is decreased (Fig. E). A simple inverse association between WMC and shape therefore does a worse job compared to particles at capturing differential effects of WMC on performance of the different problem types (Spearman's rank‐order correlation, <emph>r<subs>s</subs></emph>(<reflink idref="bib4" id="ref143">4</reflink>) = −.83, <emph>p</emph> = .06, n.s.).</p> <hd id="AN0140459439-22">Model‐fitting</hd> <p>In fitting model parameters, we compared a number of possibilities ranging from the case where all participants were constrained to share a single set of parameters (Model 1; least flexible) to the case where each participant was free to have a different set of parameters for each problem type (Model 7; most flexible). Models of intermediate complexity included the cases where two of the three free parameters {<emph>L</emph>, <emph>a</emph><subs>0</subs>, ϵ} were fixed across subjects, while the other free parameter was allowed to vary between subjects (Model 3: vary <emph>L</emph>; Model 4: vary ϵ; Model 5: vary <emph>a</emph><subs>0</subs>). Since the probability‐matching choice rule always fit better than the maximum‐probability choice rule (compare numbers without and with parentheses, respectively, in Table), we restrict our attention to the results of the former case.</p> <p>In terms of BIC, the model in which the number of particles <emph>L</emph> and shape <emph>a</emph><subs>0</subs> were fixed across subjects, while guessing rate ϵ was allowed to vary between participants (i.e., Model 4), was found to fit best (Table). By contrast, fit for the model in which the number of particles <emph>L</emph> was allowed to vary between participants, with <emph>a</emph><subs>0</subs> and ϵ fixed (Model 3), was comparatively poor. The comparison model—with a large number (<reflink idref="bib10" id="ref144">10</reflink>,000) of particles—resulted in poorer fit both when we allowed shape <emph>a</emph><subs>0</subs> and noise ϵ to vary between subjects (NLL = 41,624, BIC = 42,955), and when only noise was allowed to vary between subjects (NLL= 42,469, BIC= 43,141). The probability‐matching choice rule yielded a better fit than the maximum‐probability choice rule in all models.</p> <p>The upper panels of Fig. display the blockwise average learning curves resulting from respectively simulating from Model 3 (vary particles), Model 4 (vary noise), and Model 5 (vary shape) using the best‐fit parameters for each. All models produce similar behavior on average, recapitulating the ordering of problem types in the experimental data and the qualitative character of the learning curves (cf. Fig. C).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0007.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0007.jpg" title="SHJ model‐fitting results. Average behavior of best‐fit parameters (upper) and scatterplot of average working memory capacity (WMC) against best‐fit parameters (lower) for (A) Model 3 (vary particles L, fix a0, ϵ); (B) Model 4 (vary noise ϵ, fix a0, L); (C) Model 5 (vary shape a0, fix ϵ, L). Lower panels: line of least squares (gray); regression line for best‐fit intercept‐slope/intercept‐only model (black)." /> </p> <p></p> <p>The lower panels of Fig. plot each participant's average WMC against their best‐fit parameters for each model. When only the number of particles <emph>L</emph> was allowed to vary between participants (Model 3), WMC and <emph>L</emph> were positively correlated (<emph>r</emph> = .30, <emph>p</emph> &lt; .01), which was consistent with our initial hypothesis. Best‐fit values of the other parameters (fixed across subjects) were <emph>a</emph><subs>0</subs> = 0.5 and ϵ = 0.04. However, this model was not found to fit the data best. Furthermore, assuming that a participant's best‐fit number of particles is a (linear) function of WMC, we found that an intercept‐only model (NLL = 37,823, BIC = 39,160), with best‐fit intercept set to <emph>L</emph> = 1, fit these data better than a slope‐intercept model relating these variables (NLL = 37,823, BIC = 39,166), allowing the other parameters (<emph>a</emph><subs>0</subs> and ϵ) to vary freely in both cases.</p> <p>The best‐fitting model allowed the guessing rate ϵ to vary between participants, while fixing the remaining parameters across participants (Model 4). In this case, ϵ was found to be negatively correlated with our aggregate WMC measure (<emph>r</emph> = −.30, <emph>p</emph> &lt; .01), suggesting that high‐WMC participants tended to be less "noisy" in their choices. Best‐fit values of the remaining parameters, fixed across subjects, were <emph>a</emph><subs>0</subs> = 0.5 and <emph>L</emph> = 1. Furthermore, we found that a slope‐intercept model (NLL = 38,740, BIC = 40,083) fit these data better than an intercept‐only model (NLL = 39,188, BIC = 40,525). The best‐fit slope was β<subs>1</subs> = −0.6, supporting an inverse relationship between WMC and variability in behavior.</p> <p>In the model in which only shape <emph>a</emph><subs>0</subs> was allowed to vary between participants (Model 5), WMC and <emph>a</emph><subs>0</subs> were also significantly negatively correlated (<emph>r</emph> = −.35, <emph>p</emph> &lt; .01). Best‐fit values of the other parameters (fixed across subjects) were <emph>L</emph> = 1 and ϵ = 0.03. Here, a slope‐intercept model (NLL = 38,329, BIC = 39,671) fit better than an intercept‐only model (NLL = 39,147, BIC = 40,484), with best‐fit slope β<subs>1</subs> = −0.7.</p> <p>While model comparison did not support a model allowing a unique set of parameters (<emph>L</emph>, <emph>a</emph><subs>0</subs>, ϵ) for each participant (Model 6), this was the second‐best fitting model and it was of interest to examine how the free parameters might trade off against each other. The only significant correlations found were a negative correlation between best‐fit shape and number of particles (<emph>r</emph> = −.28, <emph>p</emph> &lt; .01), and a positive correlation between shape and guess rate (<emph>r</emph> = .19, <emph>p</emph> &lt; .05).</p> <hd id="AN0140459439-24">Knowledge restructuring</hd> <p></p> <hd id="AN0140459439-25">Simulations</hd> <p>Sewell and Lewandowsky ([<reflink idref="bib82" id="ref145">82</reflink>]) found a positive association between WMC and knowledge restructuring, as measured by an individual's ability to switch between different categorization strategies. We hypothesized that a greater number of particles would also give rise to this effect since a greater diversity of hypotheses could be represented, leading to an enhanced ability to flexibly shift between representations with changes in task demands (Fig.). As for our simulations of model performance in the SHJ task, we report results in which the probability‐matching choice rule is used, since it always yielded better fits to the data (see Table below).</p> <p>Model comparison, knowledge restructuring task</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="top"&gt;&lt;tr&gt;&lt;th align="left"&gt;Model&lt;/th&gt;&lt;th align="left"&gt;No. Free Parameters&lt;/th&gt;&lt;th align="left"&gt;NLL&lt;/th&gt;&lt;th align="left"&gt;BIC&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;1&lt;/td&gt;&lt;td align="left"&gt;4&lt;/td&gt;&lt;td align="left"&gt;24,535 (25,051)&lt;/td&gt;&lt;td align="left"&gt;24,557 (25,074)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;2&lt;/td&gt;&lt;td align="left"&gt;103&lt;/td&gt;&lt;td align="left"&gt;23,366 (23,833)&lt;/td&gt;&lt;td align="left"&gt;23,950 (24,416)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;3&lt;/td&gt;&lt;td align="left"&gt;103&lt;/td&gt;&lt;td align="left"&gt;23,837 (24,641)&lt;/td&gt;&lt;td align="left"&gt;24,421 (25,224)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;4&lt;/td&gt;&lt;td align="left"&gt;103&lt;/td&gt;&lt;td align="left"&gt;23,826 (24,421)&lt;/td&gt;&lt;td align="left"&gt;24,409 (25,005)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;5&lt;/td&gt;&lt;td align="left"&gt;103&lt;/td&gt;&lt;td align="left"&gt;22,915 (23,924)&lt;/td&gt;&lt;td align="left"&gt;23,499 (24,507)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;6&lt;/td&gt;&lt;td align="left"&gt;400&lt;/td&gt;&lt;td align="left"&gt;21,165 (21,709)&lt;/td&gt;&lt;td align="left"&gt;23,431 (23,975)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>5 Notes.</item> <item>6 We compared model fit under different constraints of the number of parameters. Model 1: single set of parameters {<emph>L</emph>, <emph>b</emph>, <emph>a</emph><subs>0</subs>, ϵ} fixed across all participants. Model 2: different number of particles <emph>L</emph> per participant, with {<emph>b</emph>, <emph>a</emph><subs>0</subs>, ϵ} fixed across participants. Model 3: different bias <emph>b</emph> per participant, with {<emph>L</emph>, <emph>a</emph><subs>0</subs>, ϵ} fixed across participants. Model 4: different shape <emph>a</emph><subs>0</subs> per participant, with {<emph>L</emph>, <emph>b</emph>, ϵ} fixed across participants. Model 5: different noise ϵ per participant, with {<emph>L</emph>, <emph>b</emph>, <emph>a</emph><subs>0</subs>} fixed across participants. Model 6: single set of parameters per participant. Values for the maximum‐probability choice rule are shown in parentheses. BIC, Bayesian information criterion; NLL, negative log likelihood.</item> </ulist> <p>Fig. shows the effect of varying the number of particles <emph>L</emph> on the degree of context sensitivity (CS) change between test sessions (all other model parameters were kept fixed: <emph>b</emph> = 0.9, <emph>a</emph><subs>0</subs> = 1, ϵ = 0). Averaging over simulation runs, we observe that the extent of CS‐change increases gradually with the number of particles, regardless of whether the model initially learns a context‐insensitive (CI; Fig. A, left) or knowledge‐partitioning (KP; Fig. A, right) strategy. This graded effect predominantly reflects the effect of averaging over CS changes which are of "all or none" character—switch or no switch—where the probability of switching increases with the number of particles (Fig. B). Interestingly, the empirical CS‐change scores also display some degree of bimodality, though this is not to the same extent, nor does the degree of bimodality notably differ between high‐ and low‐WMC participants (see Appendix, Fig. S2A). Analogous to the increase in successful switching that we observe in simulations, it is also the case that participants' probability of making a successful switch (defined as for simulations, that is, a change in CS between test sessions that crosses 0.5) increases on average with higher WMC (see Appendix, Fig. S2B).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0008.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0008.jpg" title="A greater number of particles leads to improved strategy switching. (A) In both the context‐sensitive (CI)‐first (left) and knowledge‐partitioning (KP)‐first (right) condition, increasing the number of particles L leads to a greater change in context sensitivity (CS) score on average when prompted to change strategy. Average CS scores from 1,500 simulation runs per condition. (B) The effect arises because the probability of successfully switching between strategies, P(switch), increases with more particles. A successful switch is here defined as a change in context sensitivity between test sessions, ΔCS, which &quot;crosses&quot; a score of 0.5. Lower inset: with fewer particles (L = 20), it will frequently occur that the model completely fails to switch (i.e., ΔCS = 0), as visible from the distribution over change values ΔCS. Upper inset: with more particles (L = 100), such failures are very unlikely. Switch probabilities and distributions are from 3,000 simulation runs. All other parameter values were fixed: b = 0.9, a0 = 1, ϵ = 0." /> </p> <p></p> <hd id="AN0140459439-27">Model‐fitting</hd> <p>As for the SHJ tasks, we fit models of different complexity to the data. In the knowledge restructuring task, we found that allowing each participant to have their own set of parameters fit the data better in terms of BIC than simpler, less flexible models (Table). As in the SHJ case, the comparison model, with <emph>L</emph> = 10,000 particles, always resulted in a poorer fit, and the probability‐matching choice rule yielded a better fit than the maximum‐probability choice rule in all models (Table).</p> <p>Fig. A–C show aspects of behavior of the best model using the best‐fitting parameters for each participant. Fig. A shows that the average changes in context sensitivity between transfer tests of the model qualitatively resemble the empirical data (cf. Fig. C). Similarly, Fig. B confirms that the model generalizes its categorization behavior to test stimuli in a strategy‐dependent manner that closely resembles the "ideal" response profiles (cf. Fig. B), and the generalization patterns of participants (see fig. 7 in Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref146">82</reflink>]). Splitting the best‐fit parameters according to the number of particles into top and bottom quartiles also leads to a pattern of changes in context sensitivity that qualitatively resembles that of participants when grouped by WMC scores: simulations using greater numbers of particles show larger CS changes (compare Figs. C and D).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01dec19/cogs12805-fig-0009.jpg?ephost1=dGJyMNHX8kSepq84v%2bvlOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs12805-fig-0009.jpg" title="Knowledge restructuring model‐fitting results. (A) Simulated average changes in context sensitivity (±1 SE; obscured by markers) for CI‐first (squares) and KP‐first (circles) conditions. (B) Simulated average probabilities of categorizing a test stimulus as an instance of category A in the CI‐first (upper) and KP‐first (lower) conditions in the first transfer test. Darker shading indicates a higher probability. (C) Simulated average change (+1 SE) in context sensitivity (CS) given a split into top and bottom quartiles of the best‐fit parameters for all participants ranked in terms of numbers of particles. (D–G) Scatter plots of average WMC scores vs. best‐fitting parameters, with lines of least squares (gray) and regression lines for best‐fit intercept‐slope models (black): (D) number of particles L; (E) guessing rate ϵ; (F) shape a0; (G) bias b." /> </p> <p></p> <p>When we examined the relationships between individuals' average WMC scores and best‐fitting parameters (Fig. D), we found that there was no significant correlation between WMC and best‐fit number of particles (<emph>r</emph>(<reflink idref="bib98" id="ref147">98</reflink>) = .09, <emph>p</emph> = .38, n.s.). However, this correlation analysis is affected by tradeoffs between parameters, which would likely act to reduce the correlation coefficient. A more robust analysis comes from comparing a slope‐intercept model, in which WMC is assumed to be linearly related to the number of particles, to an intercept‐only model, where the number of particles is assumed to be fixed and independent of WMC; the other parameters are free to vary, as this analysis is less affected by parameter tradeoffs (the same analysis was applied to the SHJ results, above). A slope‐intercept model (NLL = 22,045, BIC = 23,756) was found to fit this relationship better than an intercept‐only model (NLL = 22,078, BIC = 23,786), but the best‐fitting slope was small, suggesting a rather weak effect (β<subs>1</subs> = 9; black line in Fig. D).</p> <p>As in the SHJ tasks, we found a significant negative correlation between WMC and guessing rate ϵ (<emph>r</emph>(<reflink idref="bib98" id="ref148">98</reflink>) = −.26, <emph>p</emph> &lt; .01; Fig. E). A slope‐intercept model with slope β<subs>1</subs> = −0.3 (NLL = 22,153, BIC = 23,863) fit better than an intercept‐only model (NLL = 22,326, BIC = 24,032).</p> <p>We found no significant correlation between WMC and shape <emph>a</emph><subs>0</subs> (<emph>r</emph>(<reflink idref="bib98" id="ref149">98</reflink>) = −.03, <emph>p</emph> = .75, n.s.; Fig. F). A slope‐intercept model (NLL = 21,473, BIC = 23,184), with slope β<subs>1</subs> = −0.4, was found to fit this relationship better than an intercept‐only model (NLL = 21,496, BIC = 23,201).</p> <p>Finally, there was a significant correlation between WMC and bias <emph>b</emph> (<emph>r</emph>(<reflink idref="bib98" id="ref150">98</reflink>) = .27, <emph>p</emph> &lt; .01; Fig. G). A slope‐intercept model (NLL = 21,459, BIC = 23,170), with slope β<subs>1</subs> = 0.9, was found to fit this relationship better than an intercept‐only model (NLL = 21,489, BIC = 23,194).</p> <p>As in the SHJ case, it was of interest to examine how these best‐fitting parameters potentially traded off against each other. We found a negative correlation between the number of particles <emph>L</emph> and the guessing rate ϵ (<emph>r</emph>(<reflink idref="bib98" id="ref151">98</reflink>) = −.40, <emph>p</emph> &lt; .01). We also found that bias <emph>b</emph> was positively correlated with number of particles <emph>L</emph> (<emph>r</emph>(<reflink idref="bib98" id="ref152">98</reflink>) = .44, <emph>p</emph> &lt; .01), and negatively correlated with the guessing rate (<emph>r</emph>(<reflink idref="bib98" id="ref153">98</reflink>) = −.59, <emph>p</emph> &lt; .01). Other correlations were not significant.</p> <hd id="AN0140459439-29">Discussion</hd> <p>Dealing with the world's many uncertainties in a consistent and principled manner presents a formidable computational challenge. That humans routinely do so despite necessarily finite cognitive resources is an impressive feat. Algorithms for approximate Bayesian inference provide one natural source of ideas for how this may be achieved. Thus, one suggestion has been that people may approximate Bayesian computations by representing and manipulating a set of samples drawn according to the relevant probability distributions (Sanborn &amp; Chater, [<reflink idref="bib77" id="ref154">77</reflink>]), that is, by implementing Monte Carlo inference (Gelfand &amp; Smith, [<reflink idref="bib35" id="ref155">35</reflink>]; Gordon et al., [<reflink idref="bib43" id="ref156">43</reflink>]). Such methods admit a spectrum of degrees of approximation, from essentially ideal performance given plentiful computational resources (e.g., a large number of samples), to much coarser approximations when such resources are scarce (e.g., few samples). In the current work, we considered constraints on working memory capacity (WMC) in the context of probabilistic inference, asking whether parallels may be drawn between WMC limitations and resource‐constrained, approximate Bayesian inference. In particular, we hypothesized that variations in task performance that correlate with WMC would be captured by assuming that WMC directly reflects the number of samples, or "particles," available to perform inference.</p> <p>To test this, we focused on experiments that suggest a positive association between WMC and two apparently disparate aspects of categorization: (a) the ease with which novel categories are learned (Lewandowsky, [<reflink idref="bib56" id="ref157">56</reflink>]); and (b) the ability to switch between different categorization strategies (Sewell &amp; Lewandowsky, [<reflink idref="bib82" id="ref158">82</reflink>]). We saw that such categorization tasks can be considered probabilistic inference problems in which individuals seek to infer the most probable category structure(s) given their prior assumptions and what they subsequently observe. We assumed that individuals approximate inference by representing and manipulating in working memory a relatively small number of hypotheses (samples/particles) about the possible underlying category structures. The number of hypotheses an individual is able to entertain at a given time was assumed to depend on his or her WMC.</p> <p>Support for our principal hypothesis was decidedly mixed. On the one hand, we provided a "proof of concept" that increasing the number of particles in our algorithm could both hasten category learning and improve switching performance, at least on average. In simulations of the SHJ problem types, we also found that the degree to which increasing the number of particles differentially improved performance in the problem types was closely matched to the manner in which higher WMC is differentially associated with improved performance in these problem types; this pattern was not matched as well by changes in other parameters. Furthermore, when the model was fit to individuals' behavior in the knowledge restructuring experiment of Sewell and Lewandowsky ([<reflink idref="bib82" id="ref159">82</reflink>]), linear regression between WMC and number of particles suggested a positive—albeit rather weak—relationship. On the other hand, when the model was fit to individuals' performance in the SHJ tasks (Lewandowsky, [<reflink idref="bib56" id="ref160">56</reflink>]), model comparison did not support a variant in which the number of particles changes as a function of WMC. Rather, the winning model favored setting the number of particles to one, and captured individual variation in performance through the guessing‐rate, or "noise," parameter. Possible reasons for this mixed picture are discussed next.</p> <hd id="AN0140459439-30">Limitations</hd> <p>One possible reason for our failure to find a relationship in the SHJ case is the relatively weak effect of varying the number of particles on learning rate. That is, although we demonstrated that increasing the number of particles could hasten category learning in these problems, the effect was subtle—the improvement in learning was relatively small, and generally reached asymptote at a comparatively small number of particles (cf. Fig. A).</p> <p>A second contributory factor to the mixed picture—though we believe our regression analyses mitigate this—is likely the substantial correlations between model parameters. In formulating the category learning model, we included the possibility that various of its parameters—not just number of particles—would show variation when fit to behavior. As our results made clear, the parameters showed substantial correlations, making the job of disentangling their effects more difficult. In the SHJ case, the best‐fitting model had a separate noise/guessing‐rate ϵ for each participant, with other parameters fixed across participants; both correlation and regression indicated a negative relationship between WMC and ϵ, suggesting that higher WMC participants were less "noisy" in their choices. When we allowed all parameters to vary between individuals (the second best fitting model), we saw that ϵ and shape <emph>a</emph><subs>0</subs> were significantly positively correlated, as one might anticipate—recall that a higher <emph>a</emph><subs>0</subs> leads to more tolerance of category structures with mixed labels, which would lead to more errors. Furthermore, <emph>a</emph><subs>0</subs> was negatively correlated with the number of particles <emph>L</emph>, which is also expected, since an increasing number of particles tends to reduce the number of errors. However, in this case we found no significant correlation between particles <emph>L</emph> and ϵ, which we might have expected given their tendencies to decrease and increase errors, respectively. In the knowledge restructuring experiment, the best model allowed all parameters to vary between participants, and here we did indeed find that <emph>L</emph> and ϵ were negatively correlated. The fact that the bias parameter <emph>b</emph> was respectively positively and negatively correlated with <emph>L</emph> and ϵ also makes sense, since a lower bias would tend to generate more classification errors. Although we haven't demonstrated it here, we expect that <emph>b</emph> and <emph>L</emph> would also interact in strategy‐switching, in addition to the category learning phase, since a higher bias may require a larger number of particles to ensure that switching occurs reliably.</p> <p>Clearly, our model has multiple sources of variability, or "noise," that trade off in ways that unfortunately make it difficult to draw strong conclusions from our model‐fitting results. Of course, this is not an uncommon scenario, and the challenge of apportioning behavioral variability to different possible sources is a general one. In relation to the latter, it is interesting that we found in all cases that a model with a relatively low number of particles (i.e., in the range of 0–100) fit better than a model with a large number (<reflink idref="bib10" id="ref161">10</reflink>,000) of particles. The purpose of the latter was to approximate exact inference more closely, thereby providing a comparison in which noise in the inference process (as opposed to other sources of noise, such as in the choice process) was minimized. The finding therefore lends some support to the idea that inference noise plays a role in accounting for variability in participants' behavior (e.g., Wyart &amp; Koechlin, [<reflink idref="bib95" id="ref162">95</reflink>]). However, we would caution against drawing too strong a conclusion here—though we did not see much evidence of floor/ceiling effects in our fitting results, a more decisive comparison would involve an expanded range of parameters (e.g., considering ϵ on the full range [0, 1]).</p> <p>We also found that a probability‐matching choice rule always fit the data better than a maximum‐probability choice rule. Probability‐matching behavior has previously been reported in the categorization literature (Estes et al., [<reflink idref="bib34" id="ref163">34</reflink>]; Gluck &amp; Bower, [<reflink idref="bib41" id="ref164">41</reflink>]), so this result is perhaps not surprising, even if it is strictly suboptimal in this setting. However, in the context of our model, it is difficult to assign responsibility for probability matching to the inference or choice mechanism, since probability matching could conceivably arise from either separately, or both together. Indeed, since an inference mechanism based on sampling, such as the one we have described, would naturally tend to probability matching under a limited number of samples (cf. Vul, Goodman, Griffiths, &amp; Tenenbaum, [<reflink idref="bib92" id="ref165">92</reflink>]), the addition of a probability‐matching choice process makes disentangling these separate sources of variability particularly challenging.</p> <p>Why, then, did we include noise in the choice process at all? Here, the motivation was simply to improve model fit—at least some participants' behavior was more variable than even a severely resource‐constrained particle filter (i.e., a single particle). The guessing rate therefore primarily represents our ignorance about variability arising from sources distinct from sample‐based inference (e.g., attentional lapses). It is interesting that in both experiments the best‐fitting model had guessing rates that were negatively correlated with WMC. This is consistent with observations that an increase in WMC load is accompanied with what look like random responses (e.g., Adam, Vogel, &amp; Awh, [<reflink idref="bib1" id="ref166">1</reflink>]; Zhang &amp; Luck, [<reflink idref="bib97" id="ref167">97</reflink>]), since we would then expect individuals with lower WMC (as well as higher WMC individuals under increased memory load) to guess more often <emph>because</emph> their capacity is lower. However, given that our starting point was the operationalization of WMC in terms of number of particles <emph>L</emph>, the fact that we only found a negative correlation between <emph>L</emph> and ϵ in one of the two experiments is only partially consistent with this.</p> <p>Another limitation concerns our model's inability to handle particular attentional phenomena. In our presentation of the results of Sewell and Lewandowsky ([<reflink idref="bib82" id="ref168">82</reflink>]), we briefly highlighted that high‐WMC participants displayed significantly greater changes in context sensitivity (CS) in Session 1 but not in Session 2, where low‐WMC participants appeared to "catch up"; in our model, by contrast, there is no reason to expect the amount of CS change to vary for different sessions (compare Figs. D and C). At least some of the asymmetry in the human data is likely to arise due to attentional factors that are not included in our model. In particular, Sewell and Lewandowsky ([<reflink idref="bib82" id="ref169">82</reflink>]) noted that participants in the KP‐first condition generally found it easier to switch in Session 1 than participants in the CI‐first condition (compare the magnitude of CS change between transfer tests 1 and 2 for the two conditions in Fig. C). In their interpretation of this, Sewell and Lewandowsky appealed to dimensional relevance shifts, and specifically to evidence that it is easier to attend to a previously relevant dimension than to a previously ignored dimension (e.g., Kruschke, [<reflink idref="bib53" id="ref170">53</reflink>]). Thus, in the CI‐first condition, participants initially learn to <emph>ignore</emph> one of the dimensions (color, or "context"), since it is not involved in the CI strategy; this means that it will be harder to switch to the KP strategy, since the latter requires attending to the previously ignored dimension. In the KP‐first condition, by contrast, participants initially attend to all stimulus dimensions, so do not have to learn to attend to a previously ignored dimension. A modest augmentation of the current model with a prior that incorporates the assumption that only a subset of stimulus dimensions may be relevant to classification (i.e., a sparsity assumption) would conceivably address the asymmetry between KP‐first and CI‐first conditions, but presumably not the fact that low‐WMC participants appear to catch up with high‐WMC participants in Session 2.</p> <p>Finally, we should of course consider the possibility that the class of Bayesian CART models considered here may not be most appropriate. The principal motivation for focusing on this class was the intuition that a tree‐based representation, following a sequence of axis‐aligned partitions of stimulus space, was a natural fit to the task domains. It would certainly be of interest, however, to compare the performance of the current approach with those of alternative category representations (e.g., Anderson, [<reflink idref="bib3" id="ref171">3</reflink>]; Ashby et al., [<reflink idref="bib4" id="ref172">4</reflink>]; Love et al., [<reflink idref="bib59" id="ref173">59</reflink>]).</p> <hd id="AN0140459439-31">WMC and search efficiency</hd> <p>Category learning in the model proceeded quicker with more samples due to what we might refer to as increased <emph>search efficiency</emph>. Category structures that represent "good" solutions to the category learning problem were those with high posterior probability, and so the inference problem could be thought of in terms of search for such category structures in the hypothesis space (cf. Fig.). The more resources available to search this space—the more samples—then the more likely it is that (a) a good solution is discovered at all, and (b) a good solution is discovered quickly. In our simulations, we found that the marginal benefit to learning rate of increasing the number of samples was rapidly diminishing (cf. Fig. A), though we expect the point at which this occurs to depend on both the complexity of the problem and the precise details of the inference algorithm.</p> <p>In more psychological terms, the implication is that the greater the number of hypotheses that one can entertain and manipulate within working memory, the more likely that one will quickly discover good solutions. The idea of exploring a space of solutions is of course well‐established in psychology, where problem‐solving has long been cast in such terms (Newell &amp; Simon, [<reflink idref="bib65" id="ref174">65</reflink>]; Simon, [<reflink idref="bib85" id="ref175">85</reflink>]). There, however, the search problem is conventionally defined in terms of finding a path from an initial state to an explicit goal state while minimizing the path cost. This is rather different from the search in the present case, which is best described in terms of simple stochastic hill‐climbing in the absence of an explicit goal representation or, indeed, a path cost. Nevertheless, the idea that one may have greater or lesser resources with which to search may be fruitful in considering the link between WMC and problem solving more generally (Hambrick &amp; Engle, [<reflink idref="bib47" id="ref176">47</reflink>]). Other stochastic sampling algorithms that have been applied to finding action sequences in large search spaces, such as Monte Carlo tree search (Coulom, [<reflink idref="bib23" id="ref177">23</reflink>]; Gelly &amp; Silver, [<reflink idref="bib36" id="ref178">36</reflink>]), may also be a natural source of inspiration in such settings.</p> <p>Interestingly, we also found some evidence in the data of Lewandowsky ([<reflink idref="bib56" id="ref179">56</reflink>]) that WMC may interact with extent of Type II advantage in the SHJ tasks. Additional analysis of the experimental data was prompted by the observation in our model that the degree of Type II advantage appeared to be modulated by the number of particles (cf. Fig. A). This is consistent with the recent suggestion, in the context of category learning in older adults, that Type II advantage is modulated by WMC (Rabi &amp; Minda, [<reflink idref="bib73" id="ref180">73</reflink>]), though our present model does not speak to the observation that relative performance on Type II and Type IV problems may sometimes reverse (e.g., in older adults—see Badham, Sanborn, &amp; Maylor, [<reflink idref="bib11" id="ref181">11</reflink>]; Rabi &amp; Minda, [<reflink idref="bib73" id="ref182">73</reflink>]).</p> <hd id="AN0140459439-32">WMC and flexibility</hd> <p>A greater number of samples led not only to faster category learning, but also to an improved ability to switch between categorization strategies. This was due to an increase in what we might call <emph>representational adequacy</emph>. That is, with a greater number of samples, the full posterior distribution over category structures was more accurately represented, encompassing category structures that were assigned lower probability. By representing this greater plurality of category structures, the model could easily express alternative hypotheses when instructed to switch strategy, as operationalized by a reweighting of the current sample/hypothesis set (cf. Fig.).</p> <p>Again, in more psychological terms, the obvious interpretation is that the greater one's ability to entertain a variety of hypotheses, the more flexible one will be. There is evidence that individuals with higher WMC are better at solving the so‐called insight problems, and this may be because such problems are exactly those that require keeping in mind several different possibilities (Gilhooly &amp; Fioratou, [<reflink idref="bib39" id="ref183">39</reflink>]; Murray &amp; Byrne, [<reflink idref="bib64" id="ref184">64</reflink>]). Indeed, insight problems typically involve inducing task representations in participants which are not conducive to solving the problem, and so require "restructuring" of the initial task representation (Ohlsson, [<reflink idref="bib71" id="ref185">71</reflink>]; Weisberg, [<reflink idref="bib94" id="ref186">94</reflink>]).</p> <hd id="AN0140459439-33">Related work</hd> <p>The current study is framed by a number of related strands of research. Most pertinently, Lewandowsky and colleagues have themselves previously addressed the experimental results discussed here, though using a rather different modeling approach. Lewandowsky ([<reflink idref="bib56" id="ref187">56</reflink>]) found that individual differences in category learning performance could be captured by varying only the learning rate of a particular category learning model (ALCOVE; Kruschke, [<reflink idref="bib52" id="ref188">52</reflink>]), but did not establish a rationale for why WMC should be related to this parameter. Sewell and Lewandowsky ([<reflink idref="bib81" id="ref189">81</reflink>]) found that while a "single‐module" model such as ALCOVE failed to capture the general ability to fluidly switch between categorization strategies, a "multiple‐module" model, such as ATRIUM (Erickson &amp; Kruschke, [<reflink idref="bib32" id="ref190">32</reflink>])—which is able to learn more than one mapping between stimuli and category labels—could do so. However, a mechanism by which such recoordination could take place was not proposed, nor was the issue of why WMC should be related to this ability addressed. In the current work, we provide a model able to capture both experimental results using a single mechanism (i.e., variation in the number of samples), propose a simple mechanism for how recoordination could occur (i.e., importance reweighting), and offer rationales for why WMC may be associated with faster learning (search efficiency) and flexibility (representational adequacy).</p> <p>Levy et al. ([<reflink idref="bib55" id="ref191">55</reflink>]) directly anticipate our suggestion that the number of samples used for inference may be equated with WMC in their exploration of "garden path" effects in sentence processing. Briefly, garden path sentences (e.g., "The old man the boat.") are grammatical sentences that people typically fail to parse correctly, at least at first, due to early parts of the sentence tending to promote one (incorrect) interpretation over another. This initial interpretation then leads to subsequent difficulties of comprehension. Levy et al. ([<reflink idref="bib55" id="ref192">55</reflink>]) suggested that difficulties in parsing such sentences correctly—and in particular, the probability of successfully re‐parsing the sentence in light of disambiguating information arriving late in the sentence—may be explained by constraints on the resources (i.e., number of samples) available for incremental parsing. They showed that a particle filter model for performing online inference could reproduce these phenomena, with variation of the number of particles altering the strength of the effects. In particular, as the number of particles decreased, the probability that the correct interpretation of the sentence was not represented in the ensemble—leading to parse failure—increased. This is exactly analogous to the mechanism suggested to account for category switching performance in the current work: a lower number of particles makes it less likely that the alternative strategy is represented, meaning that the probability of being able to switch is decreased. However, the current work goes beyond Levy et al. ([<reflink idref="bib55" id="ref193">55</reflink>]) both in expanding the range of phenomena explained (i.e., both switching <emph>and</emph> learning effects) and in actually measuring correlations between best‐fit model parameters and WMC scores.</p> <p>The HyGene model of Dougherty and colleagues (Dougherty, Thomas, &amp; Lange, [<reflink idref="bib30" id="ref194">30</reflink>]; Thomas, Dougherty, Sprenger, &amp; Harbison, [<reflink idref="bib90" id="ref195">90</reflink>]) is also closely related to the current work. HyGene provides a general framework for diagnostic inference, incorporating processes by which hypotheses may be generated and maintained in working memory. This includes the assumption that working memory processes constrain the number of hypotheses that one can actively maintain, though to the best of our knowledge this framework has not been applied to the domain of category learning that we consider here.</p> <p>Finally, a number of previous models have considered the category learning problem in Bayesian terms (Anderson, [<reflink idref="bib3" id="ref196">3</reflink>]; Goodman et al., [<reflink idref="bib42" id="ref197">42</reflink>]; Sanborn et al., [<reflink idref="bib78" id="ref198">78</reflink>], [<reflink idref="bib79" id="ref199">79</reflink>]). Notably, both Sanborn et al. ([<reflink idref="bib78" id="ref200">78</reflink>], [<reflink idref="bib79" id="ref201">79</reflink>]) and Goodman et al. ([<reflink idref="bib42" id="ref202">42</reflink>]), despite considering rather different category representations, considered sample‐based inference to be a particularly good candidate as a psychological mechanism for approximating Bayesian inference. For example, Sanborn et al. found that they were able to replicate a wide range of category learning effects by fitting relatively few samples to experimental data, though individual differences were not explored in that work. Our use of a representation based on classification and regression trees (CART) was primarily driven by pragmatic reasons, in particular what seemed most natural for the tasks concerned, rather than a theoretical commitment to a particular way of representing categories. We expect similar results to be obtained with alternative category representations, such as those used in the Rational Model of Categorization (Anderson, [<reflink idref="bib2" id="ref203">2</reflink>]; Sanborn et al., [<reflink idref="bib79" id="ref204">79</reflink>]) and Rational Rules (Goodman et al., [<reflink idref="bib42" id="ref205">42</reflink>]), though this remains to be shown.</p> <hd id="AN0140459439-34">Future directions</hd> <p>The current work suggests a number of avenues for future investigation. One is to further explore the relative contributions of different components of the inference process. For example, search in the model effectively relies on two processes. The first is resampling, in which particles with lower probability are discarded and particles with higher probability are copied. Intuitively, this should be beneficial for learning since search is then focused on more "promising" (i.e., high probability) regions of hypothesis space. The second process is the proposal and acceptance/rejection of new hypotheses via MCMC moves, leading to local hill‐climbing in probability space. A more detailed understanding of how these processes interact, and how they may relate to various psychological phenomena, would be of interest.</p> <p>Similarly, one could consider alternative conceptualizations of the process by which participants switch between different categorization strategies. We implemented strategy‐switching as a simple reweighting operation on particles according to a new target distribution. One consequence of this modeling choice is that it may be impossible—at least for the initial time step—to switch to a new strategy if the corresponding region of hypothesis space is not represented. Though we found some hints of bimodality in the human data, the prospect of such "catastrophic failure" may not seem entirely realistic, so one could imagine exploring modifications such as allowing additional propose‐accept/reject steps during this phase.</p> <p>More generally, it is likely that there is a trade‐off between the sophistication of the processes by which individual hypotheses are maintained and manipulated, and the number of such hypotheses that one would need to support. In other words, one could presumably replace a larger number of relatively "dumb" particles/hypotheses with a smaller number of comparatively "smart" particles/hypotheses. Indeed, it has recently been suggested that, at least when considering more global hypotheses about the world where the hypothesis space becomes particularly complex, only one hypothesis would plausibly be represented (Bramley, Dayan, Griffiths, &amp; Lagnado, [<reflink idref="bib13" id="ref206">13</reflink>]). How to negotiate this spectrum of possibilities is a pressing challenge.</p> <p>Clearly, future work should also test whether the current modeling approach can be applied to other category learning tasks and beyond. As mentioned in the Introduction, it has been suggested that category learning tasks which can be solved with relatively simple, verbalizable rules ("rule‐based" tasks) are especially reliant on working memory, while tasks with solutions that generally defy description in terms of simple rules ("information‐integration" tasks) are not (Ashby &amp; Maddox, [<reflink idref="bib5" id="ref207">5</reflink>], [<reflink idref="bib6" id="ref208">6</reflink>]; Ashby &amp; O'Brien, [<reflink idref="bib7" id="ref209">7</reflink>]). However, recent results suggest rather that working memory is equally involved in these different types of task (Craig &amp; Lewandowsky, [<reflink idref="bib25" id="ref210">25</reflink>]; Lewandowsky et al., [<reflink idref="bib58" id="ref211">58</reflink>]). An obvious first step would therefore be to assess whether the current approach can be applied to tasks that are more clearly of the information‐integration type.</p> <p>A broader challenge for rational process models is to find constraints that will help determine more precisely the algorithms that underpin cognition. In the present work, we followed previous suggestions that inference algorithms based on Monte Carlo sampling are promising, but this only weakly constrains the variety of models under consideration. Determining the signatures of particular modeling choices within this larger class, and how these may succeed or fail in matching features of human cognition and behavior, is a substantial task for future research.</p> <hd id="AN0140459439-35">Acknowledgments</hd> <p>This work was supported by an EPSRC doctoral training award (KL), ESRC grant ES/K004948/1 (AS), EPSRC grant EP/I032622/1 (DL), and by the Royal Society (SL). We are extremely grateful to Rick Cooper, David Sewell, and Maarten Speekenbrink for their comments on previous versions of the manuscript.</p> <p>GRAPH: Fig. S1. Interaction of Type II advantage with working memory capacity. Average learning curves (±1 SE) for Types II and IV in the experiment of Lewandowsky (2011) for (A) all participants; (B) participants with lower‐median WMC scores; and (C) participants with upper‐median WMC scores. Only the high WMC participants show a Type II advantage (see main text for statistics).</p> <p>GRAPH: Fig. S2. Participants' context‐sensitivity changes and switch probabilities. (A) Distribution of (absolute) changes in context sensitivity (ΔCS; pooling over both test sessions) for low (left) and high (right) WMC participants. (B) The probability of making a successful switch of categorization strategy goes up with increasing WMC. Mean probabilities of a successful switch were, respectively,.64,.77,.89, and.96 for participants with WMC scores in the lower quartile (↓25), lower median (↓50), upper median (↑50), and upper quartile (↑25) of the experimental population. These scores are superimposed, for comparison, on the probability of switching as a function of the number of particles obtained from simulations (cf. Fig. 8B). As in the simulation results, a successful switch is defined as a change in context sensitivity between test sessions, ΔCS, that "crosses" a score of 0.5.</p> <ref id="AN0140459439-36"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref69" type="bt">1</bibl> <bibtext> In both experiments, <emph>p</emph>= 3, but we use the more general notation for presentation purposes.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref17" type="bt">2</bibl> <bibtext> Of course, in reality, bar position and height were much more restricted than indicated—we mean only to emphasize by the use of <ephtml> &lt;math altimg="urn:x-wiley:03640213:media:cogs12805:cogs12805-math-0073" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mi mathvariant="double-struck"&gt;R&lt;/mi&gt;&lt;/math&gt; </ephtml> that these are continuous variables.</bibtext> </blist> </ref> <ref id="AN0140459439-37"> <title> References </title> <blist> <bibtext> Adam, K., Vogel, E., &amp; Awh, E. (2017). Clear evidence for item limits in working memory. Cognitive Psychology, 97, 79 – 97.</bibtext> </blist> <blist> <bibtext> Anderson, J. (1990). The adaptive character of thought. Hillsdale, NJ : Erlbaum.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref90" type="bt">3</bibl> <bibtext> Anderson, J. (1991). The adaptive nature of human categorization. Psychological Review, 98, 409 – 429.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref91" type="bt">4</bibl> <bibtext> Ashby, F., Alfonso‐Reese, L., Turken, A., &amp; Waldron, E. (1998). A neuropsychological theory of multiple systems in category learning. Psychological Review, 105, 442 – 481.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref60" type="bt">5</bibl> <bibtext> Ashby, F., &amp; Maddox, W. (2005). Human category learning. Annual Review of Psychology, 56, 149 – 178.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref61" type="bt">6</bibl> <bibtext> Ashby, F., &amp; Maddox, W. (2011). Human category learning 2.0. Annals of the New York Academy of Sciences, 1224, 137 – 161.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref62" type="bt">7</bibl> <bibtext> Ashby, F., &amp; O'Brien, J. (2005). Category learning and multiple memory systems. Trends in Cognitive Sciences, 9 (2), 83 – 89.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref24" type="bt">8</bibl> <bibtext> Baddeley, A. (1992). Working memory. Science, 255, 556 – 559.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref25" type="bt">9</bibl> <bibtext> Baddeley, A., &amp; Hitch, G. (1974). Working memory. Psychology of Learning and Motivation, 8, 47 – 89.</bibtext> </blist> <blist> <bibtext> Baddeley, A., Thompson, N., &amp; Buchanan, M. (1975). Word length and the structure of short term memory. Journal of Verbal Learning and Verbal Behavior, 14, 575 – 589.</bibtext> </blist> <blist> <bibtext> Badham, S., Sanborn, A., &amp; Maylor, E. (2017). Deficits in category learning in older adults: Rule‐based versus clustering accounts. Psychology and Aging, 32 (5), 473 – 488.</bibtext> </blist> <blist> <bibtext> Bernardo, J., &amp; Smith, A. (1994). Bayesian theory. Chichester, UK : Wiley.</bibtext> </blist> <blist> <bibtext> Bramley, N., Dayan, P., Griffiths, T., &amp; Lagnado, D. (2017). Formalizing Neurath's ship: Approximate algorithms for online causal learning. Psychological Review, 124 (3), 301 – 338.</bibtext> </blist> <blist> <bibtext> Breiman, L., Friedman, J., Olshen, R., &amp; Stone, C. (1984). Classification and regression trees. Belmont, CA : Wadsworth.</bibtext> </blist> <blist> <bibtext> Brown, S., &amp; Steyvers, M. (2009). Detecting and predicting changes. Cognitive Psychology, 58 (58), 49 – 67.</bibtext> </blist> <blist> <bibtext> Bruner, J., Goodnow, J., &amp; Austin, G. (1956). A study of thinking. New York : Wiley.</bibtext> </blist> <blist> <bibtext> Busemeyer, J. R. (1985). Decision making under uncertainty: A comparison of simple scalability, fixed‐sample, and sequential‐sampling models. Journal of Experimental Psychology: Learning, Memory, and Cognition, 11 (3), 538 – 564.</bibtext> </blist> <blist> <bibtext> Chater, N., &amp; Oaksford, M. (Eds.) (2008). The probabilistic mind: Prospects for Bayesian cognitive science. Oxford, UK : Oxford University Press.</bibtext> </blist> <blist> <bibtext> Chipman, H., George, E., &amp; McCulloch, R. (1998). Bayesian CART model search. Journal of the American Statistical Association, 93 (443), 935 – 948.</bibtext> </blist> <blist> <bibtext> Chopin, N. (2002). A sequential particle filter method for static models. Biometrika, 89 (3), 539 – 552.</bibtext> </blist> <blist> <bibtext> Conway, A., Jarrold, C., Kane, M., Miyake, A., &amp; Towse, J. (Eds.) (2007). Variation in working memory. New York : Oxford University Press.</bibtext> </blist> <blist> <bibtext> Conway, A., Kane, M., &amp; Engle, R. (2003). Working memory capacity and its relation to general intelligence. Trends in Cognitive Sciences, 7 (12), 547 – 552.</bibtext> </blist> <blist> <bibtext> Coulom, R. (2006). Efficient selectivity and backup operators in Monte‐Carlo tree search. In H. van der Herik, P. Ciancarini, &amp; H. Donkers (Eds.), 5th International conference on computer and games (pp. 72 – 83). Berlin : Springer.</bibtext> </blist> <blist> <bibtext> Cowan, N. (2001). The magical number 4 in short‐term memory: A reconsideration of mental storage capacity. Behavioral and Brain Sciences, 24 (1), 87 – 114.</bibtext> </blist> <blist> <bibtext> Craig, S., &amp; Lewandowsky, S. (2012). Whichever way you choose to categorize, working memory helps you learn. The Quarterly Journal of Experimental Psychology, 65 (3), 439 – 464.</bibtext> </blist> <blist> <bibtext> Daneman, M., &amp; Carpenter, P. (1980). Individual differences in working memory and reading. Journal of Verbal Learning and Verbal Behavior, 19 (4), 450 – 466.</bibtext> </blist> <blist> <bibtext> Daw, N., &amp; Courville, A. (2008). The rat as particle filter. In J. Platt, D. Koller, Y. Singer, &amp; S. Roweis (Eds.), Advances in neural information processing systems 20 (pp. 369 – 376). Cambridge, MA : MIT Press.</bibtext> </blist> <blist> <bibtext> Daw, N., Courville, A., &amp; Dayan, P. (2008). Semi‐rational models of conditioning: The case of trial order. In N. Chater &amp; M. Oaksford (Eds.), The probabilistic mind: Prospects for Bayesian cognitive science (pp. 427 – 448). New York : Oxford University Press.</bibtext> </blist> <blist> <bibtext> Doucet, A., de Freitas, N., &amp; Gordon, N. (Eds.) (2001). Sequential Monte Carlo methods in practice. New York : Springer.</bibtext> </blist> <blist> <bibtext> Dougherty, M., Thomas, R., &amp; Lange, N. (2010). Toward an integrative theory of hypothesis generation, probability judgment, and hypothesis testing. In B. Ross (Ed.), The psychology of learning and motivation (vol. 52, pp. 299 – 342). Burlington, VT : Academic Press.</bibtext> </blist> <blist> <bibtext> Doya, K., Ishii, S., Pouget, A., &amp; Rao, R. (Eds.) (2007). Bayesian brain: Probabilistic approaches to neural coding. Cambridge, MA : MIT Press.</bibtext> </blist> <blist> <bibtext> Erickson, M., &amp; Kruschke, J. (1998). Rules and exemplars in category learning. Journal of Experimental Psychology: General, 127 (2), 107 – 140.</bibtext> </blist> <blist> <bibtext> Estes, W. K. (1950). Toward a statistical theory of learning. Psychological Review, 57 (2), 94 – 107.</bibtext> </blist> <blist> <bibtext> Estes, W. K., Campbell, J. A., Hatsopoulos, N., &amp; Hurwitz, J. B. (1989). Base‐rate effects in category learning: A comparison of parallel network and memory storage‐retrieval models. Journal of Experimental Psychology: Learning, Memory, and Cognition, 15 (4), 556 – 571.</bibtext> </blist> <blist> <bibtext> Gelfand, A., &amp; Smith, A. (1990). Sampling‐based approaches to calculating marginal densities. Journal of the American Statistical Association, 85 (410), 398 – 409.</bibtext> </blist> <blist> <bibtext> Gelly, S., &amp; Silver, D. (2011). Monte‐Carlo tree search and rapid action value estimation in computer Go. Artificial Intelligence, 175, 1856 – 1875.</bibtext> </blist> <blist> <bibtext> Gelman, A., Carlin, J., Stern, H., &amp; Rubin, D. (2004). Bayesian data analysis (2nd ed.). Boca Raton, FL : Chapman &amp; Hall/CRC.</bibtext> </blist> <blist> <bibtext> Gigerenzer, G., &amp; Goldstein, D. (1996). Reasoning the fast and frugal way: Models of bounded rationality. Psychological Review, 103 (4), 650 – 669.</bibtext> </blist> <blist> <bibtext> Gilhooly, K., &amp; Fioratou, E. (2009). Executive functions in insight versus noninsight problem solving: An individual differences approach. Thinking &amp; Reasoning, 15 (4), 355 – 376.</bibtext> </blist> <blist> <bibtext> Gilks, W., &amp; Berzuini, C. (2001). Following a moving target—Monte Carlo inference for dynamic Bayesian models. Journal of the Royal Statistical Society Series B, 63, 127 – 146.</bibtext> </blist> <blist> <bibtext> Gluck, M., &amp; Bower, G. (1988). From conditioning to category learning: An adaptive network model. Journal of Experimental Psychology: General, 117 (3), 227 – 247.</bibtext> </blist> <blist> <bibtext> Goodman, N., Tenenbaum, J., Feldman, J., &amp; Griffiths, T. (2008). A rational analysis of rule‐based concept learning. Cognitive Science, 32 (1), 108 – 154.</bibtext> </blist> <blist> <bibtext> Gordon, N., Salmond, D., &amp; Smith, A. (1993). Novel approach to nonlinear/non‐Gaussian Bayesian state estimation. IEE Proceedings F Radar and Signal Processing, 140 (2), 107 – 113.</bibtext> </blist> <blist> <bibtext> Griffiths, T., &amp; Tenenbaum, J. (2005). Structure and strength in causal induction. Cognitive Psychology, 51, 334 – 384.</bibtext> </blist> <blist> <bibtext> Griffiths, T., &amp; Tenenbaum, J. (2006). Optimal predictions in everyday cognition. Psychological Science, 17, 180 – 226.</bibtext> </blist> <blist> <bibtext> Griffiths, T., Vul, E., &amp; Sanborn, A. (2012). Bridging levels of analysis for probabilistic models of cognition. Current Directions in Psychological Science, 21 (4), 263 – 268.</bibtext> </blist> <blist> <bibtext> Hambrick, D., &amp; Engle, R. (2003). The role of working memory in problem solving. In J. Davidson &amp; R. Sternberg (Eds.), The psychology of problem solving (pp. 176 – 206). New York : Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Just, M., &amp; Carpenter, P. (1992). A capacity theory of comprehension: Individual differences in working memory. Psychological Review, 99, 122 – 149.</bibtext> </blist> <blist> <bibtext> Kahneman, D. (2003). A perspective on judgment and choice: Mapping bounded rationality. American Psychologist, 58 (9), 697 – 720.</bibtext> </blist> <blist> <bibtext> Kass, R., &amp; Raftery, A. (1995). Bayes factors. Journal of the American Statistical Association, 9 (430), 773 – 795.</bibtext> </blist> <blist> <bibtext> Körding, K., &amp; Wolpert, D. (2004). Bayesian integration in sensorimotor learning. Nature, 427 (6971), 244 – 247.</bibtext> </blist> <blist> <bibtext> Kruschke, J. (1992). ALCOVE: An exemplar‐based connectionist model of category learning. Psychological Review, 99, 22 – 44.</bibtext> </blist> <blist> <bibtext> Kruschke, J. (1996). Dimensional relevance shifts in category learning. Connection Science, 8 (2), 225 – 247.</bibtext> </blist> <blist> <bibtext> Kurtz, K., Levering, K., Stanton, R., Romero, J., &amp; Morris, S. (2013). Human learning of elemental category structures: Revising the classic result of Shepard, Hovland, and Jenkins (1961). Journal of Experimental Psychology: Learning, Memory, and Cognition, 39 (2), 552 – 572.</bibtext> </blist> <blist> <bibtext> Levy, R., Reali, F., &amp; Griffiths, T. (2008). Modeling the effects of memory on human online sentence processing with particle filters. In D. Koller, D. Schuurmans, Y. Bengio, &amp; L. Bottou (Eds.), Advances in neural information processing systems 21 (pp. 937 – 944). Cambridge, MA : The MIT Press.</bibtext> </blist> <blist> <bibtext> Lewandowsky, S. (2011). Working memory capacity and categorization: Individual differences and modeling. Journal of Experimental Psychology: Learning, Memory, and Cognition, 37 (3), 720 – 738.</bibtext> </blist> <blist> <bibtext> Lewandowsky, S., Griffiths, T., &amp; Kalish, M. (2009). The wisdom of individuals: Exploring people's knowledge of everyday events using iterated learning. Cognitive Science, 33, 969 – 998.</bibtext> </blist> <blist> <bibtext> Lewandowsky, S., Yang, L.‐X., Newell, B., &amp; Kalish, M. (2012). Working memory does not dissociate between different perceptual categorization tasks. Journal of Experimental Psychology: Learning, Memory, and Cognition, 38 (4), 881 – 904.</bibtext> </blist> <blist> <bibtext> Love, B., Medin, D., &amp; Gureckis, T. (2004). Sustain: A network model of category learning. Psychological Review, 111 (2), 309 – 332.</bibtext> </blist> <blist> <bibtext> Ma, W., Husain, M., &amp; Bays, P. (2014). Changing concepts of working memory. Nature Neuroscience, 17 (3), 347 – 356.</bibtext> </blist> <blist> <bibtext> MacKay, D. C. (2003). Information theory, inference, and learning algorithms. Cambridge, UK : Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Medin, D., &amp; Schaffer, M. (1978). Context theory of classification learning. Psychological Review, 85 (3), 207 – 238.</bibtext> </blist> <blist> <bibtext> Miller, G. A. (1956). The magical number seven, plus or minus two: Some limits on our capacity for processing information. Psychological Review, 63 (2), 81 – 97.</bibtext> </blist> <blist> <bibtext> Murray, M. A., &amp; Byrne, R. M. (2005). Attention and working memory in insight problem solving. In B. Bara, L. Barsalou, &amp; M. Bucciarelli (Eds.), Proceedings of the XXVII Annual Conference of the Cognitive Science Society (pp. 1571 – 1575). Mahwah, NJ : Lawrence Erlbaum.</bibtext> </blist> <blist> <bibtext> Newell, A., &amp; Simon, H. (1972). Human problem solving. Englewood Cliffs, NJ : Prentice‐Hall.</bibtext> </blist> <blist> <bibtext> Nosofsky, R. (1986). Attention, similarity, and the identification‐categorisation relationship. Journal of Experimental Psychology: General, 115 (1), 39 – 57.</bibtext> </blist> <blist> <bibtext> Nosofsky, R., Gluck, M., Palmeri, T., McKinley, S., &amp; Glauthier, P. (1994). Comparing models of rule‐based classification learning: A replication and extension of Shepard, Hovland, and Jenkins (1961). Memory and Cognition, 22 (3), 352 – 369.</bibtext> </blist> <blist> <bibtext> Nosofsky, R., Palmeri, T., &amp; McKinley, S. (1994). Rule‐plus‐exception model of classification learning. Psychological Review, 101 (1), 53 – 79.</bibtext> </blist> <blist> <bibtext> Oberauer, K., Farrell, S., Jarrold, C., &amp; Lewandowsky, S. (2016). What limits working memory capacity? Psychological Bulletin, 142 (7), 758 – 799.</bibtext> </blist> <blist> <bibtext> Oberauer, K., &amp; Kliegl, R. (2006). A formal model of capacity limits in working memory. Journal of Memory and Language, 55, 601 – 626.</bibtext> </blist> <blist> <bibtext> Ohlsson, S. (1992). Information processing explanations of insight and related phenomena. In M. Keane &amp; K. Gilhooly (Eds.), Advances in the psychology of thinking (pp. 1 – 44). London : Harvester‐Wheatsheaf.</bibtext> </blist> <blist> <bibtext> Posner, M., &amp; Keele, S. (1968). On the genesis of abstract ideas. Journal of Experimental Psychology, 77, 353 – 363.</bibtext> </blist> <blist> <bibtext> Rabi, R., &amp; Minda, J. (2016). Category learning in older adulthood: A study of the Shepard, Hovland, and Jenkins (1961) tasks. Psychology and Aging, 31, 185 – 197.</bibtext> </blist> <blist> <bibtext> Restle, F. (1962). The selection of strategies in cue learning. Psychological Review, 69 (4), 329 – 343.</bibtext> </blist> <blist> <bibtext> Robert, C., &amp; Casella, G. (2004). Monte Carlo statistical methods. New York : Springer.</bibtext> </blist> <blist> <bibtext> Rosch, E. (1973). Natural categories. Cognitive Psychology, 4, 328 – 350.</bibtext> </blist> <blist> <bibtext> Sanborn, A., &amp; Chater, N. (2016). Bayesian brains without probabilities. Trends in Cognitive Sciences, 20 (12), 883 – 893.</bibtext> </blist> <blist> <bibtext> Sanborn, A., Griffiths, T., &amp; Navarro, D. (2006). A more rational model of categorization. In R. Sun &amp; N. Miyake (Eds.), Proceedings of the 28th Annual Conference of the Cognitive Science Society (pp. 726 – 731). Mahwah, NJ : Erlbaum.</bibtext> </blist> <blist> <bibtext> Sanborn, A., Navarro, D., &amp; Griffiths, T. (2010). Rational approximations to rational models: Alternative algorithms for category learning. Psychological Review, 117 (4), 1144 – 1167.</bibtext> </blist> <blist> <bibtext> Schwarz, G. (1978). Estimating the dimension of a model. The Annals of Statistics, 6 (2), 461 – 464.</bibtext> </blist> <blist> <bibtext> Sewell, D., &amp; Lewandowsky, S. (2011). Restructuring partitioned knowledge: The role of recoordination in category learning. Cognitive Psychology, 62, 81 – 122.</bibtext> </blist> <blist> <bibtext> Sewell, D., &amp; Lewandowsky, S. (2012). Attention and working memory capacity: Insights from blocking, highlighting, and knowledge restructuring. Journal of Experimental Psychology: General, 141 (3), 444 – 469.</bibtext> </blist> <blist> <bibtext> Shepard, R. N., Hovland, C. I., &amp; Jenkins, H. M. (1961). Learning and memorization of classifications. Psychological Monographs: General and Applied, 75 (13), 1 – 42.</bibtext> </blist> <blist> <bibtext> Simon, H. (1982). Models of bounded rationality (vol. 1). Cambridge, MA : MIT Press.</bibtext> </blist> <blist> <bibtext> Simon, H. (1983). Search and reasoning in problem solving. Artificial Intelligence, 21, 7 – 29.</bibtext> </blist> <blist> <bibtext> Stewart, N., Chater, N., &amp; Brown, G. (2006). Decision by sampling. Cognitive Psychology, 53, 1 – 26.</bibtext> </blist> <blist> <bibtext> Suchow, J. W., Bourgin, D. D., &amp; Griffiths, T. L. (2017). Evolution in mind: Evolutionary dynamics, cognitive processes, and bayesian inference. Trends in Cognitive Sciences, 21 (7), 522 – 530.</bibtext> </blist> <blist> <bibtext> Suchow, J. W., Fougnie, D., Brady, T. F., &amp; Alvarez, G. A. (2014). Terms of the debate on the format and structure of visual memory. Attention, Perception, &amp; Psychophysics, 76 (7), 2071 – 2079.</bibtext> </blist> <blist> <bibtext> Tenenbaum, J., Kemp, C., Griffiths, T., &amp; Goodman, N. (2011). How to grow a mind: Statistics, structure, and abstraction. Science, 331, 1279 – 1285.</bibtext> </blist> <blist> <bibtext> Thomas, R., Dougherty, M., Sprenger, A., &amp; Harbison, J. (2008). Diagnostic hypothesis generation and human judgment. Psychological Review, 115 (1), 155 – 185.</bibtext> </blist> <blist> <bibtext> Tversky, A., &amp; Kahneman, D. (1974). Judgment under uncertainty: Heuristics and biases. Science, 185 (4157), 1124 – 1131.</bibtext> </blist> <blist> <bibtext> Vul, E., Goodman, N., Griffiths, T., &amp; Tenenbaum, J. (2014). One and done? Optimal decisions from very few samples. Cognitive Science, 38 (4), 599 – 637.</bibtext> </blist> <blist> <bibtext> Vul, E., &amp; Pashler, H. (2008). Measuring the crowd within: probabilistic representations within individuals. Psychological Science, 19 (7), 645 – 647.</bibtext> </blist> <blist> <bibtext> Weisberg, R. (1995). Prolegomena to theories of insight in problem solving: A taxonomy of problems. In R. Sternberg &amp; J. Davidson (Eds.), The nature of insight (pp. 157 – 196). Cambridge, MA : MIT Press.</bibtext> </blist> <blist> <bibtext> Wyart, V., &amp; Koechlin, E. (2016). Choice variability and suboptimality in uncertain environments. Current Opinion in Behavioral Sciences, 11, 109 – 115.</bibtext> </blist> <blist> <bibtext> Yuille, A., &amp; Kersten, D. (2006). Vision as Bayesian inference: analysis by synthesis? Trends in Cognitive Sciences, 10, 301 – 308.</bibtext> </blist> <blist> <bibtext> Zhang, W., &amp; Luck, S. (2008). Discrete fixed‐resolution representations in visual working memory. Nature, 453, 233.</bibtext> </blist> </ref> <aug> <p>By Kevin Lloyd; Adam Sanborn; David Leslie and Stephan Lewandowsky</p> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib12" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib51" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib96" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib44" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib79" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib45" firstref="ref6"></nolink> <nolink nlid="nl7" bibid="bib18" firstref="ref7"></nolink> <nolink nlid="nl8" bibid="bib77" firstref="ref8"></nolink> <nolink nlid="nl9" bibid="bib89" firstref="ref9"></nolink> <nolink nlid="nl10" bibid="bib91" firstref="ref10"></nolink> <nolink nlid="nl11" bibid="bib35" firstref="ref11"></nolink> <nolink nlid="nl12" bibid="bib75" firstref="ref12"></nolink> <nolink nlid="nl13" bibid="bib31" firstref="ref14"></nolink> <nolink nlid="nl14" bibid="bib46" firstref="ref15"></nolink> <nolink nlid="nl15" bibid="bib28" firstref="ref18"></nolink> <nolink nlid="nl16" bibid="bib38" firstref="ref19"></nolink> <nolink nlid="nl17" bibid="bib49" firstref="ref20"></nolink> <nolink nlid="nl18" bibid="bib84" firstref="ref21"></nolink> <nolink nlid="nl19" bibid="bib24" firstref="ref22"></nolink> <nolink nlid="nl20" bibid="bib63" firstref="ref23"></nolink> <nolink nlid="nl21" bibid="bib26" firstref="ref26"></nolink> <nolink nlid="nl22" bibid="bib21" firstref="ref27"></nolink> <nolink nlid="nl23" bibid="bib22" firstref="ref28"></nolink> <nolink nlid="nl24" bibid="bib10" firstref="ref29"></nolink> <nolink nlid="nl25" bibid="bib48" firstref="ref30"></nolink> <nolink nlid="nl26" bibid="bib70" firstref="ref31"></nolink> <nolink nlid="nl27" bibid="bib69" firstref="ref32"></nolink> <nolink nlid="nl28" bibid="bib60" firstref="ref33"></nolink> <nolink nlid="nl29" bibid="bib88" firstref="ref34"></nolink> <nolink nlid="nl30" bibid="bib56" firstref="ref35"></nolink> <nolink nlid="nl31" bibid="bib58" firstref="ref36"></nolink> <nolink nlid="nl32" bibid="bib81" firstref="ref37"></nolink> <nolink nlid="nl33" bibid="bib82" firstref="ref38"></nolink> <nolink nlid="nl34" bibid="bib42" firstref="ref40"></nolink> <nolink nlid="nl35" bibid="bib29" firstref="ref47"></nolink> <nolink nlid="nl36" bibid="bib17" firstref="ref49"></nolink> <nolink nlid="nl37" bibid="bib33" firstref="ref50"></nolink> <nolink nlid="nl38" bibid="bib74" firstref="ref51"></nolink> <nolink nlid="nl39" bibid="bib86" firstref="ref52"></nolink> <nolink nlid="nl40" bibid="bib87" firstref="ref54"></nolink> <nolink nlid="nl41" bibid="bib93" firstref="ref55"></nolink> <nolink nlid="nl42" bibid="bib57" firstref="ref56"></nolink> <nolink nlid="nl43" bibid="bib15" firstref="ref57"></nolink> <nolink nlid="nl44" bibid="bib55" firstref="ref58"></nolink> <nolink nlid="nl45" bibid="bib83" firstref="ref67"></nolink> <nolink nlid="nl46" bibid="bib16" firstref="ref82"></nolink> <nolink nlid="nl47" bibid="bib68" firstref="ref84"></nolink> <nolink nlid="nl48" bibid="bib72" firstref="ref85"></nolink> <nolink nlid="nl49" bibid="bib76" firstref="ref86"></nolink> <nolink nlid="nl50" bibid="bib52" firstref="ref87"></nolink> <nolink nlid="nl51" bibid="bib62" firstref="ref88"></nolink> <nolink nlid="nl52" bibid="bib66" firstref="ref89"></nolink> <nolink nlid="nl53" bibid="bib59" firstref="ref92"></nolink> <nolink nlid="nl54" bibid="bib14" firstref="ref93"></nolink> <nolink nlid="nl55" bibid="bib19" firstref="ref96"></nolink> <nolink nlid="nl56" bibid="bib27" firstref="ref105"></nolink> <nolink nlid="nl57" bibid="bib78" firstref="ref107"></nolink> <nolink nlid="nl58" bibid="bib20" firstref="ref110"></nolink> <nolink nlid="nl59" bibid="bib43" firstref="ref112"></nolink> <nolink nlid="nl60" bibid="bib40" firstref="ref114"></nolink> <nolink nlid="nl61" bibid="bib37" firstref="ref116"></nolink> <nolink nlid="nl62" bibid="bib34" firstref="ref119"></nolink> <nolink nlid="nl63" bibid="bib41" firstref="ref120"></nolink> <nolink nlid="nl64" bibid="bib80" firstref="ref122"></nolink> <nolink nlid="nl65" bibid="bib50" firstref="ref123"></nolink> <nolink nlid="nl66" bibid="bib61" firstref="ref131"></nolink> <nolink nlid="nl67" bibid="bib67" firstref="ref132"></nolink> <nolink nlid="nl68" bibid="bib54" firstref="ref135"></nolink> <nolink nlid="nl69" bibid="bib11" firstref="ref138"></nolink> <nolink nlid="nl70" bibid="bib98" firstref="ref147"></nolink> <nolink nlid="nl71" bibid="bib95" firstref="ref162"></nolink> <nolink nlid="nl72" bibid="bib92" firstref="ref165"></nolink> <nolink nlid="nl73" bibid="bib97" firstref="ref167"></nolink> <nolink nlid="nl74" bibid="bib53" firstref="ref170"></nolink> <nolink nlid="nl75" bibid="bib65" firstref="ref174"></nolink> <nolink nlid="nl76" bibid="bib85" firstref="ref175"></nolink> <nolink nlid="nl77" bibid="bib47" firstref="ref176"></nolink> <nolink nlid="nl78" bibid="bib23" firstref="ref177"></nolink> <nolink nlid="nl79" bibid="bib36" firstref="ref178"></nolink> <nolink nlid="nl80" bibid="bib73" firstref="ref180"></nolink> <nolink nlid="nl81" bibid="bib39" firstref="ref183"></nolink> <nolink nlid="nl82" bibid="bib64" firstref="ref184"></nolink> <nolink nlid="nl83" bibid="bib71" firstref="ref185"></nolink> <nolink nlid="nl84" bibid="bib94" firstref="ref186"></nolink> <nolink nlid="nl85" bibid="bib32" firstref="ref190"></nolink> <nolink nlid="nl86" bibid="bib30" firstref="ref194"></nolink> <nolink nlid="nl87" bibid="bib90" firstref="ref195"></nolink> <nolink nlid="nl88" bibid="bib13" firstref="ref206"></nolink> <nolink nlid="nl89" bibid="bib25" firstref="ref210"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1237804 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Why Higher Working Memory Capacity May Help You Learn: Sampling, Search, and Degrees of Approximation – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Lloyd%2C+Kevin%22">Lloyd, Kevin</searchLink><br /><searchLink fieldCode="AR" term="%22Sanborn%2C+Adam%22">Sanborn, Adam</searchLink><br /><searchLink fieldCode="AR" term="%22Leslie%2C+David%22">Leslie, David</searchLink><br /><searchLink fieldCode="AR" term="%22Lewandowsky%2C+Stephan%22">Lewandowsky, Stephan</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Cognitive+Science%22"><i>Cognitive Science</i></searchLink>. Dec 2019 43(12). – Name: Avail Label: Availability Group: Avail Data: Wiley-Blackwell. 350 Main Street, Malden, MA 02148. Tel: 800-835-6770; Tel: 781-388-8598; Fax: 781-388-8232; e-mail: cs-journals@wiley.com; Web site: http://www.wiley.com/WileyCDA – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 43 – Name: DatePubCY Label: Publication Date Group: Date Data: 2019 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Short+Term+Memory%22">Short Term Memory</searchLink><br /><searchLink fieldCode="DE" term="%22Bayesian+Statistics%22">Bayesian Statistics</searchLink><br /><searchLink fieldCode="DE" term="%22Cognitive+Ability%22">Cognitive Ability</searchLink><br /><searchLink fieldCode="DE" term="%22Individual+Differences%22">Individual Differences</searchLink><br /><searchLink fieldCode="DE" term="%22Inferences%22">Inferences</searchLink><br /><searchLink fieldCode="DE" term="%22Correlation%22">Correlation</searchLink><br /><searchLink fieldCode="DE" term="%22Classification%22">Classification</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Strategies%22">Learning Strategies</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/cogs.12805 – Name: ISSN Label: ISSN Group: ISSN Data: 1551-6709 – Name: Abstract Label: Abstract Group: Ab Data: Algorithms for approximate Bayesian inference, such as those based on sampling (i.e., Monte Carlo methods), provide a natural source of models of how people may deal with uncertainty with limited cognitive resources. Here, we consider the idea that individual differences in working memory capacity (WMC) may be usefully modeled in terms of the number of samples, or "particles," available to perform inference. To test this idea, we focus on two recent experiments that report positive associations between WMC and two distinct aspects of categorization performance: the ability to learn novel categories, and the ability to switch between different categorization strategies ("knowledge restructuring"). In favor of the idea of modeling WMC as a number of particles, we show that a single model can reproduce both experimental results by varying the number of particles--increasing the number of particles leads to both faster category learning and improved strategy-switching. Furthermore, when we fit the model to individual participants, we found a positive association between WMC and best-fit number of particles for strategy switching. However, no association between WMC and best-fit number of particles was found for category learning. These results are discussed in the context of the general challenge of disentangling the contributions of different potential sources of behavioral variability. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2019 – Name: AN Label: Accession Number Group: ID Data: EJ1237804 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1237804 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/cogs.12805 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 43 Subjects: – SubjectFull: Short Term Memory Type: general – SubjectFull: Bayesian Statistics Type: general – SubjectFull: Cognitive Ability Type: general – SubjectFull: Individual Differences Type: general – SubjectFull: Inferences Type: general – SubjectFull: Correlation Type: general – SubjectFull: Classification Type: general – SubjectFull: Learning Strategies Type: general – SubjectFull: Models Type: general – SubjectFull: Learning Processes Type: general Titles: – TitleFull: Why Higher Working Memory Capacity May Help You Learn: Sampling, Search, and Degrees of Approximation Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Lloyd, Kevin – PersonEntity: Name: NameFull: Sanborn, Adam – PersonEntity: Name: NameFull: Leslie, David – PersonEntity: Name: NameFull: Lewandowsky, Stephan IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 12 Type: published Y: 2019 Identifiers: – Type: issn-electronic Value: 1551-6709 Numbering: – Type: volume Value: 43 – Type: issue Value: 12 Titles: – TitleFull: Cognitive Science Type: main |
| ResultId | 1 |