The Effects of Linear Order in Category Learning: Some Replications of Ramscar et al. (2010) and Their Implications for Replicating Training Studies
Saved in:
| Title: | The Effects of Linear Order in Category Learning: Some Replications of Ramscar et al. (2010) and Their Implications for Replicating Training Studies |
|---|---|
| Language: | English |
| Authors: | Eva Viviani (ORCID |
| Source: | Cognitive Science. 2024 48(5). |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 29 |
| Publication Date: | 2024 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Symbolic Learning, Learning Processes, Artificial Intelligence, Prediction, Error Patterns, Training, Replication (Evaluation), Pictorial Stimuli, Error Correction, Prior Learning, Stimulus Generalization |
| DOI: | 10.1111/cogs.13445 |
| ISSN: | 0364-0213 1551-6709 |
| Abstract: | Ramscar, Yarlett, Dye, Denny, and Thorpe (2010) showed how, consistent with the predictions of error-driven learning models, the order in which stimuli are presented in training can affect category learning. Specifically, learners exposed to artificial language input where objects preceded their labels learned the discriminating features of categories better than learners exposed to input where labels preceded objects. We sought to replicate this finding in two online experiments employing the same tests used originally: A four pictures test (match a label to one of four pictures) and a four labels test (match a picture to one of four labels). In our study, only findings from the four pictures test were consistent with the original result. Additionally, the effect sizes observed were smaller, and participants over-generalized high-frequency category labels more than in the original study. We suggest that although Ramscar, Yarlett, Dye, Denny, and Thorpe (2010) feature-label order predictions were derived from error-driven learning, they failed to consider that this mechanism also predicts that performance in any training paradigm must inevitably be influenced by participant prior experience. We consider our findings in light of these factors, and discuss implications for the generalizability and replication of training studies. |
| Abstractor: | As Provided |
| Entry Date: | 2024 |
| Accession Number: | EJ1427185 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHQ0zOHH5ymGkTUZxyxgC8_AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDFGk9j0k_2ICD29fLAIBEICBm-qTdb1ngMRsb5hKAoMXSZUSEYft7d8DnRMRcApYjGgacccVDfC1DWEHc_SDRxGev7sibNbQxeWepm_9BRnacYW91mkWfBuPKo-iX0PNrlWVIyZvGWp5aAc6zX7EeX8nAbbS7nB9O3-4ALoZ58nvm_K4C8qx8AEBoVH07EUswoPe44KkZpU6D29XkszWuaEFT-82tSBqkhDXBvmX Text: Availability: 1 Value: <anid>AN0177511222;cgn01may.24;2024May30.06:37;v2.2.500</anid> <title id="AN0177511222-1">The Effects of Linear Order in Category Learning: Some Replications of Ramscar et al. (2010) and Their Implications for Replicating Training Studies </title> <p>Ramscar, Yarlett, Dye, Denny, and Thorpe (2010) showed how, consistent with the predictions of error‐driven learning models, the order in which stimuli are presented in training can affect category learning. Specifically, learners exposed to artificial language input where objects preceded their labels learned the discriminating features of categories better than learners exposed to input where labels preceded objects. We sought to replicate this finding in two online experiments employing the same tests used originally: A four pictures test (match a label to one of four pictures) and a four labels test (match a picture to one of four labels). In our study, only findings from the four pictures test were consistent with the original result. Additionally, the effect sizes observed were smaller, and participants over‐generalized high‐frequency category labels more than in the original study. We suggest that although Ramscar, Yarlett, Dye, Denny, and Thorpe (2010) feature‐label order predictions were derived from error‐driven learning, they failed to consider that this mechanism also predicts that performance in any training paradigm must inevitably be influenced by participant prior experience. We consider our findings in light of these factors, and discuss implications for the generalizability and replication of training studies.</p> <p>Keywords: Language learning; Discrimination; Replication; Categorization</p> <hd id="AN0177511222-2">Introduction</hd> <p>Natural languages provide shared codes that map meaning onto form in a manner that enables their users to communicate about the world. This raises the question of how the correct "scope" of usage for the linguistic forms these codes comprise are learned. How, for example, do learners identify the critical features differentiating the referents of the labels "dog" and "cat" (as shared with other users of the code)? Ramscar et al. ([<reflink idref="bib32" id="ref1">32</reflink>]) proposed that if such learning is underpinned by error‐driven processes, it must be inherently discriminative. Error‐driven learning involves acquiring and updating the predictive value (both positive and negative) of associations between cues and outcomes, which are typically strengthened when cues and outcomes <emph>co‐occur</emph>, and weakened/inhibited by cue <emph>background rates</emph> (how often a given cue occurs in the absence of a given outcome) and <emph>blocking</emph> (the prior predictability of an outcome in a context in which it co‐occurs with a cue). These factors result in <emph>cue‐competition</emph>, which enables learners to unlearn unreliable cues to the benefit of reliable cues, and reduce their uncertainty about the features associated with a given label. Accordingly, if the world provides multiple potential cues to the label "dog," and some of these are predictive (e.g., "barks"), while others are spurious or unreliable ("has fur"), then mastering label use will require learners to discriminate (and dissociate) any unreliable cues from the cues (i.e., features of the environment) that are reliably predictive of (and hence should be associated with) a label.</p> <p>A consequence of this analysis—and the key focus in Ramscar et al. ([<reflink idref="bib32" id="ref2">32</reflink>])—is that for effective unlearning to occur, the structure of events must support discriminative learning. A second consequence is that the effects of this support can be hugely affected by an individual's prior experience, since this influences background rates and blocking. This point was <emph>not</emph> considered in Ramscar et al. ([<reflink idref="bib32" id="ref3">32</reflink>]), but it is critical when it comes to predicting the <emph>outcome</emph> of learning from a specific training regime, and is an issue we will return to later. In particular, Ramscar et al. ([<reflink idref="bib32" id="ref4">32</reflink>]) suggest that because of the different functions and properties of labels versus referents, the <emph>temporal order</emph> in which stimuli are encountered can directly affect how any associative relationships between them are learned. Specifically, if referents contain many features, some of which are poor cues to label usage and some good, then in order for cue‐competition to discriminate the appropriate from the inappropriate features, it is necessary they be encountered <emph>before</emph> rather than after labels. If learners encounter numerous varied exemplars of "dogs" and "cats" followed by their labels, this will allow cue‐competition to strengthen the informative cues for the label "dog" (e.g., the feature "barks"), while uninformative cues ("has fur"), which result in prediction errors, will become disassociated. By contrast, when the order of presentation is reversed, such that each label serves as a <emph>singular</emph> predictive cue to each referent, cue‐competition over the features of the referent <emph>cannot</emph> occur. While learning will still occur in the condition, the associative values learned will simply reflect the <emph>frequency</emph> with which each feature of a referent follows each label. Critically, the cue‐competition that leads to the isolation of the discriminative (i.e., predictive) features of the referents will not be supported.</p> <p>Ramscar et al. ([<reflink idref="bib32" id="ref5">32</reflink>]) provide empirical evidence for this effect of linear order—which they refer to as "Feature‐Label‐Order" (FLO) effect —in two experiments with human learners. We focus on one of these, an artificial language learning experiment with adult participants. Participants (<emph>N</emph>=32) learned novel labels for categories of objects ("Fribbles," <ulink href="http://www.tarrlab.org">www.tarrlab.org</ulink>) under either a feature‐label (FL) learning condition—where they first saw a novel object (a Fribble) and then read a sentence presenting the category label ("That was a..." [wug/dep]); or in a label‐feature (LF) learning condition—where they first read the category label ("This is a..." [wug/dep]) and then saw the novel object.</p> <p>Fribble‐stimuli were divided into the categories shown in Fig. 1. There was one easy category—all Fribbles labeled "bim" (control Fribbles) were bright blue (i.e., a single feature mapped to a single label); the critical predictions applied to the three experimental categories. For these, each category was divided into two subsets with the most salient feature—body‐type (i.e., body shape and color)—systematically distributed such that it was not a defining feature for categorization. Instead, each label was predicted by other, more subtle discriminative features (circled for low‐frequency deps and high‐frequency tobs in Fig. 1). In addition, the frequency of training items was manipulated such that 75% of the exemplars of one category and 25% of the exemplars of another category shared the same body shape. Accordingly, to successfully learn the subcategories, participants had to learn to ignore (i.e., unlearn) the uninformative but salient body‐type feature in favor of the discriminative features. This is particularly challenging for low‐frequency items: For example, low‐frequency "deps" share the same body‐type (i.e., circular+brown) as high‐frequency "tobs," such that the label "tob" and this body‐type co‐occur frequently: Were participants to focus on the frequency of this association, they would overgeneralize and identify low‐frequency "deps" as a "tobs." This leads to the prediction that learners in the LF condition will primarily be influenced by the frequency of the co‐occurrences between labels and features and will thus show poor learning of the discriminative features compared to the FL condition.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0001.jpg" title="1 Examples and stimuli occurrence rates from Ramscar et al. ([32]). Left panel shows Fribble stimuli from Ramscar et al. ([32]) and the current study. The label bim was used for a control group of uniformly blue examples. Experimental categories (Tob, Wug, and Dep). Each category has a high‐frequency (75%) and low‐frequency (25%) subset that each contain &quot;high‐saliency&quot; feature (body type) that is shared with a subset of another category. The sets of discriminating features are circled in the low‐frequency dep and high‐frequency tob exemplars and right panel shows further exemplars from these two subcategories." /> </p> <p></p> <p>Ramscar and colleagues tested this using two four alternative force choice (4AFC) tests (administered between‐participants) where participants either matched an unseen Fribble to the labels, or matched a label to one of four new Fribbles from each category. Since no differences between test‐sets were found, the results were pooled for analysis. Fig. 2 shows the data (replotted in the format that will be used in the current paper). It can be seen that the FL condition outperforms the LF condition, and this was shown to be consistent with a computational simulation (implemented using the Rescorla–Wagner learning rule [Rescorla &amp; Wagner, [<reflink idref="bib33" id="ref6">33</reflink>]]) trained on the same input structure as the human learners. (Participants in both groups were at ceiling with the control fribbels.)</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0002.jpg" title="2 Data from Ramscar et al. ([32]). Violin plots show correct responses, split by frequency and learning‐condition. Point shows mean and error bars 95% confidence intervals." /> </p> <p></p> <p>Data were analyzed in a (2 <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$\times$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 2) ANOVA where the difference between conditions was seen in a main effect of learning condition as well as an interaction between learning condition and frequency, in the direction of a larger frequency effect in the LF condition, and a larger FL benefit for low‐frequency items than for high‐frequency items. We have since reanalyzed these original data[<reflink idref="bib1" id="ref7">1</reflink>] using logistic mixed effect models similar to those used in subsequent replications discussed below (and the current study) which better reflect the underlying binary nature of the data (Jaeger, [<reflink idref="bib16" id="ref8">16</reflink>]). These continue to reveal a strong main effect of learning condition ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 2.34, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.46, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, odds ratio = 10.4), and frequency ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 3.07, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.43, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, odds ratio = 21.5), but find no evidence of the interaction ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.17, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.84, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$=$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .83, odds ratio = 1.187).</p> <p>In sum, Ramscar et al. found a difference in performance between FL and LF trained learners, demonstrating that the order in which information is presented affects the learning of appropriately constrained generalizations. Theoretically, this provides support for a central tenet of error‐driven learning, namely, that the brain tries to predict incoming input and <emph>adjusts</emph> these predictions based on error, such that the availability of useful error is a critical aspect of learning. There are also potential educational implications: If manipulating the sequential order of information can benefit learners, this offers the possible of developing educational programs specifically designed to elicit discriminative learning (see Ramscar, Dye, Popick, &amp; O'Donnell‐McCarthy, [<reflink idref="bib26" id="ref9">26</reflink>], for an application to number learning).</p> <hd id="AN0177511222-5">Replications of the FLO effect</hd> <p>Ramscar et al. ([<reflink idref="bib32" id="ref10">32</reflink>]) has made a significant impact in Cognitive Science. At the time of writing, it has received over 360 citations on Google Scholar, averaging to <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#8764;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$\sim$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3 citations per month over the past decade. The effect of linear order in learning has also been the subject of empirical research in Education (e.g., Eitel &amp; Scheiter, [<reflink idref="bib11" id="ref11">11</reflink>]). However, as for most influential studies, no direct replication has previously been published.</p> <p>The most direct replication to date is an unpublished study by Ramscar and McClure ([<reflink idref="bib29" id="ref12">29</reflink>]) which used the same learning materials with a new group of participants (<emph>N</emph> = 3) also recruited and tested at Stanford University. The study had methodological differences, such as being part of an FMRI experiment, and participants completing both AFC tests with some counterbalancing of test order. Notably, unlike in the original study, not all participants scored perfectly on control items. Accordingly, following an unused exclusion criterion from the original study, participants who scored less than 80% on these items were excluded from the analyses (<emph>N</emph> = 7, with an additional participant excluded due to having greater than 10% missing data, leaving <emph>N</emph> = 30). The data are shown in Fig. 3, and we have again (re)analyzed the data using logistic mixed‐effect models.[<reflink idref="bib2" id="ref13">2</reflink>] Analyses confirmed that there was again a strong main effect of frequency ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.970, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.28, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, odds ratio = 7.17). In contrast to the original study, the main effect of learning condition was not significant ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.57, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.4, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .15, odds ratio = 1.77); however, there was a marginally significant interaction between frequency and learning condition ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.54, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .06, odds ratio = 3) which broken down to show no significant effect of learning condition for high‐frequency items ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.075, SE = 0.519, <emph>p</emph> = .871) but a significant effect of learning condition—in the theoretically predicted direction of an FL benefit—for low‐frequency items ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.07, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.52, <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0024" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .04, odds ratio = 2.92).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0003.jpg" title="3 Data (N = 30) replicating the FLO effect (Ramscar &amp; McClure, [29]). Violin plots show correct response rates, split by frequency and learning‐condition. Point shows mean and error bars 95% confidence intervals." /> </p> <p></p> <p>In addition to this more direct replication, numerous corollary experiments published over the years have supported the FLO hypothesis (Apfelbaum &amp; McMurray, [<reflink idref="bib1" id="ref14">1</reflink>]; Dye, Ramscar, &amp; Suh, [<reflink idref="bib10" id="ref15">10</reflink>]; Hupp, Sloutsky, &amp; Culicover, [<reflink idref="bib15" id="ref16">15</reflink>]; Hoppe, van Rij, Hendriks, &amp; Ramscar, [<reflink idref="bib14" id="ref17">14</reflink>]; Nixon, [<reflink idref="bib20" id="ref18">20</reflink>]; Ramscar, Thorpe, &amp; Denny, [<reflink idref="bib31" id="ref19">31</reflink>]; Ramscar et al., [<reflink idref="bib26" id="ref20">26</reflink>]; Ramscar, Dye, Gustafson, &amp; Klein, [<reflink idref="bib25" id="ref21">25</reflink>]; Ramscar, [<reflink idref="bib24" id="ref22">24</reflink>]; St. Clair, Monaghan, &amp; Ramscar, [<reflink idref="bib36" id="ref23">36</reflink>]; Vujović, Ramscar, &amp; Wonnacott, [<reflink idref="bib40" id="ref24">40</reflink>]). We consider the three most recent of these in detail.</p> <p>Hoppe et al. ([<reflink idref="bib14" id="ref25">14</reflink>]) studied the effect of the different linear ordering in prefixing versus suffixing on learning of the relationships between nouns and affixes in an artificial language. Under a discriminative learning account, suffixing benefits learning of abstract common dimensions from the stem nouns, that is, a benefit of the linear ordering where features (of the nouns) proceed outcomes (affixes). Their online experiment found a suffixing benefit (a FLO effect), but interestingly only for categories distinguished by multiple overlapping features; when cues were nonambiguous, both orders yield similar learning outcomes. This suggests that suffixing (i.e., as shown in Hoppe et al., [<reflink idref="bib14" id="ref26">14</reflink>]) offers a discriminative learning advantage specifically in the circumstances where cue‐competition is necessary to disambiguate cues in the environment.</p> <p>Nixon ([<reflink idref="bib20" id="ref27">20</reflink>]) found a FLO effect in an (online) study that closely matched the design of the original experiment in Ramscar et al. ([<reflink idref="bib32" id="ref28">32</reflink>]) except that the stimuli features over which generalization must occur were acoustic rather than visual. Specifically, English speakers learned to associate novel syllables (analogous to the Fribbles) with geometrical shapes (analogous to the labels), with the key discriminative feature of the syllables being lexical tone, which is not used discriminatively in English (and thus should not be salient). The design followed Ramscar et al. ([<reflink idref="bib32" id="ref29">32</reflink>]) in having low‐frequency items where there was a conflicting high‐frequency association between a salient feature (here the base syllable) and an alternative meaning. Strong learning for low‐frequency items thus involved down‐weighting the high‐frequency uninformative cue in favor of the discriminating tone cue. In the FL condition, audio syllables (multiple cues) were presented before geometric shapes (a single outcome), and in the LF, the order was reversed (geometric shapes before audio syllables). The key result was an interaction between frequency and learning condition and a simple effect of stronger performance in FL than in LF specifically for low‐frequency items (the equivalent statistic for high frequency was not reported, but the means were in the reverse direction, though the items were near ceiling). This benefit of FL for low‐frequency items is consistent with the key predictions of the discriminative account and confirms that the FLO effect is <emph>not</emph> about the ordering of visual versus linguistic stimuli per se, but rather about the arrangement of <emph>information</emph> to ensure that cue‐competition over relevant features is possible.</p> <p>Finally, online experiments by Vujović et al. ([<reflink idref="bib40" id="ref30">40</reflink>]) looked at the same advantage of <emph>suffixing</emph> over <emph>prefixing</emph> as Hoppe et al. ([<reflink idref="bib14" id="ref31">14</reflink>]), but using Fribble stimuli and the same basic class structure as Ramscar et al. ([<reflink idref="bib32" id="ref32">32</reflink>]). Four hundred participants in two experiments were exposed to novel spoken words (in contrast to the written stimuli in Ramscar et al., [<reflink idref="bib32" id="ref33">32</reflink>]) comprising a CVC stem and a CV affix, which appeared either before (prefix) or after (suffix) the stem. The stems were accompanied by visual stimuli (a subset of the Fribbles used in Ramscar et al., [<reflink idref="bib32" id="ref34">32</reflink>]) and semantic and phonological properties determined the affix use. The design mirrored the key feature of the original in that correct classification–and thus affix use—depended on unlearning the salient body shape cue in favor of discriminating cues, and this was particularly challenging for low‐frequency items due to the frequent association of the body‐type with an alternative label. Two experiments were reported that differed only in the number of individual nouns in the experiment (16 in the first, 8 in the second). Participants underwent various tests, including an AFC generalization test where they selected the best match for a novel Fribble from two audio heard stem <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0025" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$+$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> affix combinations. Analyses indicated an interaction between frequency and learning condition in both experiments, reflecting both stronger frequency effects with prefixes and a stronger suffix benefit for low‐frequency items. In contrast to the original study, there was no evidence of an FL benefit for high‐frequency items—as in Nixon, the means were in the reverse direction.</p> <p>Vujović et al. ([<reflink idref="bib40" id="ref35">40</reflink>]) also included computational simulations building on those in the original paper, which suggested that whether error‐driven learning lead to an FL specific benefit for high‐frequency items (and thus an overall "main effect" of FL) is dependent on exactly how the AFC test is simulated—that is, on how the process which evaluates the weights of associations with targets and with foils is implemented. In other words, seeing an FL benefit for high‐frequency items is not predicted by the error‐driven learning algorithm per se. This contrasts with low‐frequency items, which show an FL benefit in the model regardless of decision rule implementation, due to the specific need to "unlearn" the conflicting salient cue for these items (and the model also consistently shows a stronger frequency effect for in the LF condition regardless of implementation). Interestingly, even for low‐frequency items, Vujovic's participants showed a smaller benefit for the FL (i.e., suffixing) condition than in the original items, due to much greater overgeneralization based on body‐type. Their modeling suggested that this was consistent with learners being in an earlier stage of learning relative to those original study, since the model also exhibited overgeneralization with suffixing stimuli in the early stages of learning. They further suggested this could be due to their more complex learning paradigm, which asked learners to learn noun names and affix categorization simultaneously. Consistent with the idea that the overall complexity of the input might modulate the FLO effect, a later version of the experiment (reported as experiment 6 in Vujović, [<reflink idref="bib39" id="ref36">39</reflink>]) using a more complex artificial language with more lexical items, did <emph>not</emph> show the FLO effect—with strong overgeneralization seen in both conditions. This suggests that the ability of testing to reveal learning (and unlearning) is far more dependent on participants meeting a certain threshold of discrimination more than Ramscar et al. ([<reflink idref="bib32" id="ref37">32</reflink>]) considered.</p> <p>Overall, these later studies offer support for Ramscar et al. ([<reflink idref="bib32" id="ref38">32</reflink>])'s analysis, though they suggest that whether a "features first" benefit is seen may depend on a variety of factors such as cue ambiguity, whether there is competition between discriminating and more frequent cues, the effect of overall task complexity, and learners' experience/state of learning. It is also important to note that where a FLO effect has been seen, effect sizes have been smaller (see Fig. 11 in the General Discussion[<reflink idref="bib3" id="ref39">3</reflink>]).</p> <hd id="AN0177511222-7">Overview of the experiments in the current paper</hd> <p>Taken together, the experiments described above provide corroborating evidence in support of the FLO effect, though even in the closest replications, the effect sizes were smaller and data generally noisier than in the original study. There is also some question as to whether the FL benefit in the original paradigm should be expected to hold across high‐ and low‐frequency items, or just for low‐frequency items. Given the theoretical and potential educational implications of the original finding, and bearing in mind the increased emphasis on the need for replication in Psychology (Open Science Collaboration, [<reflink idref="bib21" id="ref40">21</reflink>]), we conducted two large scale (near) direct replications of the original study. In contrast to the original, but consistent with the three more recent studies described above, we ran these studies online. Some researchers have raised concerns regarding the validity of online versus in‐lab experiments (e.g., Finley &amp; Penningroth, [<reflink idref="bib12" id="ref41">12</reflink>]), although these primarily concern the collection of behavior data such as reaction times. On the other hand, online recruitment methods have grown in popularity due to the possibility of recruiting larger samples in short periods, thus addressing concerns about the use of low‐powered samples in Psychology experiments (e.g., Morey &amp; Lakens, [<reflink idref="bib19" id="ref42">19</reflink>]), and due to the necessity in the context of the COVID‐19 pandemic. It is thus important to establish the replicability of various experimental paradigms in online contexts. In what follows, we report these replications of Ramscar et al. ([<reflink idref="bib32" id="ref43">32</reflink>])'s experiment 1, which were conducted exclusively online and were preregistered on OSF.</p> <hd id="AN0177511222-8">Experiment 1</hd> <p></p> <hd id="AN0177511222-9">Overview and predictions</hd> <p>This experiment used the same stimuli and training and test structure as the original experiment. One extra test was added at the end of the experiment —a contingency test (where they had to judge how well a given Fribble matched a label; Fig. 4). For clarity and conciseness of presentation, the methods and results for this test (which are somewhat mixed) are provided in the online Appendices rather than the main text (see online Appendix: Contingency and additional subset analysis).</p> <p>Analysis plans were preregistered. Importantly, we preregistered that in the light of findings suggesting an effect of test‐type in this paradigm (Vujović, [<reflink idref="bib39" id="ref44">39</reflink>]), we would examine the results of the two AFC tests separately.[<reflink idref="bib4" id="ref45">4</reflink>]</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0004.jpg" title="4 Design of the experiment." /> </p> <p></p> <p>For each test, we preregistered the following predictions: We expected to find a main effect of frequency, that is, overall better performance with high‐frequency items. This is essentially a positive control (i.e., if it was not observed, it would indicate an issue with the paradigm). On the basis of Ramscar et al. ([<reflink idref="bib32" id="ref46">32</reflink>]), we also looked for a main effect of frequency in the direction of overall higher performance in FL than in LF. However, since this was not observed in the other replications described above (or originally predicted), we preregistered it as a secondary prediction.</p> <p>The most critical are our primary predictions: (i) An interaction between frequency and learning condition, reflecting a larger frequency effect in the LF condition and (ii) A simple effect of learning condition for low‐frequency items, reflecting stronger learning in the FL condition. We pre‐registered that observing either of these would support the FLO analysis.</p> <hd id="AN0177511222-11">Methods5</hd> <p></p> <hd id="AN0177511222-12">Participants</hd> <p>Two hundred and ninety‐nine participants were recruited via Prolific. Each confirmed they had native or native‐like English proficiency and were between 18 and 35 years old. They were randomly assigned to the LF and FL conditions, and to the Four Pictures + Contingency or Four Labels + Contingency test (see Fig. 4). Data from 101 participants who scored below 80% on control items and 22 with 10% or more missing trials on either test were excluded following preregistered criteria, leaving 75 in the four pictures, and 93 in the four labels task. Data collection used a preregistered optional stopping procedure, starting with 50 participants and adding in batches of 20 up to a preset maximum (see Sample size planning, and also Data loss and deviations from preregistration), with Bayes Factors checked to see if we had substantial evidence for <emph>either</emph> H1 <emph>or</emph> H0 for each of our hypotheses (in which case data collection stopped) using the process explained below. [<reflink idref="bib6" id="ref47">6</reflink>]</p> <hd id="AN0177511222-13">Stimuli</hd> <p>These were identical to Ramscar et al. ([<reflink idref="bib32" id="ref48">32</reflink>]): 75 "Fribble" pictures (<ulink href="http://www.tarrlab.org">www.tarrlab.org</ulink>) were divided into four categories, including a "control" Fribble (always blue) and the three experimental categories (comprising six subcategories) described above (Fig. 1). The control category comprised 15 exemplars, each high‐frequency subcategory 15 exemplars, and each low‐frequency subcategory 5 exemplars. Labels were presented as text. The experiment was redeveloped with the JsPsych library (De Leeuw, [<reflink idref="bib6" id="ref49">6</reflink>]) and hosted on the Gorilla platform.</p> <hd id="AN0177511222-14">Procedure</hd> <p>Participants were told that they would "learn an alien language." They saw pictures of "aliens" and heard "how they are referred to in the alien language." The training and testing structure employed is depicted in Fig. 4.</p> <hd id="AN0177511222-15">Learning</hd> <p>A total of 90 high‐frequency, 30 low‐frequency, and 30 control trials were presented in two identical blocks separated by a short break. Each trial presented a Fribble and its label, with the order and timing of Fribble/label presentations varying between conditions, as illustrated in Fig. 5.[<reflink idref="bib7" id="ref50">7</reflink>]</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0005.jpg" title="5 Temporal structure of the learning trials." /> </p> <p></p> <hd id="AN0177511222-17">Four pictures AFC task</hd> <p>Test‐trials presented a label and four previously unseen Fribbles, and participants had to pick the Fribble matching the label (Fig. 6). Responses made after 3500 ms were coded as timed‐out. There were 8 control tests (control fribble as target, and foils from both high‐ and low‐frequency categories), 24 high‐frequency tests (a high‐frequency target + two foils from the other two high‐frequency categories + one control fribble), and 24 low‐frequency trials (a low‐frequency target + two foils from the two other low‐frequency categories + a control fribble foil). Note that these last tests can be expected to be particularly difficult because one of the foils will have a body‐type that has frequently been associated with the target label (Fig. 1).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0006.jpg" title="6 Schematic representation of a single trial in both AFC tasks." /> </p> <p></p> <hd id="AN0177511222-19">Four labels AFC task</hd> <p>In each trial, participants saw a picture of a Fribble—either a control Fribble (8 trials), a high‐frequency Fribble (24 trials), or a low‐frequency fribble (24 trials)—above four labels: dep, tob, wug, and bim.</p> <hd id="AN0177511222-20">Analyses</hd> <p>We conduct equivalent analyses to the ANOVAs/<emph>t</emph>‐tests in Ramscar et al. ([<reflink idref="bib32" id="ref51">32</reflink>]) but use logistic/linear mixed‐effect models (using the package lme4; Bates, Mächler, Bolker, &amp; Walker, [<reflink idref="bib5" id="ref52">5</reflink>]) with participant as the random effect and fixed effects for learning condition, frequency, and their interactions. We run different versions of the models with different codings to capture main versus simple effects, with the fixed effect coefficients providing statistics for our hypotheses. Instead of interpreting frequentist <emph>z</emph>‐values and <emph>p</emph>‐values, we use the relevant beta and SE values from the fixed effects to calculate Bayes Factors, which provide a way of performing the Bayesian equivalent of significance testing, but have the advantage of providing information that a <emph>p</emph>‐value cannot: A "null" result (i.e., <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0026" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .05) does not tell us whether we have evidence for the null or no evidence for any conclusion at all (or even evidence against the null). This information is important, particularly in a replication study where in the case of null results we wish to know whether we do indeed have evidence against the effects which were originally found.</p> <p>Bayes Factors were computed following the approach advocated by Dienes ([<reflink idref="bib7" id="ref53">7</reflink>], [<reflink idref="bib8" id="ref54">8</reflink>], Dienes, &amp; Wonnacott, [<reflink idref="bib35" id="ref55">35</reflink>]) using Diene's calculator (Baguley &amp; Kaye, [<reflink idref="bib3" id="ref56">3</reflink>]). The calculation requires three numbers: (i) the estimate (beta), (ii) SE from the relevant coefficient in the mixed‐effect model, and (iii) a rough estimate of the predicted difference (i.e., predicted size of the beta) for the hypothesis. We based these on beta values extracted from the coefficients of equivalent models run over the data from the earlier replication by Ramscar &amp; McClure ([<reflink idref="bib29" id="ref57">29</reflink>]). In the calculator, the predicted value is used as a parameter (or the <emph>scale factor</emph>) in a model representing the plausibility of different effect sizes if H1 is true. The calculator tests whether the data summary is more likely under this model of H1 than under a model representing the null (i.e., only plausible effect is 0). The result is a ratio representing the relative strength of evidence for H1 versus the null—this is the Bayes Factor. Values above 1 indicate more evidence for H1 and below 1 more evidence for H0.[<reflink idref="bib8" id="ref58">8</reflink>] Bayes Factors are interpreted continuously; however, for hypothesis testing, we can also use discrete evidential categories. We use: BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0027" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3 indicates substantial/moderate evidence for H1 and a BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 1/3 indicates moderate/substantial evidence for H0, otherwise the evidence is ambiguous (i.e., the data are insensitive to test the hypothesis). Note that BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3 is approximately as conservative as <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0030" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .05, though alignment is not guaranteed.</p> <p>Since Bayes Factors are sensitive to the choice of values for the predicted effects, and since there is some subjectivity in this, we also calculated "robustness regions" for each BF (indicated as <emph>Robustness Region</emph> = [x:y]). These show the range of predicted values we could have used as the parameter (scale factor) for the model of H1 and still have drawn the same conclusion based on the cutoffs of BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0031" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3 or BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0032" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 1/3. That is, x and y represent how low/high a value we could have used and still obtained a BF which was greater than 3 (if the BF is <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3) lower than 1/3 (if the BF is lower than 1/3) or between 1/3 and a third (if the BF is between 1/3 and 3). Further details can be found in the Details of statistical analyses. Note we chose (and preregistered) to test one‐tailed hypotheses (since the predictions are clearly directional) and throughout the reporting beta estimates are consistently reported as positive when they are in the predicted direction and negative when they are not. We also report the (more familiar) <emph>p</emph>‐values, without interpreting them.</p> <hd id="AN0177511222-21">Results</hd> <p>Control trials were not included in analyses (average accuracy on these was 99% after participants' exclusion).</p> <hd id="AN0177511222-22">Four picture AFC task</hd> <p>Data are plotted in Fig. 7. Note that the very low performance with low‐frequency items is due to the fact that participants overgeneralized in 66% of trials (i.e., they erroneously matched based on body shape).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0007.jpg" title="7 Violin plots showing correct response rates split by frequency and learning‐condition. Point shows mean and error bars 95% confidence intervals." /> </p> <p></p> <p>There is strong evidence for higher overall performance for high than low frequency ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 4.46, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.35, tails = 1, <emph>p</emph> = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, Predicted Effect = 1.97, BF = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0037" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mn&gt;2.2&lt;/mn&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msup&gt;&lt;mn&gt;10&lt;/mn&gt;&lt;mn&gt;34&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$2.2 \times 10^{34}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , Robustness Region = [0.03: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0038" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 10]). Evidence for overall higher performance in FL than in LF tends toward the null but is ambiguous ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.13, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.32, tails = 1, <emph>p</emph> = .34, Predicted Effect = 0.58, BF = 0.67, Robustness Region = [0:1.3]). The evidence for each of our primary predictions is ambiguous: Interaction between frequency and learning condition ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.98, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.66, tails = 1, <emph>p</emph> = .07, Predicted Effect = 1, BF = 2.1, Robustness Region = [0: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 10]) evidence for greater performance in FL for low‐frequency items (simple effect) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.62, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0045" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.42, tails = 1, <emph>p</emph> = .07, Predicted Effect = 1.07, BF = 1.74, Robustness Region = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0046" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mo&gt;[&lt;/mo&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;mo&gt;:&lt;/mo&gt;&lt;mn&gt;6.89&lt;/mn&gt;&lt;mo&gt;]&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$[0:6.89]$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ). In sum, there is a strong overall effect of frequency, and the evidence that this effect is smaller in the FL condition tends in the direction of supporting the hypothesis (BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 1) but is in the ambiguous range. The evidence also tends toward supporting the hypothesis of more accurate performance with low‐frequency items in FL than in LF, but again is ambiguous.</p> <hd id="AN0177511222-24">Four labels AFC task</hd> <p>Data are in Fig. 8. Again, the very low performance with low‐frequency items is due to participants overgeneralizing based on body shape (64% of low‐frequency trials).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0008.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0008.jpg" title="8 Violin plots showing correct response rates split by frequency and learning‐condition. Point shows mean and error bars 95% confidence intervals." /> </p> <p></p> <p>There is strong evidence for higher overall performance for high than low frequency ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0048" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 4.34, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.39, tails = 1, <emph>p</emph> = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, Predicted Effect = 1.97, BF = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mn&gt;8.31&lt;/mn&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msup&gt;&lt;mn&gt;10&lt;/mn&gt;&lt;mn&gt;25&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$8.31 \times 10^{25}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , Robustness Region = [0.04: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 10]). There is substantial evidence for the null hypothesis that overall performance is higher in FL than in LF ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = –0.36, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.38, tails = 1, <emph>p</emph> = .825, Predicted Effect = 0.58, BF = 0.33, Robustness Region = [0.56:inf]). Critically, there is also evidence for the null for our primary predictions: Interaction: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = –0.7, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0056" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.66, tails = 1, <emph>p</emph> = .855, Predicted Effect = 1, BF = 0.31, Robustness Region = [0.92:inf]); Simple Effect: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0057" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = –.71, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0058" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.39, tails = 1, <emph>p</emph> = .965, Predicted Effect = 1.07, BF = 0.13, Robustness Region = [0.37:inf]).</p> <p>In sum, there is a strong overall effect of frequency and there is evidence against the hypothesis that this effect is smaller in the FL condition. There is also evidence against FL participants being more accurate with low‐frequency items than those in LF participants.</p> <hd id="AN0177511222-26">Discussion</hd> <p>Experiment 1 revealed strong frequency effects and high performance on high‐frequency items in both the LF and FL conditions. Critically, we did not see evidence for either of the key hypotheses or the secondary hypothesis. However, the pattern of evidence differed for the two tests. In the four pictures test, although evidence did not meet our criteria for substantial, there was more evidence for H1 than for the null for both the interaction and the simple effect. In contrast, in the four labels test, we found evidence for the null for both effects. Thus, overall, this experiment does not confirm the key findings of Ramscar et al. ([<reflink idref="bib32" id="ref59">32</reflink>]).</p> <p>In considering why we did not find the predicted effects, it is important to note the extremely low performance for the items from the low‐frequency subsets in both conditions and in both tests. Looking at the AFC tests, the majority of participants have scores well below 33% (which could be viewed as "chance" given that participants so rarely chose the control Fribble), which was not the case in either the original study or the unpublished replication study by Ramscar &amp; McClure ([<reflink idref="bib29" id="ref60">29</reflink>])(see Fig. 3). The poor performance in this experiment was mostly because participants overgeneralized and picked the Fribbles based on body shape, indicating that they did not learn to discriminate low‐ from high‐frequency items (there was also a significant data loss, with about one‐third of participants not meeting our criteria). Given the floor effects observed in the low‐frequency items, it is difficult to determine whether the differences in learning predicted for the FL and LF conditions are supported or not.</p> <p>To account for the reduced performance, we considered whether there might be some unintended differences in the learning phase compared to the original lab experiments. One potential difference is that originally, the stimuli had a consistent size across participants, filling a large monitor screen. By contrast, stimulus size in this experiment depended on the size of the monitor of the participants' computer, introducing variability. The possibility that smaller stimuli did not allow participants to learn the discriminative features of the objects prompted Experiment 2: A follow‐up replication with a monitor calibration phase.</p> <hd id="AN0177511222-27">Experiment 2</hd> <p></p> <hd id="AN0177511222-28">Methods</hd> <p></p> <hd id="AN0177511222-29">Participants</hd> <p>Two hundred and fifty‐five participants were recruited using the same criteria as before. However, this time participants were recruited both through Prolificand through the University of Tübingen. Prolific participants were compensated as in Experiment 1, while the university participants participated for credit. The stopping procedure described in the Participants section was employed again.</p> <hd id="AN0177511222-30">Participant exclusion</hd> <p>Of the participants, 131 completed the Four Pictures task followed by the contingency task, while 124 completed the Four Labels task followed by the contingency task. Again, as preregistered, participants scoring below 80% on control items in the 4AFC tasks were excluded (35 participants from Four Pictures and 20 from the Four Labels). Exclusions for missing data (over 10% in any task) resulted in a further 13 from the Pictures and 6 from Four Labels being excluded, leaving 83 participants in the Four pictures and 98 in the Four Labels task.</p> <hd id="AN0177511222-31">Stimuli and procedure</hd> <p>Experiment 2 was the same as Experiment 1 but added a resizing procedure to standardize stimuli size across monitors, using the JsPsych ‐resize‐ plugin (De Leeuw, [<reflink idref="bib6" id="ref61">6</reflink>]). Participants matched the on‐screen size of a container to a credit card using a slider, allowing us to set a consistent Fribble size (160x160mm) regardless of monitor.</p> <hd id="AN0177511222-32">Results</hd> <p></p> <hd id="AN0177511222-33">Four picture AFC task</hd> <p>We plot the accuracy by frequency and condition in Fig. 9. Again, the very low performance with low‐frequency items is due to participants overgeneralizing based on body shape—here in 69% of low‐frequency trials.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0009.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0009.jpg" title="9 Violin plots showing correct response rates split by frequency and learning‐condition. Point shows mean and error bars 95% confidence intervals." /> </p> <p></p> <p>There is strong evidence for higher overall performance for high than low frequency ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0059" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 4.89, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0060" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.37, tails = 1, <emph>p</emph> = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0061" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, Predicted Effect = 1.97, BF = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0062" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mn&gt;4.88&lt;/mn&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msup&gt;&lt;mn&gt;10&lt;/mn&gt;&lt;mn&gt;35&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$4.88 \times 10^{35}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , Robustness Region = [0.03: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0063" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 10]). Evidence for overall higher performance in FL than in LF tends toward the null but is ambiguous ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0064" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.004, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0065" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.28, tails = 1, <emph>p</emph> = .505, Predicted Effect = 0.58, BF = 0.44, Robustness Region = [0:0.79]). Turning to our primary predictions, critically, there is evidence for the interaction that crosses substantial criteria ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0066" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.35, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0067" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.71, tails = 1, <emph>p</emph> = .03, Predicted Effect = 1, BF = 3.58, Robustness Region = [0.64:2.26]). The evidence for the simple effect of an FL benefit for low‐frequency items was ambiguous <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0068" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.68, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0069" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.42, tails = 1, <emph>p</emph> = .05, Predicted Effect = 1.07, BF = 2.14, Robustness Region = [0:8.76]). In sum, there is a strong overall effect of frequency; however, in line with our prediction, there is evidence that this effect is smaller in the FL than in the LF condition. The evidence also tends in the direction of supporting the hypothesis of more accurate performance with low‐frequency items in FL than in LF, but it is in the ambiguous range.</p> <hd id="AN0177511222-35">Four labels AFC task</hd> <p>We plot the accuracy by frequency and condition in Fig. 10. Again, the very low performance with low‐frequency items is due to the fact that participants are overgeneralizing based on body shape (52% of low‐frequency trials). There is strong evidence for higher overall performance for high than low frequency ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0070" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 3.86, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0071" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.31, tails = 1, <emph>p</emph> = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0072" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001, Predicted Effect = 1.97, BF = <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0073" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mn&gt;4.37&lt;/mn&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msup&gt;&lt;mn&gt;10&lt;/mn&gt;&lt;mn&gt;32&lt;/mn&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$4.37 \times 10^{32}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> , Robustness Region = [0.03: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0074" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 10]). There is substantial evidence for the null hypothesis that there is higher performance in FL than in LF ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0075" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = −0.19, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0076" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.31, tails = 1, <emph>p</emph> = .735, Predicted Effect = .58, BF = 0.32, Robustness Region = [0.56:inf]). For the primary predictions, the evidence for the interaction is ambiguous ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0077" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.39, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0078" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.59, tails = 1, <emph>p</emph> = .25, Predicted Effect = 1, BF = 0.86, Robustness Region = [0:3.21]) and there is evidence for the null for the simple effect ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0079" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.003, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0080" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.34, tails = 1, <emph>p</emph> = .495, Predicted Effect = 1.07, BF = 0.31, Robustness Region = [0.97:Inf]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0010.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0010.jpg" title="10 Violin plots showing correct response rates split by frequency and learning‐condition. Point shows mean and error bars 95% confidence intervals." /> </p> <p></p> <p>In sum, there is a strong overall effect of frequency. The evidence that this effect is larger in the LF condition than in the FL tends toward the null, but is ambiguous. There is evidence against the hypothesis that FL participants are more accurate with low‐frequency items than LF participants.</p> <hd id="AN0177511222-37">Discussion</hd> <p>As in Experiment 1, performance was generally lower than in the original study, with most participants scoring at or below "chance" (33%) on low‐frequency items due to extensive overgeneralization. The pattern again differed between the two tests and in each case the direction of the evidence matched Experiment 1. However, in terms of our preregistered evidence thresholds, in the four pictures task, the evidence for H1 for the interaction between frequency and learning‐condition reaches substantial criteria (BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0081" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3), indicating that this test taps into the effects of stimuli ordering in training, as predicted. However, the evidence for the simple effect, although in direction of H1, was ambiguous. In the four labels test, evidence favored H0 for both key hypotheses (interaction and simple effect), with evidence for the null reaching the substantial criterion (BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0082" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 1/3) for the simple effect. Again, we did not see evidence for a main effect in either AFC test, aligning with the preregistered expectations and previous replications (we consider this further in the General Discussion).</p> <p>Overall, the evidence pattern suggests quantitative rather than qualitative differences between Experiments 1 and 2. However, in terms of our preregistered inference criteria, only in Experiment 2 did we find evidence for H1 for one of the tests. Accordingly, on the basis of the analyses conducted so far, it is not possible to ascertain whether this reflects a genuine difference between the experiments. To determine this, we performed additional exploratory (nonpreregistered) analyses to see whether the evidence differs.</p> <hd id="AN0177511222-38">Comparing and combining data sets: Further exploratory analyses</hd> <p>There were two differences between the experiments: (<reflink idref="bib1" id="ref62">1</reflink>) standardizing stimulus size in the current experiment, and (<reflink idref="bib2" id="ref63">2</reflink>) different recruitment methods: Solely Prolific recruited in Experiment 1, versus a mix of Prolific and university students in Experiment 2. Although (<reflink idref="bib2" id="ref64">2</reflink>) was not intentional, the original experiments were done with university students, and as we noted in the introduction, the same learning mechanisms that predict the FLO effect also predict that populations with different levels of experience can be expected to perform differently in a learning experiment (albeit this fact is not usually acknowledged when training studies are reported). Given this, we ran analyses (<reflink idref="bib1" id="ref65">1</reflink>) over data from the Prolific‐recruited students looking for a modulating effect of <emph>stimuli‐consistency‐type</emph> (i.e., inconsistent size vs. consistent size) and (<reflink idref="bib2" id="ref66">2</reflink>) on Experiment 2 data looking for a modulating effect of <emph>recruitment‐type</emph> (Prolific recruited vs. university recruited).</p> <p>We used the same approach as our previous preregistered analyses, except that we included the fixed effects of <emph>stimuli‐consistency‐type/recruitment‐type</emph> and their interactions in the models. We were interested in evidence for interactions between each of these factors and the primary tests of the FLO effect—that is, to see if there was a three‐way interaction of <emph>stimuli‐consistency‐type/recruitment‐type</emph> by <emph>learning condition</emph> by <emph>frequency</emph> and if there was an interaction of <emph>stimuli‐consistency‐type/recruitment‐type</emph> by <emph>learning condition</emph> specifically for low‐frequency items. While we again used Bayes Factors, these were two‐tailed as we did not start with predictions in one direction (see Details of statistical analyses). For both test tasks, in every case, the evidence for an interaction was around 1, that is, highly ambiguous (BFs between 0.81 and 1.12).[<reflink idref="bib9" id="ref67">9</reflink>]</p> <p>Since these analyses find no evidence that the changes we made between the experiments modulated the FLO effect, we reran the analyses for the four pictures and four label tasks using our preregistered methods but using pooled data sets from Experiments 1 and 2. For the four picture test, this exploratory analysis showed substantial evidence for H1 for both the interaction between <emph>frequency</emph> and <emph>learning condition</emph> ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0083" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.18, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0084" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.49, tails = 1, <emph>p</emph> = .005, Predicted Effect = 1, BF = 9.18, Robustness Region = [0.27: 5.74]) and the simple effect: ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0085" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.66, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0086" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.3, tails = 1, <emph>p</emph> = .007, Predicted Effect = 1.07, BF = 5.11, Robustness Region = [0.2: 2.08]). For the four labels test, there was evidence for the null for both the interaction between <emph>frequency</emph> and <emph>learning condition</emph> ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0087" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = –0.12, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0088" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.44, tails = 1, <emph>p</emph> = .695, Predicted Effect = 1, BF = 0.33, Robustness Region = [1:inf]) and the simple effect ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0089" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = –0.35, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0090" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.26, tails = 1, <emph>p</emph> = .542, Predicted Effect = 1.07, BF = 0.11, Robustness Region = [0.31:inf]).</p> <hd id="AN0177511222-39">Summary of findings from these analyses</hd> <p>Our exploratory analysis of the combined data from the experiments provides substantial evidence <emph>for</emph> both of the primary predictions in the four picture task, and <emph>against</emph> (relative evidence for the null) both primary predictions in the four labels task. In other words, we see patterns of evidence which are qualitatively in line with each of the experiments individually, but the weight of evidence in this larger sample is now sufficient to cross our "substantial" threshold in all cases. Although not preregistered, this can be interpreted bearing in mind that Bayes Factors remain a valid measure of evidence in a combined sample, with larger samples expected to yield stronger evidence. We also point out that the sample sizes in our preregistered experiments ended up being somewhat smaller than originally planned due to a data collection error (see Data loss and deviations from preregistration).[<reflink idref="bib10" id="ref68">10</reflink>]</p> <p>These analyses revealed no significant impact from changes in stimulus size or recruitment methods on the 2AFC tests between Experiments 1 and 2. Since these results are ambiguous (i.e., we do not have substantial evidence for either H1 <emph>or</emph> the null), we cannot draw strong conclusions. We also note that in the contingency data (see Contingency and additional subset analysis), we <emph>did</emph> find some evidence for an interaction with recruitment type (in the direction of stronger evidence for the simple effect for low‐frequency items in the university‐recruited than in the Prolific‐recruited individuals). However, for the AFC tests, the key conclusion must be that our paradigm and sample size are not sufficiently sensitive to test for these differences between experiments.</p> <hd id="AN0177511222-40">General discussion</hd> <p>Ramscar et al. ([<reflink idref="bib32" id="ref69">32</reflink>]) presented an analysis and simulations that predicted stimuli sequencing effects in error‐driven learning. In the current work, we set out to look for evidence of these key effects in two preregistered replications conducted online, which differed from one another in terms of stimuli size and recruitment methods. In both experiments, we saw different patterns of results for the two AFC tests. In our preregistered analyses: For the four pictures task, the evidence tended toward H1 for both the interaction and the simple effect in both experiments, crossing the substantial criteria (BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0091" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#62;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&gt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 3) for the interaction in Experiment 2. In the four labels task, in both experiments, evidence tended toward the null for both the interaction and the simple effect, crossing the substantial criteria (BF <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0092" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 1/3) for both in Experiment 1, and for the simple effect in Experiment 2. In exploratory analyses over the combined data, there was substantial evidence for both the interaction and the simple effect in the four pictures task, and evidence for the null for both effects in the four labels task.</p> <p>The evidence from the four picture task is thus (broadly) in line with that from the original study–and, critically, with the predictions of computational models implementing error‐driven learning—albeit with moderate evidence levels. This adds to the existing body of evidence supporting the theory (e.g., Nixon, [<reflink idref="bib20" id="ref70">20</reflink>]; Vujović et al., [<reflink idref="bib40" id="ref71">40</reflink>]). However, the null results in the four labels test suggest that observations of the FLO effect can be affected by the test task. We discuss this point below in the section Test‐task effects.</p> <p>Focusing on the four pictures task, while the results are overall consistent with predictions, there are important differences compared to the original experiment. One key observation is that learning of the low‐frequency items in the FL condition is much lower than in the original study, and thus the benefit over the LF condition is much smaller. This can be clearly seen in the comparison of effect sizes in Fig. 11: The effect is much smaller than in the original, though roughly consistent with the more recent studies by Vujovic and Nixon—which were also conducted online.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0011.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0011.jpg" title="11 Effect sizes for the simple effect—that is, the benefit of the FL over the LF condition specifically for low‐frequency items; beta and SE are log odds; right shows forest plot as odds ratio." /> </p> <p></p> <p>Not‐withstanding that effects are generally smaller in replications, it is worth considering why learning of low‐frequency items is not stronger in FL in our study. We saw that this poor performance reflects overgeneralization based on body‐type: Almost 70% of the time, our participants ignored the discriminating features and erroneously picked the foil Fribble whose body‐type was frequently associated with the label. Vujović et al. ([<reflink idref="bib40" id="ref72">40</reflink>]) showed that a similar pattern of overgeneralization also occurs during the early stages of training in computational simulations. Critically, in these simulations, overgeneralization in FL decreases as exposure to the input increases, due to increased opportunity to use prediction error to unlearn the erroneous cue. This suggests that (at test) our online participants may have reached an average point in learning than is different to those in the original study. We return to this point in the section Implications for conducting online training experiments.</p> <p>Another difference is that the original study found an FL benefit for both low‐ and high‐frequency items (reflected in a main effect), but in the current study that was only seen for low‐frequency items (as originally predicted by Ramscar et al., [<reflink idref="bib32" id="ref73">32</reflink>]). Indeed—though we did not analyze high‐frequency items in isolation—the data trends in the direction of an LF benefit consistent with patterns seen in Nixon ([<reflink idref="bib20" id="ref74">20</reflink>]) and Vujović et al. ([<reflink idref="bib40" id="ref75">40</reflink>]) (this is why a main effect was preregistered only as a secondary hypothesis). Further, Vujović et al. ([<reflink idref="bib40" id="ref76">40</reflink>]) pointed out that stronger LF learning for high‐frequency items is consistent with the model if participants base their choice on the picture whose features are (in sum) most positively associated with the label, assuming positive associations are mapped to raw associative weights in the model (since though negative weights for the incorrect label are greater in FL, <emph>positive</emph> weights for the correct label are greater in the LF condition). In contrast, the behavior of the original participants in showing an FL benefit with these items is more consistent with the use of a "choice" rule in the AFC task, in which in determining the probability that a label matches a particular set of features also takes into account the match with the other label (though in interpreting the findings with high‐frequency items in the original, it should also be noted that responses are close to ceiling). For a fuller discussion of these issues, see Vujović et al. ([<reflink idref="bib40" id="ref77">40</reflink>]).</p> <hd id="AN0177511222-42">Implications for conducting online training experiments</hd> <p>This paper adds to the growing body of research conducted online rather than in laboratory. We suggested above that our results (at least with the four pictures tasks) are consistent with learners being at an earlier average point in learning as compared to those in the original study (and, though to a lesser extent, in the unpublished lab‐based replication by Ramscar &amp; McClure, [<reflink idref="bib29" id="ref78">29</reflink>]). Vujović et al. ([<reflink idref="bib40" id="ref79">40</reflink>]) suggested something similar for their online study.</p> <p>Could the move to an online platform have led to this difference? Recent discussions have highlighted numerous factors affecting online versus lab experiments (Gagné &amp; Franzen, [<reflink idref="bib13" id="ref80">13</reflink>]; Rodd, [<reflink idref="bib34" id="ref81">34</reflink>]). One difference is that every online participant uses their own hardware. In this study, we explored the impact of monitor variability: Experiment 1 did not take into account the fact that monitor size could affect stimuli size, which was corrected in Experiment 2. Our subsequent analyses did <emph>not</emph> find substantial evidence that this factor modulated the effects of interest (although we also did not find substantial evidence for the null, meaning that we cannot draw very strong conclusions about this methodological aspect). However, we did not consider other hardware aspects, such as processing capacity and network quality, which could potentially affect stimuli presentation (including the critical timing differences between FL and LF). Future work could mitigate against this by restricting the experiment to participants with computers that meet specific hardware criteria.</p> <p>Another key difference in online experiments is the lack of control over participants' environments. Participants in lab experiments may have less risk of distractions and may be more motivated and attentive if they believe they are being observed. We mitigate against this to some extent via our exclusion criteria—they must have least paid sufficient attention to learn the easy control category and be sufficiently attentive during the test not to have too many "timed‐out" trials. However, we cannot rule out that the different environments led to reduced attention to the input stimuli.</p> <p>A final difference between lab‐based experiments and online experiments —which we would argue is particularly relevant for <emph>learning</emph> paradigms —is in recruitment procedures, which can change the profile of participant populations. Platforms such as Prolific provide a mechanism for wide‐scale recruitment across (relatively) diverse participants; however, while recruitment criteria can be stipulated (e.g., age range, English proficiency, lack of learning disabilities, etc.), they cannot be guaranteed. Critically, as we noted at the outset, it is well established that in error‐driven learning, while the <emph>association rates</emph> between cues and outcomes promote the learning of positive associations, what actually gets learned also depend on factors that tend to inhibit learning: The <emph>background rates</emph> of cues, and the prior predictability of outcomes in context (<emph>blocking</emph>). Not only do these factors interact to produce cue competition (which led Ramscar et al., [<reflink idref="bib32" id="ref82">32</reflink>] to predict the FLO effect), they also predict that the outcome of training in any given task is itself a function of prior experience. This is best explained by considering a far simpler learning paradigm than the one tested here. In paired‐associate learning (PAL), participants learn word pairs and then recall one word—the target, when given another—the cue. Ramscar, Hendrix, Love, and Baayen ([<reflink idref="bib27" id="ref83">27</reflink>]); Ramscar, Hendrix, Shaoul, Milin, &amp; Baayen ([<reflink idref="bib28" id="ref84">28</reflink>]); Ramscar, Sun, Hendrix, and Baayen ([<reflink idref="bib30" id="ref85">30</reflink>]) have shown that age‐related "declines" in adult PAL performance can be accurately <emph>predicted</emph> by error‐driven models that estimate the learnability of word pairs as a function of learner's previous experience with the actual words in question, modulated by the three factors that promote and inhibit learning described above. These models/factors explain why adults of all ages are better at learning "easy" PAL pairs like <emph>baby‐cries</emph> than "hard" pairs like <emph>obey‐eagle</emph>: because the high association rate of <emph>baby‐and‐cries</emph> successfully promotes learning of this association. They can explain why harder word pairs become proportionally far more difficult to learn as experience grows: because <emph>obey</emph> rarely occurs with <emph>eagle</emph>, increased knowledge of background rates inhibits learning. Further, they successfully predicted that older age‐matched L2 speakers would outperform native German speakers when asked to learn German PAL pairs (Ramscar et al., [<reflink idref="bib30" id="ref86">30</reflink>]): because on average native speakers have better knowledge of background rates than bilinguals (see also Qiu &amp; Johns, [<reflink idref="bib23" id="ref87">23</reflink>]). All these findings highlight a fact that was overlooked by Ramscar et al. ([<reflink idref="bib32" id="ref88">32</reflink>]): that the outcome of learning can <emph>never</emph> be independent of prior experience.</p> <p>To understand how experience impact the Fribbles task, consider that successful learning of the low‐frequency exemplars involves unlearning a highly salient distractor feature. This is in contrast to most categories of natural objects, which actually tend to share common features (Torralba &amp; Oliva, [<reflink idref="bib38" id="ref89">38</reflink>]). This must inevitably cause people exposed to natural categories to increasingly learn that common features are salient to category learning, and given this, it follows that learners with more experience of the world (i.e., older learners) ought to find the fribble task harder than learners with less experience (who will have had less training on actual co‐occurrences between common features and natural categories). In a similar vein, learners with more experience of taking psychological experiments will have more experience of the ways in which such experiments can contain manipulations which violate real‐world expectations. Although we did not collect age in our study, since the maximum age was 35, it seems likely that the participants are on average older than students in the original study, who were undergraduate psychology students with an average age of 19 (meanwhile, our Tübingen participants, many of whom were masters' students, were mainly linguists). These differences might begin to explain the performance of the original participants when it came to learning to discriminate the low‐frequency Fribbles.</p> <p>Speaking against this account is the fact that we did <emph>not</emph> find evidence for performance differences between our student participants and our Prolific‐recruited participants in our AFC tests. On the other hand, although we did not see evidence for this difference between students and nonstudents in the present study, we also did not see evidence for <emph>no</emph> difference. Moreover, in the analyses of the contingency data reported in Appendix C, and Fig. S6 in particular, we did see evidence of better learning of the low‐frequency discriminating features in students than in Prolific‐recruited participants. Further exploratory analyses in Appendix B are consistent with stronger learning of the discriminating features for low‐frequency items in the student group. The more general point is that there may be differences in experience between the current participants and those in the original experiment, and if there are, it is highly likely that this will impact their performance, which suggests that researchers should be more circumspect in assuming the generality of their findings when it comes to results from training paradigms (it is notable in this regard that the <emph>biggest</emph> by‐decade change in PAL performance observed across the lifespan is between adults in their 20s and 30s; Ramscar et al., [<reflink idref="bib27" id="ref90">27</reflink>]).</p> <hd id="AN0177511222-43">Test‐task effects</hd> <p>A surprising outcome in this study was that we did <emph>not</emph> find evidence for our predicted effects in the four labels task in either experiment. In fact, when combining data sets, we see evidence for the null. Assuming that participants learn as predicted—as shown by the Four Pictures task and the original study—why were the primary predictions not met in the four label task?</p> <p>It is notable here that performance with low‐frequency items was stronger in the four labels task, and particularly in the LF condition. A possible explanation for this is that the test itself might result in cue competition, and this could wipe out differences from training condition. Recall that in the four labels task, a single Fribble is presented simultaneously with four label choices. Thus, a natural way to approach the task would be to consider the picture against each label in turn, from left to right. For example, if the Fribble picture is a low‐frequency "dep," the participant will see this alongside all four labels. If they look back and forth between picture and labels, when they reach "tob," this could itself provide an error signal, helping them to discard "tob" in favor of the correct label "dep."</p> <p>Speaking against this explanation is that, in principle, something similar could occur in the four pictures task: Participants could in turn attempt to match each of the four Fribble pictures against the single label. However, since the complexity of the Fribble pictures would make it much harder to do this within the 3500 ms available, it seems that the time constraint—which was the same in both tasks—may have been more effective in blocking this process in the four pictures task. A supplementary analysis comparing the likelihood of a "timed‐out" trial for the two responses showed that this was greater for the four picture task—four picture: 4.8%, four labels: 2.5%; fixed effect of test task in the logistic mixed effect model: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0093" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.74, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0094" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.15, <emph>z</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0095" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$=$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 5.3 <emph>p</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0096" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$&lt;$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .001—supporting the claim that the time limit was indeed more constraining. Thus, it seems responses in the four pictures task may better reflect the implicit knowledge actually gleaned in training.</p> <p>How does this explanation fit with previous studies? Ramscar et al. ([<reflink idref="bib32" id="ref91">32</reflink>]) reported their results across data from the combined AFC tasks; however, we have since been able to look at their data separately for the two tasks (Figs. 12 and 13), it is clear that performance for low‐frequency items is stronger in FL than in LF for <emph>both</emph> tests. One possibility for the difference—consistent with our discussion above of how learners in the current study might be at a different point in learning—is that learners in the original study had progressed to a stage where the FL benefit was sufficiently strong that it could not be confounded by learning during the test.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0012.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0012.jpg" title="12 Data from the Four Pictures task in Ramscar et al. ([32])." /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/CGN/01may24/cogs13445-fig-0013.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="cogs13445-fig-0013.jpg" title="13 Data from the Four Labels task in Ramscar et al. ([32])." /> </p> <p></p> <p>Interestingly, Vujović et al. ([<reflink idref="bib40" id="ref92">40</reflink>]) also included different types of AFC tests and also found stronger patterns in some tests than others. However, the strongest results in that paper were in fact found in the task which, at first glance, is more similar to the current four labels test. Specifically, there was a single Fribble and two phrases to choose from. However, recall that Vujović et al. ([<reflink idref="bib40" id="ref93">40</reflink>]) used <emph>audio</emph> rather than <emph>written</emph> stimuli, and in this task, participants had to listen to the two phrases sequentially before making their choice. Therefore, participants did <emph>not</emph> have the same opportunity to deliberately consider the Fribble against a set of labels in turn.</p> <p>More generally, this result shows again that the choice of test task type can have unexpected consequences when trying to tap into these learning effects. Finally, we note that there are also further test‐type differences when we consider the results of the additional contingency test.[<reflink idref="bib11" id="ref94">11</reflink>] The results of this new test were mixed, with evidence for the null for our predictions in some cases, and for H1 in others, and exploratory analyses finding tentative evidence of differences between learner groups. This is fully discussed in the relevant online Appendix (C)—here, we note that this is again consistent with a modulating role of test‐type which affects our ability to tap into the effects which signature discriminative learning. The more general conclusion for training experiments is that effects gleaned from short periods of exposure must inevitably depend on the ability of a test to detect any effects of training.</p> <hd id="AN0177511222-46">Data loss and deviations from preregistration</hd> <p>While we view the use of preregistration as a strength of this work, it does lay bare deviations and procedural errors which must be acknowledged. We deviated from our preregistration in the size of the final sample for both Experiments 1 and 2. Our preregistration said that unless we had sufficient evidence from a smaller sample, we would continue collecting data up to a maximum of 100 participants in each of the 4AFC tasks in each <emph>net of exclusions</emph>. Exclusions were specified as depending both on participants' performance in control tasks and amount of missing data for their AFC test (no more than 10% timed‐out trials being permissible). Unfortunately, an error in the interim analyses we conducted during data collection meant that exclusions were in fact only identified on the basis of control trials, not missing data. This was not identified until after our data collection for both experiments was completed, meaning that those participants had not been replaced. Thus, our final data sets in each experiment are less than the planned maximum. Although we could theoretical rectify this by collecting more data in each experiment, in the end we were able to gain a larger sample by combining across the experiments. Since this combined data set actually led to a clear pattern of results for our key hypotheses for each AFC test, we did not feel that additional data collection in individual hypotheses was a good use of resources at this point.</p> <p>Another deviation is that we preregistered an analyses plan for Experiments 1 and 2 which included a set of specific priors for the Bayes Factors, basing these on data from Ramscar &amp; McClure ([<reflink idref="bib29" id="ref95">29</reflink>]). We later found a small error in our initial analysis of that experiment which (slightly) changed the relevant values. We decided to use the new values (i.e., those based on more accurate analyses) as the priors for the later experiments, rather than sticking with the preregistered values. Critically, the differences are small and there is no place where using these rather than the preregistered values led to qualitatively different results (i.e., there was no case where this affected the direction of the evidence as for/against our hypotheses, or where we claim substantial evidence for H1 or H0 where the use of the original values would have indicated ambiguous evidence). The reader can verify this by checking that the preregistered values (see footnote 4 above) fall within the robustness regions in each case (it is also shown in our analyses script).[<reflink idref="bib12" id="ref96">12</reflink>]</p> <hd id="AN0177511222-47">Open practises</hd> <p>Data, analysis, and preregistration of all experiments are available on OSF. All analyses are carried out in R, and all analysis scripts are on OSF.</p> <hd id="AN0177511222-48">Funding</hd> <p>This work was supported by the Leverhulme Research Project RPG‐2019‐160 "Language Learning as Expectation: a Discriminative Perspective" awarded to the last author. The second author was supported in part by a grant from the Deutsche Forschungsgemeinschaft (DFG 381713393).</p> <hd id="AN0177511222-49">Competing interests</hd> <p>We have no competing interests.</p> <hd id="AN0177511222-50">Open Research Badges</hd> <p>This article has earned Open Data and Open Materials badges. Data and materials are available at https://10.17605/OSF.IO/ZH8AJ</p> <p>GRAPH: Fig. S1: Four pictures task. Combined data from the two experiments split according to participant recruitment type.Fig. S2: Summed weights from the simulation plotted (averaged over 50 runs).Fig. S3: Example of high‐frequency "match" and low‐frequency "mismatch" trial for the category tob used in the Contingency Judgment task and in the modeling.Fig. S4: Contingency task. Experiment 1. Violin plots showing average score (between –100 +100) for match and mismatch trials for high and low frequency in the two learning conditions.Fig. S5: Contingency task. Experiment 2. Violin plots showing average score (between –100 +100) for match and mismatch trials for high and low frequency in the two learning conditions.Table S1: Results of Bayes Factors tests for each primary prediction in each experiment.Fig. S6: Contingency task. Combined data from the two experiments split according to participant recruitment type.Fig. S7: Proportion of runs of the simulations (1000 runs per N participants) where the BF met various criteria for establishing evidence for H1/H0.</p> <ref id="AN0177511222-51"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref7" type="bt">1</bibl> <bibtext> Thank you to an anonymous reviewer for suggesting this analysis.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref13" type="bt">2</bibl> <bibtext> We are not able to present separate plots/analyses for the two tests since labeling of the testing conditions in the data files was unclear and information about the order in which the tests were administered to each participant has been lost.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref39" type="bt">3</bibl> <bibtext> In this figure, we have plotted the simple effect of a benefit for FL over LF specifically for low‐frequency items (i.e., the effect most clearly predicted by the computational model which has been most consistently found; Hopp et al. is not included due to the differences in the design).</bibtext> </blist> <blist> <bibl id="bib4" idref="ref45" type="bt">4</bibl> <bibtext> We also preregistered additional analyses over the tests combined, following the original study. Since we found different patterns of results for the two tests, we do not report these, though they can be viewed on OSF.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref52" type="bt">5</bibl> <bibtext> Methods including exclusion criteria, optional stopping criteria, and analyses were preregistered on OSF. Any deviations from this plan are noted in the text.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref47" type="bt">6</bibl> <bibtext> Dienes ([9]) states that optional stopping does not affect Bayes factors as it does <emph>p‐</emph> values in frequentist analyses. Frequentist methods may yield significant <emph>p</emph>‐values under optional stopping even if the null hypothesis (H0) is true due to <emph>p</emph>‐value fluctuations. In contrast, Bayes factors, which are symmetric, will increase if H0 is false and decrease if H0 is true, providing a valid measure of evidence regardless of the stopping rule used. This respects the "stopping rule principle" which ensures that the evidence is derived solely from the data, not the method of collection. Thus, Bayes Factors remain a valid measure of evidence regardless of data collection procedure.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref50" type="bt">7</bibl> <bibtext> Rapid presentation was used to prevent explicit learning and strategizing.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref54" type="bt">8</bibl> <bibtext> Note that our approach means that we are only using predicted effect sizes as part of the hypothesis testing process, they are <emph>not</emph> influencing our <emph>estimations</emph> of effects in the current study (which are taken from the mixed effect models which do not incorporate priors into estimation).</bibtext> </blist> <blist> <bibl id="bib9" idref="ref67" type="bt">9</bibl> <bibtext> Four picture test, interaction stimuli‐consistency‐type with learning condition by frequency <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0097" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.352, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0098" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.19, tails = 2, <emph>p</emph> = .768, Predicted Effect = 0.501, BF = 0.928, Robustness Region = [0: 3.51]; four picture test, interaction stimuli‐consistency‐type by learning condition (for low‐frequency items) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0099" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = –0.643, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0100" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.682, tails = 2, <emph>p</emph> = .346, Predicted Effect = 0.537, BF = 0.931, Robustness Region = [0:3.05]; four picture test, interaction recruitment‐type with learning condition by frequency <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0101" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.028, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0102" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.426, tails = 2, <emph>p</emph> = .984, Predicted Effect = 0.501, BF = 0.984, Robustness Region = [0: 4.03]; Four picture test, interaction recruitment‐consistency‐type by learning condition (for low‐frequency items) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0103" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.258, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0104" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.807, tails = 2, <emph>p</emph> = .119, Predicted Effect = 0.537, BF = 1.21, Robustness Region = [0:8.02]; Four label test, interaction stimuli‐consistency‐type with learning condition by frequency <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0105" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.863, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0106" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.281, tails = 2, <emph>p</emph> = .5, Predicted Effect = 0.501, BF = 0.96, Robustness Region = [0: 4.56]; four label test, interaction stimuli‐consistency‐type by learning condition (for low‐frequency items) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0107" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.065, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0108" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.743, tails = 2, <emph>p</emph> = .93, Predicted Effect = 0.537, BF = 0.811, Robustness Region = [0:2.11]; four label test, interaction recruitment‐type with learning condition by frequency <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0109" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.321, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0110" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.25, tails = 2, <emph>p</emph> = .797, Predicted Effect = 0.501, BF = 0.932, Robustness Region = [0:3.65]; four label test, interaction recruitment‐consistency‐type by learning condition (for low‐frequency items) <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0111" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.992, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0112" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.724, tails = 2, <emph>p</emph> = .171, Predicted Effect = 0.537, BF = 1.12, Robustness Region = [0:5.41].</bibtext> </blist> <blist> <bibtext> We also confirmed the modulating role of test‐type in this combined sample, that is, the results are different in the Four Pictures compared with Four Labels test: There is evidence that both the frequency by learning condition interaction and the simple effect are larger in the four pictures test than in the four labels test (interaction: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0113" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.391, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0114" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.647, tails = 1, <emph>p</emph> = .016, Predicted Effect = 0.501, BF = 3.412, Robustness Region = [0.44:3.97]; simple effect: <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0115" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mi&gt;&amp;#946;&lt;/mi&gt;&lt;annotation encoding="application/x-tex"&gt;$\beta$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 1.006, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:03640213:media:cogs13445:cogs13445-math-0116" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;E&lt;/mi&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$SE$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> = 0.388, tails = 1, <emph>p</emph> = .005, Predicted Effect = 0.537, BF = 10.443, Robustness Region = [0.2: 7.33).</bibtext> </blist> <blist> <bibtext> We added it at the end of the current experiment and report in the online Appendix C: Contingency and additional subset analysis.</bibtext> </blist> <blist> <bibtext> Note that more generally, the inclusion of robustness regions throughout the paper mitigates against a potential criticism of the Bayes Factor, which is that any choice of priors is subjective. Given the regions, a researcher who thinks a different choice of prior would be more appropriate can easily check whether this would affect our conclusions.</bibtext> </blist> </ref> <ref id="AN0177511222-52"> <title> References </title> <blist> <bibtext> Apfelbaum, K. S., &amp; McMurray, B. (2017). Learning during processing: Word learning doesn't wait for word recognition to finish. Cognitive Science, 41, 706 – 747.</bibtext> </blist> <blist> <bibtext> Arppe, A., Hendrix, P., Milin, P., Baayen, R. H., Sering, T., &amp; Shaoul, C. (2018). ndl: Naive discriminative learning. R package version 0.2.18. https://CRAN.R‐project.org/package=ndl.</bibtext> </blist> <blist> <bibtext> Baguley, T., &amp; Kaye, W. (2010). Review of: Understanding psychology as a science: An introduction to scientific and statistical inference, by Z. Dienes. British Journal of Mathematical and Statistical Psychology, 63 (3), 695 – 698.</bibtext> </blist> <blist> <bibtext> Barr, D. J., Levy, R., Scheepers, C., &amp; Tily, H. J. (2013). Random effects structure for confirmatory hypothesis testing: Keep it maximal. Journal of Memory and Language, 68 (3), 255 – 278.</bibtext> </blist> <blist> <bibtext> Bates, D., Mächler, M., Bolker, B., &amp; Walker, S. (2015). Fitting linear mixed‐effects models using lme4. Journal of Statistical Software, 67 (1), 1 – 48.</bibtext> </blist> <blist> <bibtext> De Leeuw, J. R. (2015). jspsych: A javascript library for creating behavioral experiments in a web browser. Behavior Research Methods, 47 (1), 1 – 12.</bibtext> </blist> <blist> <bibtext> Dienes, Z. (2008). Understanding psychology as a science: An introduction to scientific and statistical inference. Macmillan International Higher Education.</bibtext> </blist> <blist> <bibtext> Dienes, Z. (2014). Using Bayes to get the most out of non‐significant results. Frontiers in Psychology, 5, 781.</bibtext> </blist> <blist> <bibtext> Dienes, Z. (2016). How Bayes factors change scientific practice. Journal of Mathematical Psychology, 72, 78 – 89.</bibtext> </blist> <blist> <bibtext> Dye, M., Ramscar, M., &amp; Suh, E. (2011). For the price of a song: How pitch category learning comes at a cost to absolute frequency representations. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 33.</bibtext> </blist> <blist> <bibtext> Eitel, A., &amp; Scheiter, K. (2015). Picture or text first? Explaining sequence effects when learning with pictures and text. Educational Psychology Review, 27 (1), 153 – 180.</bibtext> </blist> <blist> <bibtext> Finley, A. J., &amp; Penningroth, S. L. (2015). Online versus in‐lab: Pros and cons of an online prospective memory experiment. Advances in Psychology Research, 113, 135 – 162.</bibtext> </blist> <blist> <bibtext> Gagné, N., &amp; Franzen, L. (2021). How to run behavioural experiments online: Best practice suggestions for cognitive psychology and neuroscience.</bibtext> </blist> <blist> <bibtext> Hoppe, D. B., van Rij, J., Hendriks, P., &amp; Ramscar, M. (2020). Order matters! Influences of linear order on linguistic category learning. Cognitive Science, 44 (11), e12910.</bibtext> </blist> <blist> <bibtext> Hupp, J. M., Sloutsky, V. M., &amp; Culicover, P. W. (2009). Evidence for a domain‐general mechanism underlying the suffixation preference in language. Language and Cognitive Processes, 24 (6), 876 – 909.</bibtext> </blist> <blist> <bibtext> Jaeger, T. F. (2008). Categorical data analysis: Away from ANOVAs (transformation or not) and towards logit mixed models. Journal of Memory and Language, 59 (4), 434 – 446.</bibtext> </blist> <blist> <bibtext> Kopp, B., &amp; Wolff, M. (2000). Brain mechanisms of selective learning: Event‐related potentials provide evidence for error‐driven learning in humans. Biological Psychology, 51 (2–3), 223 – 246.</bibtext> </blist> <blist> <bibtext> Kuznetsova, A., Brockhoff, P. B., &amp; Christensen, R. H. B. (2017). lmerTest package: Tests in linear mixed effects models. Journal of Statistical Software, 82 (13), 1 – 26.</bibtext> </blist> <blist> <bibtext> Morey, R. D., &amp; Lakens, D. (2016). Why most of psychology is statistically unfalsifiable.</bibtext> </blist> <blist> <bibtext> Nixon, J. S. (2020). Of mice and men: Speech sound acquisition as discriminative learning from prediction error, not just statistical tracking. Cognition, 197, 104081.</bibtext> </blist> <blist> <bibtext> Open Science Collaboration. (2015). Estimating the reproducibility of psychological science. Science, 349, aac4716. https://doi.org/10.1126/science.aac4716</bibtext> </blist> <blist> <bibtext> Powell, M. (2009). The bound optimization by quadratic approximation (BOBYQA) algorithm for bound constrained optimization without derivatives. Technical report, Cambridge, England.</bibtext> </blist> <blist> <bibtext> Qiu, M., &amp; Johns, B. T. (2020). Semantic diversity in paired‐associate learning: Further evidence for the information accumulation perspective of cognitive aging. Psychonomic Bulletin &amp; Review, 27 (1), 114 – 121.</bibtext> </blist> <blist> <bibtext> Ramscar, M. (2013). Suffixing, prefixing, and the functional order of regularities in meaningful strings. Psihologija, 46 (4), 377 – 396.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Dye, M., Gustafson, J. W., &amp; Klein, J. (2013). Dual routes to cognitive flexibility: Learning and response‐conflict resolution in the dimensional change card sort task. Child Development, 84 (4), 1308 – 1323.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Dye, M., Popick, H. M., &amp; O'Donnell‐McCarthy, F. (2011). The enigma of number: Why children find the meanings of even small number words hard to learn and how we can help them do better. PLoS One, 6 (7), e22501.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Hendrix, P., Love, B., &amp; Baayen, R. H. (2013). Learning is not decline: The mental lexicon as a window into cognition across the lifespan. Mental Lexicon, 8 (3), 450 – 481.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Hendrix, P., Shaoul, C., Milin, P., &amp; Baayen, H. (2014). The myth of cognitive decline: Non‐linear dynamics of lifelong learning. Topics in Cognitive Science, 6 (1), 5 – 42.</bibtext> </blist> <blist> <bibtext> Ramscar, M., &amp; McClure, S. M. (2011). Manipulating information structure as a method of localizing information processing in the brain. In Poster presented at the 18th Annual Meeting of the Cognitive Neuroscience Society; April 2011.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Sun, C. C., Hendrix, P., &amp; Baayen, H. (2017). The mismeasurement of mind: Life‐span changes in paired‐associate‐learning scores reflect the "cost" of learning, not cognitive decline. Psychological Science, 28 (8), 1171 – 1179.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Thorpe, K., &amp; Denny, K. (2007). Surprise in the learning of color words. In Proceedings of the Annual Meeting of the Cognitive Science Society, volume 29.</bibtext> </blist> <blist> <bibtext> Ramscar, M., Yarlett, D., Dye, M., Denny, K., &amp; Thorpe, K. (2010). The effects of feature‐label‐order and their implications for symbolic learning. Cognitive Science, 34 (6), 909 – 957.</bibtext> </blist> <blist> <bibtext> Recorla, R. A., &amp; Wagner, A. R. (1972). A Theory of Pavlovian Conditioning: Variations in the Effectiveness of Reinforcement and Nonreinforcement. In A. H. Black, &amp; W. F. Prokasy (Eds.), Classical Conditioning II: Current Research and Theory (pp. 64 – 99). New York : Appleton‐ Century‐Crofts.</bibtext> </blist> <blist> <bibtext> Rodd, J. (2019). How to maintain data quality when you can't see your participants. APS Observer, 32.</bibtext> </blist> <blist> <bibtext> Silvey, C., Dienes, Z., &amp; Wonnacott, E. (2021). Bayes factors for mixed‐effects models. PsyArXiv. https://doi.org/10.31234/osf.io/m4hju</bibtext> </blist> <blist> <bibtext> St Clair, M. C., Monaghan, P., &amp; Ramscar, M. (2009). Relationships between language structure and language learning: The suffixing preference and grammatical categorization. Cognitive Science, 33 (7), 1317 – 1329.</bibtext> </blist> <blist> <bibtext> R Core Team. (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing, Vienna, Austria. https://<ulink href="http://www.R-project.org/">www.R-project.org/</ulink></bibtext> </blist> <blist> <bibtext> Torralba, A., &amp; Oliva, A. (2003). Statistics of natural image categories. Network: Computation in Neural Systems, 14 (3), 391.</bibtext> </blist> <blist> <bibtext> Vujović, M. (2020). Exploring language learning as uncertainty reduction using artificial language learning. PhD thesis, University College London.</bibtext> </blist> <blist> <bibtext> Vujović, M., Ramscar, M., &amp; Wonnacott, E. (2021). Language learning as uncertainty reduction: The role of prediction error in linguistic generalization and item‐learning. Journal of Memory and Language, 119, 104231.</bibtext> </blist> </ref> <aug> <p>By Eva Viviani; Michael Ramscar and Elizabeth Wonnacott</p> <p>Reported by Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib32" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib33" firstref="ref6"></nolink> <nolink nlid="nl3" bibid="bib16" firstref="ref8"></nolink> <nolink nlid="nl4" bibid="bib26" firstref="ref9"></nolink> <nolink nlid="nl5" bibid="bib11" firstref="ref11"></nolink> <nolink nlid="nl6" bibid="bib29" firstref="ref12"></nolink> <nolink nlid="nl7" bibid="bib10" firstref="ref15"></nolink> <nolink nlid="nl8" bibid="bib15" firstref="ref16"></nolink> <nolink nlid="nl9" bibid="bib14" firstref="ref17"></nolink> <nolink nlid="nl10" bibid="bib20" firstref="ref18"></nolink> <nolink nlid="nl11" bibid="bib31" firstref="ref19"></nolink> <nolink nlid="nl12" bibid="bib25" firstref="ref21"></nolink> <nolink nlid="nl13" bibid="bib24" firstref="ref22"></nolink> <nolink nlid="nl14" bibid="bib36" firstref="ref23"></nolink> <nolink nlid="nl15" bibid="bib40" firstref="ref24"></nolink> <nolink nlid="nl16" bibid="bib39" firstref="ref36"></nolink> <nolink nlid="nl17" bibid="bib21" firstref="ref40"></nolink> <nolink nlid="nl18" bibid="bib12" firstref="ref41"></nolink> <nolink nlid="nl19" bibid="bib19" firstref="ref42"></nolink> <nolink nlid="nl20" bibid="bib35" firstref="ref55"></nolink> <nolink nlid="nl21" bibid="bib13" firstref="ref80"></nolink> <nolink nlid="nl22" bibid="bib34" firstref="ref81"></nolink> <nolink nlid="nl23" bibid="bib27" firstref="ref83"></nolink> <nolink nlid="nl24" bibid="bib28" firstref="ref84"></nolink> <nolink nlid="nl25" bibid="bib30" firstref="ref85"></nolink> <nolink nlid="nl26" bibid="bib23" firstref="ref87"></nolink> <nolink nlid="nl27" bibid="bib38" firstref="ref89"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1427185 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: The Effects of Linear Order in Category Learning: Some Replications of Ramscar et al. (2010) and Their Implications for Replicating Training Studies – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Eva+Viviani%22">Eva Viviani</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1330-0585">0000-0002-1330-0585</externalLink>)<br /><searchLink fieldCode="AR" term="%22Michael+Ramscar%22">Michael Ramscar</searchLink><br /><searchLink fieldCode="AR" term="%22Elizabeth+Wonnacott%22">Elizabeth Wonnacott</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3261-7131">0000-0002-3261-7131</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Cognitive+Science%22"><i>Cognitive Science</i></searchLink>. 2024 48(5). – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 29 – Name: DatePubCY Label: Publication Date Group: Date Data: 2024 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Symbolic+Learning%22">Symbolic Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Error+Patterns%22">Error Patterns</searchLink><br /><searchLink fieldCode="DE" term="%22Training%22">Training</searchLink><br /><searchLink fieldCode="DE" term="%22Replication+%28Evaluation%29%22">Replication (Evaluation)</searchLink><br /><searchLink fieldCode="DE" term="%22Pictorial+Stimuli%22">Pictorial Stimuli</searchLink><br /><searchLink fieldCode="DE" term="%22Error+Correction%22">Error Correction</searchLink><br /><searchLink fieldCode="DE" term="%22Prior+Learning%22">Prior Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Stimulus+Generalization%22">Stimulus Generalization</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/cogs.13445 – Name: ISSN Label: ISSN Group: ISSN Data: 0364-0213<br />1551-6709 – Name: Abstract Label: Abstract Group: Ab Data: Ramscar, Yarlett, Dye, Denny, and Thorpe (2010) showed how, consistent with the predictions of error-driven learning models, the order in which stimuli are presented in training can affect category learning. Specifically, learners exposed to artificial language input where objects preceded their labels learned the discriminating features of categories better than learners exposed to input where labels preceded objects. We sought to replicate this finding in two online experiments employing the same tests used originally: A four pictures test (match a label to one of four pictures) and a four labels test (match a picture to one of four labels). In our study, only findings from the four pictures test were consistent with the original result. Additionally, the effect sizes observed were smaller, and participants over-generalized high-frequency category labels more than in the original study. We suggest that although Ramscar, Yarlett, Dye, Denny, and Thorpe (2010) feature-label order predictions were derived from error-driven learning, they failed to consider that this mechanism also predicts that performance in any training paradigm must inevitably be influenced by participant prior experience. We consider our findings in light of these factors, and discuss implications for the generalizability and replication of training studies. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2024 – Name: AN Label: Accession Number Group: ID Data: EJ1427185 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1427185 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/cogs.13445 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 29 Subjects: – SubjectFull: Symbolic Learning Type: general – SubjectFull: Learning Processes Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Prediction Type: general – SubjectFull: Error Patterns Type: general – SubjectFull: Training Type: general – SubjectFull: Replication (Evaluation) Type: general – SubjectFull: Pictorial Stimuli Type: general – SubjectFull: Error Correction Type: general – SubjectFull: Prior Learning Type: general – SubjectFull: Stimulus Generalization Type: general Titles: – TitleFull: The Effects of Linear Order in Category Learning: Some Replications of Ramscar et al. (2010) and Their Implications for Replicating Training Studies Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Eva Viviani – PersonEntity: Name: NameFull: Michael Ramscar – PersonEntity: Name: NameFull: Elizabeth Wonnacott IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 05 Type: published Y: 2024 Identifiers: – Type: issn-print Value: 0364-0213 – Type: issn-electronic Value: 1551-6709 Numbering: – Type: volume Value: 48 – Type: issue Value: 5 Titles: – TitleFull: Cognitive Science Type: main |
| ResultId | 1 |