When Two Learners Are Better than One: Using Flashcards with a Partner Improves Metacognitive Accuracy
Saved in:
| Title: | When Two Learners Are Better than One: Using Flashcards with a Partner Improves Metacognitive Accuracy |
|---|---|
| Language: | English |
| Authors: | Megan N. Imundo (ORCID |
| Source: | Metacognition and Learning. 2025 20(1). |
| Availability: | Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ |
| Peer Reviewed: | Y |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Higher Education Postsecondary Education |
| Descriptors: | Undergraduate Students, Instructional Materials, Word Recognition, Paired Associate Learning, Cooperative Learning, Learning Strategies, Recall (Psychology), Learning Processes, Metacognition, Individual Activities, Study Habits, Test Preparation, Self Management |
| DOI: | 10.1007/s11409-024-09406-w |
| ISSN: | 1556-1623 1556-1631 |
| Abstract: | We investigated the benefits of two ways to use flashcards to perform retrieval practice: alone versus with a partner. In three experiments, undergraduate students learned word-definition pairs using flashcards alone (Individual condition) or with another student (Paired condition). Participants then made global judgments of learning (gJOLs; Experiments 1-3), and item-level judgments of learning (iJOLs; Experiment 3). Finally, participants took a cued-recall test after a 5-min delay (Experiments 1-3) and a 24-h delay (Experiments 2-3). In Experiment 1, students in the Paired condition dropped flashcards less often than in the Individual condition (dropping was prohibited in Experiments 2-3). In addition, although final test performance tended to be similar across conditions, inaccurate gJOLs for the immediate test--inflated by ~ 20% relative to actual immediate test performance--were common in the Individual condition but not in the Paired condition in Experiments 1-2. In Experiment 3, we tested whether this difference in metacognitive calibration was due to the Paired condition requiring overt retrieval by instructing participants in the Individual condition to retrieve out loud. With this change, participants in the Individual and Paired conditions reported similarly accurate gJOLs and iJOLs. Taken together, these findings suggest that although performing retrieval practice with flashcards alone versus with a partner yields comparable amounts of learning, doing so with a partner can increase metacognitive accuracy, a benefit possibly driven by the facilitation of overt retrieval. Overall, these findings have implications for self-regulated learning and effective exam preparation. |
| Abstractor: | As Provided |
| Notes: | https://osf.io/hac38/?view_only=9ff44668b2184f1b8e412452f3a41640 |
| Entry Date: | 2024 |
| Accession Number: | EJ1452322 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwG1Jn5aH-kBut2-t3JDRXMPAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDO8nE8lAtuTUL1s5hwIBEICBm7uKLn_q85Opt_UIsAwogDe_ghSyARPsTvQq_SMZNvYoOWzqzGMg3DqQWptvJnqBAT6Eff20ge6EkWk-uF6wQ6vkQcQIQdrpgOl0B7HK0zVLB6jx45UNHwnRnMo2W23plngmXP8mBSzbOQBNpR0s6rC6EyHT5oREh17iR-JExIbTzhct0rsX7GbHhtBjLBnPY3K8QWjh0EbMapS_ Text: Availability: 1 Value: <anid>AN0181532824;[3d0h]01jan.25;2025Feb17.13:03;v2.2.500</anid> <title id="AN0181532824-1">When two learners are better than one: using flashcards with a partner improves metacognitive accuracy: When two learners are better than one: using flashcards with a partner improves metacognitive accuracy: M. N. Imundo et al </title> <p>We investigated the benefits of two ways to use flashcards to perform retrieval practice: alone versus with a partner. In three experiments, undergraduate students learned word-definition pairs using flashcards alone (Individual condition) or with another student (Paired condition). Participants then made global judgments of learning (gJOLs; Experiments 1–3), and item-level judgments of learning (iJOLs; Experiment 3). Finally, participants took a cued-recall test after a 5-min delay (Experiments 1–3) and a 24-h delay (Experiments 2–3). In Experiment 1, students in the Paired condition dropped flashcards less often than in the Individual condition (dropping was prohibited in Experiments 2–3). In addition, although final test performance tended to be similar across conditions, inaccurate gJOLs for the immediate test—inflated by ~ 20% relative to actual immediate test performance—were common in the Individual condition but not in the Paired condition in Experiments 1–2. In Experiment 3, we tested whether this difference in metacognitive calibration was due to the Paired condition requiring overt retrieval by instructing participants in the Individual condition to retrieve out loud. With this change, participants in the Individual and Paired conditions reported similarly accurate gJOLs and iJOLs. Taken together, these findings suggest that although performing retrieval practice with flashcards alone versus with a partner yields comparable amounts of learning, doing so with a partner can increase metacognitive accuracy, a benefit possibly driven by the facilitation of overt retrieval. Overall, these findings have implications for self-regulated learning and effective exam preparation.</p> <p>Copyright comment Springer Nature or its licensor (e.g. a society or other partner) holds exclusive rights to this article under a publishing agreement with the author(s) or other rightsholder(s); author self-archiving of the accepted manuscript version of this article is solely governed by the terms of such publishing agreement and applicable law.</p> <p>Learning scientists often recommend that students use flashcards to prepare for exams (e.g., Smith &amp; Weinstein, [<reflink idref="bib47" id="ref1">47</reflink>]). This suggestion is based on the premise that flashcards facilitate <emph>retrieval practice</emph> (i.e., practice testing), which is a potent enhancer of long-term memory (i.e., the <emph>testing effect</emph>; Pan &amp; Rickard, [<reflink idref="bib36" id="ref2">36</reflink>]; Roediger &amp; Butler, [<reflink idref="bib42" id="ref3">42</reflink>]; Rowland, [<reflink idref="bib44" id="ref4">44</reflink>] offer comprehensive reviews). Indeed, an in-depth review of popular learning techniques ranked retrieval practice as among the most effective (Dunlosky et al., [<reflink idref="bib5" id="ref5">5</reflink>]). Large surveys indicate that most undergraduate students use flashcards to prepare for their classes and often engage in retrieval practice when doing so, with the most common purpose being to learn vocabulary (Wissman et al., [<reflink idref="bib55" id="ref6">55</reflink>]; Zung et al., [<reflink idref="bib57" id="ref7">57</reflink>]). Flashcards are commonly prepared by writing a key concept or term on one side and associated information (e.g., related concepts, definitions, etc.) on the reverse, thus making it convenient to quiz oneself or others.</p> <p>Beyond its benefits for memory, retrieval practice can also aid learning in other, less obvious ways. One such benefit involves improving students' control of study behaviors (e.g., time per item, decision to stop studying) during self-regulated learning. According to prominent theories of metacognition (e.g., Nelson &amp; Narens, [<reflink idref="bib33" id="ref8">33</reflink>]), such control is commonly based on students' monitoring of their own learning (e.g., judgments of learning, confidence in retrieved answers). If a student inaccurately monitors her learning and is overconfident, then she may stop studying prematurely and be left with poor mastery of to-be-learned information. Retrieval practice can prevent that overconfidence: Miller and Geraci ([<reflink idref="bib30" id="ref9">30</reflink>]) found that a single retrieval practice opportunity, which usually provides learners with concrete evidence as to their mastery of the material (e.g., via retrieval success or failure), can lower inflated judgments of learning (also Tullis et al., [<reflink idref="bib51" id="ref10">51</reflink>]). Retrieval practice can also help students optimize their study activities: Soderstrom and Bjork ([<reflink idref="bib48" id="ref11">48</reflink>]) found that students spend more time studying difficult materials, and learn them more effectively, after engaging in retrieval practice. These findings reinforce the value of retrieval practice as not just a memory enhancer, but also a way to improve metacognitive accuracy and study decisions. It should be noted, however, that such benefits have typically been demonstrated using methods that do not involve flashcards.</p> <hd id="AN0181532824-2">Optimizing flashcard-based retrieval practice</hd> <p>Although flashcards can facilitate retrieval practice, the conditions under which they are most effective remains to be fully established (Lin et al., [<reflink idref="bib26" id="ref12">26</reflink>]; Pan et al., [<reflink idref="bib37" id="ref13">37</reflink>]; Senzaki et al., [<reflink idref="bib45" id="ref14">45</reflink>]; Zung et al., [<reflink idref="bib57" id="ref15">57</reflink>] offer additional discussion), and there is evidence that students use flashcards ineffectively and remain susceptible to illusions of competence when doing so. For instance, students may choose to download premade flashcard sets, even though generating flashcards can facilitate learning (Pan et al., [<reflink idref="bib37" id="ref16">37</reflink>]). Students also often drop flashcards before their content is well-learned: Kornell and Bjork ([<reflink idref="bib20" id="ref17">20</reflink>]) found that dropping is common after just one correct retrieval attempt, resulting in reduced learning relative to conditions wherein dropping is disallowed. Further, students prefer smaller flashcard stacks, thinking that they are more beneficial for learning (Wissman et al., [<reflink idref="bib55" id="ref18">55</reflink>]), when larger stacks enable learning to be better distributed out in time (i.e., the <emph>spacing effect</emph>; Kornell, [<reflink idref="bib19" id="ref19">19</reflink>]). Finally, one-third of students do not always check the accuracy of their responses when using flashcards (Wissman et al., [<reflink idref="bib55" id="ref20">55</reflink>]). This pattern is especially problematic when considering that students sometimes drop flashcards even before a single successful retrieval (possibly due to inadequately assessing the correctness of their responses; e.g., Kornell &amp; Bjork, [<reflink idref="bib20" id="ref21">20</reflink>], Experiment 3). Together, these findings reveal substantial room for improvement in students' use of flashcards.</p> <p>One promising method for improving flashcard use involves doing so with a partner—that is, using flashcards in pairs as opposed to individually. There are several reasons why using flashcards in pairs may be beneficial. First, is the need for overt responses. Some studies of the testing effect find that overt responding more reliably produces a testing effect than covert responding (Jönsson et al., [<reflink idref="bib17" id="ref22">17</reflink>]; Krumboltz &amp; Weisman, [<reflink idref="bib22" id="ref23">22</reflink>]; Kubik et al., [<reflink idref="bib23" id="ref24">23</reflink>]), whereas others <emph>do</emph> observe a testing effect following covert retrieval practice (Carpenter &amp; Pashler, [<reflink idref="bib4" id="ref25">4</reflink>]), or observe no difference in the testing effect between overt and covert retrieval practice (Putnam &amp; Roediger, [<reflink idref="bib39" id="ref26">39</reflink>]). There is not yet an established explanation as to why overt retrieval practice may at times lead to greater learning than covert retrieval practice. Jönsson and colleagues (2014) suggest that overt retrieval might elicit greater processing as it requires both the attempt to retrieve the target content and the overt expression of that content. Tauber and colleagues (2018) suggest that overt retrieval practice establishes accountability for the learner. This accountability may encourage more complete retrieval attempts, especially when material is complex enough (i.e., more than single words) to allow for exhaustive retrieval. Thus, overt retrieval may discourage learners from "cheating themselves" by not fully articulating a response to a given question or cue, with retrieval attempts more potent and more informative for metacognitive judgments as a result.</p> <p>Second, the presence of others may affect learners' emotional states positively. Supporting evidence comes from students' self-reports which indicate that studying with others increases motivation to learn, is more enjoyable, and improves learning relative to studying individually (McCabe &amp; Lummis, [<reflink idref="bib28" id="ref27">28</reflink>]; Wissman &amp; Rawson, [<reflink idref="bib54" id="ref28">54</reflink>]). One review of the collaborative testing literature, for example, indicates that testing with peers might reduce test anxiety (LoGuidice et al., [<reflink idref="bib27" id="ref29">27</reflink>]). Evidence suggests that experiencing positive emotions can contribute to learning outcomes (Holzer et al., [<reflink idref="bib13" id="ref30">13</reflink>]; Pekrun et al., [<reflink idref="bib38" id="ref31">38</reflink>]). Affect also has implications for a learner's metacognitive experiences (Efklides, [<reflink idref="bib6" id="ref32">6</reflink>]). The Metacognitive and Affective Model of Self-Regulated Learning (Efklides et al., [<reflink idref="bib7" id="ref33">7</reflink>]; also Hayat et al., [<reflink idref="bib12" id="ref34">12</reflink>]), for example, suggests that metacognition includes an affective component and metacognition both affects and is affected by positive and negative emotions. Further, the presence of others may increase motivation during learning (i.e., social facilitation); however, if students fear evaluation from their partner, then their learning may suffer (Geen, [<reflink idref="bib10" id="ref35">10</reflink>]).</p> <p>Third, learners may seek feedback from their partner rather than assessing the validity of their response via a sense of fluency, thus reducing susceptibility to illusions of competence. Additionally, a partner might offer explanations and correct errors, further facilitating learning (Johnson et al., [<reflink idref="bib16" id="ref36">16</reflink>]; LoGuidice et al., [<reflink idref="bib27" id="ref37">27</reflink>]). All of these reasons suggest that using flashcards with a partner—which has yet to be extensively investigated—may be beneficial.</p> <hd id="AN0181532824-3">The present study</hd> <p>We investigated the hypothesis that flashcard-based retrieval practice with a partner is better for learning than individual flashcard-based retrieval practice. Additionally, in an exploratory manner, we examined if flashcard-based retrieval practice leads participants to report more accurate post-learning metacognitive judgments when it is implemented with a partner as opposed to implemented individually. In a similarly exploratory approach, we also examined potential differences between individual and paired flashcard learning in terms of the mechanics of flashcard use (e.g., cycles through the flashcard set), associated study decisions (e.g., dropping cards), and affective states.</p> <p>Across three experiments, undergraduate students learned word-definition pairs using flashcards alone (the Individual condition) or with another student (the Paired condition), answered relevant survey questions, and then completed a final test. In Experiment 1, dropping of flashcards was allowed whereas in Experiments 2–3 it was prohibited. Additionally, whereas in Experiment 1 both Individual and Paired learners engaged in cycles of study and retrieval practice, in Experiments 2–3, all learners completed an initial study period such that Individual learners then only engaged in retrieval practice. We believe that this approach is more aligned with students' own behaviors when using flashcards in daily life (Zung et al., [<reflink idref="bib57" id="ref38">57</reflink>]). Finally, in Experiment 3 participants in the Individual condition were instructed to overtly retrieve (i.e., talk out loud) during the flashcard phase. Importantly, in all experiments and across conditions, we controlled for total learning time, used the same flashcards and learning environments, and gave similar instructions.</p> <hd id="AN0181532824-4">Experiment 1</hd> <p>In the first experiment, learners had 20 min each to study a set of vocabulary-definition pairs and to perform retrieval practice on those pairs. They were assigned to do so by themselves or with a partner. In the case of Individual learners, such learning involved 20 min of studying followed by 20 min of retrieval practice. For Paired learners, the logistics were somewhat more complex: One partner served as the "tester" and the other partner as the "testee" before the roles reversed. Hence, in the Paired condition, one partner engaged in 20 min of practice testing from the outset, whereas the other partner did so after those 20 min had elapsed.</p> <hd id="AN0181532824-5">Method</hd> <p>The study was preregistered at: https://osf.io/mqunz/?view_only=bdb8d5cce52c43a6ba400a58a70749f5</p> <hd id="AN0181532824-6">Participants</hd> <p>One hundred and fifty-two undergraduate students (Individual condition, <emph>n</emph> = 64; Paired condition, <emph>n</emph> = 88) from the participant pool at a large public research university participated in exchange for course credit. Data from two additional participants were excluded because they experienced technical malfunctions. The target sample size, 150, was determined using an a priori power analysis conducted in G*Power (Faul et al., [<reflink idref="bib8" id="ref39">8</reflink>]) in which at least 32 participants per group is needed to detect a medium effect size (Cohen's <emph>f</emph> = 0.25) in a between-participants design at 80% power and with a standard 0.05 error probability. To reach that target, data collection occurred continuously for eight weeks and concluded only with the scheduled close of the participant pool recruitment period.</p> <hd id="AN0181532824-7">Design</hd> <p>The experiment employed a 2 × 2 between-subjects factorial design with factors of Condition (Individual versus Paired) and First Learning Activity (Study First versus Test First; detailed later in this manuscript). Participants (a) learned individually or in pairs and (b) studied or tested first before switching learning activities.</p> <hd id="AN0181532824-8">Materials</hd> <p>The materials included 40 word-definition pairs, each consisting of a Graduate Record Examination (GRE) vocabulary word and its definition (e.g., <emph>monolithic</emph>: <emph>made of only one stone</emph>). The words were drawn from <emph>The Economist</emph>'s "Most Difficult GRE Words" list for 2020, whereas the definitions were drawn from Dictionary.com. The words and their definitions were 4–10 letters and 5–10 words in length, respectively; the words had a Kucera-Francis frequency of 1–3. In the case of multiple definitions, the first definition was used, and if that definition contained the GRE word, the second definition was used. All stimuli are listed in the Appendix.</p> <p>Each word-definition pair was printed on a 4 × 6 in. white index flashcard. For the <emph>standard</emph> flashcard set, which was designed for retrieval practice, each card displayed a GRE word on the front and the word and its definition on the back. For the <emph>study-only</emph> flashcard set, which was designed for studying, each card displayed a GRE word and its definition on the front and the back was blank. All text was printed in Times New Roman size 24 font, with the GRE words bolded. There were 40 cards per flashcard set, with one card per word-definition pair.</p> <hd id="AN0181532824-9">Global Judgments of Learning (gJOL)</hd> <p>Participants were asked to predict their performance on the immediate test: "If, in a few minutes, you were shown the definitions you just studied, for what percentage (%) of these definitions are you confident you could remember the corresponding word?" The gJOL was open response from 0 to 100%.</p> <hd id="AN0181532824-10">Procedure</hd> <p>The experiment was run in 2-h timeslots involving up to four participants each and using three nearly-identical laboratory testing rooms (see Fig. 1). All participants were told that they would be learning vocabulary words using flashcards, and all flashcards were randomly shuffled prior to each timeslot. Each Individual learner completed the experiment in a separate testing room, whereas the two Paired learners per timeslot did so in a shared testing room.</p> <p>Graph: Fig. 1 Procedure used in Experiment 1</p> <p>The experiment consisted of four phases. All participants first completed a flashcard phase in which they learned and practiced challenging vocabulary-definition pairs. Then they completed a series of survey questions—which included providing a global judgment of learning (gJOL)—, a distractor task, and a final cued-recall test.</p> <hd id="AN0181532824-11">Random Assignment and Counterbalancing</hd> <p>Within each timeslot, two participants were randomly assigned to the Paired condition and up to two participants were randomly assigned to the Individual condition. When fewer than four participants signed up for a timeslot, two were assigned to the Paired condition (if possible) and any others were randomly assigned to the Individual condition. The decision to prioritize filling the Paired condition occurred prior to data collection and stemmed from the inherent challenge of bringing two participants together in one timeslot to run that condition (it also maintained random assignment and was consistently applied by all experimenters, thus reducing potential bias). A moderate imbalance in sample size per condition resulted.</p> <p>Given that using flashcards in pairs entails one person being tested at a time and the other person viewing (i.e., studying) the answers while administering the tests, participants' engagement in studying or testing from the outset of the experiment (before switching activities, which resembles using flashcards across separate study and test phases) was counterbalanced. Thus, task order (i.e., First Learning Activity) was equated across both conditions.</p> <hd id="AN0181532824-12">Flashcard Phase</hd> <p></p> <hd id="AN0181532824-13">Individual condition</hd> <p>The experimenter seated each participant in a testing room, distributed the study-only or standard flashcard set and, depending on the given set, instructed them to learn the words via studying (i.e., reading) or testing (i.e., retrieval practice). Participants were permitted to cycle through the set as many times as desired and in any order for 20 min. Skipping or dropping flashcards was allowed but not specifically discussed. Afterwards, the flashcard set was replaced (i.e., the standard set was switched for the study-only set, or vice versa) and participants were instructed to use the new set for another 20 min. Hence, equal amounts of time were spent engaged in studying and testing.</p> <hd id="AN0181532824-14">Paired condition</hd> <p>Participants were seated face-to-face at a small table on which the standard flashcard set was placed. The experimenter demonstrated how the flashcards were to be used. One participant (the "tester") was to hold up each flashcard with the word-only side facing the other participant (the "testee") and read the word and definition silently as the "testee" attempted to verbally provide a definition. After the "testee" indicated that they had finished their attempt, the "tester" was to reverse the card to reveal the definition. Participants proceeded accordingly for 20 min, during which they were permitted to cycle through the set as many times as desired and in any order. As in the Individual condition, dropping was allowed by either the tester or testee but was not explicitly discussed. Verbal feedback was disallowed to minimize off-task conversations and to ensure that participants in the Paired condition did not have an unfair advantage over participants in the Individual condition due to receiving personalized or elaborative feedback. After 20 min, the experimenter directed participants to switch roles and continue for another 20 min. Thus, equal amounts of time were spent engaged in studying (as the "tester") and testing (as the "testee").</p> <hd id="AN0181532824-15">Survey and Distractor Task</hd> <p>After the flashcard phase, participants used desktop computers to (a) answer demographic questions, (b) complete the Positive and Negative Affect Schedule—Short Form (PANAS; Watson et al., [<reflink idref="bib53" id="ref40">53</reflink>]), (c) provide a global Judgment of Learning (gJOL), (d) report their level of attentional focus from 0–100%, and (e) answer questions regarding their activities during the flashcard phase and their own flashcard use in everyday learning sessions. The exact wording for each of these items and the survey items used in the subsequent experiments are available at https://osf.io/hac38/?view%5fonly=9ff44668b2184f1b8e412452f3a41640. Participants then completed a 5-min distractor task during which they solved anagrams.</p> <hd id="AN0181532824-16">Final Cued-Recall Test</hd> <p>During the final cued-recall test, each of the 40 definitions were presented individually and in a random order for 60 s. Participants attempted to type the matching GRE word (similar to Pan &amp; Rickard, [<reflink idref="bib35" id="ref41">35</reflink>]). The experiment concluded afterwards.</p> <hd id="AN0181532824-17">Results</hd> <p>Data for all experiments are available at https://osf.io/hac38/?view_only=9ff44668b2184f1b8e412452f3a41640.</p> <p>All analyses were conducted using independent samples t-tests with equal variances assumed unless otherwise noted. In all analyses, α was set at 0.05. The sample sizes per analysis differed slightly in some cases as some participants declined to answer all questions. In a parallel set of analyses not reported here, the effect of First Learning Activity—that is, whether a participant had engaged in studying prior to testing, or vice versa—was not significant on any aspect of measured behavior during the learning or final test phases. Those patterns were unsurprising given that such effects were potentially eclipsed by subsequent cycles of testing and studying. Consequently, all analyses reported here involve data collapsed across First Learning Activity. Parallel analyses that do not do so are included in the supplemental online materials https://osf.io/hac38/?view_only=9ff44668b2184f1b8e412452f3a41640.</p> <hd id="AN0181532824-18">Flashcard phase</hd> <p></p> <hd id="AN0181532824-19">Number of learning cycles</hd> <p>Participants indicated the number of learning cycles (i.e., practicing through the entire flashcard set) they completed per 20-min period. These responses were summed for a total number of cycles in the entire flashcard phase; if participants indicated an incomplete cycle, then 0.50 was added (this method, albeit somewhat imprecise, was consistently applied across conditions). Individual learners typically completed one more learning cycle (<emph>M</emph> = 5.36, <emph>SD</emph> = 1.64) across the entire flashcard phase than did Paired learners (<emph>M</emph> = 4.32, <emph>SD</emph> = 1.44). This difference was significant, <emph>t</emph> (<reflink idref="bib150" id="ref42">150</reflink>) = 4.16, <emph>p</emph> &lt; 0.001, <emph>d</emph> = 0.68, 95% CI [0.55, 1.54].</p> <hd id="AN0181532824-20">Dropping of flashcards</hd> <p>Participants reported whether they had dropped flashcards from study, and if so, why they chose to do so. These data were coded by two independent raters blind to condition (with interrater reliabilities of Cohen's <emph>κ</emph> = 0.99 and 0.85 for if they dropped and why, respectively). A Chi-square test revealed that significantly more Individual learners (53%) dropped flashcards from study than Paired learners (5%), <emph>χ</emph><sups><emph>2</emph></sups> (<reflink idref="bib2" id="ref43">2</reflink>) = 44.43, <emph>p</emph> &lt; 0.001. Fifty-eight percent of all participants who dropped a flashcard from study did so because they believed that they had learned the word-definition pair, 32% did so because they deemed the pair too difficult to learn, and 11% did so for other reasons. As only four Paired learners dropped flashcards, formal comparisons of reasons for dropping between conditions were not possible. Those four participants, however, all dropped cards because they deemed materials too difficult to learn, whereas only 24% of Individual learners dropped flashcards for that reason (most did so based on sufficient learning).</p> <hd id="AN0181532824-21">Final cued-recall test</hd> <p></p> <hd id="AN0181532824-22">Overall performance</hd> <p>Given the difficulty of the GRE words, we used an accuracy threshold wherein final test responses had to match the actual spelling by ≥ 75% to be counted as correct. Corresponding analyses under strict scoring (i.e., perfect spelling) yielded the same patterns (available in online supplemental materials at https://osf.io/hac38/?view%5fonly=9ff44668b2184f1b8e412452f3a41640). Contrary to our hypothesis that Paired flashcard learning would yield higher test performance than Individual flashcard learning, final test performance was not significantly different between the Individual and Paired conditions, <emph>t</emph> (<reflink idref="bib150" id="ref44">150</reflink>) = 1.29, <emph>p</emph> = 0.20, <emph>d</emph> = 0.21, 95% CI [−0.03, 0.13]. This result indicates that recall of the GRE words was no different shortly after individual or paired flashcard learning (Table 1 presents the descriptive statistics for each condition).</p> <p>Table 1 Cued-recall test performance in Experiments 1–3</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Condition&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Experiment 1&lt;sup&gt;a&lt;/sup&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="4"&gt;&lt;p&gt;Experiment 2&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="4"&gt;&lt;p&gt;Experiment 3&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Immediate Test&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Immediate Test&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Delayed Test&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Immediate Test&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Delayed Test&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Individual&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.48&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.24&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.49&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.28&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.40&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.28&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.42&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.26&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.36&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.23&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Paired&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.43&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.23&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.44&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.25&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.35&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.24&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.37&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.24&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.34&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.22&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p> <sups>a</sups>Only an immediate cued-recall test was administered in Experiment 1</p> <hd id="AN0181532824-23">Metacognitive judgments</hd> <p></p> <hd id="AN0181532824-24">Correlations with Final Test Performance</hd> <p>To examine whether there was a significant relationship between participants' own assessment of their learning and their actual test score, we conducted a series of exploratory bivariate correlations relating gJOL and final test performance for both conditions (see Fig. 2). Individual learners demonstrated moderate-to-large correlations between their gJOL and final test performance when learning individually, <emph>r</emph> (<reflink idref="bib61" id="ref45">61</reflink>) = 0.59, <emph>p</emph> &lt; 0.001, as did Paired learners, <emph>r</emph> (<reflink idref="bib86" id="ref46">86</reflink>) = 0.60, <emph>p</emph> &lt; 0.001. These correlations suggest that participants engaged in appropriate metacognitive monitoring, with participants who reported greater gJOLs tending to score higher on the cued-recall test afterward. Although the magnitude of the relationships between gJOL and test performance was similar between the Paired and Individual conditions, Fig. 2 clearly shows that the intercepts of the regression lines between the two conditions (computed by regressing test performance onto gJOL data) differ, prompting further analyses of participants' metacognitive calibration.</p> <p>Graph: Fig. 2 Metacognitive calibration demonstrated by those in the Individual flashcard learning and the Paired flashcard learning conditions. Each panel displays the correlation between final test performance and global judgments of learning (gJOLs). A dotted line represents the hypothetical case of perfect calibration between gJOLs and test scores; crucially, participants in the Individual condition tended to substantially overestimate their final test performance</p> <hd id="AN0181532824-25">Metacognitive calibration</hd> <p>Metacognitive calibration is a form of absolute metacognitive accuracy; i.e., the extent to which a learner can accurately estimate their learning. Here, metacognitive calibration was measured as the difference in size between their predicted test performance and their actual test performance. We computed metacognitive calibration by subtracting participants' actual test performance from their gJOLs, with positive scores indicating overconfidence and negative scores indicating underconfidence. Unlike the previous analyses, metacognitive calibration provides evidence for the direction of participants' judgment errors (e.g., if one condition tends to exhibit overestimation and the other condition tends to exhibit underestimation, then their average calibration will differ even if their correlation coefficients are similar). Thus, gJOL-test performance correlations and metacognitive calibration scores provide complementary, but distinct, information about learners' metacognitive judgments.</p> <p>An exploratory independent samples t-test compared Individual and Paired learners' metacognitive calibration scores (Fig. 2). Individual learners were overconfident (<emph>M</emph> = 0.20, <emph>SD</emph> = 0.22), whereas Paired learners were relatively accurate (<emph>M</emph> = 0.00, <emph>SD</emph> = 0.22),<emph> t</emph> (<reflink idref="bib149" id="ref47">149</reflink>) = 5.66, <emph>p</emph> &lt; 0.001, 95% CI [0.13, 0.28].</p> <hd id="AN0181532824-26">Positive and negative affect</hd> <p>We conducted separate analyses for the positive affect and negative affect subscales of the PANAS. Participants reported comparable positive affect in the Individual (<emph>M</emph> = 26.05, <emph>SD</emph> = 7.54) and Paired (<emph>M</emph> = 26.28, <emph>SD</emph> = 8.12) conditions, <emph>t</emph> (<reflink idref="bib150" id="ref48">150</reflink>) = −0.18, <emph>p</emph> = 0.86, <emph>d</emph> = 0.03, 95% CI [−2.80, 2.32]. Those who studied in pairs, however, reported significantly higher negative affect (<emph>M</emph> = 16.93, <emph>SD</emph> = 6.54) than those who studied individually (<emph>M</emph> = 14.20, <emph>SD</emph> = 3.52), <emph>t</emph> (<reflink idref="bib150" id="ref49">150</reflink>) = −3.03,<emph> p</emph> = 0.003, <emph>d</emph> = 0.50, 95% CI [−4.51, −0.95].</p> <hd id="AN0181532824-27">Attentional focus</hd> <p>Self-reported percentage time focused during the experimental tasks did not significantly differ between the Individual (<emph>M</emph> = 78.4%, <emph>SD</emph> = 16.5%) and Paired (<emph>M</emph> = 80.8%, <emph>SD</emph> = 19.1%) conditions, <emph>t</emph> (<reflink idref="bib150" id="ref50">150</reflink>) = −0.78, <emph>p</emph> = 0.44, <emph>d</emph> = 0.13, 95% CI [−8.18, 3.55].</p> <hd id="AN0181532824-28">Experiment 1 discussion</hd> <p>With respect to effects on memory, the results of Experiment 1 suggest that collaborative and individual practice of difficult vocabulary-definition pairs using flashcards yield comparable test performance after a 5-min delay. These results were contrary to our predictions. It is, however, possible that the delay between the learning and test phases was not long enough to observe the benefits of collaborative practice. In line with the framework of desirable difficulties (Bjork, 1994), the benefits of more challenging but potentially beneficial learning activities are often observed at a delay (e.g., Roediger &amp; Karpicke, 2006). It is thus possible that the positive effects of more effortful or complete retrieval encouraged by paired flashcard practice testing may emerge on a delayed test.</p> <p>There were, however, some benefits of collaborative practice that may be particularly meaningful for learners engaging in self-regulated study. Paired learners were far less likely to drop cards from study than Individual learners. Moreover, prediction errors of test performance from Paired learners did not exhibit a systematic bias whereas Individual learners on average overestimated their learning by approximately 20%. Possibly, these two results are related: If Paired learners were more metacognitively accurate during the flashcard phase of the study than Individual learners, they may have been less likely to prematurely drop cards from study. Vice versa, if Paired learners were less likely to drop cards from study than Individual learners for other reasons (perhaps because their partner was holding the flashcard deck, adding friction to the drop decision, or because they were instructed to limit discussion with their partner during the flashcard learning phase), their metacognitive judgments may have benefited from relatively equal time spent on each vocabulary term. In our view, it is crucial to ascertain whether the metacognitive calibration benefit in the Paired condition is merely a result of lower rates of dropping flashcards, which we address in the second experiment.</p> <p>Finally, the effect of First Learning Activity (i.e., whether a participant had engaged in studying prior to testing, or vice versa) did not significantly impact any aspect of behavior during the learning or final test phases, possibly because any such effects were eclipsed by subsequent cycles of testing and studying. From an ecological validity standpoint, requiring that students first study and then test themselves (or vice versa) seems at odds with the common view of flashcards as a retrieval practice tool. Additionally, the effects of collaboration on learning often have been examined within the context of testing on previously studied content, and are therefore often compared to individual testing (e.g., Barber et al., [<reflink idref="bib1" id="ref51">1</reflink>]; Gilley &amp; Clarkston, [<reflink idref="bib11" id="ref52">11</reflink>]; Imundo, [<reflink idref="bib14" id="ref53">14</reflink>]). It may therefore be more appropriate to compare the effects of paired flashcard practice to the effects of individual retrieval practice with flashcards.</p> <hd id="AN0181532824-29">Experiment 2</hd> <p>Experiment 2 again compared the effects of individual versus paired flashcard use on learning. To examine if there might be a benefit of paired practice over individual practice for long-term learning, a 24-h delayed test was added. We chose a 24-h delay because delays of one day or longer tend to yield stronger testing effects than delays occurring within the same day (Rowland, [<reflink idref="bib44" id="ref54">44</reflink>]). To rule out the possibility that Paired learners are more metacognitively accurate simply due to lower rates of dropping flashcards from study, dropping flashcards from study was explicitly prohibited in Experiment 2. Additionally, to increase participants' ease in interacting with one another in the Paired condition, a brief icebreaker activity prior to the flashcard portion was incorporated. Finally, as the effect of First Learning Activity (i.e., whether a participant had engaged in studying prior to testing, or vice versa) did not significantly impact any aspect of behavior during the flashcard phase or final test performance, First Learning Activity was removed as a factor and a period of initial study of the vocabulary-definition pairs prior to the flashcard phase was added.</p> <hd id="AN0181532824-30">Method</hd> <p>Experiment 2 was not preregistered.</p> <hd id="AN0181532824-31">Participants</hd> <p>One hundred and forty-one participants were included in this study (Individual: <emph>n</emph> = 78, Paired: <emph>n</emph> = 63). An additional thirty participants were recruited for this study but were excluded due to technical issues or experimenter error (<emph>n</emph> = 4), for failing to follow instructions (<emph>n</emph> = 11; e.g., did not practice test the entire time), or for reporting that they dropped flashcards from study during that phase (<emph>n</emph> = 15).</p> <p>Of the 141 participants in the final sample, all reported an immediate gJOL and 84 (59.6%) offered a delayed gJOL. Six (4.3%) participants did not report a delayed JOL because they did not complete the delayed test portion of the study. An additional 50 participants (35.5%) took the delayed test but chose not to offer a gJOL (in accordance with our IRB protocol, participants were not required to answer every question).[<reflink idref="bib1" id="ref55">1</reflink>]Finally, one participant (0.7%) mistakenly reported that they were participating in Session 1 (rather than Session 2) of the study when inputting their information into the delayed test link such that the page prompting participants for a gJOL did not appear.</p> <hd id="AN0181532824-32">Design</hd> <p>Experiment 2 employed a 2 × 2 mixed factorial design with Condition (Individual or Paired) as the between-subjects factor and Test Delay (5-min or 24-h) as the within-subjects factor. The 40 word-definition pairs used in this study were divided into two sets of 20 pairs (i.e., Set A and Set B): One set was used for the immediate test and one set was used for the 24-h delayed test, counterbalanced across participants by time slot. Although First Learning Activity was not manipulated for the Individual condition in this experiment and was not included in any subsequent statistical models, the nature of the Paired condition required that one member of the pair act as the tester first and one member of the pair act as the testee first.</p> <hd id="AN0181532824-33">Materials</hd> <p>The materials used in Experiment 2 were identical to the materials used in Experiment 1 except that only the standard flashcard set was used. Given a change in the software used to run the final test portion of the study (more details below), the cued-recall test was scored by two independent raters. Interrater reliability for all cued-recall test items was adequate (Cohen's <emph>κ's</emph> = 0.84 – 1.00). All disagreements were resolved by a third rater. Two gJOLs were used in this experiment, each open response from 0 to 100%. The first gJOL was to predict performance on the immediate test, "If, in a few minutes, you were shown the definitions you just studied, for what percentage (%) of these definitions are you confident you could remember the corresponding word?" and the second gJOL was to predict performance on the delayed test (and was administered immediately prior to the delayed test): "If, in a few minutes, you were shown the definitions you studied in Part 1 (yesterday using flashcards in the psychology lab), for what percentage (%) of these definitions are you confident you could remember the corresponding word?".</p> <hd id="AN0181532824-34">Procedure</hd> <p>Aside from the following changes listed below, the procedure of Experiment 2 was the same as Experiment 1 (see Fig. 3).</p> <p>Graph: Fig. 3 Procedure of Experiments 2 and 3</p> <p>The experiment was run in two sessions spaced 24 h apart. Aside from the flashcard portion, all phases of the study were run using Qualtrics (https://<ulink href="http://www.qualtrics.com/">www.qualtrics.com/</ulink>). The first session was run in 90-min timeslots involving up to six participants each and using four nearly-identical laboratory testing rooms. The session began with an initial study phase conducted individually on a desktop computer. During the initial study phase, participants studied each vocabulary-definition pair for seven seconds one-at-a-time in a random order. They did this twice, studying each vocabulary-definition pair for a total of 14 s, for an overall study time of approximately 10 min.</p> <hd id="AN0181532824-35">Flashcard Phase</hd> <p>Given that learners received approximately 10 min of initial study, the flashcard portion of the study was shortened to two 15-min periods (such that total time spent learning the materials remained approximately 40 min). During the flashcard phase, all participants solely used the standard flashcard set.</p> <hd id="AN0181532824-36">Individual condition</hd> <p>Participants were instructed to test themselves during the entirety of the flashcard phase. They were told that the experimenter would check in on them after 15 min. Dropping of flashcards was prohibited.</p> <hd id="AN0181532824-37">Paired condition</hd> <p>Given the elevated negative affect reported by Paired learners in Experiment 1, two changes were made to make learners feel more comfortable during the study and to allow for behaviors that students might engage in when collaboratively practice testing in daily life. First, between the initial study phase and the flashcard phase, Paired learners were given two min to complete an icebreaker activity. During this activity, participants were encouraged to introduce themselves to their partner and to converse with them to find one thing that they had in common (e.g., favorite color). Second, although explanations and clarifications were still disallowed during the flashcard phase to avoid an unfair benefit to the Paired condition, participants were told that they could provide brief comments (e.g., good job).</p> <hd id="AN0181532824-38">Survey and distractor task</hd> <p>As dropping flashcards from study was explicitly prohibited, participants were asked whether they dropped flashcards from study only as a compliance check; the question about why they dropped flashcards from study was removed.</p> <hd id="AN0181532824-39">Final cued-recall test</hd> <p></p> <hd id="AN0181532824-40">Immediate (5-min)</hd> <p>Twenty definitions were presented. Prior to completing the test, participants reported a gJOL (as they did in Experiment 1).</p> <hd id="AN0181532824-41">Delayed (24-h)</hd> <p>The morning after Session 1, participants were emailed the test link and were told that they had until 11:59 pm that day to complete the test on their own laptop or desktop computer in a quiet, distraction-free place. Prior to completing the test, participants again reported a gJOL.</p> <hd id="AN0181532824-42">Results</hd> <p></p> <hd id="AN0181532824-43">Flashcard phase</hd> <p></p> <hd id="AN0181532824-44">Number of practice cycles</hd> <p>Participants indicated the number of practice cycles (i.e., practicing through the entire flashcard set) they completed per 15-min period of the flashcard phase. These two numbers were again summed to compute a total number of practice cycles. Unlike in Experiment 1, Individual learners (<emph>M</emph> = 4.29, <emph>SD</emph> = 1.75) and Paired learners (<emph>M</emph> = 3.95, <emph>SD</emph> = 1.45) completed about the same number of practice cycles through the flashcard deck, <emph>t</emph> (<reflink idref="bib139" id="ref56">139</reflink>) = 1.22,<emph> p</emph> = 0.22, <emph>d</emph> = 0.21, 95% CI [−0.21, 0.88].</p> <hd id="AN0181532824-45">Final cued-recall test</hd> <p></p> <hd id="AN0181532824-46">Overall performance</hd> <p>To examine the effect of individual versus paired flashcard practice on learning, a 2 × 2 ANOVA was conducted with Condition (Individual or Paired) as the between-subjects factor, Test Delay (5-min or 24-h) as the within-subjects factor, and test performance as the dependent variable. Six participants did not complete the delayed test[<reflink idref="bib2" id="ref57">2</reflink>] and were therefore excluded from this analysis, leaving 75 Individual and 60 Paired learners in the analysis.</p> <p>Immediate test scores were higher than delayed test scores, suggesting that forgetting occurred during the 24-h delay, <emph>F</emph> (<reflink idref="bib1" id="ref58">1</reflink>, 133) = 41.80, <emph>p</emph> &lt; 0.001, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.24. Replicating the result of Experiment 1, Paired and Individual learners overall demonstrated similar test performance, <emph>F</emph> (<reflink idref="bib1" id="ref59">1</reflink>, 133) = 1.44, <emph>p</emph> = 0.23, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.01.[<reflink idref="bib3" id="ref60">3</reflink>]The nonsignificant Condition x Test Delay interaction suggests that this similarity did not change between the immediate test and the delayed test,<emph> F</emph> (<reflink idref="bib1" id="ref61">1</reflink>, 133) = 0.003, <emph>p</emph> = 0.96, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> &lt; 0.001. Performance on both the immediate and delayed tests, however, were numerically lower in the Paired condition (as was the case in Experiment 1), which suggests that there may be a modest reduction in the efficacy of learning (or the rate of learning) that occurs when using flashcards in pairs versus individually.</p> <hd id="AN0181532824-47">Equivalence test</hd> <p>To examine if the null effect obtained in Experiment 2 was equivalent with the null effect obtained in Experiment 1, we conducted a two one-sided tests (TOST) procedure to test for equivalence (Lakens et al., [<reflink idref="bib25" id="ref62">25</reflink>]) using the TOSTER package in R (Caldwell, [<reflink idref="bib3" id="ref63">3</reflink>]; Lakens, [<reflink idref="bib24" id="ref64">24</reflink>]). First, we set the smallest effect size of interest. Based on the group sizes from Experiment 1 with <emph>α</emph> = 0.05, we used G*Power (Faul et al., [<reflink idref="bib8" id="ref65">8</reflink>]) to determine that the smallest effect size Experiment 1 had 80% power to detect was <emph>d</emph> = 0.47. For that reason, we set the lower equivalence bound to <emph>d</emph> = −0.47 and the upper equivalence bound to <emph>d</emph> = 0.47. We then used the data obtained in Experiment 2 to run two Welch's one-sided t-tests. The test for the upper bound was significant, <emph>t</emph> (137.57) = −1.67, <emph>p</emph> = 0.048, as was the test for the lower bound, <emph>t</emph> (137.57) = 3.91, <emph>p</emph> &lt; 0.001. These significant t-tests indicate that we can reject the null hypothesis that the true effect was smaller than <emph>d</emph> = −0.47 or larger than <emph>d</emph> = 0.47; i.e., the effect size obtained in Experiment 2 falls within the equivalence range. Thus, we can conclude the null effect obtained in Experiment 2 is equivalent to the null effect obtained in Experiment 1.</p> <hd id="AN0181532824-48">Metacognitive judgments</hd> <p></p> <hd id="AN0181532824-49">Correlations with final test performance</hd> <p>To examine whether there was a significant relationship between participants' own assessments of their learning and their actual test score, a series of bivariate correlations related gJOL and final test performance for both conditions and for both the Immediate and Delayed tests (see Fig. 4).</p> <p>Graph: Fig. 4 Metacognitive calibration for the Immediate test (left panel) and the Delayed test (right panel). Each panel displays the correlation between test performance and global judgments of learning (gJOLs). The red and blue lines represent least squares regression fits to Individual and Paired data, respectively. A dotted line represents the hypothetical case of perfect calibration between gJOLs and test scores; again, participants in the Individual condition tended to substantially overestimate their future cued-recall test performance for the immediate test but this tendency did not extend to the delayed test</p> <p>As in Experiment 1, for the immediate test, participants demonstrated moderate-to-large correlations between their gJOL and their actual final test performance after learning individually, <emph>r</emph> (<reflink idref="bib76" id="ref66">76</reflink>) = 0.51, <emph>p</emph> &lt; 0.001, and after learning with a partner, <emph>r</emph> (<reflink idref="bib61" id="ref67">61</reflink>) = 0.57, <emph>p</emph> &lt; 0.001. These correlations were somewhat reduced when examining the relationship between delayed gJOLs and performance on the delayed test, Individual: <emph>r</emph> (<reflink idref="bib51" id="ref68">51</reflink>) = 0.43, <emph>p</emph> = 0.001; Paired: <emph>r</emph> (<reflink idref="bib29" id="ref69">29</reflink>) = 0.36, <emph>p</emph> = 0.047.</p> <hd id="AN0181532824-50">Metacognitive calibration</hd> <p>To include the maximum number of participants in the analysis of metacognitive calibration at immediate test, participants' metacognitive calibration was analyzed using separate independent samples t-tests for the Immediate and Delayed tests.</p> <hd id="AN0181532824-51">Immediate test</hd> <p>Again replicating the results of Experiment 1, Individual learners (<emph>M</emph> = 0.18, <emph>SD</emph> = 0.27) were more overconfident than Paired learners (<emph>M</emph> = 0.08, <emph>SD</emph> = 0.26), <emph>t</emph> (<reflink idref="bib139" id="ref70">139</reflink>) = 2.14, <emph>p</emph> = 0.034, <emph>d</emph> = 0.36, 95% CI [0.007, 0.19].</p> <hd id="AN0181532824-52">Delayed test</hd> <p>In contrast to the results for the immediate test, both Individual learners (<emph>M</emph> = −0.05, <emph>SD</emph> = 0.27) and Paired learners (<emph>M</emph> = −0.02, <emph>SD</emph> = 0.26) were well-calibrated, if slightly underconfident, <emph>t</emph> (<reflink idref="bib82" id="ref71">82</reflink>) = −0.39, <emph>p</emph> = 0.70, <emph>d</emph> = −0.09, 95% CI [−0.14, 0.10].</p> <hd id="AN0181532824-53">Positive and negative affect</hd> <p>As in Experiment 1, there was no difference in self-reported positive affect by Individual learners (<emph>M</emph> = 26.95, <emph>SD</emph> = 8.11) and Paired learners (<emph>M</emph> = 27.73, <emph>SD</emph> = 8.00),<emph> t</emph> (<reflink idref="bib139" id="ref72">139</reflink>) = −0.57, <emph>p</emph> = 0.57, <emph>d</emph> = −0.10, 95% CI [−3.48, 1.92]. However, in contrast to Experiment 1, self-reported negative affect also did not differ between Individual learners (<emph>M</emph> = 14.73, <emph>SD</emph> = 4.41) and Paired learners (<emph>M</emph> = 15.54, <emph>SD</emph> = 3.99), <emph>t</emph> (<reflink idref="bib130" id="ref73">130</reflink>) = −1.13, <emph>p</emph> = 0.26, <emph>d</emph> = −0.19, 95% CI [−2.22, 6.61]. It is possible that the inclusion of the icebreaker activity and the eased restrictions on verbal exchanges led to less negative affect for the Paired condition in Experiment 2.</p> <hd id="AN0181532824-54">Attentional focus</hd> <p>As in Experiment 1, self-reported percentage time focused during the experimental tasks did not significantly differ between the Individual (<emph>M</emph> = 86.3%, <emph>SD</emph> = 14.6%) and Paired (<emph>M</emph> = 87.8%, <emph>SD</emph> = 12.8%) learning conditions, <emph>t</emph> (<reflink idref="bib139" id="ref74">139</reflink>) = −0.61, <emph>p</emph> = 0.54, <emph>d</emph> = −0.10, 95% CI [−6.05, 3.20].</p> <hd id="AN0181532824-55">Experiment 2 discussion</hd> <p>Experiment 2 replicated and extended the two primary findings of Experiment 1. First, immediate cued-recall test performance was similar across those who used flashcards individually and those who used flashcards collaboratively (combining Experiment 1 and 2 data together shows this result is highly similar across the two experiments; analysis available in the online supplemental materials). This similarity was then maintained for the delayed (24-h) test. This result does not align with the suggestion that paired flashcard learning might encourage more effortful retrieval and thus act as a "desirable difficulty" with its benefits emerging after a delay (as is sometimes the case when comparing the effects of more effortful versus less effortful learning strategies, e.g., testing versus restudy, Roediger &amp; Karpicke, 2006). Second, participants in the Individual condition again demonstrated overconfidence in their learning for the immediate test whereas participants in the Paired condition again demonstrated relatively accurate metacognitive judgments. This overconfidence occurred even after dropping of flashcards was prohibited in Experiment 2, suggesting that the miscalibration observed in the Individual condition cannot simply be attributed to a lack of exposure to dropped items. At a delay, however, participants' metacognitive judgments were similarly well-calibrated. This improved metacognitive calibration at a delay is in line with prior work demonstrating that delayed JOLs tend to be more accurate than JOLs made immediately after learning (e.g., Nelson &amp; Dunlosky, [<reflink idref="bib32" id="ref75">32</reflink>]).</p> <p>In Experiment 2, we did not find evidence that learner affect, number of learning cycles, or level of focus differed between the two conditions. Consequently, in Experiment 3 we sought to identify a potential explanation for the metacognitive benefits of paired flashcard learning. Specifically, we investigated whether the use of overt retrieval in paired flashcard learning might drive its benefit. Perhaps requiring participants to overtly retrieve—regardless of whether another person is present or not—offers the learner more concrete evidence of their learning, informing more accurate metacognitive judgments.</p> <hd id="AN0181532824-56">Experiment 3</hd> <p>In Experiment 3 we examined whether learners were more metacognitively accurate following paired flashcard learning as compared to individual flashcard learning because learners in the Paired condition were required to overtly retrieve. Overt retrieval, in contrast to covert retrieval, might establish natural accountability for one's responses during retrieval practice that could enhance the benefits of testing. For example, Sumeracki and Castillo ([<reflink idref="bib49" id="ref76">49</reflink>]) observed a testing effect for students in a classroom that overtly retrieved during practice testing but not for students that covertly retrieved. However, overt and covert retrieval practice both produced a testing effect when students were first informed that one of them would be called on randomly by the teacher. Beyond learning, this accountability may extend benefits to metacognitive judgments by encouraging greater completeness of one's retrieved answers and thus enhancing the quality of evidence available when making these judgments (Tauber et al., [<reflink idref="bib50" id="ref77">50</reflink>]). In the present experiment, we instructed participants in the Individual condition to retrieve out loud during the flashcard phase. If overt retrieval was responsible for the benefits of Paired flashcard learning, then we would expect to observe no difference in metacognitive calibration between the two conditions.</p> <p>Additionally, participants in Experiment 3 reported both global JOLs, as they did in Experiments 1 and 2, and item-level JOLs (iJOLs); i.e., participants predicted both their overall performance and their likelihood of retrieving each vocabulary-definition pair. Item-level JOLs are commonly used in studies of metacognition to measure individuals' relative accuracy (<emph>metacognitive resolution</emph>; i.e., their ability to discriminate between information that will or will not be remembered; Rhodes, [<reflink idref="bib41" id="ref78">41</reflink>]; Vuorre &amp; Metcalfe, [<reflink idref="bib52" id="ref79">52</reflink>]). Until this point, we had assessed participants' absolute accuracy (i.e., <emph>metacognitive calibration</emph>), measuring the difference between participants' average/overall metacognitive judgments and their actual learning outcomes. Resolution and calibration reflect different dimensions of metacognition (Rhodes, [<reflink idref="bib41" id="ref80">41</reflink>]) and can thus at times offer divergent results that offer insight into metacognitive processes; for example, in studies of age-related differences in metacognition (Siegel &amp; Castel, [<reflink idref="bib46" id="ref81">46</reflink>]). Further, the rate of dropping flashcards from study in the Individual condition in Experiment 1 suggests that learners at least sometimes engage in spontaneous item-level judgments during flashcard learning. Thus, we incorporated both types of metacognitive judgments in Experiment 3.</p> <hd id="AN0181532824-57">Method</hd> <p>Experiment 3 was preregistered at https://aspredicted.org/4D9_DFP.</p> <hd id="AN0181532824-58">Participants</hd> <p>Four-hundred and five participants were included in this study (Individual: <emph>n</emph> = 187, Paired: <emph>n</emph> = 218). Eighty participants were recruited from the Psychology subject pool at the same large public research university as Experiments 1 and 2, and 325 participants were recruited from the Psychology subject pool at a similar large public research university in the same region. An additional 120 participants were recruited for this study but were excluded based on our preregistered criteria: dropping flashcards (<emph>n</emph> = 37), using their phone to complete the study (<emph>n</emph> = 8), failure to follow instructions (e.g., reporting that they did not retrieve during the flashcard portion of the study) (<emph>n</emph> = 67), experimenter error (<emph>n</emph> = 7), and technical issues (<emph>n</emph> = 1). We collected greater than our preregistered number of participants in this study because of an error in the set-up of the study that was not identified until midway through data collection. A subset of participants erroneously received the same set of vocabulary words to test on during both the immediate and delayed tests, rendering their delayed test scores unusable. To ensure that we were adequately powered for all our planned analyses, we collected additional data until we met our preregistered sample size for participants with usable delayed test scores (<emph>n</emph> = 210).</p> <hd id="AN0181532824-59">Design</hd> <p>Experiment 3 employed a 2 × 2 mixed factorial design with Condition (Individual or Paired) as the between-subjects factor and Test Delay (5-min or 24-h) as the within-subjects factor.</p> <hd id="AN0181532824-60">Materials</hd> <p>The materials used in Experiment 3 were identical to the materials used in Experiment 2. A subset of the cued-recall test responses (<emph>n</emph> = 1320 responses) was scored by two independent raters. Interrater reliability for all cued-recall test items was adequate (Cohen's <emph>κ's</emph> = 0.68 – 1.00). All disagreements were resolved by discussion and the remaining cued-recall test responses were scored by a single rater.</p> <hd id="AN0181532824-61">Procedure</hd> <p>Aside from the following changes listed below, the procedure of Experiment 3 was the same as in Experiment 2. All participants completed the study in nearly-identical laboratory testing rooms. Participants from one university completed the study on desktop computers and participants from the other university completed the study on their own laptop computers.</p> <hd id="AN0181532824-62">Flashcard phase</hd> <p>Participants in the Individual condition were instructed to test themselves out loud during the entirety of the flashcard phase. To ensure compliance, an audio monitor was placed in the center of the laboratory testing room and the receiver was placed in a separate area with the experimenter. Participants were informed that the audio monitor only transmitted sound to the experimenter (i.e., it did not record their audio). If the experimenter noted that the participant had stopped testing themselves aloud for more than two min they checked in on the participant and reminded them to test themselves aloud.</p> <hd id="AN0181532824-63">Survey</hd> <p>There were two additions made to the survey that was administered after the flashcard learning phase in Session 1. The first was an additional global JOL (i.e., Session 1 gJOL: Delayed Test) that queried participants about how well they believed they would do on a delayed test: "If tomorrow you were shown the definitions you just studied, what percentage (%) of these definitions are you confident you could remember the corresponding word?" This gJOL was an exploratory item to investigate if participants in each condition might differ in their tendency to predict their long-term learning, and to ensure alignment between the gJOLs and the item-by-item JOLs that participants also gave (described below). For clarity, we now refer to the gJOL for the delayed test administered in this and the previous experiment as "Session 2 gJOL: Delayed Test).</p> <p>The second addition was the inclusion of item-level JOLs (iJOLs). These iJOLs were placed at the beginning of the survey. Participants were shown the vocabulary-definition pairs one-at-a-time in a random order and asked to rate the likelihood that they would be able to type the correct vocabulary word if shown only the definition from 0% (will not be able to) to 100% (certainly will be able to). For each vocabulary-definition pair, they gave two ratings using a slider scale: one for if shown the definition "in a few minutes" and one for if shown the definition "tomorrow." Participants were instructed to report their initial judgment upon seeing the vocabulary-definition pair. If they did not report their iJOLs for a given pair within 10 s, a message appeared on the screen encouraging them to respond.</p> <hd id="AN0181532824-64">Results4</hd> <p></p> <hd id="AN0181532824-65">Flashcard phase</hd> <p></p> <hd id="AN0181532824-66">Number of practice cycles</hd> <p>Unlike Experiment 2 (but like Experiment 1), Individual learners (<emph>M</emph> = 4.55, <emph>SD</emph> = 1.82) completed about one more practice cycle through the flashcard set than Paired learners (<emph>M</emph> = 3.82, <emph>SD</emph> = 1.38), <emph>t</emph> (<reflink idref="bib401" id="ref82">401</reflink>) = 4.53,<emph> p</emph> &lt; 0.001, <emph>d</emph> = 0.45, 95% CI [0.41, 1.04].</p> <hd id="AN0181532824-67">Final cued-recall test</hd> <p></p> <hd id="AN0181532824-68">Overall performance</hd> <p>We analyzed final cued-recall test performance using a 2 (Condition: Individual or Paired) × 2 (Test Delay: Immediate or Delayed) mixed ANOVA, with Condition as a between-subjects factor and Test Delay as the within-subjects factor. This analysis only included participants who had both usable immediate and delayed test scores. Unlike in the previous experiments, the Test Delay x Condition interaction was significant, <emph>F</emph> (<reflink idref="bib1" id="ref83">1</reflink>, 208) = 5.11, <emph>p</emph> = 0.02, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.02. The pattern of the interaction suggested that the Individual condition's rate of forgetting was greater than the Paired condition's. Since the interaction was significant, follow-up independent samples t-tests were conducted. The effect of Condition at the immediate test was significant, <emph>t</emph> (<reflink idref="bib208" id="ref84">208</reflink>) = 2.20, <emph>p</emph> = 0.03, <emph>d</emph> = 0.31, 95% CI [0.007, 0.14]. The Individual condition scored significantly higher (<emph>M</emph> = 0.43, <emph>SD</emph> = 0.25) than the Paired condition (<emph>M</emph> = 0.36, <emph>SD</emph> = 0.22) at the immediate test. The effect of Condition at the delayed test was nonsignificant, <emph>t</emph> (<reflink idref="bib208" id="ref85">208</reflink>) = 0.46, <emph>p</emph> = 0.64, <emph>d</emph> = 0.07, 95% CI [−0.05, 0.08]. The Individual (<emph>M</emph> = 0.36, <emph>SD</emph> = 0.23) and Paired (<emph>M</emph> = 0.34, <emph>SD</emph> = 0.22) conditions scored similarly on the delayed test.[<reflink idref="bib5" id="ref86">5</reflink>]</p> <hd id="AN0181532824-69">Immediate test only</hd> <p>Given that a number of participants only had usable test data for the immediate test, we ran an additional independent samples t-test comparing immediate test scores for all Individual and Paired participants who had a usable immediate test score (i.e., regardless of whether they had usable delayed test data). In this analysis, which included an additional 193 participants, the difference between the Individual (<emph>M</emph> = 0.42, <emph>SD</emph> = 0.26) and Paired (<emph>M</emph> = 0.37, <emph>SD</emph> = 0.24) conditions' immediate test scores was nonsignificant, <emph>t</emph> (<reflink idref="bib401" id="ref87">401</reflink>) = 1.69, <emph>p</emph> = 0.09, <emph>d</emph> = 0.17, 95% CI [−0.007, 0.09], although numerically higher for the Individual condition.</p> <hd id="AN0181532824-70">Metacognitive judgments</hd> <p></p> <hd id="AN0181532824-71">Correlations with final test performance</hd> <p>Individual learners demonstrated moderate correlations between their gJOL and immediate final test performance when learning individually, <emph>r</emph> (<reflink idref="bib184" id="ref88">184</reflink>) = 0.51, <emph>p</emph> &lt; 0.001, as when learning in pairs, <emph>r</emph> (<reflink idref="bib213" id="ref89">213</reflink>) = 0.46, <emph>p</emph> &lt; 0.001 (see Fig. 5). The strength of these correlations was maintained when examining the relationship between Session 2 gJOL: Delayed Test judgments and performance on the delayed test for participants in the Individual condition, <emph>r</emph> (<reflink idref="bib88" id="ref90">88</reflink>) = 0.59, <emph>p</emph> &lt; 0.001, but was somewhat reduced for participants in the Paired condition, <emph>r</emph> (<reflink idref="bib118" id="ref91">118</reflink>) = 0.32, <emph>p</emph> &lt; 0.001.</p> <p>Graph: Fig. 5 Metacognitive calibration for the Immediate test and immediate global JOL (gJOL) (left panel), the Delayed test and Delayed test gJOL administered after the flashcard learning phase in Session 1 (middle panel), and the Delayed test and Delayed gJOL administered in Session 2 immediately prior to the Delayed test (right panel). Each panel displays the correlation between test performance and gJOLs. The red and blue lines represent least squares regression fits to Individual and Paired data, respectively. The dotted lines represent the hypothetical case of perfect calibration between gJOLs and test scores</p> <p>New to Experiment 3 was a gJOL in Session 1 asking participants to predict their test performance if given a test on the vocabulary-definition pairs the next day (i.e., Session 1 gJOL: Delayed Test). For this gJOL, participants' judgments were less strongly related to actual test performance than participants' immediate test gJOLs: Individual: <emph>r</emph> (<reflink idref="bib88" id="ref92">88</reflink>) = 0.36, <emph>p</emph> &lt; 0.001, Paired: <emph>r</emph> (<reflink idref="bib118" id="ref93">118</reflink>) = 0.35, <emph>p</emph> &lt; 0.001. Possible explanations for these differences in the magnitude of the gJOL-test performance correlations will be explored in the following sections.</p> <hd id="AN0181532824-72">Metacognitive calibration</hd> <p>Metacognitive calibration was again calculated by subtracting participants' actual test performance from their predicted performance (i.e., their gJOL). To include the maximum number of participants in the analysis of metacognitive calibration at immediate test, participants' metacognitive calibration was analyzed using separate independent samples t-tests for the Immediate and Delayed tests. A Bonferroni correction for multiple comparisons was used such that the standard for significance was <emph>p</emph> &lt; 0.017.</p> <hd id="AN0181532824-73">Immediate test</hd> <p>Unlike in Experiments 1 and 2, Individual learners (<emph>M</emph> = 0.06, <emph>SD</emph> = 0.25) and Paired learners (<emph>M</emph> = 0.07, <emph>SD</emph> = 0.25), had calibration scores close to 0 (i.e., they were relatively accurate, if slightly overconfident) and these scores did not significantly differ from one another, <emph>t</emph> (<reflink idref="bib399" id="ref94">399</reflink>) = −0.41, <emph>p</emph> = 0.69, <emph>d</emph> = −0.04, 95% CI [−0.06, 0.04].</p> <hd id="AN0181532824-74">Session 2 gJOL: delayed test</hd> <p>Similar to the results of Experiments 1 and 2 (and aligned with participants' calibration at immediate test), both Individual learners (<emph>M</emph> = −0.10, <emph>SD</emph> = 0.20) and Paired learners (<emph>M</emph> = −0.10, <emph>SD</emph> = 0.24) were well-calibrated, if slightly underconfident, <emph>t</emph> (<reflink idref="bib208" id="ref95">208</reflink>) = −0.02, <emph>p</emph> = 0.99, <emph>d</emph> = −0.002, 95% CI [−0.06, 0.06].</p> <hd id="AN0181532824-75">Session 1 gJOL: delayed test</hd> <p>Like the other calibration scores obtained in Experiment 3, participants in the Individual (<emph>M</emph> = −0.04, <emph>SD</emph> = 0.26) and Paired (<emph>M</emph> = −0.05, <emph>SD</emph> = 0.26) conditions were both slightly underconfident, <emph>t</emph> (<reflink idref="bib208" id="ref96">208</reflink>) = 0.26, <emph>p</emph> = 0.80, <emph>d</emph> = 0.04, 95% CI [−0.06, 0.08].</p> <p>We further compared participants' calibration scores between the two delayed test gJOLs using an ANOVA with Delayed Test gJOL Timing (Session 1 or Session 2) as the within-subjects factor and Condition (Individual or Paired) as the between-subjects factor. The interaction between these factors was nonsignificant, <emph>F</emph> (<reflink idref="bib1" id="ref97">1</reflink>, 208) = 0.19, <emph>p</emph> = 0.67, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.001, as was the main effect of Condition, <emph>F</emph> (<reflink idref="bib1" id="ref98">1</reflink>, 208) = 0.02, <emph>p</emph> = 0.89, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> &lt; 0.001. The main effect of Delayed Test gJOL Timing, however, was significant, <emph>F</emph> (<reflink idref="bib1" id="ref99">1</reflink>, 208) = 21.87, <emph>p</emph> &lt; 0.001, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.10. Session 1 Delayed Test gJOLs (<emph>M</emph> = −0.04, <emph>SD</emph> = 0.26) were significantly more accurate than Session 2 Delayed Test gJOLs (<emph>M</emph> = −0.10, <emph>SD</emph> = 0.23).</p> <hd id="AN0181532824-76">Metacognitive resolution</hd> <p>Participants' metacognitive resolution was determined by associating participants' iJOLs with their performance on that item (i.e., whether they got the item correct or incorrect on the cued-recall test) to compute a Kruskal-Goodman gamma correlation (Nelson, [<reflink idref="bib31" id="ref100">31</reflink>]). Gamma correlations offer insight into participants' ability to discriminate between vocabulary-definition pairs that ultimately were remembered and those which were not. Four participants had their data excluded from the analyses because they were missing five or more iJOLs (this exclusion criterion was preregistered) and 17 additional participants were excluded from the analyses because of a lack of variation in test scores (i.e., they scored 0% or 100%).</p> <p>A 2 (Condition: Individual or Paired) × 2 (Test Delay: Immediate or Delayed) mixed ANOVA was conducted with Condition as the between-subjects factor, Test Delay as the within-subjects factor, and the participant's gamma correlation as the dependent variable. Overall, the mean gamma correlation for each condition at each test delay suggests that students' iJOLs were moderately positively associated with their actual test performance. The Condition x Test Delay interaction was nonsignificant, <emph>F</emph> (<reflink idref="bib1" id="ref101">1</reflink>, 187) = 0.73, <emph>p</emph> = 0.40, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.004. The main effect of Condition was also nonsignificant; participants in the Individual (<emph>M</emph> = 0.49, <emph>SD</emph> = 0.49) and Paired (<emph>M</emph> = 0.51, <emph>SD</emph> = 0.49) conditions had gamma correlations of similar magnitude, <emph>F</emph> (<reflink idref="bib1" id="ref102">1</reflink>, 187) = 0.16, <emph>p</emph> = 0.69, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.001.[<reflink idref="bib6" id="ref103">6</reflink>] There was, however, a significant main effect of Test Delay such that participants' gamma correlations were higher for the delayed test (<emph>M</emph> = 0.55, <emph>SD</emph> = 0.31) than for the immediate test (<emph>M</emph> = 0.46, <emph>SD</emph> = 0.34), <emph>F</emph> (<reflink idref="bib1" id="ref104">1</reflink>, 187) = 8.93, <emph>p</emph> = 0.003, <emph>η</emph><subs><emph>p</emph></subs><sups><emph>2</emph></sups> = 0.05.</p> <hd id="AN0181532824-77">Positive and negative affect</hd> <p>As in Experiment 2, participants in the Individual (<emph>M</emph> = 25.44, <emph>SD</emph> = 7.89) and Paired (<emph>M</emph> = 26.28, <emph>SD</emph> = 8.05) conditions reported similar levels of positive affect, <emph>t</emph> (<reflink idref="bib401" id="ref105">401</reflink>) = −1.06, <emph>p</emph> = 0.29, <emph>d</emph> = −0.11, 95% CI [−2.41, 0.73]. Likewise, participants in the Individual (<emph>M</emph> = 15.84, <emph>SD</emph> = 5.35) and Paired (<emph>M</emph> = 16.03, <emph>SD</emph> = 5.77) conditions reported similar levels of negative affect, <emph>t</emph> (<reflink idref="bib401" id="ref106">401</reflink>) = −0.35, <emph>p</emph> = 0.73, <emph>d</emph> = −0.04, 95% CI [−1.29, 0.90].</p> <hd id="AN0181532824-78">Attentional focus</hd> <p>As in Experiments 1 and 2, self-reported percentage time focused during the experimental tasks in Experiment 3 did not significantly differ between the Individual (<emph>M</emph> = 88.5%, <emph>SD</emph> = 14.0%) and Paired (<emph>M</emph> = 86.4%, <emph>SD</emph> = 13.1%) conditions, <emph>t</emph> (<reflink idref="bib401" id="ref107">401</reflink>) = 1.59, <emph>p</emph> = 0.11, <emph>d</emph> = 0.16, 95% CI [−0.51, 4.80].</p> <hd id="AN0181532824-79">Self-reported flashcard use in experiments 1–3</hd> <p>Table 2 summarizes data on participants' self-reported use of flashcards for exam preparation in daily life. Most students reported using flashcards at least sometimes when studying. When studying with friends, less than half of students reported using flashcards; even if they did use flashcards when studying with friends, they did so infrequently. Overall, students' self-reported flashcard practices suggest that, while they do commonly use flashcards when studying in daily life, they are far more likely to use flashcards when studying alone versus when studying with others.</p> <p>Table 2 Frequency of self-reported flashcard use when preparing for exams</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Frequency&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="4"&gt;&lt;p&gt;When studying generally&lt;/p&gt;&lt;/th&gt;&lt;th align="left" /&gt;&lt;th align="left" /&gt;&lt;th align="left" colspan="4"&gt;&lt;p&gt;When studying with a partner&lt;/p&gt;&lt;/th&gt;&lt;th align="left" /&gt;&lt;th align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Exp. 1&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Exp. 2&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Exp. 3&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Exp. 1&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Exp. 2&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;Exp. 3&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;%&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;%&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;%&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;%&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;%&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;n&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&lt;italic&gt;%&lt;/italic&gt;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Never&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;19&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;12.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;18&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;12.8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;55&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;13.6&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;25&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;16.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;36&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;35.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;108&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;26.8&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Almost never&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;37&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;24.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;41&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;29.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;139&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;34.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;54&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;35.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;52&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;36.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;148&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;36.7&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Sometimes&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;65&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;42.8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;72&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;51.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;156&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;38.7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;59&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;38.8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;50&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;35.5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;122&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;30.3&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Almost every time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;26&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;17.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;7&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;5.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;42&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;10.4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;9&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;5.9&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;2.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;23&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;5.7&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Every time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;5&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;3.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;2.1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;11&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;2.7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;5&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;3.3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;0.0&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char"&gt;&lt;p&gt;0.5&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Total&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;152&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char" /&gt;&lt;td align="left"&gt;&lt;p&gt;141&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char" /&gt;&lt;td align="left"&gt;&lt;p&gt;403&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char" /&gt;&lt;td align="left"&gt;&lt;p&gt;152&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char" /&gt;&lt;td align="left"&gt;&lt;p&gt;141&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char" /&gt;&lt;td align="left"&gt;&lt;p&gt;403&lt;/p&gt;&lt;/td&gt;&lt;td char="." align="char" /&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0181532824-80">General discussion</hd> <p>Across three experiments, using flashcards to learn with a partner did not yield greater learning compared to using flashcards alone; in fact, in Experiment 3, the Individual condition outperformed the Paired condition at the immediate test. Despite our expectation that collaboration might encourage more effortful retrieval and thus promote long-term learning, Individual and Paired flashcard use yielded learning that was not statistically different when assessed at a 24-h delay in Experiments 2–3. Although performance did not differ significantly between the two learning conditions, we did observe two advantages of flashcard-based retrieval practice with a partner as opposed to individual retrieval practice. First, when dropping was neither explicitly allowed nor disallowed, Paired learners were far less likely to drop cards from study than Individual learners. Second, there was a striking metacognitive benefit to Paired learning observed in Experiments 1–2: Whereas Individual learners were often overconfident—overestimating learning by approximately 20% in both experiments—Paired learners exhibited more accurate judgments immediately after they had finished learning with flashcards. Instructing Individual learners to overtly retrieve in Experiment 3, however, mitigated this overconfidence. Together, these findings suggest that paired flashcard practice can offer metacognitive benefits that may be important for those using flashcards during self-regulated learning and offers evidence of a potential mechanism for these effects: the facilitation of overt retrieval.</p> <hd id="AN0181532824-81">Learning efficiency of individual versus paired flashcard learning</hd> <p>In this set of experiments, we asked learners to self-report the number of learning cycles (i.e., the number of times they were able to get through all the cards in the flashcard set). We did so because collaborative learning activities can take longer than individual learning activities (Johnson &amp; Johnson, [<reflink idref="bib15" id="ref108">15</reflink>]). Thus, although total time on task was maintained across conditions, we were interested in whether the number of learning cycles completed during the flashcard learning phase would differ between the two conditions. We found moderate evidence for paired flashcard learning being more inefficient (in terms of learning cycles completed) than individual flashcard learning. Individual learners in Experiments 1 and 3 on average completed approximately one more learning cycle than Paired learners, but learners in Experiment 2 completed a similar number of learning cycles regardless of condition (although, numerically, Individual learners completed more learning cycles than Paired learners). Reconciling with prior work, it is possible that the constraints on learners in the Paired condition (i.e., no elaborative explanations) led to a pace of study similar to the pace of study in the Individual condition, but that time spent in brief discussion or switching roles in the flashcard phase led to generally one fewer learning cycle completed by the Paired condition. Likewise, in Experiment 1, Individual learners' tendency to drop cards from study could have allowed them to complete more learning cycles than those in the Paired condition; indeed, Individual learners in Experiment 1 were the only group in the entire set of studies with an average number of cycles completed greater than five. It is also possible that learners struggled to accurately self-report the number of learning cycles. This challenge may have been more prominent for Paired learners who took on multiple roles during the learning phase (i.e., that of tester and testee) and may have given more attention to monitoring their partner than tracking their number of completed learning cycles.</p> <hd id="AN0181532824-82">Why is paired flashcard learning advantageous for metacognition?</hd> <p>Our findings appear to stem from characteristics of using flashcards with a partner; in particular, its facilitation of overt responses during retrieval practice. Unlike their counterparts in the Individual condition (in the first two experiments), Paired learners had to clearly articulate a response before feedback was provided. In Experiment 3, when Individual learners were instructed to overtly retrieve, these learners were not susceptible to overconfidence and even outperformed Paired learners on the immediate test (possibly due to their ability to engage in retrieval practice for the full 30-min session whereas Paired learners only spent half that time retrieving and the rest being the "tester" for their partner). Broadly, overt retrieval possibly resulted in more effortful retrieval processes (Pyc &amp; Rawson, [<reflink idref="bib40" id="ref109">40</reflink>] offer a discussion about the benefits of effortful retrieval) which were not shortchanged by any peeking at the answers or half-hearted attempts at retrieval.</p> <p>This suggestion is corroborated by evidence from Tauber and colleagues (2018). In two experiments, participants studied term-definition pairs from Psychology (e.g., confirmation bias). In the first experiment, participants in the retrieval practice conditions were then shown each term and either covertly or overtly (i.e., by typing) retrieved its definition and provided a judgment of knowing (i.e., a judgment of how well they knew the definition to the term). Despite the covert retrieval group scoring lower on the final recall test than the overt retrieval group, they reported significantly higher judgments of knowing during retrieval practice—a similar overconfidence to that observed in the present work using predictions of future test performance. In the second experiment, an additional "enhanced" covert retrieval group was given instructions on how to practice covert retrieval, and learners in the retrieval practice conditions also judged the completeness of their retrieval during retrieval practice. In contrast to evidence from their test performance and judgments of knowing, learners in the enhanced covert retrieval group reported retrieving more of the term definitions during retrieval practice than the overt retrieval condition. These results suggest that overt retrieval practice may encourage exhaustive retrieval and offer better evidence of one's level of learning.</p> <p>Although the facilitation of overt retrieval emerged as a key contributor to the benefits of paired flashcard learning, there are other features of paired flashcard learning that may also benefit the accuracy of metacognitive judgments. Paired learners, for example, received feedback only after a complete retrieval attempt and feedback was consistently provided. This consistent feedback from their partner obviated any issues with insufficient checking of answers (Wissman et al., 2016). Inconsistently seeking out feedback may have increased Individual learners' reliance on less diagnostic cues (e.g., ease of retrieved responses; Benjamin et al., [<reflink idref="bib2" id="ref110">2</reflink>]), yielding overconfidence. A further consideration involves the increased dropping of flashcards in the Individual condition when dropping was not explicitly prohibited. Such dropping commonly occurred because a given vocabulary-definition pair had been deemed sufficiently learned (which aligns with accounts of study-time allocation such as the Region of Proximal Learning Model; e.g., Metcalfe &amp; Kornell, [<reflink idref="bib29" id="ref111">29</reflink>]) and likely deprived learners of robust evidence of their mastery of the vocabulary-definition pairs. Consequently, Individual learners in Experiment 1 based their global judgment of learning on impoverished information relative to Paired learners.</p> <p>It should be noted that this poor metacognitive calibration in the Individual condition appeared to resolve at a 24-h delay. In line with other work highlighting that delayed JOLs tend to be more accurate than immediate JOLs (e.g., Nelson &amp; Dunlosky, [<reflink idref="bib32" id="ref112">32</reflink>]), it is possible that Individual learners were less susceptible to certain metacognitive illusions (e.g., the stability bias; Kornell &amp; Bjork, [<reflink idref="bib21" id="ref113">21</reflink>]) after the passage of time. Additionally, the experience of taking the immediate test in Session 1 of Experiments 2–3 may have offered participants insight into their learning which informed their delayed JOL, and that this information was particularly useful for Individual learners. We investigated this possibility by asking learners in Experiment 3 to report a gJOL for the delayed test in Session 1 (i.e., predict their performance if tested tomorrow) and compared that to their gJOL for the delayed test administered in Session 2 immediately prior to the test. Offering evidence contrary to the possibility that participants had used their immediate test experience to inform their delayed metacognitive judgments, the gJOLs completed right before the delayed test were less well-calibrated (i.e., significantly more underconfident) than the ones completed in Session 1, and this difference was similar between the Individual and Paired conditions, which is overall consistent with the <emph>underconfidence with practice</emph> effect (Koriat et al., [<reflink idref="bib18" id="ref114">18</reflink>]).</p> <hd id="AN0181532824-83">Potential effects of collaborative learning on affective and motivational states</hd> <p>Given classroom evidence that learning with others improves motivation and enjoyment (e.g., McCabe &amp; Lummis, [<reflink idref="bib28" id="ref115">28</reflink>]), we were surprised to observe greater negative affect in the Paired condition in Experiment 1. One possible explanation is that being quizzed by a stranger increased anxiety or embarrassment. Although logistical and privacy constraints necessitated random assignment of strangers in the Paired condition, students typically know their study partners in more authentic learning environments (although students sometimes work with strangers in large classes or in assigned groups). This explanation is supported by the lack of evidence for elevated negative affect in Paired learners in Experiments 2 and 3, which incorporated a brief icebreaker activity to facilitate participants getting to know each other (if only superficially) and eased restrictions on verbal communication during the flashcard phase. Although students would likely work with those they know if engaging in paired flashcard learning in everyday life, these findings suggest that implementation of paired flashcard learning in a structured setting (e.g., as a classroom activity) should consider methods to increase students' comfort, particularly if students are asked to work with someone that they do not know.</p> <hd id="AN0181532824-84">Limitations and future work</hd> <p>The lack of differences in final test performance may stem from several design decisions. Although participants controlled their pace of study and dropping of flashcards, they did not control when to terminate the learning session (as commonly occurs during self-regulated learning). Results may have differed if participants stopped learning once they believed that they had sufficiently mastered the material. The Paired condition may have also been negatively impacted by participants' unfamiliarity with one another and limits on verbal discussion. Learning is supported both by knowledge construction and knowledge consolidation (Roelle et al., [<reflink idref="bib43" id="ref116">43</reflink>]). Whereas retrieval practice is particularly beneficial for knowledge consolidation (Roelle et al., [<reflink idref="bib43" id="ref117">43</reflink>]), collaboration may support learning by facilitating explanations and elaborations that might be particularly beneficial for knowledge construction (Fiorella &amp; Mayer, [<reflink idref="bib9" id="ref118">9</reflink>]). Collaborative learning, for example, is often examined within the context of open-ended tasks which provide ample opportunity for knowledge construction (e.g., Zhu, [<reflink idref="bib56" id="ref119">56</reflink>]). It is possible that the carefully controlled procedure and setting of these three experiments impeded exchanges which might spontaneously occur in a real-world collaborative context, supporting knowledge-building, and thus limited the benefits of collaborative flashcard use in the present work. It is further possible that the use of less-complex materials (vocabulary-definition pairs) did not promote the use of these potentially beneficial behaviors to the extent that using more complex materials (e.g., text passages) would have—although using vocabulary as the to-be-learned content aligns with students' self-reported flashcard practices (Zung et al., [<reflink idref="bib57" id="ref120">57</reflink>]). To address some of these possibilities, future work might employ a "think-aloud" procedure (e.g., Nokes-Malach et al., [<reflink idref="bib34" id="ref121">34</reflink>]), or may recruit friends that tend to study together in more naturalistic settings (e.g., study groups).</p> <hd id="AN0181532824-85">Practical implications</hd> <p>Across three experiments, we find that using flashcards in pairs results in more accurate judgments of learning than using flashcards alone, but that this advantage is not present when participants working alone are instructed to retrieve out loud. This result has practical implications for self-regulated learning and effective exam preparation. It suggests that paired flashcard learning can encourage learners to engage in overt retrieval of content during practice testing. Although learners could retrieve out loud by themselves, studies of flashcard learning suggest that the de facto procedure when studying with flashcards is to do so with covert retrieval (Pan et al., [<reflink idref="bib37" id="ref122">37</reflink>]); in this set of studies, participants did not show benefits of individual flashcard practice in Experiments 1 and 2 when they were not instructed to overtly retrieve during study and monitored to ensure adherence. Further, learners may be more consistent and comfortable retrieving out loud in a social setting than by themselves in, for example, a library or a dorm common area. Thus, when considering the fact that undergraduate students more often use flashcards when studying alone than with a friend (which implies that flashcards are commonly regarded as a solitary tool), it appears that many students are overlooking a potentially more beneficial method of using flashcards—that is, with a partner.</p> <hd id="AN0181532824-86">Acknowledgements</hd> <p>An earlier version of this work is included in the first author's dissertation. The authors would like to thank Elizabeth Ligon Bjork, Robert Bjork, Melissa Paquette-Smith, Idan Blank, and Alan Castel for their comments and suggestions on an earlier version of this manuscript.</p> <hd id="AN0181532824-87">Author contributions</hd> <p>M.N.I., I.Z., and S.C.P. designed the experiments and developed the materials. M.N.I., I.Z., S.C.P., and M.C.W. collected experiment data. M.N.I. analyzed and interpreted the data with input from I.Z., S.C.P, and M.C.W. M.N.I drafted the manuscript and I.Z, S.C.P., and M.C.W. provided substantial manuscript edits. All authors approved the manuscript for submission.</p> <hd id="AN0181532824-88">Funding</hd> <p>No funding supported this research.</p> <hd id="AN0181532824-89">Data Availability</hd> <p>De-identified research data are publicly available via Open Science Framework https://osf.io/hac38/?view%5fonly=9ff44668b2184f1b8e412452f3a41640.</p> <hd id="AN0181532824-90">Declarations</hd> <p></p> <hd id="AN0181532824-91">Ethics Approval</hd> <p>Ethical approval was granted by UCLA's Institutional Review Board (IRB# 11-002880)..</p> <hd id="AN0181532824-92">Competing Interests</hd> <p>The authors have no conflicts of interest to declare.</p> <hd id="AN0181532824-93">Appendix</hd> <p>Table 3.</p> <p>Table 3. List of GRE vocabulary word-definition pairs</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;No.&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;GRE word&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Definition&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;1.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Abeyance&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;temporary inactivity, cessation, or suspension&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;2.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Abjure&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to renounce, repudiate, or retract, especially with formal solemnity; recant&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;3.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Anodyne&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a medicine that relieves or allays pain&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;4.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Canard&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a false or baseless, usually derogatory story, report, or rumor&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;5.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Cosset&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to treat as a pet; pamper; coddle&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;6.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Ebullient&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;overflowing with fervor, enthusiasm, or excitement; high-spirited&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;7.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Ersatz&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;serving as a substitute; synthetic; artificial&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;8.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Expiate&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to atone for; make amends or reparation for&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;9.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Fracas&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a noisy, disorderly disturbance or fight; riotous brawl; uproar&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;10.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Fusillade&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a simultaneous or continuous discharge of firearms&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;11.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Gainsay&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to deny, dispute, or contradict&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;12.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Hermetic&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;made airtight by fusion or sealing&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;13.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Impugn&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to challenge as false (another's statements, motives); cast doubt upon&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;14.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Lachrymose&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;suggestive of or tending to cause tears; mournful&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;15.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Lambaste&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to beat or whip severely&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;16.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Maelstrom&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a large, powerful, or violent whirlpool&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;17.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Monolithic&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;made of only one stone&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;18.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Munificent&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;extremely liberal in giving; very generous&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;19.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Myopic&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;unable or unwilling to act prudently; shortsighted&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;20.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Noisome&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;offensive or disgusting, as an odor&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;21.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Occlude&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to close, shut, or stop up (a passage, opening)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;22.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Paean&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;any song of praise, joy, or triumph&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;23.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Panoply&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a wide-ranging and impressive array or display&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;24.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Pellucid&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;allowing the maximum passage of light, as glass; translucent&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;25.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Polemic&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a controversial argument, as one against some opinion, doctrine&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;26.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Prosaic&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;commonplace or dull; matter-of-fact or unimaginative&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;27.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Puerile&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;of or relating to a child or to childhood&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;28.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Pundit&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a learned person, expert, or authority&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;29.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Quiescent&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;being at rest; quiet; still; inactive or motionless&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;30.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Quixotic&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;extravagantly chivalrous or romantic; visionary, impractical, or impracticable&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;31.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Redress&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;the setting right of what is wrong&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;32.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Sanguine&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;cheerfully optimistic, hopeful, or confident&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;33.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Soporific&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;causing or tending to cause sleep&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;34.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Supine&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;lying on the back, face or front upward&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;35.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Tyro&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;a beginner in learning anything; novice&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;36.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Upbraid&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to find fault with or reproach severely; censure&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;37.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Verdant&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;green with vegetation; covered with growing plants or grass&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;38.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Vitiate&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to impair the quality of; make faulty; spoil&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;39.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Vitriol&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;something highly caustic or severe in effect, as criticism&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;40.&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Welter&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;to roll, toss, or heave, as waves or the sea&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0181532824-94">Publisher's Note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0181532824-95"> <title> References </title> <blist> <bibl id="bib1" idref="ref51" type="bt">1</bibl> <bibtext> Barber SJ, Rajaram S, Aron A. When two is too many: Collaborative encoding impairs memory. Memory &amp; Cognition. 2010; 38: 255-264. 10.3758/MC.38.3.255</bibtext> </blist> <blist> <bibl id="bib2" idref="ref43" type="bt">2</bibl> <bibtext> Benjamin AS, Bjork RA, Schwartz BL. The mismeasure of memory: When retrieval fluency is misleading as a metamnemonic index. Journal of Experimental Psychology: General. 1998; 127; 1: 55-68. 10.1037/0096-3445.127.1.55</bibtext> </blist> <blist> <bibl id="bib3" idref="ref60" type="bt">3</bibl> <bibtext> Caldwell, A. R. (2022). Exploring equivalence testing with the updated TOSTER R package. PsyArXiv. https://doi.org/10.31234/osf.io/ty8de</bibtext> </blist> <blist> <bibl id="bib4" idref="ref25" type="bt">4</bibl> <bibtext> Carpenter SK, Pashler H. Testing beyond words: Using tests to enhance visuospatial map learning. Psychonomic Bulletin &amp; Review. 2007; 14: 474-478. 10.3758/BF03194092</bibtext> </blist> <blist> <bibl id="bib5" idref="ref5" type="bt">5</bibl> <bibtext> Dunlosky J, Rawson KA, Marsh EJ, Nathan MJ, Willingham DT. Improving students' learning with effective learning techniques: Promising directions from cognitive and educational psychology. Psychological Science in the Public Interest. 2013; 14; 1: 4-58. 10.1177/1529100612453266</bibtext> </blist> <blist> <bibl id="bib6" idref="ref32" type="bt">6</bibl> <bibtext> Efklides A. Metacognition and affect: What can metacognitive experiences tell us about the learning process?. Educational Research Review. 2006; 1; 1: 3-14. 10.1016/j.edurev.2005.11.001</bibtext> </blist> <blist> <bibl id="bib7" idref="ref33" type="bt">7</bibl> <bibtext> Efklides A, Schwartz BL, Brown VSchunk DH, Greene JA. Motivation and affect in self-regulated learning. Handbook of self-regulation of learning and performance. 20182; Routledge: 64-82</bibtext> </blist> <blist> <bibl id="bib8" idref="ref39" type="bt">8</bibl> <bibtext> Faul F, Erdfelder E, Lang AG, Buchner A. G*Power 3: A flexible statistical power analysis program for the social, behavioral, and biomedical sciences. Behavior Research Methods. 2007; 39; 2: 175-191. 10.3758/BF03193146</bibtext> </blist> <blist> <bibl id="bib9" idref="ref118" type="bt">9</bibl> <bibtext> Fiorella L, Mayer RE. Eight ways to promote generative learning. Educational Psychology Review. 2016; 28: 717-741. 10.1007/s10648-015-9348-9</bibtext> </blist> <blist> <bibtext> Geen RG. Evaluation apprehension and the social facilitation/inhibition of learning. Motivation and Emotion. 1983; 7; 2: 203-212. 10.1007/BF00992903</bibtext> </blist> <blist> <bibtext> Gilley BH, Clarkston B. Collaborative testing: Evidence of learning in a controlled in-class study of undergraduate students. Journal of College Science Teaching. 2014; 43; 3: 83-91. 10.2505/4/jcst14_043_03_83</bibtext> </blist> <blist> <bibtext> Hayat AA, Shateri K, Amini M, Shokrpour N. Relationships between academic self-efficacy, learning-related emotions, and metacognitive learning strategies with academic performance in medical students: A structural equation model. BMC Medical Education. 2020; 20: 76. 10.1186/s12909-020-01995-9</bibtext> </blist> <blist> <bibtext> Holzer J, Korlat S, Haider C, Mayerhofer M, Pelikan E, Schober B, Spiel C, Toumazi T, Salmela-Aro K, Käser U, Schultze-Krumbholz A, Wachs S, Dabas M, Verma S, Iliev D, Andonovska-Trajkovska D, Plichta P, Pyżalski J, Walter N, Michalek-Kwiecień J, Lewandowska-Walter A, Wright MF, Lüftenegger M. Adolescent well-being and learning in times of COVID-19—A multi-country study of basic psychological need satisfaction, learning behavior, and the mediating roles of positive emotion and intrinsic motivation. PLoS One. 2021; 16: e0251352. 10.1371/journal.pone.0251352</bibtext> </blist> <blist> <bibtext> Imundo, M. N. (2023). Testing together: Collaborative and individual practice testing can yield different patterns of learning following practice testing with varied test formats. Dissertation.</bibtext> </blist> <blist> <bibtext> Johnson DW, Johnson RT. An educational psychology success story: Social interdependence theory and cooperative learning. Educational Researcher. 2009; 38: 365-379. 10.3102/0013189X09339057</bibtext> </blist> <blist> <bibtext> Johnson DW, Johnson RT, Smith KA. Cooperative learning returns to college what evidence is there that it works?. Change: The Magazine of Higher Learning. 1998; 30; 4: 26-35. 10.1080/00091389809602629</bibtext> </blist> <blist> <bibtext> Jönsson FU, Kubik V, Sundqvist ML, Todorov I, Jonsson B. How crucial is the response format for the testing effect?. Psychological Research Psychologische Forschung. 2014; 78: 623-633. 10.1007/s00426-013-0522-8</bibtext> </blist> <blist> <bibtext> Koriat A, Sheffer L, Ma'ayan H. Comparing objective and subjective learning curves: Judgments of learning exhibit increased underconfidence with practice. Journal of Experimental Psychology: General. 2002; 131; 2: 147-162. 10.1037/0096-3445.131.2.147</bibtext> </blist> <blist> <bibtext> Kornell N. Optimising learning using flashcards: Spacing is more effective than cramming. Applied Cognitive Psychology. 2009; 23; 9: 1297-1317. 10.1002/acp.1537</bibtext> </blist> <blist> <bibtext> Kornell N, Bjork RA. Optimising self-regulated study: The benefits—and costs—of dropping flashcards. Memory. 2008; 16; 2: 125-136. 10.1080/09658210701763899</bibtext> </blist> <blist> <bibtext> Kornell N, Bjork RA. A stability bias in human memory: Overestimating remembering and underestimating learning. Journal of Experimental Psychology: General. 2009; 138; 4: 449-468. 10.1037/a0017350</bibtext> </blist> <blist> <bibtext> Krumboltz JD, Weisman RG. The effect of overt versus covert responding to programed instruction on immediate and delayed retention. Journal of Educational Psychology. 1962; 53; 2: 89-92. 10.1037/h0041100</bibtext> </blist> <blist> <bibtext> Kubik V, Jönsson FU, de Jonge M, Arshamian A. Putting action into testing: Enacted retrieval benefits long-term retention more than covert retrieval. Quarterly Journal of Experimental Psychology. 2020; 73; 12: 2093-2105. 10.1177/1747021820945560</bibtext> </blist> <blist> <bibtext> Lakens D. Equivalence tests: A practical primer for t-tests, correlations, and meta-analyses. Social Psychological and Personality Science. 2017; 8; 4: 355-362. 10.1177/1948550617697177</bibtext> </blist> <blist> <bibtext> Lakens D, Scheel AM, Isager PM. Equivalence testing for psychological research: A tutorial. Advances in Methods and Practices in Psychological Science. 2018; 1; 2: 259-269. 10.1177/2515245918770963</bibtext> </blist> <blist> <bibtext> Lin C, McDaniel MA, Miyatsu T. Effects of flashcards on learning authentic materials. Journal of Applied Research in Memory and Cognition. 2018; 7; 4: 529-539. 10.1037/h0101829</bibtext> </blist> <blist> <bibtext> LoGuidice AB, Pachai AA, Kim JA. Testing together: When do students learn more through collaborative tests?. Scholarship of Teaching and Learning in Psychology. 2015; 1; 4: 377-389. 10.1037/stl0000041</bibtext> </blist> <blist> <bibtext> McCabe JA, Lummis SN. Why and how do undergraduates study in groups?. Scholarship of Teaching and Learning in Psychology. 2018; 4; 1: 27-42. 10.1037/stl0000099</bibtext> </blist> <blist> <bibtext> Metcalfe J, Kornell N. A region of proximal learning model of study time allocation. Journal of Memory and Language. 2005; 52; 4: 463-477. 10.1016/j.jml.2004.12.001</bibtext> </blist> <blist> <bibtext> Miller TM, Geraci L. Improving metacognitive accuracy: How failing to retrieve practice items reduces overconfidence. Consciousness and Cognition: An International Journal. 2014; 29: 131-140. 10.1016/j.concog.2014.08.008</bibtext> </blist> <blist> <bibtext> Nelson TO. A comparison of current measures of the accuracy of feeling-of-knowing predictions. Psychological Bulletin. 1984; 84: 93-116. 10.1037/0033-2909.84.1.93</bibtext> </blist> <blist> <bibtext> Nelson TO, Dunlosky J. When people's judgments of learning (JOLs) are extremely accurate at predicting subsequent recall: The "delayed-JOL effect". Psychological Science. 1991; 2; 4: 267-270. 10.1111/j.1467-9280.1991.tb00147.x</bibtext> </blist> <blist> <bibtext> Nelson, T. O, &amp; Narens, L. (1990). Metamemory: A theoretical framework and new findings. In G. Bower (Ed.), The Psychology of Learning and Motivation, 26, 125–173. Academic Press.</bibtext> </blist> <blist> <bibtext> Nokes-Malach TJ, Meade ML, Morrow DG. The effect of expertise on collaborative problem solving. Thinking &amp; Reasoning. 2012; 18; 1: 32-58. 10.1080/13546783.2011.642206</bibtext> </blist> <blist> <bibtext> Pan SC, Rickard TC. Does retrieval practice enhance learning and transfer relative to restudy for term-definition facts?. Journal of Experimental Psychology: Applied. 2017; 23; 3: 278-292</bibtext> </blist> <blist> <bibtext> Pan SC, Rickard TC. Transfer of test-enhanced learning: Meta-analytic review and synthesis. Psychological Bulletin. 2018; 144; 7: 710-756. 10.1037/bul0000151</bibtext> </blist> <blist> <bibtext> Pan, S. C, Zung, I, Imundo, M. N, Zhang, X, &amp; Qiu, Y. (2022). User-generated digital flashcards yield better learning than premade flashcards. Journal of Applied Research in Memory and Cognition. https://doi.org/10.1037/mac0000083</bibtext> </blist> <blist> <bibtext> Pekrun R, Goetz T, Titz W, Perry RPFrydenberg E. Positive emotions in education. Beyond coping: Meeting goals, visions, and challenges. 2002; Oxford University Press: 149-173. 10.1093/med:psych/9780198508144.003.0008</bibtext> </blist> <blist> <bibtext> Putnam AL, Roediger HL. Does response mode affect amount recalled or the magnitude of the testing effect?. Memory and Cognition. 2013; 41: 36-48. 10.3758/s13421-012-0245-x</bibtext> </blist> <blist> <bibtext> Pyc MA, Rawson KA. Testing the retrieval effort hypothesis: Does greater difficulty correctly recalling information lead to higher levels of memory?. Journal of Memory and Language. 2009; 60: 437-447. 10.1016/j.jml.2009.01.004</bibtext> </blist> <blist> <bibtext> Rhodes MGDunlosky J, Tauber SK. Judgments of learning: Methods, data, and theory. The Oxford Handbook of Metamemory. 2016; Oxford University Press: 65-80</bibtext> </blist> <blist> <bibtext> Roediger HL, Butler AC. The critical role of retrieval practice in long-term retention. Trends in Cognitive Sciences. 2011; 15; 1: 20-27. 10.1016/j.tics.2010.09.003</bibtext> </blist> <blist> <bibtext> Roelle J, Endres T, Abel R, Obergassel N, Nückles M, Renkl A. Happy together? On the relationship between research on retrieval practice and generative learning using the case of follow-up learning tasks. Educational Psychology Review. 2023; 35; 4: 102. 10.1007/s10648-023-09810-9</bibtext> </blist> <blist> <bibtext> Rowland CA. The effect of testing versus restudy on retention: A meta-analytic review of the testing effect. Psychological Bulletin. 2014; 140; 6: 1432-1463. 10.1037/a0037559</bibtext> </blist> <blist> <bibtext> Senzaki S, Hackathorn J, Appleby DC, Gurung RA. Reinventing flashcards to increase student learning. Psychology Learning &amp; Teaching. 2017; 16; 3: 353-368. 10.1177/1475725717719771</bibtext> </blist> <blist> <bibtext> Siegel ALM, Castel AD. Age-related differences in metacognition for memory capacity and selectivity. Memory. 2019; 27; 9: 1236-1249. 10.1080/09658211.2019.1645859</bibtext> </blist> <blist> <bibtext> Smith, M, &amp; Weinstein, Y. (2016). Learn how to study using...retrieval practice. The Learning Scientists. https://<ulink href="http://www.learningscientists.org/blog/2016/6/23-1">www.learningscientists.org/blog/2016/6/23-1</ulink></bibtext> </blist> <blist> <bibtext> Soderstrom NC, Bjork RA. Testing facilitates the regulation of subsequent study time. Journal of Memory and Language. 2014; 73: 99-115. 10.1016/j.jml.2014.03.003</bibtext> </blist> <blist> <bibtext> Sumeracki MA, Castillo J. Covert and overt retrieval practice in the classroom. Translational Issues in Psychological Science. 2022; 8; 2: 282-293. 10.1037/tps0000332</bibtext> </blist> <blist> <bibtext> Tauber SK, Witherby AE, Dunlosky J, Rawson KA, Putnam AL, Roediger HL. Does covert retrieval benefit learning of key-term definitions?. Journal of Applied Research in Memory and Cognition. 2018; 7; 1: 106-115. 10.1016/j.jarmac.2016.10.004</bibtext> </blist> <blist> <bibtext> Tullis JG, Finley JR, Benjamin AS. Metacognition of the testing effect: Guiding learners to predict the benefits of retrieval. Memory &amp; Cognition. 2013; 41: 429-442. 10.3758/s13421-012-0274-5</bibtext> </blist> <blist> <bibtext> Vuorre M, Metcalfe J. Measures of relative metacognitive accuracy are confounded with task performance in tasks that permit guessing. Metacognition and Learning. 2022; 17: 269-291. 10.1007/s11409-020-09257-1</bibtext> </blist> <blist> <bibtext> Watson D, Clark LA, Tellegen A. Development and validation of brief measures of positive and negative affect: The PANAS scales. Journal of Personality and Social Psychology. 1988; 54; 6: 1063. 10.1037/0022-3514.54.6.1063</bibtext> </blist> <blist> <bibtext> Wissman KT, Rawson KA. How do students implement collaborative testing in real-world contexts?. Memory. 2016; 24; 2: 223-239. 10.1080/09658211.2014.999792</bibtext> </blist> <blist> <bibtext> Wissman KT, Rawson KA, Pyc MA. How and when do students use flashcards?. Memory. 2012; 20; 6: 568-579. 10.1080/09658211.2012.687052</bibtext> </blist> <blist> <bibtext> Zhu C. Student satisfaction, performance, and knowledge construction in online collaborative learning. Journal of Educational Technology &amp; Society. 2012; 15; 1: 127-136</bibtext> </blist> <blist> <bibtext> Zung I, Imundo MN, Pan SC. How do college students use digital flashcards during self-regulated learning?. Memory. 2022; 30; 8: 923-941. 10.1080/09658211.2022.2058553</bibtext> </blist> </ref> <ref id="AN0181532824-96"> <title> Footnotes </title> <blist> <bibtext> It is not clear why so many participants chose not to report a delayed gJOL. It is possible that, as participants were not told that the delayed portion of the study would include a test, they were surprised by the prompt for a gJOL and were unsure how to respond.</bibtext> </blist> <blist> <bibtext> Additionally, seven participants completed the delayed test late (but within 48-h of the first session of the study). A parallel analysis available in the online supplemental materials indicated that excluding these participants does not change the pattern of results.</bibtext> </blist> <blist> <bibtext> An independent samples t-test examining the effect of Condition at immediate test only (<emph>n</emph> = 141) obtained the same result.</bibtext> </blist> <blist> <bibtext> The number of participants included in each analysis varies slightly because of missing data or because a participant chose not to answer every question.</bibtext> </blist> <blist> <bibtext> Fifteen participants completed the delayed test late (but within 48-h of the first session of the study). A parallel analysis indicated that excluding these participants did not change the pattern of results (see online supplemental materials located at https://osf.io/hac38/?view_only=9ff44668b2184f1b8e412452f3a41640).</bibtext> </blist> <blist> <bibtext> Examining the gamma correlations for all participants who had an immediate test score yielded the same result: <emph>t</emph> (379) = -0.33, <emph>p</emph> = .74, <emph>d</emph> = -0.03, 95% CI [-.09,.06].</bibtext> </blist> </ref> <aug> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib47" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib36" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib42" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib44" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib55" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib57" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib33" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib30" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib51" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib48" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib26" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib37" firstref="ref13"></nolink> <nolink nlid="nl13" bibid="bib45" firstref="ref14"></nolink> <nolink nlid="nl14" bibid="bib20" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib19" firstref="ref19"></nolink> <nolink nlid="nl16" bibid="bib17" firstref="ref22"></nolink> <nolink nlid="nl17" bibid="bib22" firstref="ref23"></nolink> <nolink nlid="nl18" bibid="bib23" firstref="ref24"></nolink> <nolink nlid="nl19" bibid="bib39" firstref="ref26"></nolink> <nolink nlid="nl20" bibid="bib28" firstref="ref27"></nolink> <nolink nlid="nl21" bibid="bib54" firstref="ref28"></nolink> <nolink nlid="nl22" bibid="bib27" firstref="ref29"></nolink> <nolink nlid="nl23" bibid="bib13" firstref="ref30"></nolink> <nolink nlid="nl24" bibid="bib38" firstref="ref31"></nolink> <nolink nlid="nl25" bibid="bib12" firstref="ref34"></nolink> <nolink nlid="nl26" bibid="bib10" firstref="ref35"></nolink> <nolink nlid="nl27" bibid="bib16" firstref="ref36"></nolink> <nolink nlid="nl28" bibid="bib53" firstref="ref40"></nolink> <nolink nlid="nl29" bibid="bib35" firstref="ref41"></nolink> <nolink nlid="nl30" bibid="bib150" firstref="ref42"></nolink> <nolink nlid="nl31" bibid="bib61" firstref="ref45"></nolink> <nolink nlid="nl32" bibid="bib86" firstref="ref46"></nolink> <nolink nlid="nl33" bibid="bib149" firstref="ref47"></nolink> <nolink nlid="nl34" bibid="bib11" firstref="ref52"></nolink> <nolink nlid="nl35" bibid="bib14" firstref="ref53"></nolink> <nolink nlid="nl36" bibid="bib139" firstref="ref56"></nolink> <nolink nlid="nl37" bibid="bib25" firstref="ref62"></nolink> <nolink nlid="nl38" bibid="bib24" firstref="ref64"></nolink> <nolink nlid="nl39" bibid="bib76" firstref="ref66"></nolink> <nolink nlid="nl40" bibid="bib29" firstref="ref69"></nolink> <nolink nlid="nl41" bibid="bib82" firstref="ref71"></nolink> <nolink nlid="nl42" bibid="bib130" firstref="ref73"></nolink> <nolink nlid="nl43" bibid="bib32" firstref="ref75"></nolink> <nolink nlid="nl44" bibid="bib49" firstref="ref76"></nolink> <nolink nlid="nl45" bibid="bib50" firstref="ref77"></nolink> <nolink nlid="nl46" bibid="bib41" firstref="ref78"></nolink> <nolink nlid="nl47" bibid="bib52" firstref="ref79"></nolink> <nolink nlid="nl48" bibid="bib46" firstref="ref81"></nolink> <nolink nlid="nl49" bibid="bib401" firstref="ref82"></nolink> <nolink nlid="nl50" bibid="bib208" firstref="ref84"></nolink> <nolink nlid="nl51" bibid="bib184" firstref="ref88"></nolink> <nolink nlid="nl52" bibid="bib213" firstref="ref89"></nolink> <nolink nlid="nl53" bibid="bib88" firstref="ref90"></nolink> <nolink nlid="nl54" bibid="bib118" firstref="ref91"></nolink> <nolink nlid="nl55" bibid="bib399" firstref="ref94"></nolink> <nolink nlid="nl56" bibid="bib31" firstref="ref100"></nolink> <nolink nlid="nl57" bibid="bib15" firstref="ref108"></nolink> <nolink nlid="nl58" bibid="bib40" firstref="ref109"></nolink> <nolink nlid="nl59" bibid="bib21" firstref="ref113"></nolink> <nolink nlid="nl60" bibid="bib18" firstref="ref114"></nolink> <nolink nlid="nl61" bibid="bib43" firstref="ref116"></nolink> <nolink nlid="nl62" bibid="bib56" firstref="ref119"></nolink> <nolink nlid="nl63" bibid="bib34" firstref="ref121"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1452322 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: When Two Learners Are Better than One: Using Flashcards with a Partner Improves Metacognitive Accuracy – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Megan+N%2E+Imundo%22">Megan N. Imundo</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-4599-4777">0000-0003-4599-4777</externalLink>)<br /><searchLink fieldCode="AR" term="%22Inez+Zung%22">Inez Zung</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-0947-2309">0000-0002-0947-2309</externalLink>)<br /><searchLink fieldCode="AR" term="%22Mary+C%2E+Whatley%22">Mary C. Whatley</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-3609-5630">0000-0003-3609-5630</externalLink>)<br /><searchLink fieldCode="AR" term="%22Steven+C%2E+Pan%22">Steven C. Pan</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-9080-5651">0000-0001-9080-5651</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Metacognition+and+Learning%22"><i>Metacognition and Learning</i></searchLink>. 2025 20(1). – Name: Avail Label: Availability Group: Avail Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Undergraduate+Students%22">Undergraduate Students</searchLink><br /><searchLink fieldCode="DE" term="%22Instructional+Materials%22">Instructional Materials</searchLink><br /><searchLink fieldCode="DE" term="%22Word+Recognition%22">Word Recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Paired+Associate+Learning%22">Paired Associate Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Cooperative+Learning%22">Cooperative Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Strategies%22">Learning Strategies</searchLink><br /><searchLink fieldCode="DE" term="%22Recall+%28Psychology%29%22">Recall (Psychology)</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Metacognition%22">Metacognition</searchLink><br /><searchLink fieldCode="DE" term="%22Individual+Activities%22">Individual Activities</searchLink><br /><searchLink fieldCode="DE" term="%22Study+Habits%22">Study Habits</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Preparation%22">Test Preparation</searchLink><br /><searchLink fieldCode="DE" term="%22Self+Management%22">Self Management</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1007/s11409-024-09406-w – Name: ISSN Label: ISSN Group: ISSN Data: 1556-1623<br />1556-1631 – Name: Abstract Label: Abstract Group: Ab Data: We investigated the benefits of two ways to use flashcards to perform retrieval practice: alone versus with a partner. In three experiments, undergraduate students learned word-definition pairs using flashcards alone (Individual condition) or with another student (Paired condition). Participants then made global judgments of learning (gJOLs; Experiments 1-3), and item-level judgments of learning (iJOLs; Experiment 3). Finally, participants took a cued-recall test after a 5-min delay (Experiments 1-3) and a 24-h delay (Experiments 2-3). In Experiment 1, students in the Paired condition dropped flashcards less often than in the Individual condition (dropping was prohibited in Experiments 2-3). In addition, although final test performance tended to be similar across conditions, inaccurate gJOLs for the immediate test--inflated by ~ 20% relative to actual immediate test performance--were common in the Individual condition but not in the Paired condition in Experiments 1-2. In Experiment 3, we tested whether this difference in metacognitive calibration was due to the Paired condition requiring overt retrieval by instructing participants in the Individual condition to retrieve out loud. With this change, participants in the Individual and Paired conditions reported similarly accurate gJOLs and iJOLs. Taken together, these findings suggest that although performing retrieval practice with flashcards alone versus with a partner yields comparable amounts of learning, doing so with a partner can increase metacognitive accuracy, a benefit possibly driven by the facilitation of overt retrieval. Overall, these findings have implications for self-regulated learning and effective exam preparation. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: Note Label: Notes Group: Note Data: https://osf.io/hac38/?view_only=9ff44668b2184f1b8e412452f3a41640 – Name: DateEntry Label: Entry Date Group: Date Data: 2024 – Name: AN Label: Accession Number Group: ID Data: EJ1452322 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1452322 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s11409-024-09406-w Languages: – Text: English Subjects: – SubjectFull: Undergraduate Students Type: general – SubjectFull: Instructional Materials Type: general – SubjectFull: Word Recognition Type: general – SubjectFull: Paired Associate Learning Type: general – SubjectFull: Cooperative Learning Type: general – SubjectFull: Learning Strategies Type: general – SubjectFull: Recall (Psychology) Type: general – SubjectFull: Learning Processes Type: general – SubjectFull: Metacognition Type: general – SubjectFull: Individual Activities Type: general – SubjectFull: Study Habits Type: general – SubjectFull: Test Preparation Type: general – SubjectFull: Self Management Type: general Titles: – TitleFull: When Two Learners Are Better than One: Using Flashcards with a Partner Improves Metacognitive Accuracy Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Megan N. Imundo – PersonEntity: Name: NameFull: Inez Zung – PersonEntity: Name: NameFull: Mary C. Whatley – PersonEntity: Name: NameFull: Steven C. Pan IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 12 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1556-1623 – Type: issn-electronic Value: 1556-1631 Numbering: – Type: volume Value: 20 – Type: issue Value: 1 Titles: – TitleFull: Metacognition and Learning Type: main |
| ResultId | 1 |