Comparing Mental Effort, Difficulty, and Confidence Appraisals in Problem-Solving: A Metacognitive Perspective
Saved in:
| Title: | Comparing Mental Effort, Difficulty, and Confidence Appraisals in Problem-Solving: A Metacognitive Perspective |
|---|---|
| Language: | English |
| Authors: | Hoch, Emely (ORCID |
| Source: | Educational Psychology Review. Jun 2023 35(2). |
| Availability: | Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ |
| Peer Reviewed: | Y |
| Page Count: | 37 |
| Publication Date: | 2023 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Metacognition, Self Efficacy, Self Evaluation (Individuals), Student Behavior, Difficulty Level, Cognitive Processes, Logical Thinking, Correlation, Reaction Time, Success, Predictor Variables, Problem Solving, Cues |
| DOI: | 10.1007/s10648-023-09779-5 |
| ISSN: | 1040-726X 1573-336X |
| Abstract: | It is well established in educational research that metacognitive monitoring of performance assessed by self-reports, for instance, asking students to report their confidence in provided answers, is based on heuristic cues rather than on actual success in the task. Subjective self-reports are also used in educational research on cognitive load, where they refer to the perceived amount of mental effort invested in or difficulty of each task item. In the present study, we examined the potential underlying bases and the predictive value of mental effort and difficulty appraisals compared to confidence appraisals by applying metacognitive concepts and paradigms. In three experiments, participants faced verbal logic problems or one of two non-verbal reasoning tasks. In a between-participants design, each task item was followed by either mental effort, difficulty, or confidence appraisals. We examined the associations between the various appraisals, response time, and success rates. Consistently across all experiments, we found that mental effort and difficulty appraisals were associated more strongly than confidence with response time. Further, while all appraisals were highly predictive of solving success, the strength of this association was stronger for difficulty and confidence appraisals (which were similar) than for mental effort appraisals. We conclude that mental effort and difficulty appraisals are prone to misleading cues like other metacognitive judgments and are based on unique underlying processes. These findings challenge the accepted notion that mental effort appraisals can serve as reliable reflections of cognitive load. |
| Abstractor: | As Provided |
| Notes: | https://doi.org/10.17605/osf.io/q2p74 |
| Entry Date: | 2023 |
| Accession Number: | EJ1378898 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwE47pqT-9rFKUxhW3zxn_CZAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDEI94Kcp7BM6EZ--igIBEICBm6Gpn9Ga7CdP54CDPN4376Ci8NlpuzvAkAcZJx8QhXo4lFUJ56mIN5UU4TRWPdl50RgoS_TV2_cZxS7cOWKn8MzMZfIAMCv3oRac7zSLOrI0UKP3CzCmXPQVGBf5JyF8uIpMgbhDAILoyPrwzrY1rFxqpnc8m2dJiXeaYqHrwWtkRyJZ6BECfxto_LEmY0Ianf9U7YMbnH4PAb4H Text: Availability: 1 Value: <anid>AN0163936292;epv01jun.23;2023Jun19.13:03;v2.2.500</anid> <title id="AN0163936292-1">Comparing Mental Effort, Difficulty, and Confidence Appraisals in Problem-Solving: A Metacognitive Perspective </title> <p>It is well established in educational research that metacognitive monitoring of performance assessed by self-reports, for instance, asking students to report their confidence in provided answers, is based on heuristic cues rather than on actual success in the task. Subjective self-reports are also used in educational research on cognitive load, where they refer to the perceived amount of mental effort invested in or difficulty of each task item. In the present study, we examined the potential underlying bases and the predictive value of mental effort and difficulty appraisals compared to confidence appraisals by applying metacognitive concepts and paradigms. In three experiments, participants faced verbal logic problems or one of two non-verbal reasoning tasks. In a between-participants design, each task item was followed by either mental effort, difficulty, or confidence appraisals. We examined the associations between the various appraisals, response time, and success rates. Consistently across all experiments, we found that mental effort and difficulty appraisals were associated more strongly than confidence with response time. Further, while all appraisals were highly predictive of solving success, the strength of this association was stronger for difficulty and confidence appraisals (which were similar) than for mental effort appraisals. We conclude that mental effort and difficulty appraisals are prone to misleading cues like other metacognitive judgments and are based on unique underlying processes. These findings challenge the accepted notion that mental effort appraisals can serve as reliable reflections of cognitive load.</p> <p>Emely Hoch and Yael Sidi contributed equally to this work.</p> <hd id="AN0163936292-2">Introduction</hd> <p>Educational research on learning and instruction has developed along two distinct paths that, until recently, have acted largely in isolation (de Bruin &amp; van Merriënboer, [<reflink idref="bib24" id="ref1">24</reflink>]; de Bruin et al., [<reflink idref="bib23" id="ref2">23</reflink>]). First, research on <emph>self-regulated learning</emph> (SRL) has focused on how students plan for a learning task, monitor their performance, and then reflect on the outcome either spontaneously or with guidance (e.g., Zimmerman, [<reflink idref="bib114" id="ref3">114</reflink>]). Within this domain, <emph>metacognitive research</emph> (see Fiedler et al., [<reflink idref="bib32" id="ref4">32</reflink>], for a review; Nelson &amp; Narens, [<reflink idref="bib62" id="ref5">62</reflink>]) aims to expose conditions under which self-appraisals of knowledge may be biased, and the consequences of such bias for subsequent learning-regulation decisions (e.g., allocation of study time, use of study strategies, and help-seeking). Second, instructional design research focuses on developing learning tasks that support effective knowledge acquisition (e.g., Richter et al., [<reflink idref="bib77" id="ref6">77</reflink>]; van Gog, [<reflink idref="bib104" id="ref7">104</reflink>]). Much of this research has been conducted against the backdrop of <emph>cognitive load theory</emph> (CLT, Chandler &amp; Sweller, [<reflink idref="bib21" id="ref8">21</reflink>]), which focuses on optimizing effort investment in learning or task performance.</p> <p>The evident potential of these two massive bodies of research to fertilize each other has drawn attention in recent years (Baars et al., [<reflink idref="bib11" id="ref9">11</reflink>]; Blissett et al., [<reflink idref="bib16" id="ref10">16</reflink>]; de Bruin et al., [<reflink idref="bib23" id="ref11">23</reflink>]; Scheiter et al., [<reflink idref="bib79" id="ref12">79</reflink>]; Seufert, [<reflink idref="bib84" id="ref13">84</reflink>]; van Gog et al., [<reflink idref="bib108" id="ref14">108</reflink>]). In a recent review, Scheiter et al. ([<reflink idref="bib79" id="ref15">79</reflink>]) highlighted that metacognitive and CLT research share the use of subjective self-appraisals of the learning process and learning outcomes. In particular, in metacognitive research, participants are asked to rate (or predict) their own expected or perceived performance immediately before or after performing each task item (e.g., solving a problem). These ratings take the form of metacognitive judgments, such as ease of learning, judgments of learning, feeling of rightness, or confidence (see Ackerman &amp; Thompson, [<reflink idref="bib5" id="ref16">5</reflink>]). For consistency with CLT terminology, hereafter we refer to these judgments as <emph>appraisals</emph>. Metacognitive research has systematically shown that such appraisals are prone to biases, as they are based on heuristic cues and lay theories (Ackerman, [<reflink idref="bib2" id="ref17">2</reflink>]; Koriat et al., [<reflink idref="bib50" id="ref18">50</reflink>]). A massive body of metamemory and meta-reasoning research has exposed how inferential cues (mis)guide metacognitive monitoring (e.g., Ackerman &amp; Beller, [<reflink idref="bib6" id="ref19">6</reflink>]; Bjork et al., [<reflink idref="bib15" id="ref20">15</reflink>]; Castel, [<reflink idref="bib20" id="ref21">20</reflink>]; Koriat, [<reflink idref="bib48" id="ref22">48</reflink>]; Undorf, [<reflink idref="bib100" id="ref23">100</reflink>]). It is also well-established that metacognitive appraisals guide (and thus may mislead) self-regulation decisions (for a review, see Fiedler et al., [<reflink idref="bib32" id="ref24">32</reflink>]). Similarly, in classic studies based on CLT, participants were asked to report the amount of effort they invested and/or the difficulty they experienced as indicators of the cognitive load associated with a particular task design (i.e., load-related appraisals). These self-appraisals are taken as indicators of the effectiveness of the given instructional design.</p> <p>It seems reasonable to consider load-related appraisals as a type of metacognitive judgment (see Scheiter et al., [<reflink idref="bib79" id="ref25">79</reflink>]). From a metacognitive perspective, CLT appraisals, like other documented metacognitive appraisals, are presumably based on heuristic cues and thus are prone to biases (Ackerman, [<reflink idref="bib2" id="ref26">2</reflink>]; Koriat, [<reflink idref="bib47" id="ref27">47</reflink>]). In the present study, we examine the processes that underlie load-related appraisals with the goal of exposing whether they are prone to bias in the same way as known metacognitive appraisals, as well as their predictive value for task outcomes. Toward this end, we utilized metacognitive concepts and research methodologies.</p> <hd id="AN0163936292-3">Metacognitive Appraisals</hd> <p>Metacognitive processes accompany the full course of cognitive activities involved in self-regulated learning, taking place spontaneously in parallel to knowledge processing. Researchers distinguish between two types of metacognitive processes: metacognitive monitoring and metacognitive control. Metacognitive monitoring refers to activities aimed at tracking, reviewing, and assessing the quality of one's cognition, while metacognitive control refers to decision-making about actions to be taken based on the outputs of those monitoring operations (for a review, see Fiedler et al., [<reflink idref="bib32" id="ref28">32</reflink>]; Nelson &amp; Narens, [<reflink idref="bib62" id="ref29">62</reflink>]). For example, when solving a mathematical problem, one assesses the likely correctness of the solution that comes to mind. Based on this assessment, the solver decides whether to provide this solution or to invest more effort in searching for another solution (Efklides, [<reflink idref="bib31" id="ref30">31</reflink>]). An unreliable assessment impairs the consequent decision. Thus, to be effective, control decisions must be based on reliable monitoring (Ackerman &amp; Thompson, [<reflink idref="bib7" id="ref31">7</reflink>]). In empirical research, monitoring reliability is commonly examined by collecting subjective performance appraisals (e.g., confidence ratings for a provided solution), comparing them to objective performance measures (e.g., the participant's performance in the task) and measuring the correspondence between the two measures.</p> <p>An essential theoretical framework for understanding monitoring processes and factors influencing their accuracy is the cue utilization approach, which originated in metacognitive research focused on memorization tasks (Koriat, [<reflink idref="bib47" id="ref32">47</reflink>]). This framework suggests that people do not objectively know their knowledge level for a given task or item but infer it from a complex set of heuristic cues. Koriat ([<reflink idref="bib47" id="ref33">47</reflink>]) classified these heuristic cues as either intrinsic cues inherent to the study items (e.g., ease of processing, familiarity of items, concreteness of items) or extrinsic cues related to the learning context (e.g., number of times items were presented for study). Cue utilization is the extent to which each heuristic cue is considered when making metacognitive appraisals. Metamemory research has extensively investigated this framework, demonstrating how cues guide metacognitive monitoring both uniquely and simultaneously (Ackerman &amp; Beller, [<reflink idref="bib6" id="ref34">6</reflink>]; Castel, [<reflink idref="bib20" id="ref35">20</reflink>]; Koriat, [<reflink idref="bib48" id="ref36">48</reflink>]; Undorf et al., [<reflink idref="bib102" id="ref37">102</reflink>]). The cue utilization approach has also been extended to the domain of problem-solving (for a review and classification, see Ackerman, [<reflink idref="bib2" id="ref38">2</reflink>]; e.g., Finn &amp; Tauber, [<reflink idref="bib33" id="ref39">33</reflink>]; Metcalfe &amp; Finn, [<reflink idref="bib56" id="ref40">56</reflink>], [<reflink idref="bib57" id="ref41">57</reflink>]; Sidi et al., [<reflink idref="bib87" id="ref42">87</reflink>]). In the realm of load-related appraisals, findings suggest that monitoring of effort is an inference-based process similar to metacognitive appraisals (e.g., Dunn &amp; Risko, [<reflink idref="bib29" id="ref43">29</reflink>]; Dunn et al., [<reflink idref="bib27" id="ref44">27</reflink>]; Koriat, [<reflink idref="bib47" id="ref45">47</reflink>]; Raaijmakers et al., [<reflink idref="bib75" id="ref46">75</reflink>]).</p> <p>Notably, while some cues have been found to predict performance and effort reliably, others have been implicated in biasing the monitoring process. For example, one of the most prominent cues within the metacognitive literature is <emph>answer fluency</emph> (cf. processing fluency, Ackerman, [<reflink idref="bib2" id="ref47">2</reflink>]; Thompson et al., [<reflink idref="bib96" id="ref48">96</reflink>]). Answer fluency reflects the ease of processing and relates to the momentary experience of ease or difficulty one feels while performing each task item (Ackerman, [<reflink idref="bib2" id="ref49">2</reflink>]). Answer fluency has been primarily studied in the context of memorization tasks (for a review, see Schwartz &amp; Jemstedt, [<reflink idref="bib82" id="ref50">82</reflink>]) and more recently with reasoning and problem-solving tasks (e.g., Ackerman &amp; Beller, [<reflink idref="bib6" id="ref51">6</reflink>]; Wang &amp; Thompson, [<reflink idref="bib111" id="ref52">111</reflink>]). Answer fluency is often operationalized by measuring response time: i.e., how much time was invested in solving a particular item (question or problem). Overall, response time has been identified as a valid cue, showing inverse relationships with both performance and metacognitive appraisals in various cognitive tasks (Benjamin &amp; Bjork, [<reflink idref="bib13" id="ref53">13</reflink>]; Hertwig et al., [<reflink idref="bib39" id="ref54">39</reflink>]). However, under some conditions, response time has been found to bias metacognitive appraisals (see Finn &amp; Tauber, [<reflink idref="bib33" id="ref55">33</reflink>] for a review). For example, Kelley and Lindsay ([<reflink idref="bib41" id="ref56">41</reflink>]) primed participants with a list of words and then asked them to answer a general knowledge test. They found that the relationship between response time and confidence was similar, specifically that participants assumed quickly retrieved answers were correct, for questions that were and were not constructed to be misleading (by including in the initial list a word that was related to the question yet was not the correct answer). Benjamin et al. ([<reflink idref="bib14" id="ref57">14</reflink>]) showed that people rely on retrieval fluency, operationalized by response time, in a general knowledge task even though response time did not correspond to their actual recall. This resulted in a negative relationship between appraisals and recall performance. In the domain of problem-solving, Ackerman and Zalmanov ([<reflink idref="bib4" id="ref58">4</reflink>]) found that the association between solving time and confidence remained persistent even when the speed at which problems were solved did not predict accuracy. All these findings suggest that relying on response time can misguide the monitoring process.</p> <p>Turning to difficulty appraisals, Kelley and Jacoby ([<reflink idref="bib42" id="ref59">42</reflink>]) showed that response time serves as a potentially misleading cue here as well. In their study, participants were asked to solve anagrams and, in some conditions, were pre-exposed to some of the solution words, resulting in shorter response times for these items. The correlation between response time and difficulty appraisals was consistently high across items and conditions. Kelley and Jacoby attributed this to the biased subjective experience of difficulty cued by response time. These findings raise the question, does response time have differential relationships with confidence, difficulty, and mental effort?</p> <hd id="AN0163936292-4">Inferring Cues for Mental Effort Appraisals from Metacognitive Research</hd> <p>Cognitive load is "a multidimensional construct that represents the load that performing a particular task imposes on the cognitive system of a particular learner" (Paas &amp; Van Merriënboer, [<reflink idref="bib67" id="ref60">67</reflink>], p. 122). CLT's central premise is that the capacity of human working memory to process novel information is limited. Therefore, instructional tasks should be designed to reduce unnecessary load and promote schema acquisition, organization, and automation (van Merriënboer &amp; Kirschner, [<reflink idref="bib109" id="ref61">109</reflink>]). Although some educational research uses objective measures of cognitive load (e.g., Chen et al., [<reflink idref="bib22" id="ref62">22</reflink>]; Korbach et al., [<reflink idref="bib46" id="ref63">46</reflink>]; Szulewski et al., [<reflink idref="bib92" id="ref64">92</reflink>]), most such research has measured cognitive load via self-report measures of effort or difficulty (Naismith et al., [<reflink idref="bib60" id="ref65">60</reflink>]), as these are easier both to administer and to interpret.</p> <p>Self-report cognitive load appraisals are meant to reflect the amount of capacity or resources allocated to accommodate the task demands (Brünken et al., [<reflink idref="bib17" id="ref66">17</reflink>], [<reflink idref="bib18" id="ref67">18</reflink>]). CLT research has traditionally used effort investment and task difficulty interchangeably for this purpose (de Jong, [<reflink idref="bib25" id="ref68">25</reflink>]). Both types of appraisals are commonly elicited using 5-, 7-, or 9-point Likert scales (with the 9-point Likert scale being the one initially proposed by Paas, [<reflink idref="bib64" id="ref69">64</reflink>]). Task difficulty scales typically use wording like (e.g., "the task was very, very easy ... very, very difficult"; e.g., Ayres, [<reflink idref="bib9" id="ref70">9</reflink>]). Effort appraisals have two common variations. One focuses on the person's voluntary investment of effort (e.g., "I invested very, very low mental effort ... very, very high mental effort"; e.g., Paas, [<reflink idref="bib64" id="ref71">64</reflink>]; cf. van Gog &amp; Paas, [<reflink idref="bib106" id="ref72">106</reflink>]), and the other focuses on the task as requiring low/high effort (e.g., "The task required very, very low... very, very high effort").[<reflink idref="bib1" id="ref73">1</reflink>]</p> <p>Scheiter et al. ([<reflink idref="bib79" id="ref74">79</reflink>]) conceptualized the various CLT appraisal items from a metacognitive perspective arguing that the phrasing of effort investment items can reflect motivational and cognitive aspects related to processing and task performance. Particularly, <emph>the effort the individual decided to invest</emph> can be referred to as goal-driven effort (Koriat et al., [<reflink idref="bib49" id="ref75">49</reflink>]). Goal-driven effort relies on top-down processing, reflecting voluntary decisions made by learners. In contrast, <emph>the effort the task demands</emph> can be referred to as data-driven effort (Koriat et al., [<reflink idref="bib49" id="ref76">49</reflink>]). It is based on bottom-up processing, focusing on task characteristics that learners cannot control, similar to asking about the difficulty of the task.</p> <p>Self-appraisals of effort and difficulty are utilized in CLT research under the assumption that they reliably reflect the cognitive resources people allocate to the task. Yet evidence suggests that load-related appraisals might partly reflect biases that stem from unreliable cues (see Scheiter et al., [<reflink idref="bib79" id="ref77">79</reflink>], for a review). For instance, Raaijmakers et al. ([<reflink idref="bib75" id="ref78">75</reflink>]) examined the effects of performance feedback as an external heuristic cue for mental effort appraisals in a complex problem-solving task. In their study, half of the participants received positive feedback, and the other half received negative feedback, irrespective of actual task performance. In three experiments, feedback indeed affected effort appraisals, with the direction depending on feedback valence: positive feedback (pointing to success) was related to lower effort appraisals than negative feedback (pointing to failure). Other studies focused on the timing and frequency of load-related appraisals (Ashburner &amp; Risko, [<reflink idref="bib8" id="ref79">8</reflink>]; Schmeck et al., [<reflink idref="bib80" id="ref80">80</reflink>]; van Gog et al., [<reflink idref="bib105" id="ref81">105</reflink>]). In one study (van Gog et al., [<reflink idref="bib105" id="ref82">105</reflink>]), single delayed appraisals provided at the end of a series of tasks yielded higher cognitive load estimates than the average of appraisals provided immediately after each task item regardless of performance. This effect was particularly pronounced for the more complex tasks in the set. Relatedly, Ashburner and Risko ([<reflink idref="bib8" id="ref83">8</reflink>]) demonstrated that post-trial appraisals were associated with perceptions of greater effort compared to post-whole task appraisals regardless of objective task demands. Looking at cues inherent to the task, Dunn and Risko ([<reflink idref="bib29" id="ref84">29</reflink>]) examined effort appraisals for reading in four types of display conditions, involving rotations to either the presented words, the frame, neither, or both. Their findings showed that participants evaluated displays in which both the words and frame were rotated as more effortful to process than those in which only the words were rotated. However, these appraisals were dissociated from actual performance, as in fact the real difference was between all displays in which the words were rotated and those where only the frame was rotated, with performance being better in the latter. These findings demonstrate that load-related appraisals may rely on external cues, expressing sensitivity to task demands while being dissociated from objective measures of success.</p> <p>Might response time be another culprit biasing load-related appraisals, similar to how response time has been found to bias performance appraisals (e.g., confidence, Finn &amp; Tauber, [<reflink idref="bib33" id="ref85">33</reflink>])? While this question has not been directly investigated, related research does suggest such a link (e.g., Leppink &amp; Pérez-Fuster, [<reflink idref="bib53" id="ref86">53</reflink>]). For example, Dunn et al., ([<reflink idref="bib28" id="ref87">28</reflink>]) drew an association between perceived task effort and task time requirements. They compared effort appraisals (how "effortful" the task is) when a decision-making task presented participants with competing "costs," the time required by a task (low or high), and how error-prone the task was. While error likelihood was more strongly associated with effort appraisals, Dunn et al. reported that both costs predicted effort appraisals across several experimental conditions.</p> <p>Task complexity has also been investigated in CLT research in relation to cognitive load. In particular, tasks involving more interacting elements that need to be stored in working memory impose a higher load (Sweller et al., [<reflink idref="bib90" id="ref88">90</reflink>]). Taking a metacognitive perspective raises the question of whether load-related appraisals are guided by task complexity. Specifically, do people acknowledge the effect of variations in task complexity on the load it imposes on their working memory? Haji et al. ([<reflink idref="bib37" id="ref89">37</reflink>]) investigated the sensitivity of goal-driven effort appraisals and response time to task complexity in simulation-based surgical skill training by comparing two groups faced with low and high complexity levels. Their low-complexity group provided lower effort appraisals than the high-complexity group, with no corresponding differences in response time. This serves as an initial indication that complexity could serve as a cue for load-related appraisals. However, it is still unclear whether complexity also serves as a cue for load appraisals when complexity varies between different items within a single task. Also, notably, research has suggested that the association between item-level complexity and response time is not straightforward, as people may not be motivated to invest the required effort for solving highly complex task items (e.g., Ackerman, [<reflink idref="bib1" id="ref90">1</reflink>]; Hawkins &amp; Heathcote, [<reflink idref="bib38" id="ref91">38</reflink>]; Paas et al., [<reflink idref="bib66" id="ref92">66</reflink>]).</p> <p>Taken together, these findings expose the need to systematically investigate how response time and task (or item-level) complexity are associated with the different mental effort appraisals in order to infer their strength as heuristic cues for mental effort appraisals compared to metacognitive appraisals.</p> <hd id="AN0163936292-5">The Predictive Value of Effort and Difficulty Appraisals</hd> <p>Monitoring accuracy has been a central factor in SRL theory and research due to its causal role in guiding subsequent metacognitive control decisions (Panadero, [<reflink idref="bib68" id="ref93">68</reflink>]; Winne &amp; Perry, [<reflink idref="bib112" id="ref94">112</reflink>]). While the effectiveness of SRL is strongly dependent on the predictive value of load-related appraisals, this has yet to be systematically examined (de Bruin et al., [<reflink idref="bib23" id="ref95">23</reflink>]). As indicated above, research has shown initial evidence that load-related appraisals rely on contextual factors other than the mental effort involved or the difficulty of the task (Raaijmakers et al., [<reflink idref="bib75" id="ref96">75</reflink>]; Schmeck et al., [<reflink idref="bib80" id="ref97">80</reflink>]; van Gog et al., [<reflink idref="bib105" id="ref98">105</reflink>]). For example, Rop et al. ([<reflink idref="bib78" id="ref99">78</reflink>]) showed that mental effort ratings decreased with increasing task experience; however, the results regarding success in the task were inconsistent across two experiments. This finding suggests that the alignment between effort ratings and performance may change over time. In the present study, we aimed to delve into the predictive value of load-related appraisals, namely, their association with task success.</p> <p>As Scheiter et al. ([<reflink idref="bib79" id="ref100">79</reflink>]) suggested, one way to consider the predictive value of load-related appraisals is by looking into the predictive power of an established external criterion. In metacognitive research, the criterion used to validate subjective task appraisals, monitoring accuracy, is the success rate in performing the task at hand. This is done using two measures: calibration and resolution. <emph>Calibration</emph> represents the overall fit between subjective appraisals and actual performance. It can be biased either upwards, resulting in overconfidence, or downwards, resulting in underconfidence. Calibration bias can mislead effort regulation and result in inferior learning outcomes (Metcalfe &amp; Finn, [<reflink idref="bib57" id="ref101">57</reflink>]; Thiede et al., [<reflink idref="bib94" id="ref102">94</reflink>]). <emph>Resolution</emph> represents the extent to which people distinguish in their confidence between correct and incorrect responses. It is measured at the individual level as a within-participant correlation between confidence and success in each item. Scheiter et al. ([<reflink idref="bib79" id="ref103">79</reflink>]) maintained that of the two measures, resolution is the more relevant for load-related appraisals due to it being a relative measure (referring to the variability of ratings across task items) rather than an absolute measure (using the numerical value of each rating). Thus, in the present study, we examined the resolution of the various load-related appraisals compared to metacognitive confidence appraisals.</p> <hd id="AN0163936292-6">Research Questions and Study Overview</hd> <p>Following this review, we aimed to examine three main research questions.</p> <p></p> <ulist> <item> RQ1. Are there differences in the extent to which response time serves as a cue for mental effort appraisals, difficulty appraisals, and confidence appraisals?</item> </ulist> <p>The literature reviewed above suggests that all these types of appraisals are guided to some extent by response time. However, no research thus far has compared these relationships within one study to examine their relative strength. As fluency, by definition, relates to the momentary experience of ease or difficulty, we expected that all load-related appraisals would be found to rely more strongly on response time as a cue compared to confidence. This is supported by evidence in the metacognitive literature that confidence shows only a modest relationship with response time (Ackerman, [<reflink idref="bib1" id="ref104">1</reflink>]).</p> <p></p> <ulist> <item> RQ2. Are there differences in the extent to which mental effort appraisals, difficulty appraisals, and confidence appraisals predict actual accuracy in each task item?</item> </ulist> <p>We expected that confidence would predict task accuracy based on ample prior research. How the different mental effort appraisals relate to task accuracy is less clear. Yet this is important to examine, because there are reasons to believe that goal-driven effort and data-driven effort may show unique relationships with task accuracy. As explained above, Koriat et al.'s ([<reflink idref="bib49" id="ref105">49</reflink>]) theory suggests that goal-driven effort, operationalized by additional time invested when motivation to succeed rises, reflects the voluntary decision to invest resources into a task. This could result in investing effort in vain due to higher internal motivation to succeed, which might not necessarily result in actual higher success. However, data-driven effort, according to Koriat et al.'s theory, reflects task demands in a similar manner to confidence. Therefore, one could expect data-driven effort to better predict task accuracy compared to goal-driven effort.</p> <p>The association of perceived difficulty with task accuracy largely relies on how people interpret requests for difficulty appraisals. Do they believe they are being asked about the amount of effort they felt they personally had to invest or about difficulty as a facet of the problem? This is an open question on which our study can shed light by comparing the strength of the relationship between difficulty to task accuracy and between difficulty and response time with relations of different appraisals.</p> <p></p> <ulist> <item> RQ 3: Are there differences in the extent to which the complexity of the problem serves as a cue for mental effort appraisals, difficulty appraisals, and confidence appraisals?</item> </ulist> <p>Empirical evidence from both CLT research and metacognitive research has shown that both load-related appraisals and confidence are sensitive to differences in task elements related to complexity (e.g., Ayres, [<reflink idref="bib9" id="ref106">9</reflink>]; Paas &amp; Van Merriënboer, [<reflink idref="bib67" id="ref107">67</reflink>]; Schmeck et al., [<reflink idref="bib80" id="ref108">80</reflink>]; van Gog et al., [<reflink idref="bib105" id="ref109">105</reflink>]). However, no studies have yet compared load-related and confidence appraisals in terms of their sensitivity to complexity, leaving this an open research question.</p> <p>This study was designed to test these research questions systematically using three experiments. All three experiments were designed to examine RQ1 on response time as a cue for the different appraisals and RQ2 on the predictive value of the various appraisals for success in the task. Experiment 3 also examines RQ3 on item complexity. In all three experiments, participants completed reasoning and problem-solving tasks in a multiple-choice test format, followed by mental effort, task difficulty, or metacognitive confidence appraisals. While the tasks differed between the experiments, the designs and procedures were very similar. All experiments included four groups differing only in the type of appraisal provided immediately after completing each task item: goal-driven effort, data-driven effort, task difficulty, or metacognitive confidence. To examine our research questions, we analyzed the associations between the various appraisals and the relevant outcome of interest (response time, task accuracy, or item complexity), as appropriate. In the first experiment, we employed a verbal logic task widely used in cognitive and metacognitive research, the cognitive reflection test (CRT, Frederick, [<reflink idref="bib35" id="ref110">35</reflink>]). The task consists of misleading verbal mathematical problems (word problems) designed so that the first solution that usually comes to mind is an incorrect but predictable one, while most respondents can arrive at the correct solution with more effort investment. The task calls for heterogeneous appraisals across items, which is essential for examining within-participant correlations between appraisals on the one hand and response times or accuracy on the other.</p> <p>Notably, most empirical research on meta-memory processes has used verbal tasks like memorizing words and answering knowledge questions. When considering meta-reasoning tasks, verbal tasks dominate research as well, with the CRT, as used in experiment 1 being an example (see Ackerman &amp; Thompson, [<reflink idref="bib7" id="ref111">7</reflink>] for examples and a review). The scarce studies utilizing non-verbal tasks show some similarities in the metacognitive mechanisms involved (Lauterman &amp; Ackerman, [<reflink idref="bib51" id="ref112">51</reflink>]; Reber et al., [<reflink idref="bib76" id="ref113">76</reflink>]). To contribute to studying non-verbal reasoning processes and examine our findings' robustness, in experiment 2, we used a non-verbal problem-solving task: the missing tan task (MTT, Ackerman, [<reflink idref="bib3" id="ref114">3</reflink>]). The MTT is a challenging non-verbal reasoning task that relies on cognitive processes also involved in geometry, navigation, and design. Participants are presented with silhouettes generated from geometric pieces (called tans). The silhouettes are comprised of six pieces drawn from a pool of seven. Participants' task is to identify which piece is not needed to form each silhouette solely through mental visualization without being able to manipulate the presented pieces.</p> <p>In the third experiment, we replicated experiments 1 and 2 using yet another non-verbal task. In order to also address RQ3, for this experiment, we chose a task which offers inherent variations in item complexity: the mental rotation task (MRT, Shepard &amp; Metzler, [<reflink idref="bib85" id="ref115">85</reflink>]). Here, participants are presented with a set of rotated stimuli and must mentally rotate each one to align with a criterion stimulus in order to determine which alternative matches the criterion figure (Searle &amp; Hamm, [<reflink idref="bib83" id="ref116">83</reflink>]). Based on previous research, we used the angle of rotation as an objective measure of item complexity (Shepard &amp; Metzler, [<reflink idref="bib85" id="ref117">85</reflink>]). Complexity from a CLT perspective usually refers to the number of interacting elements of a task that are processed simultaneously in working memory; the greater the number of interacting elements, the higher the cognitive load (e.g., van Gog &amp; Sweller, [<reflink idref="bib107" id="ref118">107</reflink>]). In the present task, we assume that the larger the rotation angle (i.e., the more mental rotation required), the higher the cognitive load.</p> <p>While the three tasks rely on different reasoning skills, they are similar in several ways, supporting comparisons between the findings. (a) The tasks call for deliberate reasoning processes. (b) The tasks are in a multiple-choice format. (c) The tasks allow for the generation of a wide variety of items, resulting in different success rates, while the time spent processing the stimuli remains similar across items. This feature is important for the variability in response time to reflect its variability within participants across success rates. (d) The tasks involve uncertainty as to whether the solution provided is correct (unlike, e.g., fitting a jigsaw puzzle piece into the right spot, which usually involves no uncertainty). (e) Having more than fifteen items allows for robust within-participant statistical analyses.</p> <p>Finally, in all experiments, we collected data on individual differences in participants' self-perceptions or beliefs about their traits, abilities, or knowledge. More specifically, (a) in experiment 1, participants reported on their need for cognition (Cacioppo &amp; Petty, [<reflink idref="bib19" id="ref119">19</reflink>]). This scale reflects cognitive style, or the extent to which the individual enjoys taking part in effortful cognitive activities. In experiments 2 and 3, participants completed scales capturing (b) test anxiety (Taylor &amp; Deane, [<reflink idref="bib93" id="ref120">93</reflink>]), reflecting self-doubt about their ability to succeed in a particular task type and (c) beliefs about the malleability of intelligence (Dweck et al., [<reflink idref="bib30" id="ref121">30</reflink>]), reflecting implicit theories of intelligence as malleable (growth mindset) or unmalleable (fixed mindset). Metacognitive research has shown that these constructs are associated with confidence or similar appraisals of expected success (e.g., Jonsson &amp; Allwood, [<reflink idref="bib40" id="ref122">40</reflink>]; Kirk-Johnson et al., [<reflink idref="bib43" id="ref123">43</reflink>]; Miele et al., [<reflink idref="bib58" id="ref124">58</reflink>]; Miesner &amp; Maki, [<reflink idref="bib59" id="ref125">59</reflink>]; Petty et al., [<reflink idref="bib72" id="ref126">72</reflink>]). Thus, we sought to examine them as potential moderators for how the different appraisals are related to response time and accuracy (Scheiter et al., [<reflink idref="bib79" id="ref127">79</reflink>]).</p> <hd id="AN0163936292-7">Experiment 1</hd> <p>Experiment 1 aimed to examine the associations between the various appraisals and response time, as well as with task accuracy, by applying a verbal logic task.</p> <hd id="AN0163936292-8">Method</hd> <p></p> <hd id="AN0163936292-9">Participants and Design</hd> <p>Data were collected online through the Prolific (<ulink href="http://www.prolific.co">www.prolific.co</ulink>) participant pool. Participants were required to be at least 20 years old, to speak English fluently (to ensure they understood the instructions), and to have no learning disabilities. Participation was voluntary, anonymous, and remunerated with 2GBP. We excluded data from participants who encountered technical problems during the experiment (6 participants), did not follow instructions (e.g., admitted to having engaged in other activities while completing the tasks, 5 participants), had little variability in appraisals (i.e., SD &lt; 4; 4 participants), or provided valid responses to less than 75% of the trials (see exclusion criteria for single trials under Data Preparation, 7 participants). The final sample comprised data from 284 participants (age: <emph>M</emph> = 34.5 years, SD = 10.9; 146 females, 125 males; age and gender missing for 13 participants). Participants were randomly assigned to one of four groups that differed in the type of appraisal they were asked to provide: goal-driven effort (<emph>n</emph> = 70), data-driven effort (<emph>n</emph> = 71), task difficulty (<emph>n</emph> = 70), and confidence (<emph>n</emph> = 73).</p> <hd id="AN0163936292-10">Materials: Misleading Math and Logic Problems (CRT Tasks)</hd> <p>The original CRT (Frederick, [<reflink idref="bib35" id="ref128">35</reflink>]) contains three misleading math problems where the first solution that commonly comes to mind is a wrong but predictable one, but a little deliberative effort can lead most respondents to the correct solution. While the CRT is suitable for examining our research question, the original CRT problem set is so widely used as to raise concerns regarding participants' pre-exposure to the task. Also, to allow robust within-participant statistical analyses, it was essential to have a larger number of task items. Therefore, for the present study we used a collection of 17 misleading math and logic problems based on several resources, fitted to a multiple-choice format, and pretested (see Appendix 1). For instance, one of the items was "25 soldiers are standing in a row 3 m from each other. How long is the row?" Participants had to choose from four answers: a) 3 m, b) 69 m, c) 72 m, or d) 75 m. The answer, which is expected to jump quickly to mind, 75 m, is wrong. The correct answer is 72 m (Oldrati et al., [<reflink idref="bib63" id="ref129">63</reflink>]).</p> <hd id="AN0163936292-11">Appraisals</hd> <p>Each participant provided one of four appraisals for all items in the study depending on their experimental group: goal-driven effort, data-driven effort, task difficulty, or confidence. Participants entered their appraisal by sliding a bar on a horizontal slider using the mouse. All scales ranged from 0 to 100. In the goal-driven effort condition, the question was "How much effort did you invest in solving the problem?," and the scale ranged from <emph>very, very low effort</emph> (0) to <emph>very, very high effort</emph> (<reflink idref="bib100" id="ref130">100</reflink>). In the data-driven effort condition, the scale anchors were the same, but participants were asked "How much effort did the problem require?" In the task difficulty condition, the question was "How difficult was the problem?", and the scale anchors were <emph>very, very easy</emph> (0) and <emph>very, very difficult</emph> (<reflink idref="bib100" id="ref131">100</reflink>). Finally, participants in the confidence condition were asked "How confident are you that your solution is correct?," on a scale from <emph>a wild guess</emph> (0) to <emph>definitely sure</emph> (<reflink idref="bib100" id="ref132">100</reflink>). The scales were adapted from metacognitive research in which such scales are commonly used for different types of metacognitive appraisals. Their advantage is their receptiveness to comparison with actual task performance (also ranging from 0 to 100), which allows calculating monitoring accuracy.</p> <hd id="AN0163936292-12">Objective Measures</hd> <p>Response time was defined as the time each participant took to respond to each task item (in seconds). Our second objective measure was item-level task accuracy or providing a correct answer for each task item.</p> <hd id="AN0163936292-13">Background Variables2</hd> <p>Previous knowledge[<reflink idref="bib3" id="ref133">3</reflink>] of the task was assessed by this question: "Have you ever encountered one or more of the problems that you solved here in other studies? If you did, please write what you remember from those problems. If not, please enter "All new." Need for cognition (Cacioppo &amp; Petty, [<reflink idref="bib19" id="ref134">19</reflink>]; Lins de Holanda Coelho et al., [<reflink idref="bib54" id="ref135">54</reflink>]) was measured with the short (six-item) scale assessing the extent to which people enjoy engaging in the process of thinking (e.g., "I really enjoy a task that involves coming up with new solutions to problems"). Responses were given on a 7-point Likert scale (1 = <emph>strongly disagree</emph>, 7 = <emph>strongly agree</emph>), with two items reverse-coded (Cronbach's <emph>α</emph> = 0.84). Experience with solving puzzles ("How often do you solve puzzles or play thought-provoking games?") was assessed with a single item on a scale from 1 (<emph>almost never</emph>) to 7 (<emph>daily</emph>).</p> <p>One item serving as an attention check was presented amid the need for cognition questions using the same scale ("I like taking logic exams. Please ignore this statement, wait at least 4 s and then respond by level two"). Less than 6% of participants failed to pass the attention check. However, these participants were only excluded if they also showed another indication of inattention (failure to follow study instructions; see under the "Participants and Design" section).</p> <p>All means and standard deviations of the background variables as a function of the type of appraisal are shown in Appendix 2.</p> <hd id="AN0163936292-14">Procedure</hd> <p>Participants first received general information about the procedure and gave their consent to participate. They then received the instructions for the CRT task: to solve verbally phrased math and logic problems by choosing one out of four solution options. In the training phase, participants solved an example problem for which the correct answer was provided. With a second example, one of the four appraisals corresponding to the relevant condition (goal-driven effort, data-driven effort, task difficulty, or confidence) was introduced. Participants were told that as the problems were not trivial, ratings across the entire range of the scale, including low and intermediate rating levels, were expected. Following the examples, participants solved the 17 CRT items in random order. Work on the items was self-paced. Once a solution option was selected, it could not be changed, and an appraisal was elicited. After completing all CRT items, participants answered the background questions. Then, participants were asked whether they had pursued other activities during the experiment, had additional comments, or had encountered technical problems. Participants completed the experiment at about 15 min on average.</p> <hd id="AN0163936292-15">Data Preparation</hd> <p>Data preparation and all analyses were conducted in R (R Core Team, [<reflink idref="bib74" id="ref136">74</reflink>]). Trials with extraordinarily short (RT &lt; 2 s; 11 trials) or extraordinarily long response times (RT &gt; 180 s; 130 trials) were excluded, as were trials where participants left the experimental environment to do something else (168 trials). This left 4752 trials included in the data analyses.</p> <p>As we were interested in the strength of the relationship between different appraisals and objective measures, this association was calculated as the within-participant correlation across items between each appraisal and the objective measures of response time and accuracy. As the objective measures under investigation differed in their scale levels, Pearson correlations were used to calculate the correlation between response time (an interval scaled variable) and appraisals, while the Goodman–Kruskal γ rank correlation was used to calculate the correlations between task accuracy (a dichotomous variable) and appraisals. Note that differences between the appraisals were expected simply because of their opposing reference points: easy items should naturally yield high confidence values but low values for task difficulty and effort. Thus, to statistically address RQ1 and RQ2, confidence appraisals were reversed to match the other appraisals' direction of association with item difficulty. It should also be noted that the correlations are interpreted differently for RQ1 and RQ2. In RQ1, the correlation indicates the degree to which response time serves as a cue for the appraisal. In RQ2, the correlation reflects the extent to which the appraisal has predictive value for task accuracy.</p> <hd id="AN0163936292-16">Results and Discussion</hd> <p>Item success rates (i.e., the percentage of participants who correctly solved a given item) ranged from 26.6 to 81.9%. On average, participants needed 24.7 s (SD = 12.5) for each item, and they correctly solved 56.0% (SD = 21.7) of the items. Means and standard deviations of appraisals, response time, and success as a function of appraisal are shown in Table 1. The four appraisal groups did not differ in either their response times, <emph>F</emph>(<reflink idref="bib3" id="ref137">3</reflink>, 280) = 2.12, MSE = 154.1, <emph>p</emph> = 0.098, η<sups>2</sups> = 0.02, or in accuracy, <emph>F</emph> &lt; 1.</p> <p>Table 1 Means and standard deviations of response time, accuracy, and appraisals as a function of type of appraisal in experiments 1–3</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;Goal-driven effort&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Data-driven effort&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Task difficulty&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Confidence&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 1 (CRT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;27.33 (16.94)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;25.29 (12.12)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;23.82 (10.31)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;22.31 (8.92)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;53.66 (20.77)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;56.22 (21.56)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;58.58 (21.42)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;55.61 (23.28)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;36.64 (21.94)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;35.12 (19.14)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;29.91 (15.33)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;80.72 (9.42)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 2 (MTT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;26.46 (14.43)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;26.31 (15.84)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;27.16 (16.05)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;24.22 (13.57)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;42.59 (16.20)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;41.52 (13.34)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;45.81 (17.54)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;42.26 (16.65)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;55.28 (16.43)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;53.60 (15.23)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;55.74 (14.72)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;67.66 (12.52)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 3 (MRT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;21.93 (10.62)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;23.68 (12.33)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;21.06 (10.81)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;21.01 (12.34)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;77.23 (20.19)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;69.75 (22.82)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;72.63 (23.14)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;72.87 (23.68)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;47.06 (20.80)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;42.95 (14.08)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;42.15 (14.91)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;81.44 (12.19)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Response time gives the average time (in seconds) for answering an item. Accuracy is calculated as the percentage of correct answers. Appraisals were given on a 0–100 scale. In experiments 2 and 3, confidence appraisals were given on a 20–100 scale to account for the probability of getting the item right just by guessing</p> <p>To describe the relationships between the appraisals and the variables of interest (response time and accuracy), Table 2 presents the mean within-participant correlations as a function of appraisal. Unsurprisingly, as mentioned above, the mathematical signs for the correlations with confidence were the inverse of those for the other appraisals. To test whether these relationships are meaningful, the mean correlations were tested against 0. Adjusted <emph>p</emph>-values were calculated using Bonferroni correction to account for multiple tests (i.e., four tests for the associations between the appraisals and response time and again between the appraisals and accuracy). As seen in Table 2, all correlations were significant except for one (the correlation between goal-driven effort and accuracy). These findings indicate that response time does indeed serve as a cue for appraisals and that appraisals do predict actual accuracy in the task. Therefore, it is plausible to examine differences between the four types of appraisals in both cases.</p> <hd id="AN0163936292-17">RQ 1: Are There Differences in the Extent to Which Response Time Serves as a Cue for Mental E...</hd> <p>Table 2 Means and standard deviations of the appraisal-response time, appraisal-accuracy, and appraisal-item complexity associations as a function of type of appraisal in experiments 1–3 and across all three experiments</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;Goal-driven effort&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Data-driven effort&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Task difficulty&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Confidence&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 1 (CRT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.49 (0.32)&lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.49 (0.32)&lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.39 (0.27)&lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.18 (0.28)&lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.06 (0.39)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.16 (0.39) &lt;sup&gt;**&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.17 (0.43) &lt;sup&gt;**&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.34 (0.44) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 2 (MTT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.43 (0.30) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.37 (0.34) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.36 (0.24) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.23 (0.24) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.16 (0.31) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.24 (0.29) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.38 (0.23) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.34 (0.27) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 3 (MRT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.56 (0.29) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.52 (0.25) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.48 (0.25) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.25 (0.26) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.34 (0.46) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.38 (0.33) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.44 (0.38) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.55 (0.37) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-item complexity&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.22 (0.21) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.21 (0.23) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.28 (0.23) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.16 (0.17) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Average across experiments 1&amp;#8211;3&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-response time&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.49 (0.31) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.46 (0.31) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.41 (0.26) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.21 (0.27) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt; Appraisal-accuracy&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.18 (0.40) &lt;sup&gt;**&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.25 (0.35) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt; &amp;#8722; 0.32 (0.38) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;0.41 (0.38) &lt;sup&gt;***&lt;/sup&gt;&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Associations are calculated as within-participant correlations and thus range from − 1 to 1. Goodman and Kruskal's γ was used to determine the relationship between appraisal and accuracy. Pearson correlations were used for the relationship between appraisal and response time, as well as appraisal and item complexity. Bonferroni adjustment was applied to <emph>p</emph>-values. Asterisks indicate whether the averaged correlations significantly differ from 0: <sups>+</sups><emph>p</emph> &lt; 0.1, <sups>*</sups><emph>p</emph> &lt; 0.05, <sups>**</sups><emph>p</emph> &lt; 0.01, and <sups>***</sups><emph>p</emph> &lt; 0.001</p> <p>To evaluate whether the different appraisals rely on response time in similar ways (see the CRT columns in Fig. 1A), the strength of the appraisal-response time correlation was compared between the different appraisal types using ANOVA[<reflink idref="bib4" id="ref138">4</reflink>] (confidence appraisals were reversed for this comparison). There was a significant large effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref139">3</reflink>, 280) = 16.93, <emph>MSE</emph> = 0.09, <emph>p</emph>&lt; 0.001, η<sups>2</sups>=0.15. Post hoc pairwise comparisons (<emph>t</emph>-tests with Bonferroni adjustment for <emph>p</emph>-values) showed that the correlation with response time was significantly weaker for confidence than for each of the other three appraisal types (goal-driven effort, data-driven effort, and task difficulty, all <emph>p</emph><subs>s</subs> &lt; 0.001). All other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. These findings present initial evidence that response time serves as a cue for load-related appraisals, that is, goal- and data-driven effort appraisals as well as task difficulty appraisals.</p> <p>Graph: Fig. 1Mean size of correlations of A appraisals with response time, B appraisals with accuracy, and C appraisals with item complexity as a function of the type of appraisal for the three different tasks in the three experiments (CRT, cognitive reflection task; MTT, missing tan task; MRT, mental rotation task). Note that reversed confidence appraisals were used to calculate the correlations in the confidence group. Error bars show ± 1 standard error</p> <p>In accordance with the metacognitive literature, confidence was negatively associated with response time. However, the metacognitive literature suggests that response time can be an unreliable cue for metacognitive judgments (e.g., Ackerman, [<reflink idref="bib3" id="ref140">3</reflink>]; Finn &amp; Tauber, [<reflink idref="bib33" id="ref141">33</reflink>]). Indeed, as we expected, load-related appraisals seem to rely more strongly on response time as a cue compared to confidence appraisals. Our findings support the idea that load-related appraisals might be biased by response time, as other metacognitive judgments are (see Scheiter et al., [<reflink idref="bib79" id="ref142">79</reflink>]).</p> <hd id="AN0163936292-18">RQ 2: Are There Differences in the Extent to Which Mental Effort Appraisals, Difficulty Appra...</hd> <p>To evaluate the predictive value of appraisals for accuracy (see the CRT columns in Fig. 1B), the strength of the appraisal-accuracy correlation was compared between the different appraisal types using ANOVA (confidence appraisals were reversed for this comparison). There was a significant medium effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref143">3</reflink>, 272) = 5.61, <emph>MSE</emph> = 0.17, <emph>p</emph> &lt; 0.001, η<sups>2</sups> = 0.06. Post hoc pairwise comparisons (<emph>t</emph>-tests with Bonferroni adjustment for <emph>p</emph>-values) showed that the correlation with accuracy was significantly stronger for confidence than for goal-driven effort appraisals, <emph>p</emph> &lt; 0.001. All other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. Thus, while the findings for response time showed a clear distinction between confidence and the load-based appraisal types, at least in the CRT, this distinction was weaker when considering the association with accuracy. Confidence differed from goal-driven effort but not from the other appraisals.</p> <p>These predictive values of appraisals for accuracy align with the argument that goal-driven appraisals can reflect labor-in-vain, effort that does not yield improvement (Koriat et al., [<reflink idref="bib49" id="ref144">49</reflink>]), resulting in a weaker relationship with accuracy compared to confidence appraisals. In this case, as difficulty appraisals shared a similar relationship with accuracy relative to data-driven effort and confidence, it appears to have been interpreted by participants to reflect the difficulty of the task rather than the effort they chose to invest in a goal-driven manner.</p> <hd id="AN0163936292-19">Experiment 2</hd> <p>Experiment 2 was designed to replicate the findings of experiment 1 with a different task. Here, a non-verbal problem-solving task was applied to examine the robustness of the results across different tasks.</p> <hd id="AN0163936292-20">Method</hd> <p></p> <hd id="AN0163936292-21">Participants and Design</hd> <p>As in experiment 1, data were collected via Prolific with the same requirements and compensation for participation. Data were excluded for participants who encountered technical problems during the experiment (2 participants), who did not follow the instructions or showed signs of low effort (e.g., failure in four very easy verification items, 15), who provided data for less than 75% of the trials (<reflink idref="bib6" id="ref145">6</reflink>), or who did not consent to the use of their data (<reflink idref="bib1" id="ref146">1</reflink>). Our final sample comprised 224 participants (age: <emph>M</emph> = 31.9, <emph>SD</emph> = 10.5; 127 female, 94 male; age and gender missing for 3 participants). Again, as in experiment 1, participants were randomly assigned to one of four groups: goal-driven effort (<emph>n</emph> = 54), data-driven effort (<emph>n</emph> = 60), task difficulty (<emph>n</emph> = 60), and confidence (<emph>n</emph> = 50).</p> <hd id="AN0163936292-22">Materials: Missing Tan Task (MTT)</hd> <p>In the MTT (Ackerman, [<reflink idref="bib3" id="ref147">3</reflink>]) the silhouette of a figure was shown together with a legend including seven geometric pieces: a square, a parallelogram, two large triangles, two small triangles, and one intermediate triangle (see Fig. 2 for an example). The seven geometric pieces were marked with letters from A to E, with two each of the A and B pieces (the large and small triangles respectively). The silhouette was generated from six of the seven pieces, which could be rotated in either direction or flipped to their mirror image but could not overlap. The task was to identify which of the seven geometric pieces was not needed to reproduce the silhouette through mental visualization alone without being able to physically manipulate or move the pieces. Responses were chosen from five multiple-choice options corresponding to the labels (A to E).</p> <p>Graph: Fig. 2An example of the missing tan task from Ackerman ([<reflink idref="bib3" id="ref148">3</reflink>]). Participants were asked to indicate which of the pieces (A to E) does not fit in the silhouette. In this example, the correct answer is C since all other pieces are needed to reproduce the given silhouette</p> <p>We used the original stimuli generated, piloted, and selected by Ackerman ([<reflink idref="bib3" id="ref149">3</reflink>]), with a total of 32 silhouettes. Two of the items were used as examples during the task instructions. Two other items expected to be easier than the others, with success rates &gt; 90%, served for attention verification but were included in the analyses, nevertheless. Thus, 30 items were used for the analyses for each participant.</p> <hd id="AN0163936292-23">Appraisals</hd> <p>The appraisals (goal-driven effort, data-driven effort, task difficulty, and confidence) and appraisal procedures were the same as those for experiment 1. The only exception was that confidence was provided on a different scale (20 to 100 rather than 0 to 100). This change was intended to draw participants' attention to the fact that with five multiple-choice options, they had a 20% chance of being correct just by guessing.</p> <hd id="AN0163936292-24">Objective Measures</hd> <p>As in experiment 1, response time and accuracy were examined as objective measures that may relate to the appraisals.</p> <hd id="AN0163936292-25">Background Variables</hd> <p>Participants were asked to provide a single-item overall judgment of their performance in the MTT ("How many problems do you think you answered correctly?", from 0 to 30). A single item was also used to elicit participants' experience with solving puzzles ("How often do you solve puzzles or play thought-provoking games?") on a scale of 1 (<emph>almost never</emph>) to 4 (<emph>every day</emph>).</p> <p>Test anxiety was assessed with five statements about how the participant generally feels about exams (e.g., "During tests I feel very tense."; Taylor &amp; Deane, [<reflink idref="bib93" id="ref150">93</reflink>]). Participants were asked to rate these statements on a 4-point scale (1 = <emph>almost never</emph>, 4 = <emph>almost always</emph>; Cronbach's <emph>α</emph> = 0.87). High values on this scale indicate more substantial test anxiety. Participants were also asked when they last took an exam using a 4-point scale (1 = <emph>during the last month</emph>, 2 = <emph>between 1 and 6 months ago</emph>, 3 = <emph>between 6 and 12 months ago</emph>, and 4 = <emph>more than 12 months ago</emph>).</p> <p>To assess participants' mindsets about the malleability of intelligence (fixed vs. growth), they were asked to rate their agreement with four statements (e.g., "You can always substantially change how intelligent you are"; Dweck et al., [<reflink idref="bib30" id="ref151">30</reflink>]) on a 7-point Likert scale (1 = <emph>strongly disagree</emph>, 7 = <emph>strongly agree</emph>; Cronbach's <emph>α</emph> = 0.88), with low values indicating a fixed mindset and high values indicating a growth mindset.</p> <p>One item serving as an attention check was presented together with the test anxiety questions using the same scale ("Please describe how you generally feel regarding exams: Physical activity promotes my thinking skills. Please ignore this statement and answer by level two"). Less than 6% of participants failed to pass the attention check. However, as in experiment 1, these participants were only excluded if they showed another indication of inattention (failure to follow instructions or signs of low effort; see Participants and Design).</p> <p>Means and standard deviations of the background variables as a function of the type of appraisal are shown in Appendix 2.</p> <hd id="AN0163936292-26">Procedure</hd> <p>The overall procedure was similar to that of experiment 1. In the specific instructions, participants were told their task was to identify which of seven geometric pieces was not needed to reproduce the silhouette by clicking on the corresponding letter. As in experiment 1, instructions were given for the task with one example, and a second example was used to introduce the appraisals. Participants then worked in a self-paced manner on the 30 MTT items, which were presented randomly. After every ten items, they were told of the number of items already completed. Background questions were presented at the end. The full experiment took about 25 min.</p> <hd id="AN0163936292-27">Data Preparation</hd> <p>As in experiment 1, we excluded trials with extraordinarily short (RT &lt; 2 s, 89 trials) or long response times (RT &gt; 180 s, 68 trials), as well as trials in which participants left the experimental environment to do something else (61 trials). This left 6574 trials to be analyzed. Data preparation and all analyses were performed as in experiment 1.</p> <hd id="AN0163936292-28">Results and Discussion</hd> <p>Item success rates ranged from 8.3 to 81.2%. On average, participants needed 26.1 s (SD = 15.0) to answer each item and were able to solve 43.1% (SD = 16.0) of the items correctly. Means and standard deviations of appraisals, response time, and accuracy as a function of the type of appraisal are shown in Table 1. The four appraisal groups did not differ in either response times or accuracy, both <emph>F</emph> &lt; 1.</p> <p>As in experiment 1, the relationships between the appraisals and the response time and accuracy variables are given as mean within-participant correlations. As seen in Table 2, all mean correlations differed significantly from 0, indicating that overall, response time served as a cue for the appraisals, and that appraisals predicted accuracy in the task.</p> <hd id="AN0163936292-29">RQ 1: Are There Differences in the Extent to Which Response Time Serves as a Cue for Mental E...</hd> <p>To evaluate whether the different appraisals rely on response time in similar ways (see the MTT columns in Fig. 1A), the strength of the appraisal-response time correlation was compared between the different appraisal types using ANOVA. There was a significant medium effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref152">3</reflink>, 220) = 4.51, <emph>MSE</emph> = 0.08, <emph>p</emph>&lt; 0.005,η<sups>2</sups> = 0.06. Post hoc pairwise comparisons showed that the correlation with response time was significantly weaker for confidence than for goal-driven effort appraisals, <emph>p</emph> = 0.002. All other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. Thus, results from experiment 1 were partly replicated, as response time served as a cue for goal-driven effort more than for confidence.</p> <p>These findings serve an important contribution to the emergent research on non-verbal tasks in the metacognitive literature. In accordance with the scarce studies that investigated non-verbal tasks (e.g., Lauterman &amp; Ackerman, [<reflink idref="bib51" id="ref153">51</reflink>]; Reber et al., [<reflink idref="bib76" id="ref154">76</reflink>]), the relationship of all appraisals with response time, as well as the replication of the differences between goal-driven and confidence appraisals, demonstrates both shared and distinctive metacognitive mechanisms between verbal and non-verbal tasks.</p> <hd id="AN0163936292-30">RQ 2: Are There Differences in the Extent to Which Mental Effort Appraisals, Difficulty Appra...</hd> <p>To evaluate the predictive value of appraisals for accuracy (see the MTT column in Fig. 1B), the strength of the appraisal-accuracy correlation was compared between the different appraisal types using ANOVA. There was a significant medium to large effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref155">3</reflink>, 220) = 7.19, MSE = 0.08, <emph>p</emph> &lt; 0.001, η<sups>2</sups> = 0.09. Post hoc pairwise comparisons showed that the correlation with accuracy was significantly stronger for confidence than for goal-driven effort appraisals, <emph>p</emph> = 0.005. In addition, the correlation was significantly stronger for task difficulty appraisals compared with both goal-driven effort, <emph>p</emph> &lt; 0.001, and data-driven effort appraisals, <emph>p</emph> = 0.036. All other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. In the MTT, similar to the CRT, confidence was more strongly related to accuracy than goal-driven effort appraisals. However, unlike in experiment 1, in the MTT, task difficulty appraisals were more strongly related to accuracy than both types of effort appraisals.</p> <p>Together with the findings regarding response time, we argue that the different appraisals we considered may reflect a continuum in terms of how individuals interpret what is asked of them to report (i.e., reflection of self-regulated effort vs. reflection of task demands), in which goal-driven appraisals and confidence appraisals are the two extremes.</p> <hd id="AN0163936292-31">Experiment 3</hd> <p>Experiment 3 was designed to replicate experiments 1 and 2 using yet another task, the mental rotation task)MRT(. This task offers inherent variations in item complexity by the variation of the angle of rotation between the original shape and its rotated copy, presented among the answer options (see Fig. 3 and Materials section below). This task feature allowed us to examine RQ3 as well as RQ1 and RQ2.</p> <p>Graph: Fig. 3An example of the mental rotation task (items were generated with the Mental Rotation Stimulus Library: Peters &amp; Battista, [<reflink idref="bib71" id="ref156">71</reflink>]). Participants were asked to indicate which of the five alternatives (A to E) was identical to the criterion figure on the left. In this example, answer B is correct because it is the same figure but rotated by 80°</p> <p>Moreover, recent metacognitive studies have found that people incorporate several cues into their metacognitive judgments at once in both memorizing and problem-solving contexts (Ackerman, [<reflink idref="bib3" id="ref157">3</reflink>]; Undorf &amp; Bröder, [<reflink idref="bib101" id="ref158">101</reflink>]). In particular, these studies encouraged researchers to identify specific task characteristics that predict either success and/or metacognitive judgments for explaining sources for difficulty. Thus, in this experiment, we considered the extent to which complexity is taken into account in each of the four appraisals.</p> <hd id="AN0163936292-32">Method</hd> <p></p> <hd id="AN0163936292-33">Participants and Design</hd> <p>Again, we collected data from respondents using Prolific. Participation was remunerated with 3.25 GBP. Data were excluded from participants who did not follow instructions (3 participants), had little variability in appraisals (<reflink idref="bib6" id="ref159">6</reflink>), completed less than 75% of the trials (<reflink idref="bib4" id="ref160">4</reflink>), or did not consent to the use of their data (<reflink idref="bib1" id="ref161">1</reflink>). This left 238 participants for the analyses (age: <emph>M</emph> = 26.5, SD = 8.3; 81 females, 156 males; age and gender missing for 1 participant). As in the other experiments, participants were randomly assigned to one of four groups that differed in the appraisal they were asked to give: goal-driven effort (<emph>n</emph> = 57), data-driven effort (<emph>n</emph> = 60), task difficulty (<emph>n</emph> = 63), or confidence (<emph>n</emph> = 58).</p> <hd id="AN0163936292-34">Materials: Mental Rotation Task (MRT)</hd> <p>The MRT version used in this study was based on the mental rotation test from Vandenberg and Kuse ([<reflink idref="bib110" id="ref162">110</reflink>]). We used the mental rotation figures of the Shepard and Metzler type (taken from the Mental Rotation Stimulus Library: Peters &amp; Battista, [<reflink idref="bib71" id="ref163">71</reflink>]), comprising three-dimensional line drawings of cubes that are put together to form a figure. Each item consisted of one criterion figure to the left and five alternatives to the right (see Fig. 3 for an example). One of the presented answer options was identical to the criterion figure in structure but was shown in a rotated position around either the horizontal or vertical axis. The other alternatives (distractors) were all mirrored versions of the criterion figure and were also rotated around one of the two axes. In addition, the task allowed for systematic manipulation of task complexity at the item level, operationalized as the angle of rotation (angular disparity) between the criterion and the target figure. Previous research showed a linear relation between rotation angle and response time (Shepard &amp; Metzler, [<reflink idref="bib85" id="ref164">85</reflink>]). Accordingly, a smaller rotation angle was considered less complex because it entailed less mental rotation; therefore, less time was required to solve the item. We employed nine levels of complexity in steps of 20°, ranging from level 1 with 20° rotation (low complexity) to level 9 with 180° rotation (high complexity). In total, 72 items were used (8 items for each level of complexity), divided into two 36-item sets with 4 items per level in each. Each participant was randomly allocated one of the two sets.</p> <hd id="AN0163936292-35">Appraisals</hd> <p>The goal-driven effort, data-driven effort, task difficulty, and confidence appraisals were elicited as in the other experiments. As in experiment 2, the confidence appraisal was assessed on a scale of 20 to 100.</p> <hd id="AN0163936292-36">Objective Measures</hd> <p>Again, the objective measures examined were response time and accuracy.</p> <hd id="AN0163936292-37">Background Variables</hd> <p>Overall judgment of performance, experience with puzzle tasks, test anxiety, time since the last exam, and mindset about the malleability of intelligence (fixed vs. growth mindset) were assessed as control variables using the same questions as in experiment 2. Their means and standard deviations are shown in Appendix 2. In addition, an attention check was performed using the same procedure as in experiment 2. As before, those who did not pass the attention check (less than 6% of the sample) were only excluded if they showed another indication of inattention (see under the "Participants and Design" section).</p> <hd id="AN0163936292-38">Procedure</hd> <p>The overall procedure was the same as in the other experiments. In the specific instructions, participants were told their task was to decide which of the five objects shown to the right was the same (rotated) object as the one to the left. The entire task took about 25 min.</p> <hd id="AN0163936292-39">Data Preparation</hd> <p>Trials with extraordinarily short (RT &lt; 2 s, 82 trials) or long response times (RT &gt; 180 s, 48 trials) were excluded, as were trials in which no response was recorded because of technical problems (42 trials). The final sample comprised 8448 trials. Data preparation and all analyses were done as in the previous experiments. Complexity (at the item level) is measured on an interval scale, as complexity level corresponds to angular disparity. Hence, Pearson's correlations were used to calculate the correlation between item complexity and appraisals.</p> <hd id="AN0163936292-40">Results</hd> <p>Item success rates ranged from 45.6 to 89.7%. On average, participants needed 21.9 s (SD = 11.5) per item and correctly solved 73.1% (SD = 22.5) of the items. Means and standard deviations of appraisals, response time, and accuracy as a function of the type of appraisal are shown in Table 1. The four appraisal groups did not differ in their response times, <emph>F</emph> &lt; 1, or in accuracy, <emph>F</emph>(<reflink idref="bib3" id="ref165">3</reflink>, 234) = 1.09, MSE = 0.05, <emph>p</emph> = 0.354, η<sups>2</sups> = 0.01. As seen in Table 2, all mean correlations significantly differed from 0, indicating that both response time and item complexity served as cues for the appraisals, while the appraisals predicted accuracy in the task.</p> <hd id="AN0163936292-41">RQ 1: Are There Differences in the Extent to Which Response Time Serves as a Cue for Mental E...</hd> <p>To evaluate whether the different appraisals rely similarly on response time (see the MRT columns in Fig. 1A), the strength of the appraisal-response time correlation was compared between the different appraisal types using ANOVA. There was a significant large effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref166">3</reflink>, 234) = 16.41, MSE = 0.07, <emph>p</emph> &lt; 0.001, η<sups>2</sups> = 0.17. Post hoc pairwise comparisons showed that the correlation with response time was significantly weaker for confidence than for all three of the other appraisal types, all <emph>p</emph> &lt; 0.001, while all other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. These findings replicate those from experiment 1 and partly replicate those from experiment 2, providing further evidence that response time serves as a cue for mental effort and task difficulty appraisals more strongly than confidence appraisals.</p> <hd id="AN0163936292-42">RQ 2: Are There Differences in the Extent to Which Mental Effort Appraisals, Difficulty Appra...</hd> <p>To evaluate the predictive value of appraisals for accuracy (see the MRT columns in Fig. 1B), the strength of the appraisal-accuracy correlation was compared between the different appraisal types using ANOVA. There was a significant small to medium effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref167">3</reflink>, 224) = 3.25, MSE = 0.15, <emph>p</emph> = 0.023, η<sups>2</sups> = 0.04. Post hoc pairwise comparisons showed that the correlation with actual accuracy was significantly stronger for confidence than for goal-driven effort appraisals, <emph>p</emph> &lt; 0.024, while all other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. Thus, the present results are consistent with those of the previous experiments in that confidence appraisals were again more strongly related to accuracy than goal-driven effort appraisals.</p> <hd id="AN0163936292-43">RQ 3: Are There Differences in the Extent to Which the Complexity of the Problem Serves as a...</hd> <p>First, correlations between complexity and accuracy, as well as complexity and response time, were tested to check the operationalization of item-level complexity. There was a significant negative correlation (Goodman–Kruskal's γ correlation) between complexity and accuracy, <emph>r</emph> = <emph>− </emph>0<emph>.</emph>11, <emph>Z</emph> = <emph>− </emph>7.03, <emph>p</emph> &lt; 0.001, indicating that less complex items were more likely to be solved. Furthermore, there was a significant positive correlation (Pearson's correlation) between complexity and response time, <emph>r</emph> = 0.09, <emph>t</emph> (8446) = 8.08, <emph>p</emph> &lt; 0.001, indicating that it took participants longer to solve more complex items. Thus, we can conclude that the operationalization of item-level complexity was successful.</p> <p>To evaluate whether appraisals rely similarly on complexity (see Fig. 1C), the strength of the appraisal-complexity correlation was compared between the different types of appraisals using ANOVA. There was a significant small to medium effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref168">3</reflink>, 234) = 2.87, MSE = 0.04, <emph>p</emph> = 0.037, η<sups>2</sups> = 0.04. Post hoc pairwise comparisons showed that the correlation with complexity was significantly weaker for confidence than for difficulty appraisals, <emph>p</emph>&lt; 0.001, while all other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs>&gt;0.05.</p> <p>In sum, experiment 3 replicates in another non-verbal problem-solving task our findings regarding response time being a cue for load-related appraisals, as it is for confidence appraisals. It also provides further evidence to the notion that goal-driven effort and confidence lie at opposite ends of a continuum. However, this experiment also has a unique contribution, as it uncovers an additional cue that may guide load-related appraisals in the form of complexity. Although research has shown that all load-related appraisals and confidence are sensitive to task demands variations, our findings suggest that difficulty appraisals rely more strongly on complexity than confidence appraisals do. Notably, this finding may also shed more light on how individuals interpret the difficulty item appraisals: both difficulty and confidence appraisals seem to be perceived as relating to task demands more than individual effort one chooses to invest in the task. However, it may be that difficulty appraisals are a purer reflection of task demands associated with differences in item complexity compared to confidence appraisals.</p> <hd id="AN0163936292-44">Global Effects: Integrating Results from the Three Experiments</hd> <p>The results were remarkably consistent across the three experiments, with only slight variations in the findings. However, there were several differences between the tasks. First, while experiment 1 utilized a verbal task, experiments 2 and 3 relied on non-verbal figural tasks. Second, the tasks varied in difficulty, with the mean success rate ranging from 40% in experiment 2 to 73% in experiment 3. There were also slight framing variations between the tasks. In particular, while the CRT task had four answer options, the MTT and MRT had five; and the scale used for confidence appraisals ranged from 0 to 100 in the CRT and 20 to 100 in the MTT and MRT. Thus, apart from the analyses we conducted for each experiment, we also conducted aggregated analyses (again using ANOVA) to validate our findings beyond these differences and illuminate global effects across the tasks. As seen in Table 2, all correlations significantly differed from zero, indicating that response time served as a cue for all appraisals, and that all appraisals predicted actual accuracy in the task.</p> <hd id="AN0163936292-45">RQ 1: Are There Differences in the Extent to Which Response Time Serves as a Cue for Mental E...</hd> <p>To evaluate whether the different types of appraisals similarly rely on response times (see Exps. 1–3 in Fig. 1A), the strength of the appraisal-response time correlation was compared between the different appraisal types across the experiments while controlling for the experimental task. A two-way ANOVA was calculated with appraisal type and experimental task (CRT, MTT, and MRT) as between-subject factors, and the appraisal-response time correlation as the dependent variable. There was a significant large main effect of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref169">3</reflink>, 734) = 34.78, MSE = 0.08, <emph>p</emph>&lt;0.001, η<subs>p</subs><sups>2</sups> = 0.13, and a significant small effect of experimental task, <emph>F</emph>(<reflink idref="bib2" id="ref170">2</reflink>, 734) = 8.81, <emph>p</emph>&lt; 0.001, MSE = 0.08, η<subs>p</subs><sups>2</sups> = 0.02. Most importantly, the interaction was not significant, <emph>F</emph>(<reflink idref="bib6" id="ref171">6</reflink>, 734) = 1.17, <emph>p</emph> = 0.318, MSE =0.08, η<subs>p</subs><sups>2</sups> = 0.01, indicating that the pattern of results was similar across the tasks. Tukey multiple comparisons for appraisal type showed a significantly weaker correlation with response time for confidence than for all the other appraisals, all <emph>p</emph><subs>s</subs> &lt; 0 001, and a weaker correlation for task difficulty appraisals than for goal-driven effort appraisals, <emph>p</emph> = 0.023. The other comparisons showed no significant differences, all <emph>p</emph><subs>s</subs> &gt; 0.05. Thus, the global pattern suggests that response time is a more reliable cue for effort and difficulty appraisals than for confidence appraisals.</p> <hd id="AN0163936292-46">RQ 2: Are There Differences in the Extent to Which Mental Effort Appraisals, Difficulty Appra...</hd> <p>To evaluate the predictive value of appraisals for accuracy (see Exps. 1–3 in Fig. 1B), the strength of the appraisal-accuracy correlation was compared between the appraisal types while controlling for the experimental task. Again, a two-way ANOVA was conducted. There were significant medium main effects of appraisal type, <emph>F</emph>(<reflink idref="bib3" id="ref172">3</reflink>, 716) = 13.07, MSE = 0.13, <emph>p</emph> &lt; 0.001, η<subs>p</subs><sups>2</sups> = 0.05, and experimental task, <emph>F</emph>(<reflink idref="bib2" id="ref173">2</reflink>, 716) = 26.54, MSE = 0.13, <emph>p</emph> &lt; 0.001, η<subs>p</subs><sups>2</sups> = 0.07. As in RQ1, the interaction was not significant, <emph>F</emph> &lt; 1, indicating that the pattern of results was similar across the tasks. Tukey multiple comparisons for appraisal type showed accuracy to be correlated significantly more strongly with confidence compared with both types of effort appraisals (goal driven and data driven), both <emph>p</emph> &lt; 0.001. Accuracy was also correlated significantly less with goal-driven effort appraisals compared with task difficulty appraisals, <emph>p</emph> &lt; 0.001, while the other comparisons showed no significant differences, all <emph>p</emph>s &gt; 0.10. Thus, overall, confidence appraisals predicted accuracy in the task more reliably than effort appraisals.</p> <hd id="AN0163936292-47">General Discussion</hd> <p>Educational research often relies on people's self-reported subjective appraisals of their experience with learning or problem-solving tasks. However, evidence from both the metacognitive and CLT research domains indicates that such subjective appraisals are prone to biases. Those findings highlight the need for a systematic investigation of the inference processes underlying different appraisals (Scheiter et al., [<reflink idref="bib79" id="ref174">79</reflink>]). In the present study, we integrate these two major theoretical approaches, examining the underlying bases and the predictive value of mental effort, task difficulty, and metacognitive confidence appraisals in three cognitively demanding problem-solving tasks by using metacognitive concepts, paradigms, and measures.</p> <p>Our first research question concerned response time as a cue for subjective appraisals. We were particularly interested in whether response time would emerge as a cue for the load-related appraisals, namely, effort and difficulty, as it has for confidence appraisals (e.g., Baars et al., [<reflink idref="bib11" id="ref175">11</reflink>]). Across all experiments, we found that, indeed, response time was significantly associated with all appraisals. Yet, interestingly, these associations were stronger for the load-related appraisals than for confidence: the former had moderate to strong relationships with response time, and the latter only had a weak relationship.</p> <p>Using response time as a cue can be informative for mental effort: research has shown that invested effort is related to time investment (Baars et al., [<reflink idref="bib11" id="ref176">11</reflink>]). However, metacognitive research has robustly shown that using response time inflexibly as a cue for confidence appraisals in problem-solving tasks can be misleading (e.g., Finn &amp; Tauber, [<reflink idref="bib33" id="ref177">33</reflink>]; Thompson et al., [<reflink idref="bib95" id="ref178">95</reflink>]; Thompson &amp; Morsanyi, [<reflink idref="bib97" id="ref179">97</reflink>]). For example, Ackerman and Zalmanov ([<reflink idref="bib4" id="ref180">4</reflink>]) examined the association between response time and confidence in problem-solving tasks in conditions where response time was either a valid or invalid predictor of performance. Across conditions and regardless of its validity, confidence varied as a function of response time, with participants reporting more confidence in solutions provided quickly than in those which took longer. As mentioned above, recent studies encourage considering multiple sources for actual and perceived difficulty beyond response time (Undorf &amp; Bröder, [<reflink idref="bib101" id="ref181">101</reflink>]). Notably, even when controlling for various such sources, still, response time as a cue has been shown to result in a monitoring bias (Ackerman, [<reflink idref="bib3" id="ref182">3</reflink>]). Under the metacognitive framework, such biased monitoring of learning is problematic, as it may misguide subsequent regulatory decisions (e.g., Metcalfe &amp; Finn, [<reflink idref="bib57" id="ref183">57</reflink>]). It remains to be seen whether such harmful effects on subsequent regulatory decisions also arise from bias in load-related appraisals. This question opens fertile ground for follow-up research.</p> <p>In addition to testing the power of response time to influence subjective appraisals, we also examined the effect of item-level complexity (our third research question). We found support for item complexity as a cue for perceived difficulty, which had a stronger correlation with complexity compared with confidence appraisals but was similar to the two effort appraisals. This finding is in line with the assumption that our operationalization of complexity, as incrementally larger (or smaller) rotation angles, did indeed impose different levels of additional cognitive load by requiring incrementally more (or less) mental rotation to correctly solve the item. Future research might extend this investigation to other sources of load and to other tasks for delving further into the question of why complexity was found here to be a significantly stronger cue only for difficulty compared with confidence and not for either of the effort appraisals.</p> <p>Notably, response time and item complexity are only two of the many cues already uncovered as underlying metacognitive appraisals. Other cues for confidence appraisals in problem-solving tasks identified in the meta-reasoning literature (see Ackerman, [<reflink idref="bib2" id="ref184">2</reflink>] for a review) include accessibility (the number of associations that come to mind when answering a question, e.g., Ackerman &amp; Beller, [<reflink idref="bib6" id="ref185">6</reflink>]), self-consistency (the consistency with which different considerations support the chosen answer, Bajšanski et al., [<reflink idref="bib12" id="ref186">12</reflink>]), and cardinality (the number of considered answer options, Bajšanski et al., [<reflink idref="bib12" id="ref187">12</reflink>]). In addition, recent research indicates that cue integration, namely, exposing and analyzing multiple cues inherent in the task, has the potential to afford a more thorough understanding of the mechanisms underlying metacognitive appraisals (e.g. Ackerman, [<reflink idref="bib3" id="ref188">3</reflink>]; Undorf et al., [<reflink idref="bib102" id="ref189">102</reflink>]). We call future research to use our methodology to examine other cues and their potential interactive role in load-related appraisals.</p> <p>Our second research question focused on the predictive value of the various appraisals for item-level success (correct answers). While all appraisals were significantly associated with success, the strength of this association was stronger for both confidence and difficulty appraisals (which were similar and with moderate to large effects) than for effort appraisals (small to medium effects). Taken together with our finding (RQ1) that response time was more strongly correlated with the effort appraisals than with difficulty or confidence, these findings support the notion that subjective mental effort appraisals, whether goal driven or data driven, reflect fluency and are therefore not a good basis for predicting actual success compared to difficulty and confidence appraisals in the examined tasks. Moreover, Ackerman ([<reflink idref="bib3" id="ref190">3</reflink>]) succeeded in improving success and attenuating biases in confidence judgments with instructions. It is worth investigating whether such instructional design features affect cues underlying effort appraisals as well.</p> <p>Overall, our findings support questioning the reliability of the commonly used load-related appraisals as reflections of cognitive load. In addition, the results indicate an important distinction between effort and difficulty appraisals. It has been suggested in CLT research that these scales can be used interchangeably under the assumption that despite their varied phrasing they all reflect cognitive load differences stemming from instructional procedures (e.g., de Jong, [<reflink idref="bib25" id="ref191">25</reflink>]; Sweller et al., [<reflink idref="bib89" id="ref192">89</reflink>]). Our findings suggest that effort appraisals are less accurate than difficulty appraisals in predicting performance. Schmeck et al. ([<reflink idref="bib80" id="ref193">80</reflink>]) similarly distinguished between effort and difficulty appraisals in relation to performance. In two experiments, they compared single delayed mental effort and difficulty appraisals at the end of a series of tasks to the average of mental effort and difficulty appraisals after each of those tasks. They found that performance was predicted only by mental effort appraisals in one experiment, while difficulty appraisals were more strongly (though not significantly) associated with performance in the other experiment. Our findings, along with those of Schmeck, offer empirical support for the notion that the two measurements reflect distinct constructs (Ayres &amp; Youssef, [<reflink idref="bib10" id="ref194">10</reflink>]; van Gog &amp; Paas, [<reflink idref="bib106" id="ref195">106</reflink>]).</p> <p>This study also contributes to the developing field of meta-reasoning (Ackerman &amp; Thompson, [<reflink idref="bib7" id="ref196">7</reflink>]). Specifically, very little is known about cue utilization and the predictive value of appraisals in non-verbal reasoning tasks (see Lauterman &amp; Ackerman, [<reflink idref="bib51" id="ref197">51</reflink>]). Here, we show initial evidence for similar patterns of relationships linking confidence with response time and accuracy in both a well-studied verbal task (the CRT) and two non-verbal tasks (the MTT and MRT). Future studies are called to shed more light on these relationships by examining different tasks, variations in instructional materials, and different populations while using preregistered hypotheses.</p> <p>Notably, the present study focused on reasoning and problem-solving tasks. However, both metacognitive and CLT research have been grounded in more typical learning contexts (e.g., memorization and comprehension tasks). Although similarities in monitoring appraisals have been demonstrated between these different cognitive processes (Ackerman, [<reflink idref="bib2" id="ref198">2</reflink>]), it is imperative to examine the generalizability of our findings to more typical learning situations by replicating the study with such tasks and in actual educational settings. Boundary conditions such as prior knowledge also need to be exposed. Finally, further insights into people's reasoning when making appraisals could be gleaned by assessing process data, for instance through think-aloud studies.</p> <hd id="AN0163936292-48">Limitations</hd> <p>A first possible limitation of our research is that our samples consisted solely of crowd workers who performed the experiments online. This could have affected the measured response times. However, there is little evidence in the data that the experiments are problematic in that respect. In addition, Prolific is an online research platform that has been empirically found to provide access to more diverse and naïve and less dishonest populations, producing higher data quality with less noise compared to other research platforms (e.g., Gupta et al., [<reflink idref="bib36" id="ref199">36</reflink>]; Peer et al., [<reflink idref="bib69" id="ref200">69</reflink>], [<reflink idref="bib70" id="ref201">70</reflink>]). Nonetheless, replications under more controlled conditions in the laboratory may serve to reaffirm our findings.</p> <p>Second, we selected tasks designed to have a certain range of difficulty, where additional effort would improve performance. That is, we chose our tasks so that they should be solvable given sufficient time. Future research should also consider tasks where additional effort does not necessarily pay off in improved performance. These might be simpler tasks in which additional effort merely increases efficiency but not performance or more difficult tasks that participants might not be able to solve even by expending effort.</p> <p>Third, we used a 0–100 scale to assess the appraisals. While this is common practice in the metacognitive literature, cognitive load research typically employs 5-, 7-, or the original 9-point scales to assess mental effort (Paas, [<reflink idref="bib64" id="ref202">64</reflink>]; Paas et al., [<reflink idref="bib65" id="ref203">65</reflink>]). On the one hand, a scale from 0 to 100 offers more sensitivity in detecting changes in participants' appraisals. On the other hand, in cognitive load research, it is a subject of debate whether one can distinguish between even nine levels of effort (Paas et al., [<reflink idref="bib65" id="ref204">65</reflink>]).</p> <p>Further, one might argue that our results are limited because of the correlational nature of our approach. Of course, correlational research has limitations like the third-variable problem or that correlations only describe relationships but causality cannot be inferred. However, it has to be noted that correlations themselves are not the core of our analyses. Rather, we use within-person correlations as dependent variables in an experimental design. Our interest lies in the varying strength of the relation of varying appraisals with response time, accuracy, or complexity level. Thus, correlations are used as dependent variables in subsequent analyses. These final analyses are the comparison of the correlations as a function of the experimentally varied type of appraisal, which are at the heart of our contribution.</p> <p>Finally, as noted in the introduction, recent research has begun to delve deeper into developing and validating unique self-report measures for different types of cognitive load, intrinsic, extraneous, and germane load (e.g., Klepsch &amp; Seufert, [<reflink idref="bib45" id="ref205">45</reflink>]; Klepsch et al., [<reflink idref="bib44" id="ref206">44</reflink>]; Leppink &amp; Pérez-Fuster, [<reflink idref="bib53" id="ref207">53</reflink>]; Leppink et al., [<reflink idref="bib52" id="ref208">52</reflink>]). While examining these different types of cognitive load was not within the scope of our study, future research could examine how the different types of load relate to response time, performance, and complexity.</p> <hd id="AN0163936292-49">Conclusion</hd> <p>While CLT research has assumed that individuals' subjective appraisals of mental load reflect the cognitive resources allocated to achieve task goals, a metacognitive approach suggests that load-related appraisals, like metacognitive appraisals, are potentially susceptible to bias and in need of thorough investigation (Scheiter et al., [<reflink idref="bib79" id="ref209">79</reflink>]). To this end, the present study employs a metacognitive framework to offer novel empirical evidence for the underlying processes on which load-related appraisals rely. The results highlight the tight relationships of load-related appraisals with response time and item-level complexity, as well as their weaker relationship with actual task performance compared to metacognitive confidence appraisals. These findings, which replicate across several unique tasks, imply that load-related appraisals are indeed susceptible to bias.</p> <p>These findings have powerful implications for research and practice in education. Effective regulation and resource allocation in everyday tasks, and especially in educational contexts, rely heavily on people's ability to accurately appraise the demands of cognitive tasks; and, as has been consistently shown in metacognitive research, biased monitoring can mislead future regulatory decisions. Thus, it is important to understand the bases and validity of cues that learners use for effort appraisals and regulation. In practice, the findings can guide the design of different instructional frameworks. For example, designers of adaptive learning environments which rely on self-reports to select the next task should consider which type of appraisal have the strongest associations with their desired learning outcome.</p> <p>This study should be perceived as a starting point for exposing the underlying processes at the heart of load-related appraisals and to inspire a new stream of future research. However, the findings already indicate that vigilance is required when collecting and interpreting subjective self-reports of effort and task difficulty. Finally, by relating subjective appraisals of cognitive load to metacognitive appraisals, the present study contributes to bridging the CLT and metacognitive research paradigms (de Bruin et al., [<reflink idref="bib23" id="ref210">23</reflink>]; Scheiter et al., [<reflink idref="bib79" id="ref211">79</reflink>]).</p> <hd id="AN0163936292-50">Acknowledgements</hd> <p>This research was supported by and inspired by discussions of the EARLI Emerging Field Group Monitoring and Regulation of Effort.</p> <hd id="AN0163936292-51">Author Contribution</hd> <p>All authors contributed to the studies' conception and design. Material preparation, data collection, and analysis were mainly performed by Yael Sidi (experiment 1), Rakefet Ackerman (experiment 2), and Emely Hoch (experiment 3). The first draft of the manuscript was written by Emely Hoch and Yael Sidi and commented on by all authors. All authors read and approved the final version of the manuscript.</p> <hd id="AN0163936292-52">Funding</hd> <p>Open Access funding enabled and organized by Projekt DEAL. This research was funded by the Jacobs Foundation in the context of the EARLI Emerging Field Group Monitoring and Regulation of Effort and by the Israel Science Foundation.</p> <hd id="AN0163936292-53">Data Availability</hd> <p>The datasets generated and/or analyzed during the current study along with the corresponding analysis syntax are available in the OSF repository (https://doi.org/10.17605/osf.io/q2p74).</p> <hd id="AN0163936292-54">Declarations</hd> <p></p> <hd id="AN0163936292-55">Ethics Approval</hd> <p>The ethical conduct of the studies was reviewed and approved by the Behavioral Sciences Research Ethics Committee of the Technion–Israel Institute of Technology (2020–015) and the Leibniz-Institut für Wissensmedien Institutional Review Board (protocol LEK 2021/011).</p> <hd id="AN0163936292-56">Consent to Participate</hd> <p>Informed consent was obtained from all individual participants included in the study.</p> <hd id="AN0163936292-57">Conflict of Interest</hd> <p>Rakefet Ackerman is an editorial board member of Educational Psychology Review. Otherwise, the authors have no competing interests to declare relevant to this article's content.</p> <hd id="AN0163936292-58">Appendix 1. Development and Pretesting of the CRT Items</hd> <p>Initially, 31 open-ended problems were compiled from different studies and publications (De Neys et al., [<reflink idref="bib26" id="ref212">26</reflink>]; Finucane &amp; Gullion, [<reflink idref="bib34" id="ref213">34</reflink>]; National Institute for Testing and Evaluation [<reflink idref="bib61" id="ref214">61</reflink>]; Oldrati et al., [<reflink idref="bib63" id="ref215">63</reflink>]; Primi et al., [<reflink idref="bib73" id="ref216">73</reflink>]; Shtulman &amp; McCallum, [<reflink idref="bib86" id="ref217">86</reflink>]; Sirota et al., [<reflink idref="bib88" id="ref218">88</reflink>]; Toplak et al., [<reflink idref="bib98" id="ref219">98</reflink>]; Trippas et al., [<reflink idref="bib99" id="ref220">99</reflink>]; Valerjev, [<reflink idref="bib103" id="ref221">103</reflink>]; Young et al., [<reflink idref="bib113" id="ref222">113</reflink>]).Thirty-two participants completed a pretest of these items via Prolific for 2GBP monetary compensation. Items were presented to participants in random order. Following analysis of success rates and response times, we excluded eight items for the following reasons: very high success rates with quick response times; previously known to many participants; very low success rates with long response times and minimal number of expected answers with long response times. This analysis left us with 23 items. We then fitted the problems to a multiple-choice format. The multiple-choice items included four answer options generated such that each item had a correct answer, one misleading answer (i.e., an answer that was incorrect but predictable), and two distractors. The two distractors were developed based on the pretest, and comprised the two (incorrect) answers provided by the highest proportion of pretest participants. If all given answers appeared at the same rates, we selected those that seemed most misleading. We generated new distractors for five items which did not result in enough wrong answers in the pretest. Thirty-four participants then completed a second pretest using the multiple-choice format. One problem was excluded for being too easy, two problems were selected as training items, and three very easy problems were selected as attention check items, leaving 17 problems in the final set.</p> <hd id="AN0163936292-59">Appendix 2. Descriptive Statistics of Background Variables</hd> <p>Table 3 Means and Standard Deviations of Background Variables as a Function of Type of Appraisal in Experiments 1–3</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;Goal-driven effort&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Data-driven effort&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Task difficulty&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Confidence&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 1 (CRT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Need for cognition&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.57 (1.25)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.68 (1.13)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.96 (1.05)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;5.03 (0.86)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Experience with puzzles&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.05 (1.58)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.84 (1.72)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.93 (1.60)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.39 (1.58)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 2 (MTT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Judgment of performance&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;50.12 (22.51)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;46.27 (21.95)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;45.23 (21.89)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;45.44 (19.35)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Experience with puzzles&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.81 (0.72)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.81 (0.69)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.77 (0.71)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.92 (0.70)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Test anxiety&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.67 (0.81)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.48 (0.70)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.62 (0.82)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.36 (0.79)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Time since last exam&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.29 (1.07)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.05 (1.00)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.86 (1.23)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;3.10 (1.13)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Mindset of intelligence&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.54 (1.21)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.37 (1.24)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.60 (1.31)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.77 (1.21)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" colspan="5"&gt;&lt;p&gt;Experiment 3 (MRT)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Judgment of performance&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;74.61 (20.64)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;70.28 (19.25)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;71.21 (21.65)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;73.13 (22.40)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Experience with puzzles&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.00 (0.80)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.87 (0.70)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;1.87 (0.68)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.10 (0.77)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Test anxiety&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.25 (0.68)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.43 (0.72)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.37 (0.74)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.31 (0.80)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Time since last exam&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.51 (1.15)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.45 (1.08)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.30 (1.29)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;2.52 (1.20)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#160;&amp;#160;Mindset of intelligence&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.62 (1.25)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.42 (1.28)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.88 (1.14)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;4.68 (1.26)&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Need for cognition was assessed on a 7-point Likert scale, with low scores indicating low need for cognition. Experience with puzzles was assessed on a 7-point scale in Experiment 1 and on a 4-point scale in Experiment 2 and Experiment 3,with low scores always indicating little experience. Judgment of performance was assessed as the number of correctly solved tasks and is given as a percentage of the number of tasks in the relevant experiment. Test anxiety and mindset about the malleability of intelligence were assessed on 4-point scales, with low scores indicating low test anxiety or a fixed mindset, respectively. Time since the last exam was assessed on a scale from 1 (<emph>during the last month</emph>) to 4 (<emph>more than 12 months ago</emph>)</p> <p>3</p> <hd id="AN0163936292-60">Publisher's Note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0163936292-61"> <title> References </title> <blist> <bibl id="bib1" idref="ref73" type="bt">1</bibl> <bibtext> Ackerman, R. (2014). The diminishing criterion model for metacognitive regulation of time investment. Journal of Experimental Psychology: General, 143(3), 1349–1368. https://doi.org/10.1037/a0035098</bibtext> </blist> <blist> <bibl id="bib2" idref="ref17" type="bt">2</bibl> <bibtext> Ackerman, R. (2019). Heuristic cues for meta-reasoning judgments. Psihologijske Teme, 28(1), 1–20. https://doi.org/10.31820/pt.28.1.1</bibtext> </blist> <blist> <bibl id="bib3" idref="ref114" type="bt">3</bibl> <bibtext> Ackerman, R. (2023). Bird's-eye view of cue integration: A methodology for exposing multiple cues underlying metacognitive judgments. Educational Psychology Review, 32:55. https://doi.org/10.1007/s10648-023-09771-z</bibtext> </blist> <blist> <bibl id="bib4" idref="ref58" type="bt">4</bibl> <bibtext> Ackerman, R., &amp; Zalmanov, H. (2012). The persistence of the fluency–confidence association in problem solving. Psychonomic Bulletin &amp; Review, 19(6), 1187–1192. https://doi.org/10.3758/s13423-012-0305-z</bibtext> </blist> <blist> <bibl id="bib5" idref="ref16" type="bt">5</bibl> <bibtext> Ackerman, R., &amp; Thompson, V. A. (2015). Meta-reasoning: What can we learn from meta-memory. In A. Feeney &amp; V. A. Thompson (Eds.), Reasoning as Memory (pp. 164–178). Psychology Press.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref19" type="bt">6</bibl> <bibtext> Ackerman, R., &amp; Beller, Y. (2017). Shared and distinct cue utilization for metacognitive judgements during reasoning and memorisation. Thinking &amp; Reasoning, 23(4), 376–408. https://doi.org/10.1080/13546783.2017.1328373</bibtext> </blist> <blist> <bibl id="bib7" idref="ref31" type="bt">7</bibl> <bibtext> Ackerman, R., &amp; Thompson, V. A. (2017). Meta-reasoning: Monitoring and control of thinking and reasoning. Trends in Cognitive Sciences, 21(8), 607–617. https://doi.org/10.1016/j.tics.2017.05.004</bibtext> </blist> <blist> <bibl id="bib8" idref="ref79" type="bt">8</bibl> <bibtext> Ashburner M, Risko EF. Judgements of effort as a function of post-trial versus post-task elicitation. Quarterly Journal of Experimental Psychology. 2021; 74; 6: 991-1006. 10.1177/17470218211005759</bibtext> </blist> <blist> <bibl id="bib9" idref="ref70" type="bt">9</bibl> <bibtext> Ayres P. Using subjective measures to detect variations of intrinsic cognitive load within problems. Learning and Instruction. 2006; 16; 5: 389-400. 10.1016/j.learninstruc.2006.09.001</bibtext> </blist> <blist> <bibtext> Ayres, P., &amp; Youssef, A. (2008). Investigating the influence of transitory information and motivation during instructional animations. In P. A. Kirschner, F. Prins, V. Jonker, &amp;G. Kanselaaer (Eds.), Proceedings of the 8th International Conference for the Learning Sciences (pp. 68–75). ICLS</bibtext> </blist> <blist> <bibtext> Baars M, Wijnia L, de Bruin A, Paas F. The relation between students' effort and monitoring judgments during learning: A meta-analysis. Educational Psychology Review. 2020; 32; 4: 979-1002. 10.1007/s10648-020-09569-3</bibtext> </blist> <blist> <bibtext> Bajšanski I, Žauhar V, Valerjev P. Confidence judgments in syllogistic reasoning: The role of consistency and response cardinality. Thinking &amp; Reasoning. 2019; 25; 1: 14-47. 10.1080/13546783.2018.1464506</bibtext> </blist> <blist> <bibtext> Benjamin AS, Bjork RA Reder L. Retrieval fluency as a metacognitive index. Metacognition and implicit memory. 1996; Erlbaum: 309-338</bibtext> </blist> <blist> <bibtext> Benjamin AS, Bjork RA, Schwartz BL. The mismeasure of memory: When retrieval fluency is misleading as a metamnemonic index. Journal of Experimental Psychology: General. 1998; 127; 1: 55-68. 10.1037/0096-3445.127.1.55</bibtext> </blist> <blist> <bibtext> Bjork RA, Dunlosky J, Kornell N. Self-regulated learning: Beliefs, techniques, and illusions. Annual Review of Psychology. 2013; 64; 1: 417-444. 10.1146/annurev-psych-113011-143823</bibtext> </blist> <blist> <bibtext> Blissett S, Sibbald M, Kok E, van Merriënboer J. Optimizing self-regulation of performance: Is mental effort a cue?. Advances in Health Sciences Education. 2018; 23; 5: 891-898. 10.1007/s10459-018-9838-x</bibtext> </blist> <blist> <bibtext> Brünken R, Plass JL, Leutner D. Direct measurement of cognitive load in multimedia learning. Educational Psychologist. 2003; 38; 1: 53-61. 10.1207/S15326985EP3801_7</bibtext> </blist> <blist> <bibtext> Brünken, R., Seufert, T., &amp; Paas, F. (2006). Measuring cognitive load. In J. L. Plass, R. Moreno, &amp; R. Brünken (Eds.), Cognitive Load Theory (pp. 181–202). Cambridge University Press. https://doi.org/10.1017/CBO9780511844744.011</bibtext> </blist> <blist> <bibtext> Cacioppo JT, Petty RE. The need for cognition. Journal of Personality and Social Psychology. 1982; 42; 1: 116-131. 10.1037/0022-3514.42.1.116</bibtext> </blist> <blist> <bibtext> Castel AD. Metacognition and learning about primacy and recency effects in free recall: The utilization of intrinsic and extrinsic cues when making judgments of learning. Memory &amp; Cognition. 2008; 36; 2: 429-437. 10.3758/MC.36.2.429</bibtext> </blist> <blist> <bibtext> Chandler P, Sweller J. Cognitive load theory and the format of instruction. Cognition and Instruction. 1991; 8; 4: 293-332. 10.1207/s1532690xci0804_2</bibtext> </blist> <blist> <bibtext> Chen S, Epps J, Paas F. Pupillometric and blink measures of diverse task loads: Implications for working memory models. British Journal of Educational Psychology. 2022; 00: 1-21. 10.1111/bjep.12577</bibtext> </blist> <blist> <bibtext> de Bruin ABH, Roelle J, Carpenter SK, Baars M. Synthesizing cognitive load and self-regulation theory: A theoretical framework and research agenda. Educational Psychology Review. 2020; 32; 4: 903-915. 10.1007/s10648-020-09576-4</bibtext> </blist> <blist> <bibtext> de Bruin ABH, van Merriënboer JJG. Bridging cognitive load and self-regulated learning research: A complementary approach to contemporary issues in educational research. Learning and Instruction. 2017; 51: 1-9. 10.1016/j.learninstruc.2017.06.001</bibtext> </blist> <blist> <bibtext> de Jong T. Cognitive load theory, educational research, and instructional design: Some food for thought. Instructional Science. 2010; 38; 2: 105-134. 10.1007/s11251-009-9110-0</bibtext> </blist> <blist> <bibtext> De Neys W, Rossi S, Houdé O. Bats, balls, and substitution sensitivity: Cognitive misers are no happy fools. Psychonomic Bulletin &amp; Review. 2013; 20; 2: 269-273. 10.3758/s13423-013-0384-5</bibtext> </blist> <blist> <bibtext> Dunn TL, Gaspar C, Risko EF. Cue awareness in avoiding effortful control. Neuropsychologia. 2019; 123: 77-91. 10.1016/j.neuropsychologia.2018.05.011</bibtext> </blist> <blist> <bibtext> Dunn TL, Inzlicht M, Risko EF. Anticipating cognitive effort: Roles of perceived error-likelihood and time demands. Psychological Research Psychologische Forschung. 2019; 83; 5: 1033-1056. 10.1007/s00426-017-0943-x</bibtext> </blist> <blist> <bibtext> Dunn TL, Risko EF. Toward a metacognitive account of cognitive offloading. Cognitive Science. 2016; 40; 5: 1080-1127. 10.1111/cogs.12273</bibtext> </blist> <blist> <bibtext> Dweck CS, Chiu C, Hong Y. Implicit theories and their role in judgments and reactions: A word from two perspectives. Psychological Inquiry. 1995; 6; 4: 267-285. 10.1207/s15327965pli0604_1</bibtext> </blist> <blist> <bibtext> Efklides A. Metacognition: Defining its facets and levels of functioning in relation to self- and co-regulation. European Psychologist. 2008; 13; 4: 277-287. 10.1027/1016-9040.13.4.277</bibtext> </blist> <blist> <bibtext> Fiedler, K., Ackerman, R., &amp; Scarampi, C. (2019). Metacognition: Monitoring and controlling one's own knowledge, reasoning and decisions. In R. J. Sternberg &amp; J. Funke (Eds.), Introduction to the psychology of human thought (pp. 89–111). Heidelberg University Publishing. https://doi.org/10.17885/heiup.470.c6669</bibtext> </blist> <blist> <bibtext> Finn B, Tauber SK. When confidence is not a signal of knowing: How students' experiences and beliefs about processing fluency can lead to miscalibrated confidence. Educational Psychology Review. 2015; 27; 4: 567-586. 10.1007/s10648-015-9313-7</bibtext> </blist> <blist> <bibtext> Finucane ML, Gullion CM. Developing a tool for measuring the decision-making competence of older adults. Psychology and Aging. 2010; 25; 2: 271-288. 10.1037/a0019106</bibtext> </blist> <blist> <bibtext> Frederick S. Cognitive reflection and decision making. Journal of Economic Perspectives. 2005; 19; 4: 25-42. 10.1257/089533005775196732</bibtext> </blist> <blist> <bibtext> Gupta, N., Rigotti, L., &amp; Wilson, A. (2021). The experimenters' dilemma: Inferential preferences over populations. ArXiv Preprint. <ulink href="http://arxiv.org/abs/2107.05064">http://arxiv.org/abs/2107.05064</ulink></bibtext> </blist> <blist> <bibtext> Haji FA, Rojas D, Childs R, de Ribaupierre S, Dubrowski A. Measuring cognitive load: Performance, mental effort and simulation task complexity. Medical Education. 2015; 49; 8: 815-827. 10.1111/medu.12773</bibtext> </blist> <blist> <bibtext> Hawkins GE, Heathcote A. Racing against the clock: Evidence-based versus time-based decisions. Psychological Review. 2021; 128; 2: 222-263. 10.1037/rev0000259</bibtext> </blist> <blist> <bibtext> Hertwig R, Herzog SM, Schooler LJ, Reimer T. Fluency heuristic: A model of how the mind exploits a by-product of information retrieval. Journal of Experimental Psychology: Learning, Memory, and Cognition. 2008; 34; 5: 1191-1206. 10.1037/a0013025</bibtext> </blist> <blist> <bibtext> Jonsson A-C, Allwood CM. Stability and variability in the realism of confidence judgments over time, content domain, and gender. Personality and Individual Differences. 2003; 34; 4: 559-574. 10.1016/S0191-8869(02)00028-4</bibtext> </blist> <blist> <bibtext> Kelley CM, Lindsay DS. Remembering mistaken for knowing: Ease of retrieval as a basis for confidence in answers to general knowledge questions. Journal of Memory and Language. 1993; 32; 1: 1-24. 10.1006/jmla.1993.1001</bibtext> </blist> <blist> <bibtext> Kelley CM, Jacoby LL. Adult egocentrism: Subjective experience versus analytic bases for judgment. Journal of Memory and Language. 1996; 35; 2: 157-175. 10.1006/jmla.1996.0009</bibtext> </blist> <blist> <bibtext> Kirk-Johnson A, Galla BM, Fraundorf SH. Perceiving effort as poor learning: The misinterpreted-effort hypothesis of how experienced effort and perceived learning relate to study strategy choice. Cognitive Psychology. 2019; 115: 101237. 10.1016/j.cogpsych.2019.101237</bibtext> </blist> <blist> <bibtext> Klepsch M, Schmitz F, Seufert T. Development and validation of two instruments measuring intrinsic, extraneous, and germane cognitive load. Frontiers in Psychology. 2017; 8: 1-18. 10.3389/fpsyg.2017.01997</bibtext> </blist> <blist> <bibtext> Klepsch M, Seufert T. Understanding instructional design effects by differentiated measurement of intrinsic, extraneous, and germane cognitive load. Instructional Science. 2020; 48; 1: 45-77. 10.1007/s11251-020-09502-9</bibtext> </blist> <blist> <bibtext> Korbach A, Brünken R, Park B. Differentiating different types of cognitive load: A comparison of different measures. Educational Psychology Review. 2018; 30; 2: 503-529. 10.1007/s10648-017-9404-8</bibtext> </blist> <blist> <bibtext> Koriat A. Monitoring one's own knowledge during study: A cue-utilization approach to judgments of learning. Journal of Experimental Psychology: General. 1997; 126; 4: 349-370. 10.1037/0096-3445.126.4.349</bibtext> </blist> <blist> <bibtext> Koriat A. Easy comes, easy goes? The link between learning and remembering and its exploitation in metacognition. Memory and Cognition. 2008; 36; 2: 416-428. 10.3758/MC.36.2.416</bibtext> </blist> <blist> <bibtext> Koriat A, Ma'ayan H, Nussinson R. The intricate relationships between monitoring and control in metacognition: Lessons for the cause-and-effect relation between subjective experience and behavior. Journal of Experimental Psychology: General. 2006; 135; 1: 36-69. 10.1037/0096-3445.135.1.36</bibtext> </blist> <blist> <bibtext> Koriat A, Nussinson R, Bless H, Shaked N Dunlosky J, Bjork RA. Information-based and experience-based metacognitive judgments: Evidence from subjective confidence. Handbook of memory and metamemory. 2008; Psychology Press: 117-135</bibtext> </blist> <blist> <bibtext> Lauterman, T., &amp; Ackerman, R. (2019). Initial judgment of solvability in non-verbal problems – a predictor of solving processes. Metacognition and Learning, 14(3), 365–383. https://doi.org/10.1007/s11409-019-09194-8</bibtext> </blist> <blist> <bibtext> Leppink J, Paas F, Van der Vleuten CPM, Van Gog T, Van Merriënboer JJG. Development of an instrument for measuring different types of cognitive load. Behavior Research Methods. 2013; 45; 4: 1058-1072. 10.3758/s13428-013-0334-1</bibtext> </blist> <blist> <bibtext> Leppink J, Pérez-Fuster P. Mental effort, workload, time on task, and certainty: Beyond linear models. Educational Psychology Review. 2019; 31; 2: 421-438. 10.1007/s10648-018-09460-2</bibtext> </blist> <blist> <bibtext> Lins de Holanda Coelho, G., Hanel, P. H. P., &amp; Wolf, L. J. (2020). The very efficient assessment of need for cognition: Developing a six-item version. Assessment, 27(8), 1870–1885https://doi.org/10.1177/1073191118793208</bibtext> </blist> <blist> <bibtext> Lunney GH. Using analysis of variance with a dichotomous dependent variable: An empirical study. Journal of Educational Measurement. 1970; 7; 4: 263-269. 10.1111/j.1745-3984.1970.tb00727.x</bibtext> </blist> <blist> <bibtext> Metcalfe J, Finn B. Familiarity and retrieval processes in delayed judgments of learning. Journal of Experimental Psychology: Learning, Memory, and Cognition. 2008; 34; 5: 1084-1097. 10.1037/a0012580</bibtext> </blist> <blist> <bibtext> Metcalfe J, Finn B. Evidence that judgments of learning are causally related to study choice. Psychonomic Bulletin &amp; Review. 2008; 15; 1: 174-179. 10.3758/PBR.15.1.174</bibtext> </blist> <blist> <bibtext> Miele DB, Finn B, Molden DC. Does easily learned mean easily remembered?. Psychological Science. 2011; 22; 3: 320-324. 10.1177/0956797610397954</bibtext> </blist> <blist> <bibtext> Miesner MT, Maki RH. The role of test anxiety in absolute and relative metacomprehension accuracy. European Journal of Cognitive Psychology. 2007; 19; 4–5: 650-670. 10.1080/09541440701326196</bibtext> </blist> <blist> <bibtext> Naismith LM, Cheung JJH, Ringsted C, Cavalcanti RB. Limitations of subjective cognitive load measures in simulation-based procedural training. Medical Education. 2015; 49; 8: 805-814. 10.1111/medu.12732</bibtext> </blist> <blist> <bibtext> National Institute for Testing and Evaluation. (n.d.). The psychometric entrance test-practice tests.https://<ulink href="http://www.nite.org.il/psychometric-entrance-test/preparation/?lang=en">www.nite.org.il/psychometric-entrance-test/preparation/?lang=en</ulink></bibtext> </blist> <blist> <bibtext> Nelson, T. O., &amp; Narens, L. (1990). Metamemory: A theoretical framework and new findings. In G. H. Bower (Ed.), The psychology of learning and motivation (Vol. 26, Issue C, pp. 125–173). Academic Press. https://doi.org/10.1016/S0079-7421(08)60053-5</bibtext> </blist> <blist> <bibtext> Oldrati V, Patricelli J, Colombo B, Antonietti A. The role of dorsolateral prefrontal cortex in inhibition mechanism: A study on cognitive reflection test and similar tasks through neuromodulation. Neuropsychologia. 2016; 91: 499-508. 10.1016/j.neuropsychologia.2016.09.010</bibtext> </blist> <blist> <bibtext> Paas F. Training strategies for attaining transfer of problem-solving skill in statistics: A cognitive-load approach. Journal of Educational Psychology. 1992; 84; 4: 429-434. 10.1037/0022-0663.84.4.429</bibtext> </blist> <blist> <bibtext> Paas F, Tuovinen JE, Tabbers H, Van Gerven PWM. Cognitive load measurement as a means to advance cognitive load theory. Educational Psychologist. 2003; 38; 1: 63-71. 10.1207/S15326985EP3801_8</bibtext> </blist> <blist> <bibtext> Paas F, Tuovinen JE, van Merriënboer JJG, AubteenDarabi A. A motivational perspective on the relation between mental effort and performance: Optimizing learner involvement in instruction. Educational Technology Research and Development. 2005; 53; 3: 25-34. 10.1007/BF02504795</bibtext> </blist> <blist> <bibtext> Paas F, Van Merriënboer JJG. Instructional control of cognitive load in the training of complex cognitive tasks. Educational Psychology Review. 1994; 6; 4: 351-371. 10.1007/BF02213420</bibtext> </blist> <blist> <bibtext> Panadero E. A review of self-regulated learning: Six models and four directions for research. Frontiers in Psychology. 2017; 8; 422: 1-28. 10.3389/fpsyg.2017.00422</bibtext> </blist> <blist> <bibtext> Peer E, Brandimarte L, Samat S, Acquisti A. Beyond the Turk: Alternative platforms for crowdsourcing behavioral research. Journal of Experimental Social Psychology. 2017; 70: 153-163. 10.1016/j.jesp.2017.01.006</bibtext> </blist> <blist> <bibtext> Peer E, Rothschild D, Gordon A, Evernden Z, Damer E. Data quality of platforms and panels for online behavioral research. Behavior Research Methods. 2021; 54; 4: 1643-1662. 10.3758/s13428-021-01694-3</bibtext> </blist> <blist> <bibtext> Peters M, Battista C. Applications of mental rotation figures of the Shepard and Metzler type and description of a mental rotation stimulus library. Brain and Cognition. 2008; 66; 3: 260-264. 10.1016/j.bandc.2007.09.003</bibtext> </blist> <blist> <bibtext> Petty RE, Briñol P, Loersch C, McCaslin MJ Leary MR, Hoyle R. The need for cognition. Handbook of individual differences in social behavior. 2009; Guilford Press: 318-329</bibtext> </blist> <blist> <bibtext> Primi C, Morsanyi K, Chiesi F, Donati MA, Hamilton J. The development and testing of a new version of the cognitive reflection test applying item response theory (IRT). Journal of Behavioral Decision Making. 2016; 29; 5: 453-469. 10.1002/bdm.1883</bibtext> </blist> <blist> <bibtext> R Core Team. (2021). R: A language and environment for statistical computing [Software]. R Foundation for Statistical Computing. https://<ulink href="http://www.r-project.org/">www.r-project.org/</ulink></bibtext> </blist> <blist> <bibtext> Raaijmakers SF, Baars M, Schaap L, Paas F, van Gog T. Effects of performance feedback valence on perceptions of invested mental effort. Learning and Instruction. 2017; 51: 36-46. 10.1016/j.learninstruc.2016.12.002</bibtext> </blist> <blist> <bibtext> Reber R, Brun M, Mitterndorfer K. The use of heuristics in intuitive mathematical judgment. Psychonomic Bulletin &amp; Review. 2008; 15; 6: 1174-1178. 10.3758/PBR.15.6.1174</bibtext> </blist> <blist> <bibtext> Richter, J., Scheiter, K., &amp; Eitel, A. (2016). Signaling text-picture relations in multimedia learning: A comprehensive meta-analysis. Educational Research Review, 17, 19–36. https://doi.org/10.1016/j.edurev.2015.12.003</bibtext> </blist> <blist> <bibtext> Rop, G., Schüler, A., Verkoeijen, P. P. J. L., Scheiter, K., &amp; Gog, T. (2018). Effects of task experience and layout on learning from text and pictures with or without unnecessary picture descriptions. Journal of Computer Assisted Learning, 34(4), 458–470. https://doi.org/10.1111/jcal.12287</bibtext> </blist> <blist> <bibtext> Scheiter, K., Ackerman, R., &amp; Hoogerheide, V. (2020). Looking at mental effort appraisals through a metacognitive lens: Are they biased? Educational Psychology Review, 32(4), 1003–1027. https://doi.org/10.1007/s10648-020-09555-9</bibtext> </blist> <blist> <bibtext> Schmeck A, Opfermann M, van Gog T, Paas F, Leutner D. Measuring cognitive load with subjective rating scales during problem solving: Differences between immediate and delayed ratings. Instructional Science. 2015; 43; 1: 93-114. 10.1007/s11251-014-9328-3</bibtext> </blist> <blist> <bibtext> Schmider E, Ziegler M, Danay E, Beyer L, Bühner M. Is it really robust? Reinvestigating the robustness of ANOVA against violations of the normal distribution assumption. Methodology. 2010; 6; 4: 147-151. 10.1027/1614-2241/a000016</bibtext> </blist> <blist> <bibtext> Schwartz, B. L., &amp; Jemstedt, A. (2021). The role of fluency and dysfluency in metacognitive experiences. In D. Moraitou &amp; P. Metallidou (Eds.), Trends and prospects in metacognition research across the life span (pp. 25–40). Springer International Publishing. https://doi.org/10.1007/978-3-030-51673-4_2</bibtext> </blist> <blist> <bibtext> Searle JA, Hamm JP. Mental rotation: An examination of assumptions. Wires Cognitive Science. 2017; 8; 6: 701-703. 10.1002/wcs.1443</bibtext> </blist> <blist> <bibtext> Seufert T. Building bridges between self-regulation and cognitive load—an invitation for a broad and differentiated attempt. Educational Psychology Review. 2020; 32; 4: 1151-1162. 10.1007/s10648-020-09574-6</bibtext> </blist> <blist> <bibtext> Shepard RN, Metzler J. Mental rotation of three-dimensional objects. Science. 1971; 171; 3972: 701-703. 10.1126/science.171.3972.701</bibtext> </blist> <blist> <bibtext> Shtulman, A., &amp; McCallum, K. (2014). Cognitive reflection predicts science understanding. Proceedings of the Annual Meeting of the Cognitive Science Society, 36, 2937–2942.</bibtext> </blist> <blist> <bibtext> Sidi, Y., Shpigelman, M., Zalmanov, H., &amp; Ackerman, R. (2017). Understanding metacognitive inferiority on screen by exposing cues for depth of processing. Learning and Instruction, 51, 61–73. https://doi.org/10.1016/j.learninstruc.2017.01.002</bibtext> </blist> <blist> <bibtext> Sirota, M., Dewberry, C., Juanchich, M., Kostovičová, L., &amp; Marshall, A. C. (2018). Measuring cognitive reflection without maths: Developing and validating the verbal cognitive reflection test. PsyArXiv. https://doi.org/10.31234/osf.io/pfe79</bibtext> </blist> <blist> <bibtext> Sweller, J., Ayres, P., &amp; Kalyuga, S. (2011). Measuring cognitive load. In Cognitive load theory (pp. 71–85). Springer New York. https://doi.org/10.1007/978-1-4419-8126-4_6</bibtext> </blist> <blist> <bibtext> Sweller J, Van Merrienboer JJG, Paas F. Cognitive architecture and instructional design. Educational Psychology Review. 1998; 10; 3: 251-296. 10.1023/A:1022193728205</bibtext> </blist> <blist> <bibtext> Sweller J, van Merriënboer JJG, Paas F. Cognitive architecture and instructional design: 20 years later. Educational Psychology Review. 2019; 31; 2: 261-292. 10.1007/s10648-019-09465-5</bibtext> </blist> <blist> <bibtext> Szulewski A, Kelton D, Howes D. Pupillometry as a tool to study expertise in medicine. Frontline Learning Research. 2017; 5; 3: 55-65. 10.14786/flr.v5i3.256</bibtext> </blist> <blist> <bibtext> Taylor J, Deane FP. Development of a short form of the test anxiety inventory (TAI). The Journal of General Psychology. 2002; 129; 2: 127-136. 10.1080/00221300209603133</bibtext> </blist> <blist> <bibtext> Thiede KW, Anderson MCM, Therriault D. Accuracy of metacognitive monitoring affects learning of texts. Journal of Educational Psychology. 2003; 95; 1: 66-73. 10.1037/0022-0663.95.1.66</bibtext> </blist> <blist> <bibtext> Thompson, V. A., Evans, J. S. B. T., &amp; Campbell, J. I. D. (2013a). Matching bias on the selection task: It's fast and feels good. Thinking &amp; Reasoning,19(3–4), 431–452. https://doi.org/10.1080/13546783.2013.820220</bibtext> </blist> <blist> <bibtext> Thompson, V. A., Turner, J. A. P., Pennycook, G., Ball, L. J., Brack, H., Ophir, Y., &amp; Ackerman, R. (2013b). The role of answer fluency and perceptual fluency as metacognitive cues for initiating analytic thinking. Cognition, 128(2), 237–251. https://doi.org/10.1016/j.cognition.2012.09.012</bibtext> </blist> <blist> <bibtext> Thompson, V. A., &amp; Morsanyi, K. (2012). Analytic thinking: Do you feel like it? Mind &amp; Society,11(1), 93–105. https://doi.org/10.1007/s11299-012-0100-6</bibtext> </blist> <blist> <bibtext> Toplak ME, West RF, Stanovich KE. Assessing miserly information processing: An expansion of the cognitive reflection test. Thinking &amp; Reasoning. 2014; 20; 2: 147-168. 10.1080/13546783.2013.844729</bibtext> </blist> <blist> <bibtext> Trippas D, Handley SJ, Verde MF, Morsanyi K. Logic brightens my day: Evidence for implicit sensitivity to logical validity. Journal of Experimental Psychology: Learning, Memory, and Cognition. 2016; 42; 9: 1448-1457. 10.1037/xlm0000248</bibtext> </blist> <blist> <bibtext> Undorf, M. (2020). Fluency illusions in metamemory. In A. M. Cleary &amp; B. L. Schwartz (Eds.), Memory quirks: The study of odd phenomena in memory (pp. 150–174). Routledge. https://doi.org/10.4324/9780429264498-12</bibtext> </blist> <blist> <bibtext> Undorf M, Bröder A. Cue integration in metamemory judgements is strategic. Quarterly Journal of Experimental Psychology. 2020; 73; 4: 629-642. 10.1177/1747021819882308</bibtext> </blist> <blist> <bibtext> Undorf M, Söllner A, Bröder A. Simultaneous utilization of multiple cues in judgments of learning. Memory &amp; Cognition. 2018; 46; 4: 507-519. 10.3758/s13421-017-0780-6</bibtext> </blist> <blist> <bibtext> Valerjev, P. (2019). Chronometry and meta-reasoning in a modified cognitive reflection test. In K. Damnjanović, O. Tošković, &amp; S. Marković (Eds.), Proceedings of the XXV Scientific Conference: Empirical Studies in Psychology (pp. 31–34).</bibtext> </blist> <blist> <bibtext> van Gog, T. (2022). The signaling (or cueing) principle in multimedia learning. In R. E. Mayer &amp; L. Fiorella (Eds.), The Cambridge handbook of multimedia learning (3rd ed., pp. 221–230). Cambridge University Press. https://doi.org/10.1017/9781108894333.022</bibtext> </blist> <blist> <bibtext> van Gog T, Kirschner F, Kester L, Paas F. Timing and frequency of mental effort measurement: Evidence in favour of repeated measures. Applied Cognitive Psychology. 2012; 26; 6: 833-839. 10.1002/acp.2883</bibtext> </blist> <blist> <bibtext> van Gog T, Paas F. Instructional efficiency: Revisiting the original construct in educational research. Educational Psychologist. 2008; 43; 1: 16-26. 10.1080/00461520701756248</bibtext> </blist> <blist> <bibtext> van Gog T, Sweller J. Not new, but nearly forgotten: The testing effect decreases or even disappears as the complexity of learning materials increases. Educational Psychology Review. 2015; 27; 2: 247-264. 10.1007/s10648-015-9310-x</bibtext> </blist> <blist> <bibtext> van Gog, T., Hoogerheide, V., &amp; van Harsel, M. (2020). The role of mental effort in fostering self-regulated learning with problem-solving tasks. Educational Psychology Review, 32(4), 1055–1072. https://doi.org/10.1007/s10648-020-09544-y</bibtext> </blist> <blist> <bibtext> van Merriënboer, J. J. G., &amp; Kirschner, P. A. (2017). Ten steps to complex learning: A systematic approach to four-component instructional design (3rd ed.). Routledge. https://doi.org/10.4324/9781315113210</bibtext> </blist> <blist> <bibtext> Vandenberg SG, Kuse AR. Mental rotations, a group test of three-dimensional spatial visualization. Perceptual and Motor Skills. 1978; 47; 2: 599-604. 10.2466/pms.1978.47.2.599</bibtext> </blist> <blist> <bibtext> Wang S, Thompson V. Fluency and feeling of rightness: The effect of anchoring and models. Psihologijske Teme. 2019; 28; 1: 37-72. 10.31820/pt.28.1.3</bibtext> </blist> <blist> <bibtext> Winne PH, Perry NE Boekaerts M, Pintrich PR, Zeidner M. Measuring self-regulated learning. Handbook of self-regulation. 2000; Academic Press: 531-566. 10.1016/B978-012109890-2/50045-7</bibtext> </blist> <blist> <bibtext> Young, A. G., Powers, A., Pilgrim, L., &amp; Shtulman, A. (2018). Developing a cognitive reflection test for school-age children. In T. T. Rogers, M. Rau, X. Zhu, &amp; C. W. Kalish (Eds.), Proceedings of the 40th Annual Conference of the Cognitive Science Society (pp. 1232–1237). Cognitive Science Society.</bibtext> </blist> <blist> <bibtext> Zimmerman BJ. Becoming a self-regulated learner: An overview. Theory into Practice. 2002; 41; 2: 64-70. 10.1207/s15430421tip4102_2</bibtext> </blist> </ref> <ref id="AN0163936292-62"> <title> Footnotes </title> <blist> <bibtext> CLT also distinguishes between three types of cognitive load (intrinsic load, extraneous load, and germane load), for each of which recent research has developed and validated unique measures (e.g., Klepsch and Seufert [45]; Klepsch et al., [44]; Leppink et al., [52]). However, investigating the different types of cognitive load was not the focus of this study. Moreover, the distinction itself, as well as the construct of germane load as a distinguishable category of cognitive load, is under debate in the CLT literature (Sweller et al., [91]). As the present study is an initial investigation of the bases of self-reported effort and difficulty, we focus on the classic framing of mental effort and rely on the most common measures used in CLT research.</bibtext> </blist> <blist> <bibtext> Background variables were also exploratively examined as possible moderators in all three experiments. However, since few significant results and, in particular, no consistent patterns emerged, the results are not reported for the sake of brevity.</bibtext> </blist> <blist> <bibtext> Four participants stated that they had already encountered most or all of the items. However, since excluding them from the analyses did not change the pattern of results, these participants were kept in the sample to maintain statistical power.</bibtext> </blist> <blist> <bibtext> The assumption of normally distributed data was violated in all three experiments. However, in such cases, the <emph>F</emph>-statistic in fixed effects models is still considered robust when group sizes are equal (Lunney [55]; Schmider et al., [81]). Furthermore, the assumption of homogeneity of variances was violated when answering RQ1 in experiment 2. Since data transformation did not resolve this issue and using non-parametric alternatives to ANOVA did not change the pattern of results, we report the results of ANOVA throughout the manuscript.</bibtext> </blist> </ref> <aug> <p>Reported by Author; Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib24" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib23" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib114" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib32" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib62" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib77" firstref="ref6"></nolink> <nolink nlid="nl7" bibid="bib104" firstref="ref7"></nolink> <nolink nlid="nl8" bibid="bib21" firstref="ref8"></nolink> <nolink nlid="nl9" bibid="bib11" firstref="ref9"></nolink> <nolink nlid="nl10" bibid="bib16" firstref="ref10"></nolink> <nolink nlid="nl11" bibid="bib79" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib84" firstref="ref13"></nolink> <nolink nlid="nl13" bibid="bib108" firstref="ref14"></nolink> <nolink nlid="nl14" bibid="bib50" firstref="ref18"></nolink> <nolink nlid="nl15" bibid="bib15" firstref="ref20"></nolink> <nolink nlid="nl16" bibid="bib20" firstref="ref21"></nolink> <nolink nlid="nl17" bibid="bib48" firstref="ref22"></nolink> <nolink nlid="nl18" bibid="bib100" firstref="ref23"></nolink> <nolink nlid="nl19" bibid="bib47" firstref="ref27"></nolink> <nolink nlid="nl20" bibid="bib31" firstref="ref30"></nolink> <nolink nlid="nl21" bibid="bib102" firstref="ref37"></nolink> <nolink nlid="nl22" bibid="bib33" firstref="ref39"></nolink> <nolink nlid="nl23" bibid="bib56" firstref="ref40"></nolink> <nolink nlid="nl24" bibid="bib57" firstref="ref41"></nolink> <nolink nlid="nl25" bibid="bib87" firstref="ref42"></nolink> <nolink nlid="nl26" bibid="bib29" firstref="ref43"></nolink> <nolink nlid="nl27" bibid="bib27" firstref="ref44"></nolink> <nolink nlid="nl28" bibid="bib75" firstref="ref46"></nolink> <nolink nlid="nl29" bibid="bib96" firstref="ref48"></nolink> <nolink nlid="nl30" bibid="bib82" firstref="ref50"></nolink> <nolink nlid="nl31" bibid="bib111" firstref="ref52"></nolink> <nolink nlid="nl32" bibid="bib13" firstref="ref53"></nolink> <nolink nlid="nl33" bibid="bib39" firstref="ref54"></nolink> <nolink nlid="nl34" bibid="bib41" firstref="ref56"></nolink> <nolink nlid="nl35" bibid="bib14" firstref="ref57"></nolink> <nolink nlid="nl36" bibid="bib42" firstref="ref59"></nolink> <nolink nlid="nl37" bibid="bib67" firstref="ref60"></nolink> <nolink nlid="nl38" bibid="bib109" firstref="ref61"></nolink> <nolink nlid="nl39" bibid="bib22" firstref="ref62"></nolink> <nolink nlid="nl40" bibid="bib46" firstref="ref63"></nolink> <nolink nlid="nl41" bibid="bib92" firstref="ref64"></nolink> <nolink nlid="nl42" bibid="bib60" firstref="ref65"></nolink> <nolink nlid="nl43" bibid="bib17" firstref="ref66"></nolink> <nolink nlid="nl44" bibid="bib18" firstref="ref67"></nolink> <nolink nlid="nl45" bibid="bib25" firstref="ref68"></nolink> <nolink nlid="nl46" bibid="bib64" firstref="ref69"></nolink> <nolink nlid="nl47" bibid="bib106" firstref="ref72"></nolink> <nolink nlid="nl48" bibid="bib49" firstref="ref75"></nolink> <nolink nlid="nl49" bibid="bib80" firstref="ref80"></nolink> <nolink nlid="nl50" bibid="bib105" firstref="ref81"></nolink> <nolink nlid="nl51" bibid="bib53" firstref="ref86"></nolink> <nolink nlid="nl52" bibid="bib28" firstref="ref87"></nolink> <nolink nlid="nl53" bibid="bib90" firstref="ref88"></nolink> <nolink nlid="nl54" bibid="bib37" firstref="ref89"></nolink> <nolink nlid="nl55" bibid="bib38" firstref="ref91"></nolink> <nolink nlid="nl56" bibid="bib66" firstref="ref92"></nolink> <nolink nlid="nl57" bibid="bib68" firstref="ref93"></nolink> <nolink nlid="nl58" bibid="bib112" firstref="ref94"></nolink> <nolink nlid="nl59" bibid="bib78" firstref="ref99"></nolink> <nolink nlid="nl60" bibid="bib94" firstref="ref102"></nolink> <nolink nlid="nl61" bibid="bib35" firstref="ref110"></nolink> <nolink nlid="nl62" bibid="bib51" firstref="ref112"></nolink> <nolink nlid="nl63" bibid="bib76" firstref="ref113"></nolink> <nolink nlid="nl64" bibid="bib85" firstref="ref115"></nolink> <nolink nlid="nl65" bibid="bib83" firstref="ref116"></nolink> <nolink nlid="nl66" bibid="bib107" firstref="ref118"></nolink> <nolink nlid="nl67" bibid="bib19" firstref="ref119"></nolink> <nolink nlid="nl68" bibid="bib93" firstref="ref120"></nolink> <nolink nlid="nl69" bibid="bib30" firstref="ref121"></nolink> <nolink nlid="nl70" bibid="bib40" firstref="ref122"></nolink> <nolink nlid="nl71" bibid="bib43" firstref="ref123"></nolink> <nolink nlid="nl72" bibid="bib58" firstref="ref124"></nolink> <nolink nlid="nl73" bibid="bib59" firstref="ref125"></nolink> <nolink nlid="nl74" bibid="bib72" firstref="ref126"></nolink> <nolink nlid="nl75" bibid="bib63" firstref="ref129"></nolink> <nolink nlid="nl76" bibid="bib54" firstref="ref135"></nolink> <nolink nlid="nl77" bibid="bib74" firstref="ref136"></nolink> <nolink nlid="nl78" bibid="bib71" firstref="ref156"></nolink> <nolink nlid="nl79" bibid="bib101" firstref="ref158"></nolink> <nolink nlid="nl80" bibid="bib110" firstref="ref162"></nolink> <nolink nlid="nl81" bibid="bib95" firstref="ref178"></nolink> <nolink nlid="nl82" bibid="bib97" firstref="ref179"></nolink> <nolink nlid="nl83" bibid="bib12" firstref="ref186"></nolink> <nolink nlid="nl84" bibid="bib89" firstref="ref192"></nolink> <nolink nlid="nl85" bibid="bib10" firstref="ref194"></nolink> <nolink nlid="nl86" bibid="bib36" firstref="ref199"></nolink> <nolink nlid="nl87" bibid="bib69" firstref="ref200"></nolink> <nolink nlid="nl88" bibid="bib70" firstref="ref201"></nolink> <nolink nlid="nl89" bibid="bib65" firstref="ref203"></nolink> <nolink nlid="nl90" bibid="bib45" firstref="ref205"></nolink> <nolink nlid="nl91" bibid="bib44" firstref="ref206"></nolink> <nolink nlid="nl92" bibid="bib52" firstref="ref208"></nolink> <nolink nlid="nl93" bibid="bib26" firstref="ref212"></nolink> <nolink nlid="nl94" bibid="bib34" firstref="ref213"></nolink> <nolink nlid="nl95" bibid="bib61" firstref="ref214"></nolink> <nolink nlid="nl96" bibid="bib73" firstref="ref216"></nolink> <nolink nlid="nl97" bibid="bib86" firstref="ref217"></nolink> <nolink nlid="nl98" bibid="bib88" firstref="ref218"></nolink> <nolink nlid="nl99" bibid="bib98" firstref="ref219"></nolink> <nolink nlid="nl100" bibid="bib99" firstref="ref220"></nolink> <nolink nlid="nl101" bibid="bib103" firstref="ref221"></nolink> <nolink nlid="nl102" bibid="bib113" firstref="ref222"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1378898 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Comparing Mental Effort, Difficulty, and Confidence Appraisals in Problem-Solving: A Metacognitive Perspective – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Hoch%2C+Emely%22">Hoch, Emely</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-6534-1506">0000-0002-6534-1506</externalLink>)<br /><searchLink fieldCode="AR" term="%22Sidi%2C+Yael%22">Sidi, Yael</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-2503-9166">0000-0003-2503-9166</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ackerman%2C+Rakefet%22">Ackerman, Rakefet</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-9583-8014">0000-0001-9583-8014</externalLink>)<br /><searchLink fieldCode="AR" term="%22Hoogerheide%2C+Vincent%22">Hoogerheide, Vincent</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-4176-0973">0000-0003-4176-0973</externalLink>)<br /><searchLink fieldCode="AR" term="%22Scheiter%2C+Katharina%22">Scheiter, Katharina</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-9397-7544">0000-0002-9397-7544</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Psychology+Review%22"><i>Educational Psychology Review</i></searchLink>. Jun 2023 35(2). – Name: Avail Label: Availability Group: Avail Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 37 – Name: DatePubCY Label: Publication Date Group: Date Data: 2023 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Metacognition%22">Metacognition</searchLink><br /><searchLink fieldCode="DE" term="%22Self+Efficacy%22">Self Efficacy</searchLink><br /><searchLink fieldCode="DE" term="%22Self+Evaluation+%28Individuals%29%22">Self Evaluation (Individuals)</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Behavior%22">Student Behavior</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Cognitive+Processes%22">Cognitive Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Logical+Thinking%22">Logical Thinking</searchLink><br /><searchLink fieldCode="DE" term="%22Correlation%22">Correlation</searchLink><br /><searchLink fieldCode="DE" term="%22Reaction+Time%22">Reaction Time</searchLink><br /><searchLink fieldCode="DE" term="%22Success%22">Success</searchLink><br /><searchLink fieldCode="DE" term="%22Predictor+Variables%22">Predictor Variables</searchLink><br /><searchLink fieldCode="DE" term="%22Problem+Solving%22">Problem Solving</searchLink><br /><searchLink fieldCode="DE" term="%22Cues%22">Cues</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1007/s10648-023-09779-5 – Name: ISSN Label: ISSN Group: ISSN Data: 1040-726X<br />1573-336X – Name: Abstract Label: Abstract Group: Ab Data: It is well established in educational research that metacognitive monitoring of performance assessed by self-reports, for instance, asking students to report their confidence in provided answers, is based on heuristic cues rather than on actual success in the task. Subjective self-reports are also used in educational research on cognitive load, where they refer to the perceived amount of mental effort invested in or difficulty of each task item. In the present study, we examined the potential underlying bases and the predictive value of mental effort and difficulty appraisals compared to confidence appraisals by applying metacognitive concepts and paradigms. In three experiments, participants faced verbal logic problems or one of two non-verbal reasoning tasks. In a between-participants design, each task item was followed by either mental effort, difficulty, or confidence appraisals. We examined the associations between the various appraisals, response time, and success rates. Consistently across all experiments, we found that mental effort and difficulty appraisals were associated more strongly than confidence with response time. Further, while all appraisals were highly predictive of solving success, the strength of this association was stronger for difficulty and confidence appraisals (which were similar) than for mental effort appraisals. We conclude that mental effort and difficulty appraisals are prone to misleading cues like other metacognitive judgments and are based on unique underlying processes. These findings challenge the accepted notion that mental effort appraisals can serve as reliable reflections of cognitive load. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: Note Label: Notes Group: Note Data: https://doi.org/10.17605/osf.io/q2p74 – Name: DateEntry Label: Entry Date Group: Date Data: 2023 – Name: AN Label: Accession Number Group: ID Data: EJ1378898 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1378898 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10648-023-09779-5 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 37 Subjects: – SubjectFull: Metacognition Type: general – SubjectFull: Self Efficacy Type: general – SubjectFull: Self Evaluation (Individuals) Type: general – SubjectFull: Student Behavior Type: general – SubjectFull: Difficulty Level Type: general – SubjectFull: Cognitive Processes Type: general – SubjectFull: Logical Thinking Type: general – SubjectFull: Correlation Type: general – SubjectFull: Reaction Time Type: general – SubjectFull: Success Type: general – SubjectFull: Predictor Variables Type: general – SubjectFull: Problem Solving Type: general – SubjectFull: Cues Type: general Titles: – TitleFull: Comparing Mental Effort, Difficulty, and Confidence Appraisals in Problem-Solving: A Metacognitive Perspective Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Hoch, Emely – PersonEntity: Name: NameFull: Sidi, Yael – PersonEntity: Name: NameFull: Ackerman, Rakefet – PersonEntity: Name: NameFull: Hoogerheide, Vincent – PersonEntity: Name: NameFull: Scheiter, Katharina IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Type: published Y: 2023 Identifiers: – Type: issn-print Value: 1040-726X – Type: issn-electronic Value: 1573-336X Numbering: – Type: volume Value: 35 – Type: issue Value: 2 Titles: – TitleFull: Educational Psychology Review Type: main |
| ResultId | 1 |