Does Timed Testing Affect the Interpretation of Efficiency Scores?--A GLMM Analysis of Reading Components

Saved in:
Bibliographic Details
Title: Does Timed Testing Affect the Interpretation of Efficiency Scores?--A GLMM Analysis of Reading Components
Language: English
Authors: Frank Goldhammer (ORCID 0000-0003-0289-9534), Ulf Kroehne (ORCID 0000-0002-0412-169X), Carolin Hahnel (ORCID 0000-0003-2394-3944), Johannes Naumann (ORCID 0000-0002-2625-9630), Paul De Boeck (ORCID 0000-0002-0884-2582)
Source: Journal of Educational Measurement. 2024 61(3):349-377.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 29
Publication Date: 2024
Document Type: Journal Articles
Reports - Research
Descriptors: Timed Tests, Efficiency, Scores, Test Interpretation, Test Items, Performance Based Assessment, Difficulty Level, Word Recognition, Reading Skills, Construct Validity, Semantics
DOI: 10.1111/jedm.12393
ISSN: 0022-0655
1745-3984
Abstract: The efficiency of cognitive component skills is typically assessed with speeded performance tests. Interpreting only effective ability or effective speed as efficiency may be challenging because of the within-person dependency between both variables (speed-ability tradeoff, SAT). The present study measures efficiency as effective ability conditional on speed by controlling speed experimentally. Item-level time limits control the stimulus presentation time and the time window for responding (timed condition). The overall goal was to examine the construct validity of effective ability scores obtained from untimed and timed condition by comparing the effects of theory-based item properties on item difficulty. If such effects exist, the scores reflect how well the test-takers were able to cope with the theory-based requirements. A German subsample from PISA 2012 completed two reading component skills tasks (i.e., word recognition and semantic integration) with and without item-level time limits. Overall, the included linguistic item properties showed stronger effects on item difficulty in the timed than the untimed condition. In the semantic integration task, item properties explained the time required in the untimed condition. The results suggest that effective ability scores in the timed condition better reflect how well test-takers were able to cope with the theoretically relevant task demands.
Abstractor: As Provided
Entry Date: 2024
Accession Number: EJ1449104
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHnIRlZuOt1G59gpyA2g_OXAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDOKFu5gW4MPsyjbj-gIBEICBm_kBkOyengZUHpAh9ZpO-qjvLMouqaOXxVO33vjopkqkFRZzCytjB0ruh4So-lPATz6SE5SqFYKCArX6ruG8T5CQBaBB6gboIUM39xktt2Y_os6CSG_TQp0wsO4EY826zq5LqIvJ12pSZlNtg_Yp4ObbUXHMzCmoMnEHvXKQr_L6t5eFqbPFZAfO54lSyPpX5lXnI8p5VN2Fb2RN
Text:
  Availability: 1
  Value: <anid>AN0180987481;mea01sep.24;2024Nov22.02:20;v2.2.500</anid> <title id="AN0180987481-1">Does Timed Testing Affect the Interpretation of Efficiency Scores?—A GLMM Analysis of Reading Components </title> <p>The efficiency of cognitive component skills is typically assessed with speeded performance tests. Interpreting only effective ability or effective speed as efficiency may be challenging because of the within‐person dependency between both variables (speed‐ability tradeoff, SAT). The present study measures efficiency as effective ability conditional on speed by controlling speed experimentally. Item‐level time limits control the stimulus presentation time and the time window for responding (timed condition). The overall goal was to examine the construct validity of effective ability scores obtained from untimed and timed condition by comparing the effects of theory‐based item properties on item difficulty. If such effects exist, the scores reflect how well the test‐takers were able to cope with the theory‐based requirements. A German subsample from PISA 2012 completed two reading component skills tasks (i.e., word recognition and semantic integration) with and without item‐level time limits. Overall, the included linguistic item properties showed stronger effects on item difficulty in the timed than the untimed condition. In the semantic integration task, item properties explained the time required in the untimed condition. The results suggest that effective ability scores in the timed condition better reflect how well test‐takers were able to cope with the theoretically relevant task demands.</p> <p>Cognitive efficiency can be defined as the product in a cognitive task (work output, e.g., accuracy) in relation to its costs (work input, e.g., time taken). For example, for a given work output, efficiency is higher if the work input was lower, and conversely. Cognitive efficiency tests play an important role in various educational assessment contexts. They are used in research studies (e.g., Kim et al., [<reflink idref="bib28" id="ref1">28</reflink>]), in national (e.g., NAEP's oral reading fluency test; White et al., [<reflink idref="bib53" id="ref2">53</reflink>]) and international (e.g., PISA's reading fluency test; Avvisati, [<reflink idref="bib4" id="ref3">4</reflink>]) evaluation studies as well as in classroom assessments (e.g., Merrell & Tymms, [<reflink idref="bib36" id="ref4">36</reflink>]). Cognitive efficiency tests typically assess subprocesses or foundational skills (e.g., word recognition) that enable and facilitate higher‐order cognitive processes (e.g., reading comprehension). Efficiency of executing subprocesses frees cognitive resources for learning and may even be required for higher‐order cognitive tasks. Therefore, efficiency tests can provide valuable information about learning progress and indicate potential underlying deficits.</p> <p>Typically, cognitive efficiency is measured by performance tests in which test‐takers must solve relatively simple tasks correctly and quickly. However, considering only speed or accuracy gives an incomplete picture because of their within‐person dependency (e.g., Wickelgren, [<reflink idref="bib54" id="ref5">54</reflink>]): Test‐takers can trade accuracy for speed and vice versa. One way to deal with individual differences in the tradeoff between speed and accuracy is to control response speed experimentally at item level, that is, the test‐taker's criterion on how much time to spend to produce a valid answer. This can be achieved by controlling the stimulus presentation time and the time window for responding at item‐level.</p> <p>The general goal of the present study is to investigate whether the construct interpretation of effective ability scores obtained from this kind of timed testing is valid in terms of cognitive efficiency or even improved compared to effective ability scores from untimed testing ("effective" ability refers to the accuracy rate demonstrated by the test‐taker given the chosen speed). To this end, we investigate the effects of theory‐based item properties on item difficulty within an item‐response theory (IRT) framework. If effects of item properties exist, they provide evidence for the validity of the construct interpretation of test scores. If such effects are found to be stronger in timed testing compared to untimed testing, the effective ability scores from timed testing would better reflect how well the test‐takers were able to cope with the theoretically relevant task requirements. The validation is done for two reading component skills tasks, visual word recognition and sentence‐level semantic integration, administered in an untimed and a timed testing condition. To compare the effect of item properties between timed and untimed testing, the interaction effect between item property and condition is examined.</p> <hd id="AN0180987481-2">Timed Testing to Measure the Efficiency of Component Skills</hd> <p></p> <hd id="AN0180987481-3">The tradeoff between effective speed and effective ability</hd> <p>The efficiency of cognitive component skills is typically assessed with speeded performance tests. The test‐taker is required to either correctly respond to as many items as possible in a limited amount of time or spend as little time as possible to correctly respond to a fixed number of items (speed test, Gulliksen, [<reflink idref="bib23" id="ref6">23</reflink>]). Observed individual differences in item responses and item‐response times can be explained by latent person variables of effective ability and effective speed, respectively (van der Linden, [<reflink idref="bib51" id="ref7">51</reflink>]). Furthermore, conditional dependencies within items may exist meaning that for an item the relationship between item response and item‐response time cannot be fully represented by the correlation of effective speed and effective ability. "Effective" refers to the actual balance between speed and accuracy a person chooses when completing the test. A test‐taker may change speed and accuracy across situations or conditions. The within‐person dependency of effective ability and effective speed is referred to as the speed‐ability tradeoff (SAT) function (van der Linden, [<reflink idref="bib51" id="ref8">51</reflink>]), in line with the speed‐accuracy tradeoff function investigated in experimental cognitive psychology (Luce, [<reflink idref="bib35" id="ref9">35</reflink>]; Wickelgren, [<reflink idref="bib54" id="ref10">54</reflink>]). Thus, a test‐taker completing a test may increase the effective speed at the cost of the effective ability resulting in a negative within‐person relation of effective ability and effective speed. A main difference between the speed‐accuracy tradeoff function and the SAT function is that the latter refers to latent variables (i.e., effective ability and effective speed) while the former to observable performance variables (i.e., average of accuracy and reaction time). Nevertheless, both kinds of variables are linked in measurement models separating the effects of person and item on performance.</p> <p>Obviously, the actual balance chosen by a test‐taker has consequences for ability estimates (based on response accuracy) and speed estimates (based on response times). Most importantly, the interpretation of observed differences in effective ability without considering or controlling for effective speed is ambiguous because they may be due to differences in the individual SAT function (i.e., true ability differences), differences in the decision on effective speed (i.e., the chosen tradeoff between effective speed and effective ability), or both (Goldhammer, [<reflink idref="bib17" id="ref11">17</reflink>]; Goldhammer et al., [<reflink idref="bib21" id="ref12">21</reflink>]).</p> <hd id="AN0180987481-4">Measuring cognitive efficiency conditional on speed</hd> <p>To deal with the interpretative ambiguity induced by individual differences in the chosen effective speed level, the general approach of the present study is to measure effective ability conditional on a fixed speed level. This is done by implementing a single speed condition, which is realized at the item level by controlling the presentation time of the stimulus, followed by a fixed time window for the required response (timed condition). Note that in this timed condition, cognitive efficiency is only reflected by observed differences in response accuracy because response time has become an experimentally controlled variable. Thus, for the timed condition we interpret effective ability derived from response accuracy <emph>as</emph> a measure of cognitive efficiency. In contrast, in the untimed condition without any time constraints, cognitive efficiency is always represented by effective ability derived from response accuracy <emph>and</emph> effective speed derived from response times (see e.g., Müller et al., [<reflink idref="bib37" id="ref13">37</reflink>], proposing a ratio of mean accuracy and mean log‐transformed response time).</p> <p>The suggested approach with one speed condition at the item level does not provide the full SAT function, which requires multiple speed conditions ranging from slow to fast. Nevertheless, individual differences in effective speed can already be controlled with a single speed condition, and ambiguity in the interpretation of effective ability estimates obtained for this single speed condition is reduced. As the test‐taker's decision criterion on how much time to take to produce a response to a stimulus is controlled experimentally, the response speed is not subject to choice (e.g., Davison et al., [<reflink idref="bib10" id="ref14">10</reflink>]; Goldhammer & Kroehne, [<reflink idref="bib19" id="ref15">19</reflink>]; Lohman, [<reflink idref="bib34" id="ref16">34</reflink>]; Salthouse & Hedden, [<reflink idref="bib47" id="ref17">47</reflink>]).</p> <p>Basically, we assume that individual differences in cognitive efficiency are reflected in how quickly and correctly a task can be solved. For the proposed timed testing procedure, we fix speed experimentally and observe accuracy only (i.e., responding in time correctly). The stimulus presentation time needs to be set in a way that empirical evidence can be elicited to infer cognitive efficiency. Obviously, a moderate time pressure (speededness) is needed to distinguish more or less efficient test‐takers based on their accuracies. Too strong time pressure would give rise to rapid random guessing regardless of the individual's cognitive efficiency (Goldhammer & Kroehne, [<reflink idref="bib19" id="ref18">19</reflink>]).</p> <p>In timed testing, accurate performance in an item depends on the item's difficulty and required processing time (time intensity, usually positively correlated with difficulty), the test‐taker's cognitive efficiency, and the available time for processing the stimulus and answering. As we want to control the test‐taker's criterion on how much time to spend to produce a valid answer, the stimulus presentation time needs to be predictable and is therefore constant across all items regardless of the items' time intensities. As a consequence, items with relative low time intensity will not contribute much to the measurement of highly efficient test‐takers. However, they are expected to distinguish between less efficient test‐takers whereas items with higher time intensity distinguish between efficient test‐takers. Note that efficient test‐takers may respond in the response window as required but have actually finished cognitive processing earlier (i.e., evidence accumulation and preparing a choice; Heitz, [<reflink idref="bib25" id="ref19">25</reflink>]). This is not problematic as long as this happens without being at the expense of the quality of the decision (i.e., accuracy). Actually, the issue of fit between person and item is comparable to (non‐adaptive) ability testing with easy items that do not discriminate between highly able test‐takers but between less able test‐takers and vice versa.</p> <p>Note that cognitive efficiency tests have been developed with time limits at test‐level (e.g., Test of Silent Reading Efficiency and Comprehension, TOSREC; Wagner et al., [<reflink idref="bib52" id="ref20">52</reflink>]). Although limiting the overall time is a way to manipulate overall time pressure and difficulty, in such tests it is still up to the test‐taker how to use the available time across items. In principle, engaged test‐takers could underestimate or overestimate the target speed required by the given time limit and the time intensities of items. As a consequence, obtained patterns of responses (correct vs. incorrect items) and missing responses (in particular not reached items) are not directly comparable between test‐takers if the decision on effective speed differs (i.e., the speed‐ability tradeoff). When using item‐level time constraints, we expect behavioral differences between test‐takers to be more comparable because the decision about effective speed is controlled, with test‐takers knowing in advance how much time is available to complete an item. However, it still has to be demonstrated that effective ability scores obtained from the proposed timed testing procedure can be interpreted validly as cognitive efficiency. A concern could be that time pressure at item level introduces construct‐irrelevant variance (e.g., test anxiety, perceived strain).</p> <hd id="AN0180987481-5">Validation Approach for the Interpretation of Cognitive Efficiency Scores</hd> <p></p> <hd id="AN0180987481-6">Construct validation based on item properties</hd> <p>The present study aims to validate the construct interpretation of effective ability scores obtained from the timed testing procedure in terms of cognitive efficiency. To this end, we examine whether item difficulty (i.e., the accuracy of item responses) is a function of item properties that define the construct‐specific challenges the test‐taker needs to meet to give a correct response.</p> <p>A valid construct interpretation requires that individual differences in scores are causally determined by the theoretical construct that the test score is intended to measure (AERA et al., [<reflink idref="bib1" id="ref21">1</reflink>]; Kane, [<reflink idref="bib27" id="ref22">27</reflink>]). Typically, construct validation considers person‐level differences by investigating the relationship of the test score variable to other variables representing related constructs (convergent evidence based on relations to other variables, AERA et al., [<reflink idref="bib1" id="ref23">1</reflink>]; nomological network, Cronbach & Meehl, [<reflink idref="bib9" id="ref24">9</reflink>]; nomothetic span, Embretson, [<reflink idref="bib15" id="ref25">15</reflink>]). However, in the present study we investigate between‐item differences as a source of validity evidence supporting the construct interpretation. Evidence is provided by examining the effects of item properties that can be assumed to explain task performance based on substantive theory. Following the construct representation approach (Embretson, [<reflink idref="bib15" id="ref26">15</reflink>]), item properties are identified based on a cognitive model or theory defining the measured ability construct and underlying information processing. It is assumed that, in particular, stimulus properties determine the required components of information processing and thus account for the difficulty of an item. In a theory‐based test development process, these properties are systematically considered and determine the construction of the items. This supports not only the validity of the construct interpretation of test scores but also the validity of the generalization inference (content validity).</p> <p>Let us assume that item difficulty can be empirically explained by item properties as expected (e.g., there is empirical evidence that in word recognition items, as used in the present study, a lower word frequency makes the correct recognition of the word more difficult). Correctly solving the item means that the test‐taker was able to cope with the construct‐related requirements contained in the item and vice versa. Accordingly, the test score reflects how well a test‐taker was able to master the construct‐related requirements presented by the items. Strong test‐takers with high test scores are those who are able to successfully perform the information processes required in items with challenging properties and vice versa. Thus, if item difficulty can be explained by theory‐based item properties, one may expect that effective ability scores reflect how well the test‐takers were able to cope with construct‐related task requirements.</p> <hd id="AN0180987481-7">Effects of item properties on difficulty in timed and untimed testing</hd> <p>Basically, items in a timed condition with fixed stimulus presentation time and a response window imposing moderate time pressure are expected to be more difficult than those in an untimed condition. As predicted by the within‐person SAT function, test‐takers who need to increase effective speed to meet the item‐level time constraints will make more errors and show lower effective ability. The opposite might also happen, but overall, we assume that the time pressure is increased for most test‐takers. Furthermore, we can expect that less efficient test‐takers show a stronger decrease in their effective ability due to the introduced time pressure. They make more errors because they no longer have the time to execute the cognitive processes required by item properties. This is especially the case with more difficult items, whose properties represent a more difficult task. From these considerations also follows the assumption that individual differences in effective ability observed in the timed and untimed conditions are not highly correlated.</p> <p>In the untimed condition, test‐takers are able to adapt the amount of invested time to deal with higher demands in more difficult items. In an untimed test including relatively easy items, this adaptation can be expected to be successful to some extent and fewer errors are made. Overall, this means that in the untimed condition the respective item property has less influence on item difficulty because the difference in response accuracy between high and low demands is reduced. Related to this argument, Thurstone ([<reflink idref="bib50" id="ref27">50</reflink>]) presents a model describing for a fixed person how the probability of obtaining a correct response to an item depends on the time to respond and the difficulty of the item. If difficulty is increased (e.g., in our case by manipulating an item property) while the response time remains the same, the model suggests a decrease in the probability of success. That is what we expect for the timed condition. This decrease would be smaller or even eliminated if the allowed response time is increased which is possible in the untimed condition.</p> <p>For a valid measure of cognitive efficiency, we expect that more demanding item properties (being theoretically‐motivated) are reflected in lower response accuracy. We expect this for the timed condition controlling the stimulus presentation time and response window but to a lesser extent for the untimed condition because in the untimed condition test‐takers may trade response accuracy for speed (for the lexical‐decision task, see e.g., Antos, [<reflink idref="bib3" id="ref28">3</reflink>]). Thus, as a main goal, the present study investigates whether there is an interaction effect on item difficulty between item property and the experimental condition (timed vs. untimed). If the item property is more strongly related to difficulty in the timed condition, this is taken as evidence that the effective ability scores obtained in the timed condition can be interpreted more validly in terms of efficiency than the untimed scores.</p> <hd id="AN0180987481-8">Linguistic item properties</hd> <p>In the present study, we applied the timed testing procedure to tasks of reading efficiency at word and sentence level. For validation, we investigated the effect of linguistic properties of the stimulus, that is, word or sentence, on item difficulty. The word recognition task requires the test‐taker to correctly identify words and non‐words (e.g., "glurp" is a non‐word). Similarly, in the sentence‐level semantic integration task the test‐taker needs to decide whether a presented sentence is true or false (e.g., "The grass is green" is true).</p> <p>For the word recognition task used in the present study, we considered the following linguistic properties: In the case of words, a higher <emph>frequency</emph> of words was expected to facilitate the identification of words as words (e.g., Rayner & Duffy, [<reflink idref="bib43" id="ref29">43</reflink>]). As another property of words, we considered the <emph>number of orthographic neighbors</emph> (i.e., the number of words that can be formed by substituting a single letter in a given word). A facilitating role might be expected for a higher number of orthographic neighbors (e.g., Andrews, [<reflink idref="bib2" id="ref30">2</reflink>]); however, conditions for inhibitory effects have also been discussed in previous research (e.g., Perea & Rosa, [<reflink idref="bib39" id="ref31">39</reflink>]). In the case of non‐words, we had different expectations. A non‐word was generally obtained by distorting an existing word (referred to as base word of the non‐word). Given these manipulations, we assumed that the frequency and the number of orthographic neighbors of the corresponding base word do not show a facilitating effect in non‐words as opposed to words. For non‐words, previous research even suggests an inhibitory effect of the number of orthographic neighbors; that is, a larger neighborhood makes the correct recognition of a non‐word more difficult (Yap et al., [<reflink idref="bib55" id="ref32">55</reflink>]). Non‐words showing high <emph>word similarity</emph>, particularly when being <emph>pseudohomophones</emph> (i.e., non‐words that are pronounced like real words; e.g., Goswami et al., [<reflink idref="bib22" id="ref33">22</reflink>]), were also assumed to make the detection of non‐words as non‐words more difficult. Typically, correctly identifying pseudohomophones is harder than other types of non‐words.</p> <p>For the semantic integration task used in the present study, we expected the following linguistic properties to affect item difficulty: In terms of the semantic complexity of a sentence, a higher <emph>number of propositions</emph> (e.g., Kintsch & Keenan, [<reflink idref="bib30" id="ref34">30</reflink>]) and a higher <emph>number of objects</emph> in correct and incorrect sentences make their evaluation more difficult. High <emph>predictability</emph> (e.g., Rayner et al., [<reflink idref="bib42" id="ref35">42</reflink>]) of correct sentences facilitates identifying correct sentences as correct. Further details on linguistic properties will be provided in the Method section below.</p> <hd id="AN0180987481-9">Research Questions and Hypotheses</hd> <p>The first research question was whether timed testing at item level (i.e., controlling the stimulus presentation time and the time window for responding, with moderate time pressure) provides effective ability scores that can be interpreted validly as cognitive efficiency. As second research question, we investigated whether timed testing provides more valid scores than untimed testing. The construct validation strategy was to investigate the effects of theory‐based item properties on item difficulty. If such effects can be shown, the effective ability scores reflect how well the test‐takers were able to cope with the theoretically relevant task demands.</p> <p>To address these research questions, we validated the construct interpretation of scores from two reading component skills tasks, visual word recognition and sentence‐level semantic integration (Richter et al., [<reflink idref="bib45" id="ref36">45</reflink>]). Overall, we assumed that task performance is influenced by construct‐related item properties that facilitate or impede the successful use of the respective reading component skill. However, test‐takers may trade speed for accuracy in an untimed condition suggesting that untimed and timed testing differ in the extent to which linguistic item properties are related to item difficulty. Therefore, we assumed an interaction effect between item property and condition (timed vs. untimed), such that the effect is stronger in the timed condition than in the untimed condition.</p> <p>As said, efficiency in the untimed condition is represented by both effective speed and effective ability. Accordingly, we also investigated the effect of linguistic item properties on the item time intensity for the untimed condition (Goldhammer, Hahnel, et al., [<reflink idref="bib18" id="ref37">18</reflink>]; Klein Entink et al., [<reflink idref="bib31" id="ref38">31</reflink>]). Given that time intensity is usually positively related to difficulty, we hypothesized that linguistic features that are supposed to make an item more difficult increase the time intensity of items. If such effects can be revealed, the test‐taker's speed reflects how fast the test‐taker dealt with the theory‐based task requirements. Note that in the timed condition item time intensity does not vary anymore across items. As time intensity and difficulty are typically positively correlated, the relation of linguistic item properties to item difficulty is expected to become stronger in the timed condition (as argued above). For example, in the timed condition, an item with properties making it more difficult will show a greater decrease in accuracy than in the untimed condition, in which the time intensity of the item can be adapted to the difficulty of the item.</p> <p>Overall, the aim of our study is to make a conceptual contribution to the measurement of cognitive efficiency that has practical implications: We develop a theoretical argument for how cognitive efficiency can be measured (more validly) by response accuracy alone, while response speed is controlled experimentally. In terms of practical relevance, we argue that the proposed approach is important for valid comparisons between individuals or valid diagnostic decisions.</p> <p>The present study focusing on item‐level differences for validation is based on a reanalysis of data from the study by Goldhammer et al. ([<reflink idref="bib18" id="ref39">18</reflink>]), who investigated person‐level differences in visual word recognition and sentence‐level semantic integration and their relation to reading comprehension.</p> <hd id="AN0180987481-10">Method</hd> <p></p> <hd id="AN0180987481-11">Sample</hd> <p>A total sample of <emph>N</emph> = 888 students participated in both the Programme for International Student Assessment (PISA) 2012 main study in Germany and an add‐on study including among others the two reading component skills tasks. In the sample, 46.17% were female, 49.21% were male, and 4.62% did not specify, aged 15.33 to 16.33 years. The sampling procedure for the main study consisted of two stages in which PISA‐eligible schools were first sampled. For the add‐on study, a subset of 77 schools were included with up to 14 students sampled from the group of PISA students. At school, the test instruments were administered by a proctor in a group setting using standardized bring‐in laptops. For the add‐on study, a test design with 16 booklets was implemented with random assignment of test‐takers to booklets. In 12 booklets, the test‐takers also completed the visual word recognition task and the semantic integration task in both an untimed and a timed condition. Thus, the effective sample size was <emph>N</emph> = 637.</p> <p>As part of PISA 2012, the present study followed the rules and procedures of the PISA 2012 main study in Germany. Prior to the implementation of the PISA 2012 data collection, a content and data protection review as well as approval by the Ministries of Education and Cultural Affairs of all federal states took place. The participation of drawn PISA schools and students was mandatory (although the completion of the student questionnaire was partly voluntary). Further details can be found in the national report on PISA 2012 (Prenzel et al., [<reflink idref="bib40" id="ref40">40</reflink>]).</p> <hd id="AN0180987481-12">Instruments</hd> <p> <bold>Visual word recognition</bold>. The lexical‐decision task (Richter et al., [<reflink idref="bib45" id="ref41">45</reflink>]) required the participants to distinguish between words and non‐words (e.g., "Mele" is a non‐word) by pressing the corresponding response button on the keyboard. All the words were nouns, with their length varying from three to 10 letters and one to three syllables. Words and non‐words were matched according to length (number of letters and syllables), word frequency and number of orthographic neighbors (in the case of non‐words of the base words they were derived from). More detailed information on the linguistic properties of words and non‐words can be found below (see section Linguistic Item Properties and Table 1).</p> <p>1 Table Descriptive Statistics of Item Properties by Experimental Condition</p> <p> <ephtml> <table><thead><tr><th>Task</th><th align="center">Property</th><th align="center">Condition</th><th align="center">Trials</th><th align="center"><italic>N</italic></th><th align="center"><italic>M</italic></th><th align="center"><italic>SD</italic></th><th align="center"><italic>Min</italic></th><th align="center"><italic>Max</italic></th></tr></thead><tbody><tr><td>Word recognition</td><td>Word frequency</td><td>Untimed</td><td>Words</td><td>16</td><td>4.06</td><td>1.06</td><td>3</td><td>6</td></tr><tr><td /><td /><td>Timed</td><td>Words</td><td>16</td><td>3.94</td><td>1.00</td><td>2</td><td>5</td></tr><tr><td /><td>Number of orthographic neighbors</td><td>Untimed</td><td>Words</td><td>16</td><td>15.56</td><td>13.02</td><td>1</td><td>49</td></tr><tr><td /><td /><td>Timed</td><td>Words</td><td>16</td><td>13.88</td><td>14.04</td><td>1</td><td>44</td></tr><tr><td /><td>Number of orthographic neighbors (base word)</td><td>Untimed</td><td>Non‐words</td><td>16</td><td>14.38</td><td>18.06</td><td>0</td><td>58</td></tr><tr><td /><td /><td>Timed</td><td>Non‐words</td><td>16</td><td>14.38</td><td>16.23</td><td>1</td><td>53</td></tr><tr><td /><td>Word similar</td><td>Untimed</td><td>Non‐words</td><td>16 (12)<ext-link /><sup>a</sup></td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td></tr><tr><td /><td /><td>Timed</td><td>Non‐words</td><td>16 (12)<ext-link /><sup>a</sup></td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td></tr><tr><td /><td>Pseudohomophone</td><td>Untimed</td><td>Non‐words</td><td>16 (4)<ext-link /><sup>b</sup></td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td></tr><tr><td /><td /><td>Timed</td><td>Non‐words</td><td>16 (5)<ext-link /><sup>b</sup></td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td></tr><tr><td>Semantic integration</td><td>Number of propositions</td><td>Untimed</td><td>(In)correct sentences</td><td>24</td><td>2.00</td><td>.83</td><td>1</td><td>3</td></tr><tr><td /><td /><td>Timed</td><td>(In)correct sentences</td><td>24</td><td>2.00</td><td>.83</td><td>1</td><td>3</td></tr><tr><td /><td>Number of objects</td><td>Untimed</td><td>(In)correct sentences</td><td>24</td><td>1.17</td><td>.82</td><td>0</td><td>3</td></tr><tr><td /><td /><td>Timed</td><td>(In)correct sentences</td><td>24</td><td>.88</td><td>.90</td><td>0</td><td>3</td></tr><tr><td /><td>Unpredictable</td><td>Untimed</td><td>Correct sentences</td><td>12 (6)<ext-link /><sup>c</sup></td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td></tr><tr><td /><td /><td>Timed</td><td>Correct sentences</td><td>12 (6)<ext-link /><sup>c</sup></td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td><td align="center">n/a</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note: N</emph> = number of trials.</p> <ulist> <item>2 a Number of word‐like non‐words.</item> <item>3 b Number of pseudohomophones.</item> <item>4 c Number of unpredictable sentences.</item> </ulist> <p> <bold>Sentence‐level semantic integration</bold>. Similarly, the sentence verification task (Richter et al., [<reflink idref="bib45" id="ref42">45</reflink>]) required the participants to distinguish between true sentences and false sentences (e.g., "Snails are fast" is a false sentence) by pressing the corresponding response button. Sentences were constructed under systematic variation of linguistic features; more detailed information about the design principles is presented below (see section Linguistic Item Properties and Table 1).</p> <p>The language of the tests employed in the study was German.</p> <hd id="AN0180987481-13">Experimental design</hd> <p>The within‐subject design included an untimed condition followed by a timed condition for visual word recognition and sentence‐level semantic integration. To avoid carryover effects, timed and untimed conditions used different item materials (which, however, were parallel in terms of stimulus properties used in the item design; see section Linguistic Item Properties and Table 1).</p> <p>Each trial began with a 500‐ms presentation of a centered fixation cross. When it disappeared, the stimulus was presented. For the <emph>untimed condition</emph>, participants decided individually when to respond. After responding, a blank screen appeared for 500 ms (interstimulus interval). In the <emph>timed condition</emph> (see Figure 1), the stimulus disappeared once the predefined presentation time elapsed, indicated by the response signal (see Reed, [<reflink idref="bib44" id="ref43">44</reflink>]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01sep24/jedm12393-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12393-fig-0001.jpg" title="1 Trial of the word recognition task in the timed condition with a stimulus presentation time of 741 ms (from Goldhammer, Kroehne, et al., [18])." /> </p> <p></p> <p>For visual word recognition, the stimulus presentation time was 741 ms, and for sentence‐level semantic integration, it was 1,500 ms. For word recognition, the stimulus presentation time was derived from a response time distribution obtained in a previous study without time constraints (i.e., 60th percentile) and further trialed and adjusted to the target population in a coglab study (for further details, see Goldhammer & Kroehne, [<reflink idref="bib19" id="ref44">19</reflink>]). For semantic integration, a moderate time limit was determined in a coglab study. Our goal was to choose a moderate time limit across all items and all test‐takers. How strict the selected time limit for a certain test‐taker actually is depends on the time intensity of the respective item (i.e., the average time spent on an item in an untimed condition). That is, for more time‐intensive items the same time limit is stricter and vice versa. Thereby, the range of item difficulty could be manipulated. Items that are more time‐intensive relative to the time limit represent more difficult items that can only be completed correct and in time by more efficient test‐takers, while items that are less time‐intensive relative to the time limit can also be successfully completed by less efficient test‐takers.</p> <p>The participants did not see a timer but were familiar with the amount of available time through timed practice trials. The participants had a window of 300 ms to give a response as soon as they heard the signal (a beep) via earphones. Feedback on timing was given on all timed trials. If the response was given on time, a happy face was shown for 800 ms; if the answer was too early (i.e., before the response window) or too late (i.e., after the response window), an unhappy face with the message "too early" or "too late" was shown for 1,200 ms. After presenting the feedback, a blank screen appeared for 500 ms.</p> <p>To manipulate the test‐taker's response speed in the timed condition, the experimental procedure controlled the stimulus presentation time and the time window for responding. However, whether the given time was fully used for showing maximum performance is still under control of the test‐taker, meaning the experimental procedure allows only for partial control (see also Discussion section).</p> <hd id="AN0180987481-15">Procedure</hd> <p>First, participants completed the visual word recognition test in the untimed condition, which consisted of 32 trials (16 words and 16 non‐words). Then, they completed the visual word recognition test in a timed condition (again, 16 words and 16 non‐words). After that, the participants took the sentence‐level semantic integration test in the untimed condition, including 24 trials (12 true sentences, 12 false sentences). Finally, they completed the sentence‐level semantic integration test in the timed condition (again, 12 true sentences, 12 false sentences). In both reading component skill tests, stimuli appeared in a random order, which was the same for all participants.</p> <p>Each condition began with a block of practice trials to familiarize participants with the particular stimulus presentation and the required response mode. In the timed condition, participants learned to respond in a timely manner in 12 practice trials. The goal was to adjust the test‐taker's decision criterion of how quickly to give a correct response to a stimulus so that he or she would respond in time.</p> <p>In the untimed condition, the participants were instructed to work as quickly as possible and avoid making errors. In the timed condition, participants were required to press the response button of their choice once the response signal was presented and to avoid making an error.</p> <p>In the timed conditions, responses given before the onset of the stimulus on the screen could not be responses to the stimulus and were treated as not attempted items with a missing response and a missing response time (2.57% of the expected total across items and test‐takers); there were no responses before the onset of the stimulus in the untimed conditions. In the timed condition, omissions could occur (3.76% of the expected total across items and test‐takers). For these items, the response time was treated as not available (i.e., missing) and the response as incorrect since cognitive efficiency is indicated by responding in time (and correctly). There were no omitted responses in the untimed conditions.</p> <hd id="AN0180987481-16">Linguistic Item Properties</hd> <p>Table 1 gives an overview of the item properties and their descriptive statistics by experimental condition. Overall, the distributions were comparable between the timed and untimed testing conditions.</p> <p> <bold>Visual word recognition</bold>. <emph>Word frequency</emph> reflects how often a word is used in contemporary language corpora. The information was obtained from the digital dictionary of the German language (Digitales Wörterbuch der deutschen Sprache, DWDS),[<reflink idref="bib1" id="ref45">1</reflink>] classifying a word's frequency on a seven‐level logarithmic scale ranging from 0 (seldom) to 6 (frequent).</p> <p>The <emph>number of orthographic neighbors</emph> of a word represents the number of words obtained by changing a single letter in the target word while holding the other letters constant (e.g., cat: oat, hat, car, eat, etc.) (Coltheart et al., [<reflink idref="bib8" id="ref46">8</reflink>]). The number of orthographic neighbors of a non‐word refers to the respective base word that was manipulated to create the non‐word.</p> <p>The property <emph>word similarity</emph> (yes vs. no) indicates whether a non‐word is phonologically and/or visually similar to a word.</p> <p>The item property <emph>pseudohomophone</emph> (yes vs. no) indicates that the non‐word was derived by changing the orthography while retaining the phonology (e.g., word: blue, pseudohomophone: bloo).</p> <p> <bold>Sentence‐level semantic integration</bold>. The <emph>number of propositions</emph> was used as a measure of a sentence's semantic complexity, that is, the number of elementary units that can be used to describe the semantic structure of a sentence. A proposition represents the smallest unit of knowledge that contains an independent statement and can be described as a structure consisting of a relation and a number of ordered arguments (Kintsch, [<reflink idref="bib29" id="ref47">29</reflink>]).</p> <p>The <emph>number of objects</emph> in a sentence was used as another measure of a sentence's complexity. An object is a sentence element required by a verb or preposition as a complement. To obtain this measure, the objects were counted per sentence. Although more propositions may suggest more objects, this is not necessarily true as not only objects do instantiate propositions. Accordingly, both item properties are moderately correlated, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>r</mi><mspace width="0.28em" /><mo>=</mo><mspace width="0.28em" /><mo>.</mo><mn>69</mn></mrow><annotation encoding="application/x-tex">$r\; = \;.69$</annotation></semantics></math> </ephtml> . As an example, the German sentence "Giraffen haben lange Hälse" (Giraffes have long necks) includes two propositions (have[agent: giraffes, object: necks], long[necks]) and one object (necks).</p> <p>The <emph>predictability</emph> (yes vs. no) applies to true sentences and refers to whether words in a sentence can be predicted based on the preceding sentence context. The preceding sentence context may be a term to be defined and the following words form a definition. Predictability reflects how reliably the latter can be inferred from the former (e.g., in the sentence "Most birds can fly" showing predictability "fly" can be clearly inferred from "birds"). Predictability of one word by another depends on the cooccurrence of these words in language use. For obtaining this measure, the cooccurrence of words was assessed by means of a Web search (for details see Richter & Naumann, [<reflink idref="bib46" id="ref48">46</reflink>]).</p> <hd id="AN0180987481-17">Statistical Analyses</hd> <p>To address the research questions, generalized linear mixed models (GLMMs) were estimated (De Boeck et al., [<reflink idref="bib12" id="ref49">12</reflink>]). The descriptive baseline GLMM with the logit of a correct item response as explained variable was specified as follows: 1.1 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>logit</mi><mfenced separators="" open="(" close=")"><mrow><mi>P</mi><mfenced separators="" open="(" close=")"><mrow><msub><mi>Y</mi><mrow><mi>p</mi><mi>i</mi></mrow></msub><mo>=</mo><mn>1</mn></mrow></mfenced></mrow></mfenced><mo linebreak="badbreak">=</mo><msub><mi>β</mi><mn>0</mn></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><mo>.</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{\mathrm{logit}}\left({P\left({{Y_{pi}} = 1} \right)} \right) = {\beta _0} + {b_{0p}} + {b_{0s}} + {b_{0i}}.\end{equation}$$</annotation></semantics></math> </ephtml> The model includes random intercepts of persons, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><annotation encoding="application/x-tex">${b_{0p}}$</annotation></semantics></math> </ephtml> , schools, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">${b_{0s}}$</annotation></semantics></math> </ephtml> , and items, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>b</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><annotation encoding="application/x-tex">${b_{0i}}$</annotation></semantics></math> </ephtml> , as well a fixed intercept, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>β</mi><mn>0</mn></msub><annotation encoding="application/x-tex">${\beta _0}$</annotation></semantics></math> </ephtml> . Thus, the model represents a cross‐classified multilevel structure with observation <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>Y</mi><mrow><mi>p</mi><mi>i</mi></mrow></msub><annotation encoding="application/x-tex">${Y_{pi}}$</annotation></semantics></math> </ephtml> (level 1) nested into both person and item (level 2) with person nested into school (level 3). The variance components of the random intercepts were used to determine, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ICC</mi><mo>(</mo><mi>k</mi><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}(k)$</annotation></semantics></math> </ephtml> as a measure of reliability of the random intercepts for persons, and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ICC</mi><mo>(</mo><mrow><mi>s</mi><mi>c</mi><mi>h</mi><mi>o</mi><mi>o</mi><mi>l</mi></mrow><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}({school})$</annotation></semantics></math> </ephtml> . The <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ICC</mi><mo>(</mo><mi>k</mi><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}(k)$</annotation></semantics></math> </ephtml> coefficient as the reliability of the sum of all items was computed (see De Boeck, [<reflink idref="bib11" id="ref50">11</reflink>]; Semmes et al., [<reflink idref="bib48" id="ref51">48</reflink>]) as <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>I</mi><mi>C</mi><mi>C</mi><mrow><mo>(</mo><mi>k</mi><mo>)</mo></mrow><mo>=</mo><mi>Var</mi><mrow><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo>)</mo></mrow><mo>/</mo><mrow><mo>(</mo><mrow><mi>Var</mi><mrow><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo>)</mo></mrow><mo>+</mo><mi>Var</mi><mrow><mo>(</mo><mi>ε</mi><mo>)</mo></mrow><mo>/</mo><mi>n</mi></mrow><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">$ICC(k) = {\mathrm{Var}}({{b_{0p}}})/({{\mathrm{Var}}({{b_{0p}}}) + {\mathrm{Var}}(\varepsilon)/n})$</annotation></semantics></math> </ephtml> , where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>Var</mi><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{Var}}({{b_{0p}}})$</annotation></semantics></math> </ephtml> is the variance of the person intercept, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>Var</mi><mo>(</mo><mi>ε</mi><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{Var}}(\varepsilon)$</annotation></semantics></math> </ephtml> is the error variance which is for the logit link function <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>Var</mi><mrow><mo>(</mo><mi>ε</mi><mo>)</mo></mrow><mo>=</mo><mfrac><msup><mi>π</mi><mn>2</mn></msup><mn>3</mn></mfrac><mo>=</mo><mn>3.29</mn></mrow><annotation encoding="application/x-tex">${\mathrm{Var}}(\varepsilon) = \frac{{{\pi ^2}}}{3} = 3.29$</annotation></semantics></math> </ephtml> , and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>n</mi><annotation encoding="application/x-tex">$n$</annotation></semantics></math> </ephtml> is the number of items. The proportion of variance due to school differences was computed as follows: <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ICC</mi><mrow><mo>(</mo><mrow><mi>s</mi><mi>c</mi><mi>h</mi><mi>o</mi><mi>o</mi><mi>l</mi></mrow><mo>)</mo></mrow><mo>=</mo><mi>Var</mi><mrow><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo>)</mo></mrow><mo>/</mo><mrow><mo>(</mo><mrow><mi>Var</mi><mrow><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo>)</mo></mrow><mo>+</mo><mi>Var</mi><mrow><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo>)</mo></mrow></mrow><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}({school}) = {\mathrm{Var}}({{b_{0s}}})/({{\mathrm{Var}}({{b_{0s}}}) + {\mathrm{Var}}({{b_{0p}}})})$</annotation></semantics></math> </ephtml> . Model 1.1 was tested for the accuracy data obtained from the untimed condition. The same was done for the accuracy data from the timed condition (model 1.2 which is technically identical to model 1.1).</p> <p>The baseline model was extended to an explanatory item‐response model as follows: 2.1 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>logit</mi><mfenced separators="" open="(" close=")"><mrow><mi>P</mi><mfenced separators="" open="(" close=")"><mrow><msub><mi>Y</mi><mrow><mi>p</mi><mi>i</mi></mrow></msub><mo>=</mo><mn>1</mn></mrow></mfenced></mrow></mfenced><mo linebreak="badbreak">=</mo><msub><mi>β</mi><mn>0</mn></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>k</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>u</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><mo>.</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{\mathrm{logit}}\left({P\left({{Y_{pi}} = 1} \right)} \right) = {\beta _0} + {b_{0p}} + {b_{0s}} + {\beta _{1k}}{X_{i,k}} + {u_{0i}}.\end{equation}$$</annotation></semantics></math> </ephtml> The fixed effect, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><annotation encoding="application/x-tex">${\beta _{1k}}$</annotation></semantics></math> </ephtml> , represents the effect of the item‐level covariate, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>k</mi></mrow></msub><annotation encoding="application/x-tex">${X_{i,k}}$</annotation></semantics></math> </ephtml> , indicating item property <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>k</mi><annotation encoding="application/x-tex">$k$</annotation></semantics></math> </ephtml> . The random item intercept <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>u</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><annotation encoding="application/x-tex">${u_{0i}}$</annotation></semantics></math> </ephtml> in (2.1) replaces <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>b</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><annotation encoding="application/x-tex">${b_{0i}}$</annotation></semantics></math> </ephtml> in (1.1) and represents the residual variance of item easiness that is not explained by the item‐level covariate (see linear logistic test models with error term, LLTM+e, Janssen et al., [<reflink idref="bib26" id="ref52">26</reflink>]). The error term relaxes the assumption of the original LLTM (Fischer, [<reflink idref="bib16" id="ref53">16</reflink>]) that all item variance could be explained by the linear combination of item‐level covariates.</p> <p>For each item property, the explanatory model (2.1) was tested for the accuracy data obtained from the untimed condition serving as a benchmark test procedure. The same was done for the accuracy data from the timed condition to provide validity evidence for the timed test procedure (model 2.2, which is technically identical to model 2.1). Furthermore, to directly test whether the effect of the item‐level covariate on item difficulty is changed by timed testing, data from the timed and untimed condition were modeled jointly. To this end, the GLMM (2.1) and (2.2), respectively, was extended by adding the experimental condition (timed vs. untimed) and the interaction of the respective item‐level covariate and condition. Distinct but correlated person random intercepts were assumed for the two experimental conditions, turning the model into a two‐dimensional (1PL) model: 2.3 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>logit</mi><mfenced separators="" open="(" close=")"><mrow><mi>P</mi><mfenced separators="" open="(" close=")"><mrow><msub><mi>Y</mi><mrow><mi>p</mi><mi>i</mi><mi>c</mi></mrow></msub><mo>=</mo><mn>1</mn></mrow></mfenced></mrow></mfenced><mo linebreak="badbreak">=</mo><msub><mi>β</mi><mn>0</mn></msub><mo linebreak="goodbreak">+</mo><munderover><mo>∑</mo><mrow><mi>c</mi><mspace width="0.28em" /><mo>=</mo><mspace width="0.28em" /><mn>1</mn></mrow><mn>2</mn></munderover><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi><mi>c</mi></mrow></msub><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>c</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>k</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>β</mi><mn>2</mn></msub><msub><mi>X</mi><mi>c</mi></msub><mo linebreak="goodbreak">+</mo><msub><mi>β</mi><mn>3</mn></msub><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>k</mi></mrow></msub><msub><mi>X</mi><mi>c</mi></msub><mo linebreak="goodbreak">+</mo><msub><mi>u</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><mo>,</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{\mathrm{logit}}\left({P\left({{Y_{pic}} = 1} \right)} \right) = {\beta _0} + \mathop \sum \limits_{c{\mathrm{\;}} = {\mathrm{\;}}1}^2 {b_{0pc}}{X_{i,c}} + {b_{0s}} + {\beta _{1k}}{X_{i,k}} + {\beta _2}{X_c} + {\beta _3}{X_{i,k}}{X_c} + {u_{0i}},\end{equation}$$</annotation></semantics></math> </ephtml> where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0024" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>X</mi><mi>c</mi></msub><annotation encoding="application/x-tex">${X_c}$</annotation></semantics></math> </ephtml> denotes the experimental condition with timed as reference condition (i.e., timed: <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0025" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>X</mi><mi>c</mi></msub><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">${X_c} = 0$</annotation></semantics></math> </ephtml> , untimed: <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0026" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>X</mi><mi>c</mi></msub><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">${X_c} = 1$</annotation></semantics></math> </ephtml> ), <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0027" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>β</mi><mn>2</mn></msub><annotation encoding="application/x-tex">${\beta _2}$</annotation></semantics></math> </ephtml> the main effect of the experimental condition, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>β</mi><mn>3</mn></msub><annotation encoding="application/x-tex">${\beta _3}$</annotation></semantics></math> </ephtml> the interaction effect of experimental condition and the item‐level covariate, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi><mi>c</mi></mrow></msub><annotation encoding="application/x-tex">${b_{0pc}}$</annotation></semantics></math> </ephtml> the random person intercept by condition, and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0030" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>c</mi></mrow></msub><annotation encoding="application/x-tex">${X_{i,c}}$</annotation></semantics></math> </ephtml> indicates whether item <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0031" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>i</mi><annotation encoding="application/x-tex">$i$</annotation></semantics></math> </ephtml> contributes to the measurement of the condition‐specific dimension <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0032" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>c</mi><annotation encoding="application/x-tex">$c$</annotation></semantics></math> </ephtml> ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>c</mi></mrow></msub><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">${X_{i,c}} = 1$</annotation></semantics></math> </ephtml> ) or not ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>c</mi></mrow></msub><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">${X_{i,c}} = 0$</annotation></semantics></math> </ephtml> ).</p> <p>For the log‐transformed response time, the following explanatory linear mixed model (LMM) model was estimated: 3 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>log</mi><mfenced separators="" open="(" close=")"><mrow><mi>R</mi><msub><mi>T</mi><mrow><mi>p</mi><mi>i</mi></mrow></msub></mrow></mfenced><mo linebreak="badbreak">=</mo><msub><mi>β</mi><mn>0</mn></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><msub><mi>X</mi><mrow><mi>i</mi><mo>,</mo><mi>k</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>u</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><mo linebreak="goodbreak">+</mo><msub><mi>u</mi><mrow><mi>p</mi><mi>i</mi></mrow></msub><mo>.</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{\mathrm{log}}\left({R{T_{pi}}} \right) = {\beta _0} + {b_{0p}} + {b_{0s}} + {\beta _{1k}}{X_{i,k}} + {u_{0i}} + {u_{pi}}.\end{equation}$$</annotation></semantics></math> </ephtml> The lognormal RT model also includes a response time residual, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>u</mi><mrow><mi>p</mi><mi>i</mi></mrow></msub><annotation encoding="application/x-tex">${u_{pi}}$</annotation></semantics></math> </ephtml> . This is not needed for the logit model because the logit translates in a probability so that the randomness is included in the probability being the parameter for the Bernoulli distribution of responses. All random effects and the residual of response time are modeled assuming a normal distribution. The response time model (<reflink idref="bib3" id="ref54">3</reflink>) was applied to the response time data from the untimed condition, but not to the response time data from the timed condition where response time was turned into an experimentally controlled variable.</p> <p>To investigate the unique effects of item‐level covariates on response accuracy, we included the linear combination of all the relevant item covariates simultaneously into models 2.1 and 2.2, resulting in models 4.1 and 4.2. The presented (G)LMMs were tested for both visual word recognition and sentence‐level semantic integration.</p> <p>To evaluate the adequacy of the fitted (G)LMMs, we did a residual analysis for each model by graphical checking and statistical testing following the recommendations by Bolker et al. ([<reflink idref="bib6" id="ref55">6</reflink>]). Specifically, we examined the scaled residuals (Hartig, [<reflink idref="bib24" id="ref56">24</reflink>]) which range from 0 to 1 and were formed for each observation. The residual represents the value of the cumulative density function (for simulated observation data based on the fitted model) at the value of the observed data point. Regarding the expected uniform distribution of scaled residuals, we visually inspected the QQ plot for deviations and did a Kolmogorov‐Smirnov (KS) test. Additionally, we tested for over/underdispersion, that is, whether the residual variance is larger or smaller than expected given the fitted model. We also tested for outliers, that is, whether the number of observations outside the range of simulated data are larger or smaller than expected. Finally, we graphically inspected the residuals against the predicted value to check whether the assumed linearity of the predictor's effect is justifiable. The relationship of residuals and predicted value is represented as smooth spline around the theoretical expectation (i.e., a straight line at .5). Alternatively, if the predictor is a factor, or if the number of observations on the predictor is small, a boxplot is shown by factor level. The distribution for each factor level should be uniform. Due to the high number of observations or residuals (number of test‐takers times number of items) making even small deviations significant, we weighted the graphical checking more strongly than statistical checking.</p> <p>Most of the linguistic properties are only applicable to a subset of trials (e.g., word frequency applies to trials with words but not to trials with non‐words). In this case only this subset of trials was included into the analysis (for an overview see Table 1, see also Discussion section).</p> <p>All analyses were conducted using the R environment (R Core Team, [<reflink idref="bib41" id="ref57">41</reflink>]) with the lme4 (Bates et al., [<reflink idref="bib5" id="ref58">5</reflink>]) and lmerTest (Kuznetsova et al., [<reflink idref="bib32" id="ref59">32</reflink>]) packages. For residual analysis we used the R package DHARMa (Hartig, [<reflink idref="bib24" id="ref60">24</reflink>]) with the recommended (default) settings. Marginal <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0037" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msup><mi>R</mi><mn>2</mn></msup><annotation encoding="application/x-tex">${R^2}$</annotation></semantics></math> </ephtml> indicating the proportion of variance explained by fixed effects in GLMMs (Nakagawa & Schielzeth, [<reflink idref="bib38" id="ref61">38</reflink>]) were calculated with the R package partR2 (Stoffel et al., [<reflink idref="bib49" id="ref62">49</reflink>]). Note that the used binomial models give relatively small <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0038" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><annotation encoding="application/x-tex">$R_{{\mathrm{marginal}}}^2$</annotation></semantics></math> </ephtml> values because much of the variance originates from the random effects and the residual variance (additive dispersion and distribution‐specific variance; Nakagawa & Schielzeth, [<reflink idref="bib38" id="ref63">38</reflink>]).</p> <hd id="AN0180987481-18">Results</hd> <p></p> <hd id="AN0180987481-19">Model Fit of Tested GLMMs</hd> <p>Table 2 shows the summary of the scaled residual analyses to assess the adequacy of the fitted models. For both tasks, word recognition and semantic integration, the models for accuracy data (models 1.1‐2.3, 4.1, 4.2) did not show signs of misfit. The KS tests turned out to be significant for some models (2.2, 2.3, 4.2), but the QQ plot clearly did not point to a relevant deviation from the expected uniform distribution of the scaled residuals. For both tasks, word recognition and semantic integration, the models for response time data (model 3) showed some misfit. Specifically, there was a slight deviation from the expected uniform distribution in all models suggesting underdispersion. That is, there were too many residuals around .5, and not as many residuals in the tail of the distribution as one would expect under the fitted model. Even though the dispersion tests were not significant, the results for the response time models should be interpreted with more caution.</p> <p>2 Table Summary of Scaled Residual Analyses to Assess Model Fit</p> <p> <ephtml> <table><thead><tr valign="bottom"><th>Task</th><th align="center">Model</th><th>Number of Models</th><th>QQ Plot Deviation</th><th>KS Test Sig.</th><th>Dispersion Test Sig.</th><th>Outlier Test Sig.</th><th>Residual vs. Predicted Deviation</th></tr></thead><tbody><tr><td>Word recognition</td><td>1.1</td><td>1</td><td>0</td><td>0</td><td>0</td><td>0</td><td>Na</td></tr><tr><td /><td>1.2</td><td>1</td><td>0</td><td>0</td><td>0</td><td>0</td><td>Na</td></tr><tr><td /><td>2.1</td><td>5</td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td>2.2</td><td>5</td><td>0</td><td>5</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td>2.3</td><td>5</td><td>0</td><td>5</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td align="center">3</td><td>5</td><td>5</td><td>5</td><td>0</td><td>5</td><td>5</td></tr><tr><td /><td>4.1</td><td>2</td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td>4.2</td><td>2</td><td>0</td><td>2</td><td>0</td><td>0</td><td>0</td></tr><tr><td>Semantic integration</td><td>1.1</td><td>1</td><td>0</td><td>0</td><td>0</td><td>0</td><td>Na</td></tr><tr><td /><td>1.2</td><td>1</td><td>0</td><td>1</td><td>0</td><td>0</td><td>Na</td></tr><tr><td /><td>2.1</td><td>3</td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td>2.2</td><td>3</td><td>0</td><td>3</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td>2.3</td><td>3</td><td>0</td><td>3</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td align="center">3</td><td>3</td><td>3</td><td>3</td><td>0</td><td>3</td><td>3</td></tr><tr><td /><td>4.1</td><td>1</td><td>0</td><td>0</td><td>0</td><td>0</td><td>0</td></tr><tr><td /><td>4.2</td><td>1</td><td>0</td><td>1</td><td>0</td><td>0</td><td>0</td></tr></tbody></table> </ephtml> </p> <ulist> <item>5 <emph>Note</emph>: The numbers indicate how many models showed signs of misfit for the respective criterion.</item> <item>6 QQ = quantile‐quantile; KS = Kolmogorov‐Smirnov; sig. = significant; Na = not applicable because model does not include covariates.</item> </ulist> <hd id="AN0180987481-20">Descriptive Results and Baseline Modeling</hd> <p>Table 3 provides a descriptive summary of the overall performance for each task and condition. The average proportion correct and average response time was computed based on all available observations (i.e., across all available person‐item combinations) by task and condition. As expected, the timed condition was more difficult in both the word recognition task and the sentence verification task, as evidenced by a 14% decrease in the proportion of correct responses in each case. In the timed conditions the proportion of correct responses was still above 50%, which is the rate for random rapid guessing. Thus, the speed level was moderate. Actually, as shown by the average response time, the experimentally imposed speed was overall comparable to the self‐selected speed. However, comparison of the standard deviation of response time between the timed and untimed conditions makes it clear that introducing item‐level time constraints was effective in reducing heterogeneity in response speed in the timed conditions.</p> <p>3 Table Performance across Items and Persons by Task and Condition</p> <p> <ephtml> <table><thead><tr><th /><th /><th align="center">Proportion Correct</th><th align="center">Response Time (ms)</th></tr><tr><th>Task</th><th align="center">Condition</th><th align="center"><italic>M</italic></th><th align="center"><italic>SD</italic></th><th align="center"><italic>Min</italic></th><th align="center"><italic>Max</italic></th><th align="center"><italic>M</italic></th><th align="center"><italic>SD</italic></th><th align="center"><italic>Min</italic></th><th align="center"><italic>Max</italic></th></tr></thead><tbody><tr><td>Word recognition</td><td>Untimed</td><td>.90</td><td>.06</td><td>.71</td><td>.95</td><td>1,397</td><td>615</td><td>503</td><td>26,170</td></tr><tr><td /><td>Timed</td><td>.76</td><td>.08</td><td>.51</td><td>.86</td><td>1,461</td><td>216</td><td>511</td><td>3,242</td></tr><tr><td>Semantic integration</td><td>Untimed</td><td>.87</td><td>.06</td><td>.72</td><td>.93</td><td>2,361</td><td>1,329</td><td>503</td><td>45,300</td></tr><tr><td /><td>Timed</td><td>.73</td><td>.08</td><td>.57</td><td>.90</td><td>2,214</td><td>331</td><td>501</td><td>3,726</td></tr></tbody></table> </ephtml> </p> <p>In the timed conditions the empirical range of response times is broader than the response window (word recognition: 1,241‐1,541 ms, semantic integration: 2,000‐2,300 ms). This is because we did not exclude the trials where the response to the stimulus was too early (i.e., before the response window) or too late (i.e., after the response window) in all analyses. For word recognition, across all items, 66.11% of the observed responses were on time, 20.70% were too early, and 13.20% were too late. For sentence‐level semantic integration across all items, 58.21% of the observations were on time, 17.35% were too early, and 24.44% were too late.</p> <p>Table 4 shows the model parameters and derived ICC coefficients from baseline modeling; see Equation 1.1, by task and condition. The fixed effects <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>β</mi><mn>0</mn></msub><annotation encoding="application/x-tex">${\beta _0}$</annotation></semantics></math> </ephtml> again reflect that the timed condition is more difficult than the untimed condition. The item and person variance components are comparable between conditions with the exception that for both tasks the variance of the random school intercept is much larger for the timed than for the untimed condition. Accordingly, the <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ICC</mi><mo>(</mo><mrow><mi>s</mi><mi>c</mi><mi>h</mi><mi>o</mi><mi>o</mi><mi>l</mi></mrow><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}({school})$</annotation></semantics></math> </ephtml> is higher for the timed than for the untimed condition. This suggests that differences at school‐level (including different school tracks) have a stronger relation to timed performance of students than the untimed performance. Finally, the reliability measure <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ICC</mi><mo>(</mo><mi>k</mi><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}(k)$</annotation></semantics></math> </ephtml> proved to be comparably high for the timed and untimed conditions.</p> <p>4 Table Model Parameters and Derived ICC Coefficients from Baseline Modeling by Task and Condition</p> <p> <ephtml> <table><thead><tr><th>Model</th><th align="center">Task</th><th align="center">Condition</th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>β</mi><mn>0</mn></msub><annotation encoding="application/x-tex">${\beta _0}$</annotation></semantics></math></p></th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mrow><mi>Var</mi><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>i</mi></mrow></msub><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{Var}}({{b_{0i}}})$</annotation></semantics></math></p></th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mrow><mi>Var</mi><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>p</mi></mrow></msub><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{Var}}({{b_{0p}}})$</annotation></semantics></math></p></th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0045" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mrow><mi>Var</mi><mo>(</mo><msub><mi>b</mi><mrow><mn>0</mn><mi>s</mi></mrow></msub><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{Var}}({{b_{0s}}})$</annotation></semantics></math></p></th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0046" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mrow><mi>ICC</mi><mo>(</mo><mi>k</mi><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}(k)$</annotation></semantics></math></p></th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mrow><mi>ICC</mi><mo>(</mo><mrow><mi>s</mi><mi>c</mi><mi>h</mi><mi>o</mi><mi>o</mi><mi>l</mi></mrow><mo>)</mo></mrow><annotation encoding="application/x-tex">${\mathrm{ICC}}({school})$</annotation></semantics></math></p></th></tr></thead><tbody><tr><td>1.1</td><td>Word recognition</td><td>Untimed</td><td>2.75</td><td>.37</td><td>1.17</td><td>.21</td><td>.92</td><td>.15</td></tr><tr><td>1.2</td><td /><td>Timed</td><td>1.54</td><td>.30</td><td>1.17</td><td>.65</td><td>.92</td><td>.36</td></tr><tr><td>1.1</td><td>Semantic integration</td><td>Untimed</td><td>2.35</td><td>.28</td><td>1.04</td><td>.18</td><td>.88</td><td>.15</td></tr><tr><td>1.2</td><td /><td>Timed</td><td>1.21</td><td>.27</td><td>.91</td><td>.52</td><td>.87</td><td>.36</td></tr></tbody></table> </ephtml> </p> <hd id="AN0180987481-21">Visual Word Recognition</hd> <p>The Tables 5–8 summarize the effects of different item properties on individuals' performance in the visual word recognition and sentence‐level semantic integration tasks. Each table informs about the included trials (words or non‐words; incorrect and/or correct sentences) and conditions (untimed and/or timed) on which the analyses were based.</p> <p>5 Table Effects of Linguistic Item Properties in the Word Recognition Task</p> <p> <ephtml> <table><thead><tr valign="bottom"><th>Model</th><th align="center">Dependent Variable</th><th align="center">Item‐Level Covariate</th><th align="center">Included Trials</th><th align="center">Condition</th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0048" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><annotation encoding="application/x-tex">${\beta _{1k}}$</annotation></semantics></math></p></th><th align="center">(<italic>SE</italic>)</th></tr></thead><tbody><tr><td>2.1</td><td>Logit(P+)</td><td>Word frequency</td><td>Words</td><td>Untimed</td><td>.26<ext-link /><sup>*</sup></td><td>(.13)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Word frequency</td><td>Words</td><td>Timed</td><td>.59<ext-link /><sup>***</sup></td><td>(.11)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Word frequency × untimed</td><td>Words</td><td>Timed, untimed</td><td>–.32†</td><td>(.18)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Word frequency</td><td>Words</td><td>Untimed</td><td>–.02</td><td>(.02)</td></tr><tr><td>2.1</td><td>Logit(P+)</td><td>Number of orthographic neighbors</td><td>Words</td><td>Untimed</td><td>.18</td><td>(.11)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Number of orthographic neighbors</td><td>Words</td><td>Timed</td><td>.15</td><td>(.13)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Number of orthographic neighbors × untimed</td><td>Words</td><td>Timed, untimed</td><td>.03</td><td>(.17)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Number of orthographic neighbors</td><td>Words</td><td>Untimed</td><td>–.02</td><td>(.01)</td></tr><tr><td>2.1</td><td>Logit(P+)</td><td>Number of orthographic neighbors (base word)</td><td>Non‐words</td><td>Untimed</td><td>–.09</td><td>(.09)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Number of orthographic neighbors (base word)</td><td>Non‐words</td><td>Timed</td><td>–.12<ext-link /><sup>*</sup></td><td>(.06)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Number of orthographic neighbors (base word) × untimed</td><td>Non‐words</td><td>Timed, untimed</td><td>.03</td><td>(.11)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Number of orthographic neighbors</td><td>Non‐words</td><td>Untimed</td><td>–.01</td><td>(.01)</td></tr><tr><td>2.1</td><td>Logit(P+)</td><td>Word similar</td><td>Non‐words</td><td>Untimed</td><td>–.55</td><td>(.36)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Word similar</td><td>Non‐words</td><td>Timed</td><td>–.16</td><td>(.24)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Word similar × untimed</td><td>Non‐words</td><td>Timed, untimed</td><td>–.40</td><td>(.44)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Word similar</td><td>Non‐words</td><td>Untimed</td><td>–.03</td><td>(.04)</td></tr><tr><td>2.1</td><td>Logit(P+)</td><td>Pseudohomophone</td><td>Non‐words</td><td>Untimed</td><td>–.34</td><td>(.37)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Pseudohomophone</td><td>Non‐words</td><td>Timed</td><td>–.01</td><td>(.23)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Pseudohomophone × untimed</td><td>Non‐words</td><td>Timed, untimed</td><td>–.34</td><td>(.43)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Pseudohomophone</td><td>Non‐words</td><td>Untimed</td><td>–.05</td><td>(.04)</td></tr></tbody></table> </ephtml> </p> <ulist> <item>7 <emph>Note</emph>: Model 2.3 is a two‐dimensional 1 PL model (with latent variables for timed and untimed ability).</item> <item>8 †<emph>p</emph> < .10.</item> <item>9 * <emph>p</emph> < .05.</item> <item>10 *** <emph>p</emph> < .001 (two‐sided tests).</item> <item>11 <emph>SE</emph> = standard error; RT = log‐transformed response time; P+ = probability to identify the stimulus correctly.</item> </ulist> <p>As expected, <emph>word frequency</emph> proved to be a facilitating factor (Table 5). Most importantly, the positive effect on accuracy and the amount of explained variance, respectively, was larger in the timed (model 2.2, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>.</mo><mn>59</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0306</mn></mrow><annotation encoding="application/x-tex">${\beta _{11}} =.59,\;p <.001,\,R_{\mathrm{marginal}}^{2} =.0306$</annotation></semantics></math> </ephtml> ) than in the untimed condition (model 2.1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>.</mo><mn>26</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>05</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0053</mn></mrow><annotation encoding="application/x-tex">${\beta _{11}} =.26,\;p <.05,\,R_{\mathrm{marginal}}^{2} =.0053$</annotation></semantics></math> </ephtml> ), which was also reflected by the interaction effect included in the extended model combining timed and untimed condition (model 2.3, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>3</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>32</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mspace width="0.28em" /><mo>=</mo><mo>.</mo><mn>07</mn></mrow><annotation encoding="application/x-tex">${\beta}_{3} = -.32,\;p\; =.07$</annotation></semantics></math> </ephtml> ). However, the interaction effect did not achieve statistical significance. There was no effect of word frequency on time intensity in the untimed condition (model 3).</p> <p>As shown in Table 5, other than expected, the facilitating effect for the <emph>number of orthographic neighbors</emph> when evaluating words was not significant, neither for the untimed condition (model 2.1) nor for the timed condition (model 2.2). Also, the interaction between the number of orthographic neighbors and condition was not significant (model 2.3). Finally, there was no effect of the number of orthographic neighbors on time intensity in the untimed condition (model 3).</p> <p>For non‐word trials, the assumed inhibitory effect of the <emph>number of orthographic neighbors (base word)</emph> could not be found in the untimed condition (model 2.1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>09</mn><mo>,</mo><mspace width="2.79999pt" /><mi>p</mi><mspace width="2.79999pt" /><mo>=</mo><mo>.</mo><mn>34</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0019</mn></mrow><annotation encoding="application/x-tex">${\beta}_{11}=-.09,\hspace*{0.28em}p\hspace*{0.28em}=.34,\, {R}_{\mathrm{marginal}}^{2}=.0019$</annotation></semantics></math> </ephtml> ) but in the timed condition (model 2.2, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>12</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>05</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0033</mn></mrow><annotation encoding="application/x-tex">${\beta _{11}} = -.12,\;p <.05,\,R_{\mathrm{marginal}}^{2} =.0033$</annotation></semantics></math> </ephtml> ). The corresponding interaction effect, however, did not achieve statistical significance (model 2.3, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>3</mn></msub><mo>=</mo><mspace width="0.28em" /><mo>.</mo><mn>03</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mspace width="0.28em" /><mo>=</mo><mo>.</mo><mn>78</mn></mrow><annotation encoding="application/x-tex">${\beta}_{3} = \;.03,\;p\; =.78$</annotation></semantics></math> </ephtml> ). Also, the time intensity in the untimed condition was not related to the number of orthographic neighbors (base word) (model 3).</p> <p>The remaining part of Table 5 shows the <emph>word similarity</emph> and <emph>pseudohomophone</emph> effects. Different from what we expected, there were no significant negative effects of these linguistic properties neither in the timed nor in the untimed condition.</p> <p>Finally, we included the item properties jointly as item‐level covariates into the model to investigate their unique effects. The upper part of Table 6 shows the results for the two item properties being applicable to words by condition. The overall picture does not change: A significant word frequency effect could be revealed only for the timed condition (model 4.2). The lower part of Table 6 shows the obtained result pattern for non‐words by condition. We included all three linguistic item properties being applicable to non‐words. The results revealed the expected significant inhibitory effect for the base word's number of orthographic neighbors in the timed condition (model 4.2); however, it was not significant in the untimed condition (model 4.1).</p> <p>6 Table Unique Effects of Linguistic Item Properties in the Word Recognition Task</p> <p> <ephtml> <table><thead><tr valign="bottom"><th>Model</th><th align="center">Dependent Variable</th><th align="center">Item‐Level Covariate</th><th align="center">Included Trials</th><th align="center">Condition</th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><annotation encoding="application/x-tex">${\beta _{1k}}$</annotation></semantics></math></p></th><th align="center">(<italic>SE</italic>)</th></tr></thead><tbody><tr><td>4.1</td><td>Logit(P+)</td><td>Word frequency</td><td>Words</td><td>Untimed</td><td>.21</td><td>(.18)</td></tr><tr><td /><td /><td>Number of orthographic neighbors</td><td /><td /><td>.06</td><td>(.15)</td></tr><tr><td>4.2</td><td>Logit(P+)</td><td>Word frequency</td><td>Words</td><td>Timed</td><td>.59<ext-link /><sup>***</sup></td><td>(.12)</td></tr><tr><td /><td /><td>Number of orthographic neighbors</td><td /><td /><td>–.01</td><td>(.07)</td></tr><tr><td>4.1</td><td>Logit(P+)</td><td>Number of orthographic neighbors (base word)</td><td>Non‐words</td><td>Untimed</td><td>–.04</td><td>(.09)</td></tr><tr><td /><td /><td>Word similar</td><td /><td /><td>–.41</td><td>(.41)</td></tr><tr><td /><td /><td>Pseudohomophone</td><td /><td /><td>–.20</td><td>(.37)</td></tr><tr><td>4.2</td><td>Logit(P+)</td><td>Number of orthographic neighbors (base word)</td><td>Non‐words</td><td>Timed</td><td>–.14<ext-link /><sup>*</sup></td><td>(.07)</td></tr><tr><td /><td /><td>Word similar</td><td /><td /><td>.12</td><td>(.28)</td></tr><tr><td /><td /><td>Pseudohomophone</td><td /><td /><td>–.18</td><td>(.26)</td></tr></tbody></table> </ephtml> </p> <ulist> <item>12 * <emph>p</emph> < .05.</item> <item>13 *** <emph>p</emph> < .001 (two‐sided tests).</item> <item>14 <emph>SE</emph> = standard error; P+ = probability to identify the stimulus correctly.</item> </ulist> <p>Note that model 2.3 also informs about the overall condition effect on difficulty and the correlation of the person random intercepts between conditions (not presented in the results tables). All models with word‐/non‐word‐related linguistic properties demonstrated that the timed condition was significantly more difficult than the untimed condition. The person random intercepts were only moderately correlated (models including word trials: .32; models including non‐word trials: .41) suggesting that the two latent variables representing individual differences in word recognition by condition have different interpretations.</p> <hd id="AN0180987481-22">Semantic Integration</hd> <p>The upper part of Table 7 presents the effects of the <emph>number of propositions</emph> on the performance in the semantic integration task. Other than expected, there was no relation of the number of propositions to item difficulty in the untimed condition (model 2.1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0056" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>.</mo><mn>09</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mspace width="0.28em" /><mo>=</mo><mo>.</mo><mn>51</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0005</mn></mrow><annotation encoding="application/x-tex">${\beta}_{11} =.09,\;p\; =.51,\,R_{\mathrm{marginal}}^{2} =.0005$</annotation></semantics></math> </ephtml> ). However, as assumed, the number of propositions made the task significantly harder in the timed condition (model 2.2, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0057" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>32</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>01</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0103</mn></mrow><annotation encoding="application/x-tex">${\beta}_{11} = -.32,\;p <.01,\,R_{\mathrm{marginal}}^{2} =.0103$</annotation></semantics></math> </ephtml> ). The significant interaction effect (model 2.3, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0058" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>3</mn></msub><mo>=</mo><mo>.</mo><mn>41</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>05</mn></mrow><annotation encoding="application/x-tex">${\beta}_{3} =.41,\;p <.05$</annotation></semantics></math> </ephtml> ) also indicated a difference between conditions, suggesting that the number of propositions had a less impairing effect on performance in the untimed condition. As hypothesized for the untimed condition, there was a significant positive effect of number of propositions on item time intensity (model 3), meaning that sentences with more propositions took more time.</p> <p>7 Table Effects of Linguistic Item Properties in the Semantic Integration Task</p> <p> <ephtml> <table><thead><tr valign="bottom"><th>Model</th><th align="center">Dependent Variable</th><th align="center">Item‐Level Covariate</th><th align="center">Included Trials</th><th align="center">Condition</th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0059" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><annotation encoding="application/x-tex">${\beta _{1k}}$</annotation></semantics></math></p></th><th align="center">(<italic>SE</italic>)</th></tr></thead><tbody><tr><td>2.1</td><td>Logit(P+)</td><td>Number of propositions</td><td>(In)correct sentences</td><td>Untimed</td><td>.09</td><td>(.13)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Number of propositions</td><td>(In)correct sentences</td><td>Timed</td><td>–.32<ext-link /><sup>**</sup></td><td>(.11)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Number of propositions × untimed</td><td>(In)correct sentences</td><td>Timed, untimed</td><td>.41<ext-link /><sup>*</sup></td><td>(.18)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Number of propositions</td><td>(in)correct sentences</td><td>Untimed</td><td>.14<ext-link /><sup>***</sup></td><td>(.03)</td></tr><tr><td>2.1</td><td>Logit(P+)</td><td>Number of objects</td><td>(In)correct sentences</td><td>Untimed</td><td>–.09</td><td>(.14)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Number of objects</td><td>(In)correct sentences</td><td>Timed</td><td>–.38<ext-link /><sup>***</sup></td><td>(.09)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Number of objects × untimed</td><td>(In)correct sentences</td><td>Timed, untimed</td><td>.28†</td><td>(.17)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Number of objects</td><td>(In)correct sentences</td><td>Untimed</td><td>.15<ext-link /><sup>***</sup></td><td>(.03)</td></tr><tr><td>2.1</td><td>Logit(P+)</td><td>Unpredictable</td><td>Correct sentences</td><td>Untimed</td><td>–.19</td><td>(.34)</td></tr><tr><td>2.2</td><td>Logit(P+)</td><td>Unpredictable</td><td>Correct sentences</td><td>Timed</td><td>–.50<ext-link /><sup>*</sup></td><td>(.22)</td></tr><tr><td>2.3</td><td>Logit(P+)</td><td>Unpredictable × untimed</td><td>Correct sentences</td><td>Timed, untimed</td><td>.31</td><td>(.41)</td></tr><tr><td>3</td><td>Log(RT)</td><td>Unpredictable</td><td>Correct sentences</td><td>Untimed</td><td>.12</td><td>(.08)</td></tr></tbody></table> </ephtml> </p> <ulist> <item>15 <emph>Note</emph>: Model 2.3 is a two‐dimensional 1 PL model (with latent variables for timed and untimed ability).</item> <item>16 †<emph>p</emph> < .10.</item> <item>17 * <emph>p</emph> < .05.</item> <item>18 ** <emph>p</emph> < .01.</item> <item>19 *** <emph>p</emph> < .001 (two‐sided tests).</item> <item>20 <emph>SE</emph> = standard error; RT = log‐transformed response time; P+ = probability to identify the stimulus correctly.</item> </ulist> <p>As shown in the middle part of Table 7, the results for the property <emph>number of objects</emph> were quite similar. There was a non‐significant effect in the untimed condition (model 2.1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0060" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>09</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>50</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0005</mn></mrow><annotation encoding="application/x-tex">${\beta}_{11} = -.09,\;p <.50,\,R_{\mathrm{marginal}}^{2} =.0005$</annotation></semantics></math> </ephtml> ) and a significant effect in the timed condition (model 2.2, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0061" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>38</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0163</mn></mrow><annotation encoding="application/x-tex">${\beta _{11}} = -.38,\;p <.001,\,R_{\mathrm{marginal}}^{2} =.0163$</annotation></semantics></math> </ephtml> ). The model integrating both conditions provided evidence that effects differed between conditions (model 2.3, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0062" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>3</mn></msub><mo>=</mo><mo>.</mo><mn>28</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mspace width="0.28em" /><mo>=</mo><mo>.</mo><mn>09</mn></mrow><annotation encoding="application/x-tex">${\beta}_{3} =.28,\;p\; =.09$</annotation></semantics></math> </ephtml> ) although the interaction effect did not become significant. The effect on time intensity in the untimed condition was positive and significant (model 3).</p> <p>The results obtained for <emph>predictability</emph>, summarized in the lower part of Table 7, follow the same pattern. Low predictability does not make the semantic integration task harder in the untimed condition (model 2.1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0063" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>19</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mspace width="0.28em" /><mo>=</mo><mo>.</mo><mn>59</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0009</mn></mrow><annotation encoding="application/x-tex">${\beta}_{11} = -.19,\;p\; =.59,\,R_{\mathrm{marginal}}^{2} =.0009$</annotation></semantics></math> </ephtml> ), but this was the case in the timed condition (model 2.2, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0064" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>11</mn></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>50</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mo><</mo><mo>.</mo><mn>05</mn><mo>,</mo><mspace width="0.16em" /><msubsup><mi>R</mi><mi>marginal</mi><mn>2</mn></msubsup><mo>=</mo><mo>.</mo><mn>0075</mn></mrow><annotation encoding="application/x-tex">${\beta}_{11} = -.50,\;p <.05,\,R_{\mathrm{marginal}}^{2} =.0075$</annotation></semantics></math> </ephtml> ). However, the interaction effect proved not to be significant (model 2.3, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0065" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>β</mi><mn>3</mn></msub><mo>=</mo><mo>.</mo><mn>31</mn><mo>,</mo><mspace width="0.28em" /><mi>p</mi><mspace width="0.28em" /><mo>=</mo><mo>.</mo><mn>44</mn></mrow><annotation encoding="application/x-tex">${\beta}_{3} =.31,\;p\; =.44$</annotation></semantics></math> </ephtml> ). The expected positive effect of predictability on time intensity also did not reach statistical significance (model 3).</p> <p>Finally, Table 8 presents the unique effects of linguistic properties for (in)correct sentences. In the untimed condition (model 4.1), the effects for both item properties were not significant. In the timed condition (model 4.2), the negative effect of the number of objects was still significant whereas the unique effect of the correlated number of propositions was no longer significant. Note that the linguistic property of predictability was the only variable applicable to correct sentences. Therefore, this property is not included in the analyses of unique effects.</p> <p>8 Table Unique Effects of Linguistic Item Properties in the Semantic Integration Task</p> <p> <ephtml> <table><thead><tr valign="bottom"><th>Model</th><th align="center">Dependent Variable</th><th align="center">Item‐Level Covariate</th><th align="center">Included Trials</th><th align="center">Condition</th><th align="center"><p><math display="inline" altimg="urn:x-wiley:00220655:media:jedm12393:jedm12393-math-0066" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>β</mi><mrow><mn>1</mn><mi>k</mi></mrow></msub><annotation encoding="application/x-tex">${\beta _{1k}}$</annotation></semantics></math></p></th><th align="center">(<italic>SE</italic>)</th></tr></thead><tbody><tr><td>4.1</td><td>Logit(P+)</td><td>Number of propositions</td><td>(In)correct sentences</td><td>Untimed</td><td>.30†</td><td>(.18)</td></tr><tr><td /><td /><td>Number of objects</td><td /><td /><td>–.31†</td><td>(.18)</td></tr><tr><td>4.2</td><td>Logit(P+)</td><td>Number of propositions</td><td>(In)correct sentences</td><td>Timed</td><td>–.08</td><td>(.14)</td></tr><tr><td /><td /><td>Number of objects</td><td /><td /><td>–.33*</td><td>(.13)</td></tr></tbody></table> </ephtml> </p> <ulist> <item>21 †<emph>p</emph> < .10.</item> <item>22 *<emph>p</emph> < .05</item> <item>23 (two‐sided tests).</item> <item>24 <emph>SE</emph> = standard error; P+ = probability to identify the stimulus correctly.</item> </ulist> <p>As for word recognition, model 2.3 provides information about the condition effect on difficulty and the correlation of the person random intercepts between conditions (not presented in the results tables). In all models, the timed condition was significantly more difficult than the untimed condition except in the model including number of propositions as item‐level covariate. The person random intercepts were moderately correlated (models including correct and incorrect sentences: .49; model including correct sentences: .43) suggesting that the two latent variables representing semantic integration have to be interpreted differently.</p> <hd id="AN0180987481-23">Discussion</hd> <p>The present study collected and evaluated validity evidence (AERA et al., [<reflink idref="bib1" id="ref64">1</reflink>]) based on response processes to support the construct interpretation of effective ability scores obtained from untimed and timed testing in terms of reading efficiency at word and sentence level. The validation strategy employed was to investigate the effects of theory‐based linguistic item properties on item difficulty. Overall, the results provided mixed support for the hypothesis that item properties would have stronger effects on item difficulty in the timed than the untimed condition. To this extent the findings suggest that the effective ability scores of the timed condition better reflect how well the test‐takers were able to cope with the theoretically relevant task requirements than the effective ability scores of the untimed condition. With regard to conclusions for application, it follows that timed items may have a higher diagnostic value due to their higher sensitivity to construct‐related item properties. For instance, if a student performs equally well on timed items with familiar and unfamiliar words, this suggests mastery of reading with both types of words. In contrast, untimed items are less sensitive to the difference in familiar and unfamiliar words, and thus are less informative about reading (unfamiliar) words efficiently. The improved construct validity evidence provided by timed testing is expected to have a beneficial impact on diagnostic decisions in applied assessment contexts, such as identifying weak readers for additional training programs to improve their reading fluency or further distinguishing within the group of skilled readers. From a school accountability perspective, timed testing may also be useful if one is interested in a stronger differentiation between schools, because the ICC is higher than for untimed testing.</p> <p>Overall, for the <emph>word recognition task</emph>, only few of the selected linguistic stimulus properties were related to item difficulty. However, the revealed significant effects supported our expectation that relationships in the timed condition are more pronounced than in the untimed condition. Specifically, word frequency showed a stronger facilitating effect in timed testing than in untimed testing. We also found evidence that the number of orthographic neighbors was inhibitory for non‐words in the timed condition while this was not the case in the untimed condition. Note that, although there was no clear support for our hypothesis about the linguistic item properties other than the effect of word frequency, this does not necessarily speak against it. It may be that the way in which the linguistic item properties were manipulated has no effect in the student sample, regardless of whether it is an untimed or a timed testing approach.</p> <p>The findings for the <emph>semantic integration task</emph> clearly supported our hypotheses. For all three linguistic properties—number of propositions, number of objects, and unpredictability—the expected negative effect was significant in the timed condition. In the untimed condition, however, the respective effect on difficulty was much smaller and not significant. This pattern was partly reflected by a significant interaction effect between condition and the respective item property. Finally, the effect on time intensity was positive and significant as expected with the exception of predictability. Although we found some significant effects of item properties for the timed condition, the picture of how to interpret the effective ability scores obtained remains incomplete, especially for the word recognition task, where only one item property showed the expected effects.</p> <p>The revealed differences in the effect of item‐level covariates on item difficulty between timed and untimed conditions might also be understood as violation of measurement invariance. These differences suggest that different abilities are measured in the two conditions (for a related finding, see De Boeck et al., [<reflink idref="bib13" id="ref65">13</reflink>]; Semmes et al., [<reflink idref="bib48" id="ref66">48</reflink>]). This conclusion is also supported by the only moderate correlation of person intercepts between conditions. One interpretation could be that under time pressure the test‐taker relies or needs to rely more on automated information processing, whereas with less time pressure—if needed—a more deliberate information processing including also other solution strategies is used.</p> <p>The automaticity theory (LaBerge & Samuels, [<reflink idref="bib33" id="ref67">33</reflink>]) assumes that advanced readers can recognize word meanings automatically, whereas poor readers require attentional control and effort for processing at different stages. The timed condition requiring to rely on automated information processing thus better helps to distinguish good and poor readers. Correspondingly, the item properties considered in the present study determine whether automatic processing is facilitated (e.g., frequency of a word) or hampered (e.g., number of propositions in a sentence). Actually, the word frequency effect, that is, frequent words are processed faster, is typically interpreted as a learning effect (Brysbaert et al., [<reflink idref="bib7" id="ref68">7</reflink>]). Thus, on average, the meaning of more frequent visual words can be retrieved more automatically.</p> <p>The present study used data from the study by Goldhammer et al. ([<reflink idref="bib18" id="ref69">18</reflink>]) and thus directly complements their findings. The former study investigated score interpretations based on person‐level differences (nomothetic span approach). Particularly, PISA reading comprehension was used as external criterion which was predicted by the reading component skills visual word recognition and sentence‐level semantic integration measured under untimed and timed conditions. A main finding was that effective ability measures obtained from timed testing of reading component skills better explain reading comprehension as compared to corresponding effective ability and effective speed measures obtained from untimed testing. The present study explains item‐level differences in reading component skills through theory‐based linguistic item properties, thus extending the interpretation of the test scores in terms of underlying cognitive processes. Both studies suggest, from different perspectives, that measurements controlling response speed at item level enable a clearer interpretation of effective ability scores in terms of cognitive efficiency than untimed measurements.</p> <p>A limitation of the present study is that our hypotheses about the effects of linguistic item properties are based on studies, some of which were conducted with students of a different age (e.g., university students, 8‐ and 9‐year‐old children) than the students in our study (15‐year‐old students). Thus, our hypotheses rely on the implicit assumption that the assumed effects are more or less invariant across age groups. Another limitation is that the order of the conditions and the item material presented was fixed. Thus, the experimental design could be further improved by balancing the position of the tests, the position of conditions, and the test content. For example, we cannot rule out the possibility that the differences between the two conditions are also due to the fixed order, that is, that behavior in the timed condition might be influenced by experience gained in the preceding untimed condition. Furthermore, although the timed experimental conditions were carefully designed and implemented, some subjects may not have changed or were unable to change their response speed in the timed condition as required. Furthermore, we cannot identify any test‐takers who finished the task processing too early but responded on time after waiting for the signal. Another potential limitation for a study focusing item‐level differences is the low number of items per experimental condition (i.e., 32 for word recognition, and 24 for semantic integration). The respective number of items included in a model was further reduced if only a subset was applicable for a given item property. Thus, the small number of items may have compromised the chance of detecting significant effects, and it also prevented us from considering more complex models that involve theoretically important interactions among item characteristics. Further research is needed to assess the robustness of the effects obtained. Note that with the employed modeling approach— LLTM with random items (i.e., error component)— it is much more difficult to reach statistically significant effects of item‐level covariates. Nevertheless, this approach is preferable because it provides more accurate standard errors of the effects, since it takes into account the full variability of the items (De Boeck et al., [<reflink idref="bib14" id="ref70">14</reflink>]). Another limitation that could be addressed in future research is to directly evaluate whether construct‐irrelevant characteristics (e.g., the test‐taker's test anxiety) explain differences between experimental conditions (i.e., increased item difficulty in timed condition, moderate correlation of condition‐specific latent variables). This would require, for instance, to assess state anxiety using a self‐report measure.</p> <p>The present study demonstrates how the experimental control of response speed at item level can extend educational measurement practices. The results provide one source of validity evidence supporting the conclusion that item‐level time limits controlling the stimulus presentation time and the time window for responding may provide a more valid construct interpretation of test scores in terms of underlying cognitive processes. The proposed testing procedure is applicable to constructs of cognitive efficiency when tests can be used that present a sequence of stimuli to which each must be responded to (e.g., two‐choice response paradigm). The time limit should be determined empirically by administering the items in an untimed condition. Time limits could then be chosen based on percentile values of the observed response time distributions such that a certain proportion of test‐takers of the target population—for example, 50%—are able to solve a given item correctly and in the given time. Goldhammer and Kroehne ([<reflink idref="bib19" id="ref71">19</reflink>]) showed experimentally how the choice of the item‐level time limit affects the performance (in the word recognition task). Based on that, we would assume that faster conditions would increase random responding and in turn the linguistic item properties would no longer explain item difficulty. Slower conditions should provide results being more and more similar to those obtained for the untimed condition. In the timed condition the goal is to manipulate the test‐taker's decision criterion on how much time to take to produce a response to a stimulus. This adaptation becomes easier for the test‐taker if the item‐level time constraints are predictable across items. Therefore, no item‐specific time limit for the stimulus presentation (e.g., depending on the time intensity) should be used but the same for all items within a task.</p> <p>Taken together, the study demonstrates that measurements of cognitive efficiency can be validly implemented by employing item‐level time limits controlling the stimulus presentation time and the time window for responding. As shown by our findings on reading component skills, this approach also leads to test scores that can be better interpreted in terms of underlying cognitive processes test‐takers are engaged in.</p> <hd id="AN0180987481-24">Acknowledgments</hd> <p>This research was funded and supported by the Centre for International Student Assessment (ZIB). The authors would like to thank Johannes Naumann and Tobias Richter for providing the stimulus material for the word recognition and semantic integration task for this study. Furthermore, we are thankful to three anonymous reviewers for their valuable insights and constructive comments.</p> <p>Open access funding enabled and organized by Projekt DEAL.</p> <ref id="AN0180987481-25"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref21" type="bt">1</bibl> <bibtext> https://<ulink href="http://www.dwds.de/">www.dwds.de/</ulink></bibtext> </blist> <blist> <bibl id="bib2" idref="ref30" type="bt">2</bibl> <bibtext> A previous version of the article was presented at the 2022 annual meeting of the National Council on Measurement in Education (NCME), San Diego, USA, April 21–24, 2022.</bibtext> </blist> </ref> <ref id="AN0180987481-26"> <title> References </title> <blist> <bibtext> AERA, APA, NCME, & Joint Committee on Standards for Educational Psychological Testing. (2014). Standards for educational and psychological testing. American Educational Research Association.</bibtext> </blist> <blist> <bibtext> Andrews, S. (1997). The effect of orthographic similarity on lexical retrieval: Resolving neighborhood conflicts. Psychonomic Bulletin & Review, 4 (4), 439 – 461. https://doi.org/10.3758/BF03214334</bibtext> </blist> <blist> <bibl id="bib3" idref="ref28" type="bt">3</bibl> <bibtext> Antos, S. J. (1979). Processing facilitation in a lexical decision task. Journal of Experimental Psychology: Human Perception and Performance, 5 (3), 527 – 545. https://doi.org/10.1037/0096‐1523.5.3.527</bibtext> </blist> <blist> <bibl id="bib4" idref="ref3" type="bt">4</bibl> <bibtext> Avvisati, F. (2023). What can we learn from the PISA reading‐fluency test? (121). OECD Publishing. https://www.oecd‐ilibrary.org/content/paper/c698b19a‐en</bibtext> </blist> <blist> <bibl id="bib5" idref="ref58" type="bt">5</bibl> <bibtext> Bates, D., Mächler, M., Bolker, B., & Walker, S. (2015). Fitting linear mixed‐effects models using lme4. Journal of Statistical Software, 67 (1), 1 – 48. https://doi.org/10.18637/jss.v067.i01</bibtext> </blist> <blist> <bibl id="bib6" idref="ref55" type="bt">6</bibl> <bibtext> Bolker, B. M., Brooks, M. E., Clark, C. J., Geange, S. W., Poulsen, J. R., Stevens, M. H. H., & White, J.‐S. S. (2009). Generalized linear mixed models: A practical guide for ecology and evolution. Trends in Ecology & Evolution, 24 (3), 127 – 135 Article 3. https://doi.org/10.1016/j.tree.2008.10.008</bibtext> </blist> <blist> <bibl id="bib7" idref="ref68" type="bt">7</bibl> <bibtext> Brysbaert, M., Mandera, P., & Keuleers, E. (2018). The word frequency effect in word processing: An updated review. Current Directions in Psychological Science, 27 (1), 45 – 50. https://doi.org/10.1177/0963721417727521</bibtext> </blist> <blist> <bibl id="bib8" idref="ref46" type="bt">8</bibl> <bibtext> Coltheart, M., Davelaar, E. J., Jonasson, J. T., & Besner, D. (1977). Access to the Internal Lexicon. In S. Dornic (Ed.), Attention and performance (Vol. VI, pp. 535 – 555). Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref24" type="bt">9</bibl> <bibtext> Cronbach, L. J., & Meehl, P. E. (1955). Construct validity in psychological tests. Psychological Bulletin, 52 (4), 281 – 302. https://doi.org/10.1037/h0040957</bibtext> </blist> <blist> <bibtext> Davison, M. L., Semmes, R., Huang, L., & Close, C. N. (2012). On the reliability and validity of a numerical reasoning speed dimension derived from response times collected in computerized testing. Educational and Psychological Measurement, 72 (2), 245 – 263. https://doi.org/10.1177/0013164411408412</bibtext> </blist> <blist> <bibtext> De Boeck, P. (2008). Random item IRT models. Psychometrika, 73 (4), 533 – 559. https://doi.org/10.1007/s11336‐008‐9092‐x</bibtext> </blist> <blist> <bibtext> Boeck, P. D., Bakker, M., Zwitser, R., Nivard, M., Hofman, A., Tuerlinckx, F., & Partchev, I. (2011). The estimation of item response models with the lmer function from the lme4 package in R. Journal of Statistical Software, 39 (12), 1 – 28. https://doi.org/10.18637/jss.v039.i12</bibtext> </blist> <blist> <bibtext> De Boeck, P., Chen, H., & Davison, M. (2017). Spontaneous and imposed speed of cognitive test responses. British Journal of Mathematical and Statistical Psychology, 70 (2), 225 – 237. https://doi.org/10.1111/bmsp.12094</bibtext> </blist> <blist> <bibtext> De Boeck, P., Cho, S., & Wilson, M. (2016). Explanatory item response models. In A. A. Rupp & J. P. Leighton (Eds.), The Wiley handbook of cognition and assessment: Frameworks, methodologies, and applications (pp. 247 – 266). Wiley Online Library.</bibtext> </blist> <blist> <bibtext> Whitely, S. E. (1983). Construct validity: Construct representation versus nomothetic span. Psychological Bulletin, 93 (1), 179 – 197. https://doi.org/10.1037/0033‐2909.93.1.179</bibtext> </blist> <blist> <bibtext> Fischer, G. H. (1973). The linear logistic test model as an instrument in educational research. Acta Psychologica, 37 (6), 359 – 374. https://doi.org/10.1016/0001‐6918(73)90003‐6</bibtext> </blist> <blist> <bibtext> Goldhammer, F. (2015). Measuring ability, speed, or both? Challenges, psychometric solutions, and what can be gained from experimental control. Measurement: Interdisciplinary Research and Perspectives, 13 (3–4), 133 – 164. https://doi.org/10.1080/15366367.2015.1100020</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Hahnel, C., Kroehne, U., & Zehner, F. (2021). From by product to design factor: On validating the interpretation of process indicators based on log data. Large‐Scale Assessments in Education, 9 (1), 20. https://doi.org/10.1186/s40536‐021‐00113‐5</bibtext> </blist> <blist> <bibtext> Goldhammer, F., & Kroehne, U. (2014). Controlling Individuals' Time Spent on Task in Speeded Performance Measures: Experimental Time Limits, Posterior Time Limits, and Response Time Modeling. Applied Psychological Measurement, 38, 255 – 267. https://doi.org/10.1177/0146621613517164</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Kroehne, U., Hahnel, C., & De Boeck, P. (2021). Controlling speed in component skills of reading improves the explanation of reading comprehension. Journal of Educational Psychology, 113 (5), 861 – 878. https://doi.org/10.1037/edu0000655</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Steinwascher, M. A., Kroehne, U., & Naumann, J. (2017). Modelling individual response time effects between and within experimental speed conditions: A GLMM approach for speeded tests. British Journal of Mathematical and Statistical Psychology, 70 (2), 238 – 256. https://doi.org/10.1111/bmsp.12099</bibtext> </blist> <blist> <bibtext> Goswami, U., Ziegler, J. C., Dalton, L., & Schneider, W. (2001). Pseudohomophone effects and phonological recoding procedures in reading development in English and German. Journal of Memory and Language, 45 (4), 648 – 664. https://doi.org/10.1006/jmla.2001.2790</bibtext> </blist> <blist> <bibtext> Gulliksen, H. (1950). Theory of mental tests. Wiley. https://doi.org/10.4324/9780203052150</bibtext> </blist> <blist> <bibtext> Hartig, F. (2022). DHARMa: Residual Diagnostics for Hierarchical (Multi‐Level /Mixed) Regression Models. R package version 0.4.6. [Computer software]. <ulink href="http://florianhartig.github.io/DHARMa/">http://florianhartig.github.io/DHARMa/</ulink></bibtext> </blist> <blist> <bibtext> Heitz, R. P. (2014). The Speed‐Accuracy Tradeoff: History, Physiology, Methodology, and Behavior. Frontiers in Neuroscience, 8. https://doi.org/10.3389/fnins.2014.00150</bibtext> </blist> <blist> <bibtext> Janssen, R., Schepers, J., & Peres, D. (2004). Models with item and item group predictors. In P. De Boeck, & M. Wilson (Eds.), Explanatory Item Response Models: A Generalized Linear and Nonlinear Approach (pp. 189 – 212). Springer New York. https://doi.org/10.1007/978‐1‐4757‐3990‐9_6</bibtext> </blist> <blist> <bibtext> Kane, M. T. (2001). Current Concerns in Validity Theory. Journal of Educational Measurement, 38 (4), 319 – 342. https://doi.org/10.1111/j.1745‐3984.2001.tb01130.x</bibtext> </blist> <blist> <bibtext> Kim, Y.‐S., Wagner, R. K., & Lopez, D. (2012). Developmental relations between reading fluency and reading comprehension: A longitudinal study from Grade 1 to Grade 2. Journal of Experimental Child Psychology, 113 (1), 93 – 111. https://doi.org/10.1016/j.jecp.2012.03.002</bibtext> </blist> <blist> <bibtext> Kintsch, W. (1974). The representation of meaning in memory. Erlbaum.</bibtext> </blist> <blist> <bibtext> Kintsch, W., & Keenan, J. (1973). Reading rate and retention as a function of the number of propositions in the base structure of sentences. Cognitive Psychology, 5 (3), 257 – 274. https://doi.org/10.1016/0010‐0285(73)90036‐4</bibtext> </blist> <blist> <bibtext> Klein Entink, R. H., Kuhn, J.‐T., Hornke, L. F., & Fox, J.‐P. (2009). Evaluating cognitive theory: A joint modeling approach using responses and response times. Psychological Methods, 14 (1), 54 – 75, Article 1. https://doi.org/10.1037/a0014877</bibtext> </blist> <blist> <bibtext> Kuznetsova, A., Brockhoff, P. B., & Christensen, R. H. B. (2017). lmerTest Package: Tests in linear mixed effects models. Journal of Statistical Software, 82 (13), 1 – 26. https://doi.org/10.18637/jss.v082.i13</bibtext> </blist> <blist> <bibtext> Laberge, D., & Samuels, S. J (1974). Toward a theory of automatic information processing in reading. Cognitive Psychology, 6 (2), 293 – 323. https://doi.org/10.1016/0010‐0285(74)90015‐2</bibtext> </blist> <blist> <bibtext> Lohman, D. F. (1989). Individual differences in errors and latencies on cognitive tasks. Learning and Individual Differences, 1 (2), 179 – 202. https://doi.org/10.1016/1041‐6080(89)90002‐2</bibtext> </blist> <blist> <bibtext> Luce, R. D. (1986). Response times: Their roles in inferring elementary mental organization. Oxford University Press.</bibtext> </blist> <blist> <bibtext> Merrell, C., & Tymms, P. (2007). Identifying reading problems with computer‐adaptive assessments. Journal of Computer Assisted Learning, 23 (1), 27 – 35. https://doi.org/10.1111/j.1365‐2729.2007.00196.x</bibtext> </blist> <blist> <bibtext> Müller, B., Richter, T., & Karageorgos, P. (2020). Syllable‐based reading improvement: Effects on word reading and reading comprehension in Grade 2. Learning and Instruction, 66, 101304. https://doi.org/10.1016/j.learninstruc.2020.101304</bibtext> </blist> <blist> <bibtext> Nakagawa, S., & Schielzeth, H. (2013). A general and simple method for obtaining R2 from generalized linear mixed‐effects models. Methods in Ecology and Evolution, 4 (2), 133 – 142. https://doi.org/10.1111/j.2041‐210x.2012.00261.x</bibtext> </blist> <blist> <bibtext> Perea, M., & Rosa, E. (2000). The effects of orthographic neighborhood in reading and laboratory word identification tasks: A review. Psicológica, 21 (2), 327 – 340.</bibtext> </blist> <blist> <bibtext> Prenzel, M., Sälzer, C., Klieme, E., & Köller, O. (2013). PISA 2012. Fortschritte und Herausforderungen in Deutschland [PISA 2012. Progress and challenges in Germany]. Waxmann.</bibtext> </blist> <blist> <bibtext> R Core Team. (2021). R: A language and environment for statistical computing. R Foundation for Statistical Computing. https://www.R‐project.org/</bibtext> </blist> <blist> <bibtext> Rayner, K., Ashby, J., Pollatsek, A., & Reichle, E. D. (2004). The effects of frequency and predictability on eye fixations in reading: Implications for the EZ Reader model. Journal of Experimental Psychology: Human Perception and Performance, 30 (4), 720 – 732. https://doi.org/10.1037/0096-1523.30.4.72</bibtext> </blist> <blist> <bibtext> Rayner, K., & Duffy, S. A. (1986). Lexical complexity and fixation times in reading: Effects of word frequency, verb complexity, and lexical ambiguity. Memory & Cognition, 14 (3), 191 – 201. https://doi.org/10.3758/BF03197692</bibtext> </blist> <blist> <bibtext> Reed, A V. (1973). Speed‐accuracy trade‐off in recognition memory. Science, 181 (4099), 574 – 576. https://doi.org/10.1126/science.181.4099.574</bibtext> </blist> <blist> <bibtext> Richter, T., Isberner, M.‐B., Naumann, J., & Kutzner, Y. (2012). Prozessbezogene Diagnostik von Lesefähigkeiten bei Grundschulkindern [Process‐based measurement of reading skills in primary school children]. Zeitschrift Für Pädagogische Psychologie, 26 (4), 313 – 331. https://doi.org/10.1024/1010‐0652/a000079</bibtext> </blist> <blist> <bibtext> Richter, T., & Naumann, J. (2009). Was misst der ELVES‐Subtest Satzverifikation? Analysen von Mess‐und Itemeigenschaften mit hierarchisch‐linearen Modellen [What does the ELVES sentence verification subtest measure? Analyses of measurement and item properties with hierarchical linear models]. In W. Lenhard, & W. Schneider (Eds.), Diagnostik und Förderung des Leseverständnisses (Vol. 7, pp. 131 – 149). Hogrefe & Huber Publishers.</bibtext> </blist> <blist> <bibtext> Salthouse, T. A., & Hedden, T. (2002). Interpreting reaction time measures in between‐group comparisons. Journal of Clinical and Experimental Neuropsychology, 24 (7), 858 – 872. https://doi.org/10.1076/jcen.24.7.858.8392</bibtext> </blist> <blist> <bibtext> Semmes, R., Davison, M. L., & Close, C. (2011). Modeling individual differences in numerical reasoning speed as a random effect of response time limits. Applied Psychological Measurement, 35 (6), 433 – 446. https://doi.org/10.1177/0146621611407305</bibtext> </blist> <blist> <bibtext> Stoffel, M. A., Nakagawa, S., & Schielzeth, H. (2020). partR2: Partitioning R2 in generalized linear mixed models. [Computer software]. PeerJ, 9, e11414. https://doi.org/10.1101/2020.07.26.221168 bioRxiv</bibtext> </blist> <blist> <bibtext> Thurstone, L. L. (1937). Ability, motivation, and speed. Psychometrika, 2 (4), 249 – 254. https://doi.org/10.1007/BF02287896</bibtext> </blist> <blist> <bibtext> Van Der Linden, W. J. (2009). Conceptual issues in response‐time modeling. Journal of Educational Measurement, 46 (3), 247 – 272. https://doi.org/10.1111/j.1745‐3984.2009.00080.x</bibtext> </blist> <blist> <bibtext> Wagner, R. K., Torgesen, J. K., Rashotte, C. A., & Pearson, N. A. (2010). Test of silent reading efficiency and comprehension. Pro‐Ed.</bibtext> </blist> <blist> <bibtext> White, S., Sabatini, J., Park, B. J., Chen, J., Bernstein, J., & Li, M. (2021). The 2018 NAEP Oral Reading Fluency Study. NCES 2021–025. National Center for Education Statistics, U.S. Department of Education. https://nces.ed.gov/pubsearch/pubsinfo.asp?pubid=2021026</bibtext> </blist> <blist> <bibtext> Wickelgren, W. A. (1977). Speed‐accuracy tradeoff and information processing dynamics. Acta Psychologica, 41, 67 – 85. https://doi.org/10.1016/0001‐6918(77)90012‐9</bibtext> </blist> <blist> <bibtext> Yap, M. J., Sibley, D. E., Balota, D. A., Ratcliff, R., & Rueckl, J. (2015). Responding to nonwords in the lexical decision task: Insights from the English Lexicon Project. Journal of Experimental Psychology. Learning, Memory, and Cognition, 41 (3), 597 – 613. PubMed. https://doi.org/10.1037/xlm0000064</bibtext> </blist> </ref> <aug> <p>By Frank Goldhammer; Ulf Kroehne; Carolin Hahnel; Johannes Naumann and Paul De Boeck</p> <p>Reported by Author; Author; Author; Author; Author</p> <p></p> <p>FRANK GOLDHAMMER is Professor at DIPF | Leibniz Institute for Research and Information in Education, Rostocker Str. 6, 60323 Frankfurt, Germany and Centre for International Student Assessment (ZIB), Marsstr. 20−22, 80335 Munich, Germany; f.goldhammer@dipf.de. His primary research interests include technology‐based assessment of learning outcomes and processes, international large‐scale assessment, behavioral process data including response speed, digital competencies and reading skills.</p> <p>ULF KROEHNE is Head of unit at DIPF | Leibniz Institute for Research and Information in Education, Rostocker Str. 6, 60323 Frankfurt, Germany; u.kroehne@dipf.de. His primary research interests include psychometrics and latent variable models, methods for analyzing log data, technology‐based assessment, computer‐based adaptive testing, causal inference and program evaluation.</p> <p>CAROLIN HAHNEL is Professor at Ruhr University Bochum, Universitätsstraße 150, 44801 Bochum, Germany; carolin.hahnel@ruhr‐uni‐bochum.de. Her primary research interests include reading and learning in digital environments, large‐scale assessment and the modeling and interpretation of process‐related behavioral data.</p> <p>JOHANNES NAUMANN is Professor at University of Wuppertal, Gaußstraße 20, 42119 Wuppertal, Germany; j.naumann@uni‐wuppertal.de. His primary research interests include the assessment, modeling, development and fostering of basic cognitive processes and literacy in reading and listening comprehension, problem solving, ICTs, navigation and text comprehension with digital texts.</p> <p>PAUL DE BOECK is Professor Emeritus at Ohio State University, 1827 Neil Ave., Columbus, OH 43210; deboeck.2@osu.edu. His primary research interests include explanatory measurement and psychometric modeling (e.g., IRT, GLMM), item responses and the underlying processes, and the basis of the replication crisis in psychology with focus on the robustness of findings.</p> </aug> <nolink nlid="nl1" bibid="bib28" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib53" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib36" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib54" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib23" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib51" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib35" firstref="ref9"></nolink> <nolink nlid="nl8" bibid="bib17" firstref="ref11"></nolink> <nolink nlid="nl9" bibid="bib21" firstref="ref12"></nolink> <nolink nlid="nl10" bibid="bib37" firstref="ref13"></nolink> <nolink nlid="nl11" bibid="bib10" firstref="ref14"></nolink> <nolink nlid="nl12" bibid="bib19" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib34" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib47" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib25" firstref="ref19"></nolink> <nolink nlid="nl16" bibid="bib52" firstref="ref20"></nolink> <nolink nlid="nl17" bibid="bib27" firstref="ref22"></nolink> <nolink nlid="nl18" bibid="bib15" firstref="ref25"></nolink> <nolink nlid="nl19" bibid="bib50" firstref="ref27"></nolink> <nolink nlid="nl20" bibid="bib43" firstref="ref29"></nolink> <nolink nlid="nl21" bibid="bib39" firstref="ref31"></nolink> <nolink nlid="nl22" bibid="bib55" firstref="ref32"></nolink> <nolink nlid="nl23" bibid="bib22" firstref="ref33"></nolink> <nolink nlid="nl24" bibid="bib30" firstref="ref34"></nolink> <nolink nlid="nl25" bibid="bib42" firstref="ref35"></nolink> <nolink nlid="nl26" bibid="bib45" firstref="ref36"></nolink> <nolink nlid="nl27" bibid="bib18" firstref="ref37"></nolink> <nolink nlid="nl28" bibid="bib31" firstref="ref38"></nolink> <nolink nlid="nl29" bibid="bib40" firstref="ref40"></nolink> <nolink nlid="nl30" bibid="bib44" firstref="ref43"></nolink> <nolink nlid="nl31" bibid="bib29" firstref="ref47"></nolink> <nolink nlid="nl32" bibid="bib46" firstref="ref48"></nolink> <nolink nlid="nl33" bibid="bib12" firstref="ref49"></nolink> <nolink nlid="nl34" bibid="bib11" firstref="ref50"></nolink> <nolink nlid="nl35" bibid="bib48" firstref="ref51"></nolink> <nolink nlid="nl36" bibid="bib26" firstref="ref52"></nolink> <nolink nlid="nl37" bibid="bib16" firstref="ref53"></nolink> <nolink nlid="nl38" bibid="bib24" firstref="ref56"></nolink> <nolink nlid="nl39" bibid="bib41" firstref="ref57"></nolink> <nolink nlid="nl40" bibid="bib32" firstref="ref59"></nolink> <nolink nlid="nl41" bibid="bib38" firstref="ref61"></nolink> <nolink nlid="nl42" bibid="bib49" firstref="ref62"></nolink> <nolink nlid="nl43" bibid="bib13" firstref="ref65"></nolink> <nolink nlid="nl44" bibid="bib33" firstref="ref67"></nolink> <nolink nlid="nl45" bibid="bib14" firstref="ref70"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1449104
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Does Timed Testing Affect the Interpretation of Efficiency Scores?--A GLMM Analysis of Reading Components
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Frank+Goldhammer%22">Frank Goldhammer</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0289-9534">0000-0003-0289-9534</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ulf+Kroehne%22">Ulf Kroehne</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0412-169X">0000-0002-0412-169X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Carolin+Hahnel%22">Carolin Hahnel</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-2394-3944">0000-0003-2394-3944</externalLink>)<br /><searchLink fieldCode="AR" term="%22Johannes+Naumann%22">Johannes Naumann</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-2625-9630">0000-0002-2625-9630</externalLink>)<br /><searchLink fieldCode="AR" term="%22Paul+De+Boeck%22">Paul De Boeck</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-0884-2582">0000-0002-0884-2582</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2024 61(3):349-377.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 29
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Timed+Tests%22">Timed Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Efficiency%22">Efficiency</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Interpretation%22">Test Interpretation</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Performance+Based+Assessment%22">Performance Based Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Word+Recognition%22">Word Recognition</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Skills%22">Reading Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Construct+Validity%22">Construct Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Semantics%22">Semantics</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/jedm.12393
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0655<br />1745-3984
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The efficiency of cognitive component skills is typically assessed with speeded performance tests. Interpreting only effective ability or effective speed as efficiency may be challenging because of the within-person dependency between both variables (speed-ability tradeoff, SAT). The present study measures efficiency as effective ability conditional on speed by controlling speed experimentally. Item-level time limits control the stimulus presentation time and the time window for responding (timed condition). The overall goal was to examine the construct validity of effective ability scores obtained from untimed and timed condition by comparing the effects of theory-based item properties on item difficulty. If such effects exist, the scores reflect how well the test-takers were able to cope with the theory-based requirements. A German subsample from PISA 2012 completed two reading component skills tasks (i.e., word recognition and semantic integration) with and without item-level time limits. Overall, the included linguistic item properties showed stronger effects on item difficulty in the timed than the untimed condition. In the semantic integration task, item properties explained the time required in the untimed condition. The results suggest that effective ability scores in the timed condition better reflect how well test-takers were able to cope with the theoretically relevant task demands.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1449104
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1449104
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/jedm.12393
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 29
        StartPage: 349
    Subjects:
      – SubjectFull: Timed Tests
        Type: general
      – SubjectFull: Efficiency
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Test Interpretation
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Performance Based Assessment
        Type: general
      – SubjectFull: Difficulty Level
        Type: general
      – SubjectFull: Word Recognition
        Type: general
      – SubjectFull: Reading Skills
        Type: general
      – SubjectFull: Construct Validity
        Type: general
      – SubjectFull: Semantics
        Type: general
    Titles:
      – TitleFull: Does Timed Testing Affect the Interpretation of Efficiency Scores?--A GLMM Analysis of Reading Components
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Frank Goldhammer
      – PersonEntity:
          Name:
            NameFull: Ulf Kroehne
      – PersonEntity:
          Name:
            NameFull: Carolin Hahnel
      – PersonEntity:
          Name:
            NameFull: Johannes Naumann
      – PersonEntity:
          Name:
            NameFull: Paul De Boeck
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 0022-0655
            – Type: issn-electronic
              Value: 1745-3984
          Numbering:
            – Type: volume
              Value: 61
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Journal of Educational Measurement
              Type: main
ResultId 1