Flexible Item Response Modeling for Timed Reading Comprehension Assessment

Saved in:
Bibliographic Details
Title: Flexible Item Response Modeling for Timed Reading Comprehension Assessment
Language: English
Authors: Boris Forthmann (ORCID 0000-0001-9755-7304), Wolfgang Lenhard (ORCID 0000-0002-8184-6889), Alexandra Lenhard (ORCID 0000-0001-8680-4381), Natalie Förster (ORCID 0000-0003-0634-5993)
Source: Journal of Experimental Education. 2025 93(4):770-786.
Availability: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
Peer Reviewed: Y
Page Count: 17
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Descriptors: Item Response Theory, Reading Comprehension, Reading Rate, Models, Psychometrics, Reading Tests, Timed Tests, Raw Scores
DOI: 10.1080/00220973.2024.2367162
ISSN: 0022-0973
1940-0683
Abstract: While a rich methodology for analyzing response patterns for accuracy and time-on-task is at hand via Item Response Theory (IRT), tests with time cutoffs are so far harder to handle. Given that this test mode is widely applied, especially in the context of paper-and-pencil testing, there is a lack of psychometric techniques for a relevant number of tests. In this context, the original work of Rasch and his Rasch Poisson Counts model indeed offers an approach for this scenario that is adequate to solve the problem but which leads to model violations in many cases. Recent developments in statistical modeling -- the so-called Conway Maxwell Poisson Counts Model (CMPCM) -- can solve the problem of under- and overdispersion. We apply this model to the norm data of the ELFE II reading comprehension test and analyze patterns of over- and underdispersion with regard to speededness and mode effects. CMPCM with subtest-specific dispersion was adequate to model the raw test data, with underdispersion occurring mainly in highly speeded subtests with low difficulty and overdispersion in less speeded subtests with high difficulty. Thus, the CMPCM could contribute to psychometric methodology to appropriately model tests with time cutoffs on the subtest level.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1501045
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHGjXGuZd7PqNaNtJZPQ8FuAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDDTDrQ_RR4r2T_Sx8gIBEICBmxUnchvwsPchzeGJGwSyTo4XU87I_3kybjwFp4fcp-0XdqiOat0_KMP-mAfyttkz5_OK9F-TCtddf-f1irPhFE_Bx8FLWF3pgt-fo0tmKvPYWiuVjQQrv36e4E6oJUsqr4EIuLGYhat3znK5TCze1UolTid4ykSFaklHBPfLZ-Bv4gW4xOCeeDnR6zZQ-JjKUtNlOBQVCMxHvjxw
Text:
  Availability: 1
  Value: <anid>AN0187097618;jxe01oct.25;2025Aug05.05:46;v2.2.500</anid> <title id="AN0187097618-1">Flexible Item Response Modeling for Timed Reading Comprehension Assessment </title> <p>While a rich methodology for analyzing response patterns for accuracy and time-on-task is at hand via Item Response Theory (IRT), tests with time cutoffs are so far harder to handle. Given that this test mode is widely applied, especially in the context of paper-and-pencil testing, there is a lack of psychometric techniques for a relevant number of tests. In this context, the original work of Rasch and his Rasch Poisson Counts model indeed offers an approach for this scenario that is adequate to solve the problem but which leads to model violations in many cases. Recent developments in statistical modeling – the so-called Conway Maxwell Poisson Counts Model (CMPCM) – can solve the problem of under- and overdispersion. We apply this model to the norm data of the ELFE II reading comprehension test and analyze patterns of over- and underdispersion with regard to speededness and mode effects. CMPCM with subtest-specific dispersion was adequate to model the raw test data, with underdispersion occurring mainly in highly speeded subtests with low difficulty and overdispersion in less speeded subtests with high difficulty. Thus, the CMPCM could contribute to psychometric methodology to appropriately model tests with time cutoffs on the subtest level.</p> <p>Keywords: Conway–Maxwell Poisson Counts Model (CMPCM); Item Response Theory; mode effects; Reading Comprehension; speededness</p> <p>The aim of psychometric tests is to convert empirically observable differences between persons into numbers in such a way that the numbers reflect the relations between persons (de Ayala, [<reflink idref="bib11" id="ref1">11</reflink>]; Lenhard & Lenhard, [<reflink idref="bib38" id="ref2">38</reflink>]). This undertaking is not at all trivial. First, items are needed that capture the personality trait to be measured in all its different facets. Second, the set of items used for a test must not only cover all facets but also all levels of the targeted personality trait. Since the beginning of intelligence research, it has been assumed that the ability of a person is reflected in both the correctness of a response, but also in the speed of the solution process (Goldhammer, [<reflink idref="bib23" id="ref3">23</reflink>]; Thorndike et al., [<reflink idref="bib51" id="ref4">51</reflink>], p. 33). In line with this assumption, many intelligence tests and the measurement of academic performance at least partially rely on speeded subtests.</p> <p>Including a speed component in the assessment, however, should result in an adjustment of the statistical model when estimating the personality trait. While the joint analysis of item accuracy and response time is generally challenging, it becomes especially difficult for paper-and-pencil (PP) tests, in which response times are not available for every single item because discontinue rules like time cutoffs are an integral part of the measurement process (e.g., not all items are necessarily reached by a test taker). In our study, we demonstrate the application of flexible count data models as a means of scaling a reading comprehension test that comprises three different subtests with varying degrees of speededness. Moreover, we analyze whether already established effects of the presentation mode are also reflected in the parameters of the models and the reliability of ability estimates.</p> <hd id="AN0187097618-2">Problems in modelling speeded tests</hd> <p>Both classical test theory (CTT) and item response theory (IRT) provide a range of methods for the task of optimal item selection. However, not all methods are equally suitable for all personality traits that one would like to capture. On one side of the aisle, the true score model of CTT focuses on observed raw scores, that is, the methods are generally suited for speeded and unspeeded tests alike, but the CTT framework fails to establish the conceptual link between the observed data and the latent trait of a person (de Ayala, [<reflink idref="bib11" id="ref5">11</reflink>]). On the other side of the aisle, IRT models establish the links between observed raw data and the latent trait of a person.</p> <p>Though IRT as a framework for test construction has become more and more central in recent decades, there is unfortunately still a relatively large gap in terms of the method inventory concerning speededness. Often, however, a test designer needs speeded tasks (or even both speeded and unspeeded tasks) to cover a construct in all its facets. This for example applies to reading fluency, which can be defined as precise and at the same time fast decoding. To focus only on either reading speed or accuracy would thus neglect essential information. Recently, IRT models have been proposed, which not only include accuracy but also take the working speed into account (Fox et al., [<reflink idref="bib20" id="ref6">20</reflink>]; van der Linden et al., [<reflink idref="bib52" id="ref7">52</reflink>]). Using these models, the shift in the relationship between speed and accuracy can be analyzed regarding the complexity of the item material. However, the models require the measurement of accuracy and response time at the item level and thus are very difficult to apply on PP tests or on subtests with discontinue rules or time cutoffs, where items not being reached are an integral part of the measurement process. Notably, there are scoring approaches to combine accuracy and reaction time into various forms of efficiency measures, e.g., by simply dividing accuracy by log-transformed time-on-task per item (for an overview, see Roskam, [<reflink idref="bib46" id="ref8">46</reflink>]). Nevertheless, these weighting procedures are again not scaling frameworks and they share the same drawbacks as inherent in observed sum scores (i.e., the conceptual link between data and latent trait is missing; see above). We would like to elaborate on the construct relevance in setting time cutoffs using the example of reading comprehension and its different aspects.</p> <hd id="AN0187097618-3">Facets of reading comprehension on different linguistic levels</hd> <p>Reading is a complex process that requires many different abilities. During the initial stages of reading acquisition, the main task is to correctly decode written words into phonetic sequences and also to match the correct meanings to these series of phonemes. Although the acquisition of this skill varies in difficulty and time in different languages depending on the depth of orthography, it can be noted that the final result of the learning process is roughly comparable independently of the transparency of the orthographic system (Landerl, Wimmer & Frith, [<reflink idref="bib37" id="ref9">37</reflink>]; Wimmer & Goswami, [<reflink idref="bib53" id="ref10">53</reflink>]). Eventually, the reader can automatically identify the correct phonetic form, while simultaneously accessing the semantic content of word meanings within a very short amount of time (so-called direct or semantic route; Coltheart et al., [<reflink idref="bib8" id="ref11">8</reflink>]).</p> <p>This skill of rapid and error-free decoding of written material is referred to as reading fluency and it plays a significant role for several reasons. Reading fluency is regarded as the essential factor that bridges the gap between the early stages of reading development (Ehri, [<reflink idref="bib14" id="ref12">14</reflink>]; Pikulski & Chard, [<reflink idref="bib44" id="ref13">44</reflink>]), where the phonetic structure of words is cumbersomely recoded on word, syllable, or grapheme level (e.g., the indirect route of the Dual Route Cascaded Modell of Coltheart et al., [<reflink idref="bib8" id="ref14">8</reflink>]), and higher processing levels like reading comprehension. According to these theories, fluency develops by forming connections between the visual presentation of words and their pronunciation and meanings, storing this connection in memory systems and thus reducing mental effort and speeding up the reading process. Automatization, reflected by the speed of information retrieval, is a key indicator of the extent to which these memory systems have developed. Consequently, reading fluency is also a very robust indicator for distinguishing proficient from poor readers (Fuchs et al., [<reflink idref="bib22" id="ref15">22</reflink>]; Stanovich, [<reflink idref="bib50" id="ref16">50</reflink>]), which makes it a favorable diagnostic instrument. At the same time, it exhibits very high interindividual stability as demonstrated by longitudinal reading research (Klicpera & Gasteiger-Klicpera, [<reflink idref="bib36" id="ref17">36</reflink>], p. 54), with almost parallel, by and large, non-intersecting developmental trajectories of the performances of students on different aptitude levels over the first eight years of schooling. Finally, because of its importance, reading fluency has been a major goal of reading interventions, especially in poor readers to improve reading comprehension and literacy by speeding up basic reading processes (e.g., NRP, [<reflink idref="bib42" id="ref18">42</reflink>]).</p> <p>Reading fluency is therefore a particularly important source of information for measuring reading ability. Given that reading errors occur only rarely once decoding has been automatized (Jenkins et al., [<reflink idref="bib32" id="ref19">32</reflink>]), interindividual differences in reading fluency can be measured mainly by recording the time needed to solve the items or by imposing a time limit. A traditional one-parameter-logistic IRT scaling would not correctly capture the latent ability because item difficulty at the level of individual words is restricted (i.e., early items are too easy, whereas late items are too difficult). At the same time, more sophisticated approaches such as, for example, log-normal reaction time modeling (Fox et al., [<reflink idref="bib20" id="ref20">20</reflink>]; package LNIRT; Fox, Klotzke & Klein Entink, [<reflink idref="bib21" id="ref21">21</reflink>]) are not compatible with standard methods used to assess reading fluency, like having children read out loud lists of words for one minute and counting, how many words they have been able to read correctly (e.g., Jenkins et al., [<reflink idref="bib32" id="ref22">32</reflink>]).</p> <p>However, the ability to decipher words correctly alone does not make a good reader. At the sentence and text level, there are additional processes that must be successfully mastered by the reader. For example, at the sentence level, syntactic structures must be correctly decoded, personal pronouns have to be assigned correctly, the meaning of prepositions must be correctly grasped, and much more (Schindler et al., [<reflink idref="bib47" id="ref23">47</reflink>]). Finally, at the text level, a reader must be able to distinguish between important and unimportant aspects of a text. Subsequently, the reader must integrate the important aspects into a coherent representation of the gist of this text, the so-called situational model (Kintsch, [<reflink idref="bib33" id="ref24">33</reflink>]). Among other things, working memory and fluid reasoning play an important role in this process as well (Cain, Oakhill, & Bryant, [<reflink idref="bib6" id="ref25">6</reflink>]; Cain, Oakhill, & Lemmon, [<reflink idref="bib7" id="ref26">7</reflink>]). In contrast to sight word reading, not all processes required for reading comprehension can be automatized on sentence and text level. As a consequence of the high cognitive demands that come into play at the text level, reading speed is no longer a suitable measure to differentiate between good and bad readers at this particular stage of reading acquisition (Klauda & Guthrie, [<reflink idref="bib35" id="ref27">35</reflink>]). Instead, reading performance at the text level can be quantified very well by simply looking at which items a child was able to solve, quite independently of the time needed. Interestingly, reading comprehension on sentence level plays an intermediate role between word and text level not only with regard to cognitive requirements but also with regard to measurement requirements. Given that the effects of speed and accuracy are rather balanced at this level, it constitutes a very firm and comprehensive measure of reading comprehension (Kirschmann et al., [<reflink idref="bib34" id="ref28">34</reflink>]). Notably, the mode of presentation interacts with both the complexity of the item material and aptitude (Lenhard, Schroeders, et al., [<reflink idref="bib40" id="ref29">40</reflink>]). Reading speed is higher when working on screen at the expense of accuracy. This effect is more pronounced in younger and poorer-performing children, and it more strongly affects basic reading processes.</p> <p>In summary, while the availability of IRT frameworks for jointly modeling time on task and accuracy is evolving (Fox et al., [<reflink idref="bib20" id="ref30">20</reflink>]), IRT methods for scaling tests with time cutoffs are largely missing so far. This contrasts with the widespread use of tests with time cutoffs for assessing some narrow facets of cognitive abilities like, for example, reading fluency within the broader concept of reading comprehension, or because of practical necessities in the case of PP assessments. Indeed, a unifying IRT framework to model speeded (and unspeeded) tasks as a function of one latent variable seems vital for such assessment contexts.</p> <hd id="AN0187097618-4">The Conway–Maxwell Poisson counts model (CMPCM)</hd> <p>Interestingly, Rasch's (1960) pioneering work that paved the way for the success of IRT included a general framework for dealing with speeded tasks. What is even more, Rasch's model explicitly targeted the reading process. To this purpose, Rasch used a Poisson process instead of a simple Bernoulli process, that is, he focused on the number of events (in Rasch's case: reading errors) within a certain amount of time. Rasch's Poisson counts model (RPCM) and its extensions were later used to scale various other cognitive tasks such as intelligence (Ogasawara, [<reflink idref="bib43" id="ref31">43</reflink>]), attention (Baghaei et al., [<reflink idref="bib2" id="ref32">2</reflink>]), mental speed (Doebler et al., [<reflink idref="bib13" id="ref33">13</reflink>]; Holling et al., [<reflink idref="bib27" id="ref34">27</reflink>]), memory (Jendryczko et al., [<reflink idref="bib31" id="ref35">31</reflink>]), verbal fluency (Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref36">19</reflink>]), and divergent thinking tasks (Forthmann et al., [<reflink idref="bib17" id="ref37">17</reflink>]). This clearly illustrates the wide applicability of the RPCM.</p> <p>However, the RPCM is based on the Poisson distribution, which expresses the probability of the occurrence of a given number of events <emph>y</emph>, such as the number of reading errors or, in this work, the number of correctly solved items, in an interval of time or space. This distribution assumes that each event occurs independently of the preceding event and with a known constant mean rate. The distribution is specified by only one parameter, usually labeled λ, which represents the expected value and variance of the Poisson. To put it in other words, mean and variance (notably even skewness is a function of λ) in Poisson distributions are equal and this fact is called <emph>equidispersion</emph>. The probability mass function of the Poisson is</p> <p>Graph</p> <p> <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mi>P</mi><mo stretchy="true">(</mo><mrow><mi>Y</mi><mo>=</mo><mi>y</mi></mrow><mo stretchy="true">|</mo><mrow><mi>λ</mi></mrow><mo stretchy="true">)</mo><mo>=</mo><mrow><mfrac><mrow><mrow><msup><mrow><mi>λ</mi></mrow><mrow><mi>y</mi></mrow></msup></mrow></mrow><mrow><mi>y</mi><mo>!</mo></mrow></mfrac></mrow><mi mathvariant="normal">e</mi><mtext mathvariant="normal">xp</mtext><mo>⁡</mo><mo>(</mo><mo>−</mo><mi>λ</mi><mo>)</mo><mo>.</mo></math> </ephtml> (<reflink idref="bib1" id="ref38">1</reflink>)</p> <p>In the RPCM for person <emph>j</emph> and item <emph>i</emph>, the expected value λ<emph><subs>ji</subs></emph> = μ<emph><subs>ji</subs></emph> (i.e., the parameter of a conditional Poisson distribution) is modeled as a log-linear function of an ability parameter θ<emph><subs>j</subs></emph> and item easiness parameter β<emph><subs>i</subs></emph>, i.e.</p> <p>Graph</p> <p> <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mi>λ</mi></mrow><mrow><mtext mathvariant="italic">ji</mtext></mrow></msub></mrow><mo>=</mo><mrow><msub><mrow><mi>μ</mi></mrow><mrow><mtext mathvariant="italic">ji</mtext></mrow></msub></mrow><mo>=</mo><mi>μ</mi><mo stretchy="true">(</mo><mrow><mrow><msub><mrow><mi>θ</mi></mrow><mrow><mi>j</mi></mrow></msub></mrow><mo>,</mo><mrow><msub><mrow><mi>β</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow></mrow><mo stretchy="true">)</mo><mo>=</mo><mrow><mrow><mtext mathvariant="normal">exp</mtext></mrow><mo /><mrow><mo stretchy="true">(</mo><mrow><mrow><msub><mrow><mrow><msub><mrow><mi>β</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow><mo>+</mo><mi>θ</mi></mrow><mrow><mi>j</mi></mrow></msub></mrow></mrow><mo stretchy="true">)</mo></mrow></mrow><mo>.</mo></math> </ephtml> The following basic assumption of the Poisson distribution applies to these conditional distributions in the RPCM as well: expected value and variance are equal (i.e., equidispersion).</p> <p>For most empirical data, however, this assumption is too restrictive because the variance may be greater than the mean (i.e., <emph>overdispersion</emph>) or smaller than the mean (i.e., <emph>underdispersion</emph>). If the Poisson model is used in cases of equidispersion violations, it can have serious consequences for the standard errors of all estimated parameters in the model. Thus, accurate statistical inference is therefore closely related to the issue of dispersion in count data models. When data are overdispersed, standard errors are too small in Poisson models compared to models that explicitly model overdispersion (Hilbe, [<reflink idref="bib26" id="ref39">26</reflink>]). Additionally, standard errors in Poisson models are too large for underdispersed data (Faddy & Bosch, [<reflink idref="bib15" id="ref40">15</reflink>]). Therefore, when using the Poisson model, statistical inference can be too liberal in the case of overdispersion or too conservative in the case of underdispersion. Statisticians and researchers may be more concerned with liberal inference resulting in increased type-I errors. Therefore, overdispersion is more widely known, and statistical approaches to handle it are more mature (e.g., Hilbe, [<reflink idref="bib26" id="ref41">26</reflink>]) compared to underdispersion.</p> <p>To further illustrate dispersion and its relevance to the context of this work, let's consider a scenario where the same person has taken a reading test multiple times with constant easiness. Test scores from the first 10 runs were: 25, 20, 28, 34, 27, 35, 20, 25, 22, and 25. The average score is 26.10 and the variance is 26.77, which aligns with the concept of equidispersion under the RPCM. Assuming the second test score is 10 instead of 20 and the fourth test score is 44 instead of 34, the mean score remains the same. However, the variance increases to 80.10, indicating overdispersion in the conditional distribution. This set of tests violates the assumption of equidispersion and thus contradicts the application of the RPCM. It is also possible to construct an underdispersed set of test scores by replacing the second score of 20 with a score of 24 and the fourth score of 34 with a score of 30. The test scores have a mean of 26.10 and a variance of 17.88, indicating underdispersion and a violation of the assumptions of conditional Poisson distributions under the RPCM.</p> <p>It is apparent from the repeated assessments given in the three examples that there is a difference in their reliability. The data suggest that the least reliable assessment is the one with overdispersion in the conditional distribution and that underdispersed items are more reliable than overdispersed items. Such non-equidispersed data pose significant challenges for the RPCM, as it lacks mechanisms to account for dispersion issues. In psychometric testing, deviations from equidispersion, manifesting as either over- or underdispersion, can skew the reliability estimates of latent abilities. When the RPCM is applied under these conditions, the precision of measurement can be inaccurately assessed—either overestimated or underestimated. This discrepancy underscores the necessity of precise evaluation methods and accurate measurement precision quantification. Hung ([<reflink idref="bib30" id="ref42">30</reflink>]) extended the RPCM to a negative binomial counts model to address overdispersion. This model prevents overestimation of person parameter reliability when items are overdispersed (cf. Forthmann & Doebler, [<reflink idref="bib16" id="ref43">16</reflink>]).</p> <p>Modeling of underdispersion, however, has been technically more challenging. For example, flexible distributions such as the generalized Poisson distribution (Consul & Famoye, [<reflink idref="bib9" id="ref44">9</reflink>]) or the Conway-Maxwell-Poisson (CMP; Guikema & Coffelt, [<reflink idref="bib25" id="ref45">25</reflink>]; Sellers & Shmueli, [<reflink idref="bib49" id="ref46">49</reflink>]) distribution support only a limited range of parameter values or were not parameterized in a way that facilitates interpretation, respectively. For example, the probability mass function of the generalized Poisson distribution may become undefined for underdispersed data (Czado et al., [<reflink idref="bib10" id="ref47">10</reflink>]). Recent developments related to the CMP distribution (Huang, [<reflink idref="bib29" id="ref48">29</reflink>]), however, resulted in a mean parameterized CMP (CMP<subs>μ</subs>) distribution that allows deriving a flexible IRT count model equally applicable to items with equidispersion, overdispersion, and underdispersion. The central assumptions of the CMP counts model (CMPCM) are outlined in the following.</p> <p>Similar to the RPCM, the CMPCM models the expected value μ<emph><subs>ji</subs></emph> of the number of relevant responses <emph>y<subs>ji</subs></emph> for person <emph>j</emph> and item <emph>i</emph> as log-linear</p> <p>Graph</p> <p> <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mi>μ</mi></mrow><mrow><mtext mathvariant="italic">ji</mtext></mrow></msub></mrow><mo>=</mo><mi>μ</mi><mo stretchy="true">(</mo><mrow><mrow><msub><mrow><mi>θ</mi></mrow><mrow><mi>j</mi></mrow></msub></mrow><mo>,</mo><mrow><msub><mrow><mi>β</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow></mrow><mo stretchy="true">)</mo><mo>=</mo><mrow><mrow><mtext mathvariant="normal">exp</mtext></mrow><mo /><mrow><mo stretchy="true">(</mo><mrow><mrow><msub><mrow><mrow><msub><mrow><mi>β</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow><mo>+</mo><mi>θ</mi></mrow><mrow><mi>j</mi></mrow></msub></mrow></mrow><mo stretchy="true">)</mo></mrow></mrow><mo>,</mo></math> </ephtml> (<reflink idref="bib2" id="ref49">2</reflink>)</p> <p>with ability parameter θ<emph><subs>j</subs></emph> and item easiness parameter β<emph><subs>i</subs></emph>. Please note that the term "item" here refers to the vote count of a whole (sub-)test instead of the single elements of a test scale. Analogous to extensions of the RPCM, Equation (<reflink idref="bib2" id="ref50">2</reflink>) can further include a known time parameter</p> <p>Graph</p> <p> <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mi>ω</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow></math> </ephtml> for a priori known differences between items in terms of the allotted time-on-task</p> <p>Graph</p> <p> <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msub><mrow><mi>μ</mi></mrow><mrow><mtext mathvariant="italic">ji</mtext></mrow></msub></mrow><mo>=</mo><mi>μ</mi><mo stretchy="true">(</mo><mrow><mrow><msub><mrow><mi>θ</mi></mrow><mrow><mi>j</mi></mrow></msub></mrow><mo>,</mo><mrow><msub><mrow><mi>β</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow><mo>,</mo><mrow><msub><mrow><mi>ω</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow></mrow><mo stretchy="true">)</mo><mo>=</mo><mrow><mrow><mtext mathvariant="normal">exp</mtext></mrow><mo /><mrow><mo stretchy="true">(</mo><mrow><mrow><msub><mrow><mrow><msub><mrow><mi>β</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow><mo>+</mo><mrow><msub><mrow><mi>ω</mi></mrow><mrow><mi>i</mi></mrow></msub></mrow><mo>+</mo><mi>θ</mi></mrow><mrow><mi>j</mi></mrow></msub></mrow></mrow><mo stretchy="true">)</mo></mrow></mrow><mo>.</mo></math> </ephtml> (<reflink idref="bib3" id="ref51">3</reflink>)</p> <p>The CMPCM further incorporates an item-specific dispersion parameter ν<emph><subs>i</subs></emph>.</p> <p>Then, the number of relevant responses, i.e., the random variable <emph>Y<subs>ji</subs></emph> follows a CMP<subs>μ</subs>(μ(θ<emph><subs>j</subs></emph>, β<emph><subs>i</subs></emph>, ω<emph><subs>i</subs></emph>), ν<emph><subs>i</subs></emph>) distribution (i.e., Huang's mean parameterization of the CMP distribution). It is further assumed that ability parameter θ<emph><subs>j</subs></emph> follows a normal distribution with mean zero and variance</p> <p>Graph</p> <p> <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mrow><mi>σ</mi></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow><mo>.</mo></math> </ephtml> The higher the ability parameter θ<emph><subs>j</subs></emph> the higher the expected value on a given item. The item parameter indicates easiness (i.e., it is easier to score high on items with comparably larger β<emph><subs>i</subs></emph>). The known time parameter ω<emph><subs>i</subs></emph> reflects log-time (e.g., the logarithm of time in minutes) and can be interpreted analogous to the easiness parameter. Higher values on the parameter ω<emph><subs>i</subs></emph> are associated with more time-on-task.</p> <p>In addition, the dispersion parameter is transformed to enhance its interpretability by τ<emph><subs>i</subs></emph> = log(1/ν<emph><subs>i</subs></emph>). Underdispersion is indicated when τ<emph><subs>i</subs></emph> is smaller than zero, whereas a τ<emph><subs>i</subs></emph> that is larger than zero reflects overdispersion. Finally, the CMPCM reduces to the RPCM when τ<emph><subs>i</subs></emph> equals zero. In other words, when the RPCM is estimated, τ<emph><subs>i</subs></emph> is fixed to a value of zero. Item-specific dispersion in the CMPCM reflects the error variance or local reliability of an item, with greater reliability as τ<emph><subs>i</subs></emph> decreases. As illustrated by the third example of repeated assessments (see above), underdispersed items are always more reliable than overdispersed items. Notably, the conditional variance in the CMPCM is a function of the conditional expected value μ<emph><subs>ji</subs></emph> and dispersion parameter τ<emph><subs>i</subs></emph> which cannot be given by a simple formula (Huang, [<reflink idref="bib29" id="ref52">29</reflink>]). Finally, it should be noted that the conditional expected value μ<emph><subs>ji</subs></emph> and dispersion parameter τ<emph><subs>i</subs></emph> are orthogonal (cf. Huang, [<reflink idref="bib29" id="ref53">29</reflink>]).</p> <p>The CMPCM is a generalized linear mixed model and can be estimated by the package glmmTMB (Brooks et al., [<reflink idref="bib3" id="ref54">3</reflink>]) for the statistical software R. Hence, well in accordance with numerous IRT applications, parameter estimation is performed by marginal maximum likelihood estimation. The model displayed good parameter recovery with consistent estimates already for rather small sample sizes such as <emph>N</emph><bold></bold>=<bold></bold>100 (Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref55">19</reflink>]). In several studies, the CMPCM was found to outperform the RPCM when verbal fluency tasks, tasks for the measurement of general language ability (Forthmann, Grotjahn, et al., [<reflink idref="bib18" id="ref56">18</reflink>]), or researcher capacity as indicated by bibliometric indicators (Forthmann & Doebler, [<reflink idref="bib16" id="ref57">16</reflink>]) were measured. In all of these studies, underdispersed items were found which highlights the importance of focusing on all possible deviations from equidispersion (i.e., not only overdispersion).</p> <hd id="AN0187097618-5">Possible causes for under- and overdispersion</hd> <p>To date, it is not fully understood, what leads to under- and overdispersion, but several factors have previously been identified to cause these phenomena (Winkelmann, [<reflink idref="bib54" id="ref58">54</reflink>]). For example, a specific requirement for equidispersion within the Poisson process is that the intensity λ, that is, the number of correct answers per time unit, remains constant over time. An increase of λ over time leads to underdispersion whereas a decrease leads to overdispersion. This is especially relevant for tests with an item ordering with increasing difficulty in combination with time cutoffs.</p> <p>Additionally, it has been shown that highly speeded tests—e.g., measures of verbal fluency (Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref59">19</reflink>]) or attention (Baghaei et al., [<reflink idref="bib2" id="ref60">2</reflink>])—display under- instead of overdispersion. In these cases, the items are usually rather easy and of comparable item difficulty, which means that errors occur rarely and that the order of items is less relevant. Therefore, one could expect equidispersion at first sight. To answer the question of why underdispersion as compared to equidispersion occurs in these cases, let us consider the response to a single item instead of the hit rate across the whole subtest. The stationary (i.e., equidispersed) Poisson process assumes a purely random process. For example, imagine that it is your task to press a key at random time intervals. Because the key presses are otherwise not subject to any rules, the probability of a key press occurring at a given time is independent of whether a key press occurred shortly before. The so-called hazard function, that is the probability with which an event occurs at time <emph>T</emph> under the condition that it has not yet occurred before, is therefore constant.</p> <p>Transferred to word reading, this would mean that the probability of having to wait 10 ms until the correct response would be just as high immediately after the item presentation as after 600 ms or after 3 hours. However, word reading (just as any other cognitive task) is not a random but a deliberate intellectual process that takes a certain processing time to come to a certain conclusion at the end. Thus, the probability that a correct answer occurs 10 ms after the item presentation is much lower than the probability that the correct answer occurs within the next 10 ms if 600 ms of processing time have already elapsed. To put it differently: Within each item, the hazard function usually increases continuously, that is, the correct answers are not evenly distributed across the whole spectrum of possible response times, but they accumulate at the end of the intellectual processing time. Thus, the variance of the response times is restricted, which leads to the fact that the number of correct responses within a larger time interval is also strongly restricted, i.e. underdispersion occurs.</p> <hd id="AN0187097618-6">Aim of the current study</hd> <p>The overall target of the current study was to establish a flexible IRT model to capture speed as well as ability components as different facets of reading comprehension, that is, under the umbrella of one single latent variable. Specifically, we aimed to test the applicability of the Poisson count framework on a specific test of reading comprehension, namely, the ELFE II reading comprehension test (Lenhard, Lenhard, et al., [<reflink idref="bib39" id="ref61">39</reflink>]), to assess the effects of subtests with different speededness and to capture mode effects that affect working styles of the test takers. As we will later on describe in greater depth, this test measures reading comprehension at word, sentence, and text level using one subtest for each of these levels.</p> <p>The first question was whether a simple RPCM model would suffice, or whether the data would be better represented by a more complex CMPCM model, that is, a model in which deviations from equidispersion are considered.</p> <p>Second, in many psychometric tests, items are ordered according to item difficulty with more difficult items at the end of the test. Consequently, one would suppose that the number of correct answers per time unit decreases during these subtests. With respect to the Poisson count framework, this would lead to overdispersion. In contrast, in scales with mainly easy items and a high speededness, underdispersion should occur. In summary, different features of cognitive performance tests can lead to either overdispersion or underdispersion. We thus expected that the dispersions would be different in the different subtests. In particular, we expected a lower dispersion parameter for the subtest on the word level than for the subtest on the text level because of the increasing item difficulties that occur mainly in the latter. For the subtext on sentence level, we expected a position between these two extremes, that is, an intermediate dispersion. As a result, we assumed that a CMPCM with a subtest-specific dispersion parameter would fit the data better than a CMPCM with a constant dispersion parameter.</p> <p>Third, we analyzed the effects of moderators via CMPCM, specifically whether and how the presentation mode would affect the model parameters. ELFE II can be administered as a PP test or as a computer-based (CB) test. Previous research (Lenhard, Schroeders, et al., 2017) has shown that the presentation mode affects the raw score distributions. Relevant differences mainly showed up on word level with test takers producing higher raw scores in the CB mode. Moreover, different kinds of analyses also demonstrated higher error rates (which are not the focus of the current work) in the CB mode for all subtests, that is, for all complexity levels of reading comprehension. This result was also the most pronounced and most stable across grades for the word subtest. Consequently, reliability estimates were slightly lower in the CB as compared to the PP mode. Therefore, we expected the Poisson count models to replicate this reliability difference and a more pronounced underdispersion in the word level subtest when working on screen, since in this mode, the speed effect is considerably higher.</p> <hd id="AN0187097618-7">Method</hd> <p></p> <hd id="AN0187097618-8">Participants</hd> <p>The sample in this study was comprised of the norming sample of the ELFE II reading comprehension test (Lenhard, Lenhard, et al., [<reflink idref="bib39" id="ref62">39</reflink>]). The data collection of the norming study took place in 2015 and 2016 and included 5,073 individual cases, from which a representative norming sample of <emph>N</emph><bold></bold>=<bold></bold>3,720 participants was stratified and eventually entered into this analysis. The test consists of two parallel versions with <emph>n</emph><bold></bold>=<bold></bold>1,637 completing the CB test and <emph>n</emph><bold></bold>=<bold></bold>2,047 the PP version. The test can be administered from the end of grade one to the beginning of grade seven. The age of the cohort ranged from 5.5 to 14.73 with <emph>M</emph><bold></bold>=<bold></bold>10.1 (<emph>SD</emph><bold></bold>=<bold></bold>1.66). The sample included 49% boys. A total of 3,135 of the participants were monolingual German speakers, 881 were bilingual, and 237 grew up with a family language other than German.</p> <p>The parents of all children who participated in the study provided written informed consent. All participants were informed in advance about the purpose and the process of the study. Participation was voluntary. The data was collected anonymously. The study was approved on 23 April, 2015 by the School Ethics Board of the government of the district of Lower Franconia.</p> <hd id="AN0187097618-9">The ELFE II reading comprehension test</hd> <p>ELFE II has three different subtests, which assess reading comprehension at word, sentence, and text levels. The word subtest mainly focuses on reading fluency and therefore has a strong speed component but only low difficulty, whereas the text subtest concentrates on the accuracy of the responses with only a marginal speed component. The sentence subtest is a balanced mixture of both speed and accuracy. All three subtests are administered with time cutoff criteria and are, therefore, viable candidates for the application of the Poisson count framework. The subtest on word level includes 76 items in the form of a picture word-matching task with four-word alternatives for one image. It is administered with a time limit of three minutes. In this subtest, all items are relatively easy, with a range in item difficulty from <emph>min</emph> =.72 to <emph>max</emph> =.99 (<emph>M</emph> =.94, <emph>SD</emph> =.05) in the pilot sample, i.e., there is not much difference in item difficulty. To put it in other words: The word comprehension subtest is almost a pure speed test. The retest correlation (pooled over both administration modes) after four weeks is <emph>r<subs>tt</subs></emph> =.85 (partial correlation, controlled for schooling duration) and the odd even split half correlation is <emph>r<subs>tt</subs></emph> =.95. The test on sentence level includes 36 items in the form of a gap fill task with five alternatives and is as well administered with a time limit of three minutes. The variance of item difficulties is larger than in the word comprehension subtest, with a range in item difficulty from <emph>min</emph> =.47 to <emph>max</emph> =.95 (<emph>M</emph> =.85, <emph>SD</emph> =.10). The items are presented with increasing item difficulty. The odd even split half reliability for this subtest amounts to <emph>r<subs>tt</subs></emph> =.94 with a retest reliability of <emph>r<subs>tt</subs></emph> =.91 after four weeks. Finally, the test on text level consists of 26 items, consisting of a short text, a question prompt and four alternatives and the time limit is seven minutes. The variance in item difficulty is larger than in the sentence comprehension subtest. Items are also presented with increasing difficulty, with a range in item difficulty from <emph>min</emph> =.40 to <emph>max</emph> =.92 (<emph>M</emph> =.70, <emph>SD</emph> =.13). Odd even split half reliability for this subtest amounts to <emph>r<subs>tt</subs></emph> =.88 with a retest reliability of <emph>r<subs>tt</subs></emph> =.85.</p> <hd id="AN0187097618-10">Analytic strategy</hd> <p>Importantly, the vote counts cannot be applied to one single item but only to the subtest as a whole. However, it has become established practice to use the usual IRT nomenclature when applying Poisson count models and to speak of item easiness and item dispersion when referencing count data (e.g., Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref63">19</reflink>]). Here, we will maintain this common nomenclature but we would like to clarify that the terms <emph>item easiness</emph> and <emph>item dispersion</emph> always refer to one whole subtest in the following.</p> <p>The covariate-adjusted frequency plot (CAFP; Holling et al., [<reflink idref="bib27" id="ref64">27</reflink>], [<reflink idref="bib28" id="ref65">28</reflink>]) was used to assess absolute model fit of the various count data IRT models. The CAFP is a graphical tool that compares the observed counts and the counts as implied by a given model. The model-implied frequency counts of possible values should be fairly close to the empirical frequency counts if the data fit reasonably well to a given model.</p> <p>Relative model fit was examined by means of likelihood ratio tests (Forthmann, Grotjahn, et al., [<reflink idref="bib18" id="ref66">18</reflink>]; Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref67">19</reflink>]). All tested candidate models in this work are nested in each other. Hence, the RPCM is compared to the CMPCM with a global dispersion parameter. Next, the CMPCM with global dispersion parameter is compared to the most general model, the CMPCM with item-specific dispersion parameters. To further take model parsimony into account, we compared all models based on the Akaike Information Criterion (AIC; Akaike, [<reflink idref="bib1" id="ref68">1</reflink>]) and the Bayesian Information Criterion (BIC; Schwarz, [<reflink idref="bib48" id="ref69">48</reflink>]). Smaller values on information criteria imply better model fit while at the same time a penalty for a model's number of parameters is considered.</p> <p>Finally, we analyzed mode effects through measurement invariance testing of the CMPCM parameter estimates within a multi-group model (PP vs. CB). Please note, however, that this is not identical with measurement invariance in the context of multi-group confirmatory factor analysis sensu Meredith ([<reflink idref="bib41" id="ref70">41</reflink>]). The empirical reliability was estimated as a summary statistic to inform about the measurement precision of ability estimates (Brown & Croudace, [<reflink idref="bib4" id="ref71">4</reflink>]; Green et al., [<reflink idref="bib24" id="ref72">24</reflink>]). Empirical reliability is based on the estimated trait variance</p> <p>Graph</p> <p> <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mrow><mrow><mrow><mrow><mover accent="true"><mrow><mi>σ</mi></mrow><mo>̂</mo></mover></mrow></mrow></mrow></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow></math> </ephtml> and the squared average standard error</p> <p>Graph</p> <p> <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow><msubsup><mrow><mrow><mrow><mrow><mover accent="false"><mrow><mtext mathvariant="italic">SE</mtext></mrow><mo>¯</mo></mover></mrow></mrow></mrow></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow></math> </ephtml> of the ability estimates:</p> <p>Graph</p> <p> <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mtext mathvariant="italic">Rel</mtext><mo stretchy="true">(</mo><mrow><mi>θ</mi></mrow><mo stretchy="true">)</mo><mo>=</mo><mn>1</mn><mo>‐</mo><mrow><msubsup><mrow><mrow><mrow><mrow><mover accent="false"><mrow><mtext mathvariant="italic">SE</mtext></mrow><mo>¯</mo></mover></mrow></mrow></mrow></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow><mo>/</mo><mrow><msubsup><mrow><mrow><mrow><mrow><mover accent="true"><mrow><mi>σ</mi></mrow><mo>̂</mo></mover></mrow></mrow></mrow></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow><mo>.</mo></math> </ephtml> (<reflink idref="bib1" id="ref73">1</reflink>)</p> <p>Empirical reliability estimates the squared correlation between the observed and true ability estimates (Brown & Maydeu-Olivares, [<reflink idref="bib5" id="ref74">5</reflink>]).</p> <p>All models were estimated by means of the package glmmTMB (Brooks et al., [<reflink idref="bib3" id="ref75">3</reflink>]) for the statistical software R (R Core Team, [<reflink idref="bib45" id="ref76">45</reflink>]). The glmmTMB package applies marginal maximum likelihood estimation for the ability distribution (De Boeck et al., [<reflink idref="bib12" id="ref77">12</reflink>]). For identification purposes, the mean of the ability distribution was fixed to zero (Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref78">19</reflink>]). Ability estimates are based on the <emph>maximum a posteriori</emph> (MAP) estimator. Standard errors of the MAP estimates were calculated as implemented in the glmmTMB package (Brooks et al., [<reflink idref="bib3" id="ref79">3</reflink>]). All scripts necessary to run the analyses are openly available at https://t.ly/FS7os. The raw data are available upon request to the second author.</p> <hd id="AN0187097618-11">Results</hd> <p>The CAFP in Figure 1 shows the results in terms of model-data fit. As expected, the distribution of observed frequencies of counts was bimodal. This results from differing complexity of the tasks on word, sentence and text level. Consequently, the observed frequencies of counts were more strongly influenced by word, compared to sentence, compared to text level. In the modeling, however, this imbalance is addressed by controlling for time-on-task per item and modeling easiness parameters for the different subtests. This modeling approach worked reasonably well because all models captured this general shape of the distribution. However, the CMPCM with item-specific dispersion parameters (green dot-dashed line) yielded model-implied count frequencies closer to the observed frequencies as compared to the RPCM and the CMPCM with global dispersion parameters. The latter two models were indistinguishable in terms of model-data fit in the CAFP (see Figure 1). Overall, the data fitted large portions of the CMPCM with item-specific dispersion parameters well.</p> <p>PHOTO (COLOR): Figure 1. Covariate-adjusted frequency plot for the RPCM, CMPCM with global dispersion (CMPCM-G), and CMPCM with item-specific dispersion (CMPCM). Model-implied counts are depicted by blue, red, or green lines and they should be as close as possible to the observed frequency counts to indicate model fit (Holling et al., [<reflink idref="bib27" id="ref80">27</reflink>]).</p> <p>Furthermore, the CMPCM with global dispersion parameter improved the fit between model and data when compared to the RPCM. This improvement was demonstrated by the results of the likelihood ratio test, AIC, and BIC. It is worth noting that the estimated dispersion parameter was slightly smaller than zero, although still significant. It is also important to mention that the model-data fit for these two models was indistinguishable, as shown in Figure 1. However, allowing the item dispersion parameters to vary across word, sentence, and text scores significantly improved the model-data fit, as shown by all model comparison statistics (see Table 1).</p> <p>Table 1. Model estimation results for RPCM and CMPCMs.</p> <p> <ephtml> <table><thead><tr><td /><td /><td>RPCM</td><td>CMPCM with global dispersion</td><td>CMPCM with item-specific dispersion</td></tr></thead><tbody valign="top"><tr><td>Fixed effects</td><td><italic>ω</italic></td><td><italic>β</italic> (<p><graphic href="vjxe_a_2367162_ilm0006.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mtext mathvariant="italic">SE</mtext></mrow><mrow><mi>β</mi></mrow></msub></mrow></math></p>)</td><td><italic>β</italic> (<p><graphic href="vjxe_a_2367162_ilm0007.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mtext mathvariant="italic">SE</mtext></mrow><mrow><mi>β</mi></mrow></msub></mrow></math></p>)</td><td><italic>β</italic> (<p><graphic href="vjxe_a_2367162_ilm0008.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mtext mathvariant="italic">SE</mtext></mrow><mrow><mi>β</mi></mrow></msub></mrow></math></p>)</td></tr><tr><td> Words</td><td char=".">1.099</td><td char="(">2.731 (0.008)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">2.730 (0.008)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">2.735 (0.007)<xref ref-type="table-fn" rid="tfn2">***</xref></td></tr><tr><td> Sentences</td><td char=".">1.099</td><td char="(">1.864 (0.008)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">1.863 (0.008)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">1.868 (0.008)<xref ref-type="table-fn" rid="tfn2">***</xref></td></tr><tr><td> Texts</td><td char=".">1.946</td><td char="(">0.572 (0.009)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">0.571 (0.008)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">0.575 (0.009)<xref ref-type="table-fn" rid="tfn2">***</xref></td></tr><tr><td>Random effects</td><td /><td /><td /><td /></tr><tr><td><p><graphic href="vjxe_a_2367162_ilm0009.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msubsup><mrow><mrow><mrow><mrow><mover accent="true"><mrow><mi>σ</mi></mrow><mo>̂</mo></mover></mrow></mrow></mrow></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow></math></p></td><td /><td char="(">0.193</td><td char="(">0.196</td><td char="(">0.186</td></tr><tr><td><p><graphic href="vjxe_a_2367162_ilm0010.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msubsup><mrow><mrow><mrow><mrow><mover accent="false"><mrow><mtext mathvariant="italic">SE</mtext></mrow><mo>¯</mo></mover></mrow></mrow></mrow></mrow><mrow><mi>θ</mi></mrow><mrow><mn>2</mn></mrow></msubsup></mrow></math></p></td><td /><td char="(">0.013</td><td char="(">0.012</td><td char="(">0.009</td></tr><tr><td> Empirical reliability</td><td /><td char="(">0.932</td><td char="(">0.938</td><td char="(">0.951</td></tr><tr><td>Dispersion</td><td /><td><italic>τ</italic></td><td><italic>τ</italic> (<p><graphic href="vjxe_a_2367162_ilm0011.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mtext mathvariant="italic">SE</mtext></mrow><mrow><mi>τ</mi></mrow></msub></mrow></math></p>)</td><td><italic>τ</italic> (<p><graphic href="vjxe_a_2367162_ilm0012.gif" content-type="Graph" /><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><mrow xmlns=""><msub><mrow><mtext mathvariant="italic">SE</mtext></mrow><mrow><mi>τ</mi></mrow></msub></mrow></math></p>)</td></tr><tr><td> Global</td><td /><td>0</td><td char="(">−0.091 (0.017)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td /></tr><tr><td> Words</td><td /><td /><td /><td char="(">−0.593 (0.083)<xref ref-type="table-fn" rid="tfn2">***</xref></td></tr><tr><td> Sentences</td><td /><td /><td /><td char="(">−0.152 (0.033)<xref ref-type="table-fn" rid="tfn2">**</xref>*</td></tr><tr><td> Texts</td><td /><td /><td /><td char="(">0.200 (0.030)<xref ref-type="table-fn" rid="tfn2">***</xref></td></tr><tr><td>Model comparison</td><td /><td /><td /><td /></tr><tr><td> Δχ<sup>2</sup>(<italic>df</italic>)</td><td /><td>–</td><td char="(">27.66 (1)<xref ref-type="table-fn" rid="tfn2">***</xref></td><td char="(">143.95 (2)<xref ref-type="table-fn" rid="tfn2">***</xref></td></tr><tr><td> AIC</td><td /><td>75,722</td><td char="(">75,696</td><td char="("><bold>75,556</bold></td></tr><tr><td> BIC</td><td /><td>75,751</td><td char="(">75,733</td><td char="("><bold>75,608</bold></td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Notes. N</emph> = 3720 (11,160 observations); τ = dispersion parameter (CMPCMs: <emph>τ</emph> < 0 indicates underdispersion; <emph>τ</emph> = 0 indicates equidispersion; and <emph>τ</emph> > 0 indicates overdispersion). Likelihood ratio tests for CMPCMs: CMPCM with global dispersion is compared with RPCM and CMPCM with item-specific dispersion is compared with CMPCM with global dispersion. Δχ<sups>2</sups> = likelihood ratio statistic; AIC = Akaike's Information Criterion; BIC = Bayesian Information Criterion. Lower values of information criteria imply better model fit when also model parsimony is taken into account. BIC values model parsimony more than AIC. Values in bold indicate the best-fitting model based on AIC and BIC, respectively.</p> <p>2 ***<emph>p</emph> <.001.</p> <p>As could be expected, the word score had the highest item easiness followed by the sentence score, whereas the text items were the most difficult (see Table 1). Importantly, these item easiness estimates were controlled for varying time-on-task. The item easiness parameter estimates were highly comparable across all models. Specifically, for the best fitting model, the word-level easiness estimate of 2.731 implies that for a child of average ability (i.e., θ = 0) one would expect exp(2.731) ≈ 15 correctly solved items within one minute, whereas the sentence-level (exp(1.868) ≈ 6) and text-level estimates (exp(1.868) ≈ 2) imply much lower numbers of correctly solved items for average ability and 1 min time on task. The conditional distributions of the word and sentence score were underdispersed. Underdispersion was stronger for word scores as compared to sentence scores, while the conditional distributions of the text score were found to be overdispersed. This implies that local reliability was highest for the word-level subtest and lowest for the text-level subtest.</p> <p>In particular, the underdispersion observed for the sentence-level subtest and the overdispersion observed for the text-level subtest were very similar in their deviation from equidispersion. This implies further that the gain in empirical reliability for the CMPCM with item-specific dispersion can be attributed to the comparably stronger underdispersion for the word-level subtest. The standard deviation of the ability parameters was estimated to be 0.186 which means that 95% of the ability estimates are to be expected in the interval from −0.365 to 0.365 (i.e., 0 ± 1.96 × 0.186).</p> <hd id="AN0187097618-12">Measurement invariance testing</hd> <p>In both groups (PP vs. CB), the CMPCM with item-specific dispersion parameters fit best (see Table 2). In a next step, the CMPCM with item-specific dispersion was fit as a multi-group model that allowed group-specific easiness and dispersion parameters for each ELFE II sub-score. Ability variance was also estimated in each of the groups. This model fitted better as compared to a model that constrained item-easiness parameters to be equal (strong invariance; see Table 3). Further fixing item-dispersion parameters to be equal across groups (strict invariance) worsened model fit (see Table 3). Hence, ELFE II displayed configural measurement invariance.</p> <p>Table 2. Model comparisons for both groups (CB and PP version of ELFE II).</p> <p> <ephtml> <table><thead><tr><td /><td>CB</td><td>PP</td></tr><tr><td /><td>RPCM</td><td>CMPCM with global dispersion</td><td>CMPCM with item-specific dispersion</td><td>RPCM</td><td>CMPCM with global dispersion</td><td>CMPCM with item-specific dispersion</td></tr></thead><tbody valign="top"><tr><td>AIC</td><td>33,190</td><td>33,024</td><td><bold>32,813</bold></td><td>42,058</td><td>42,057</td><td><bold>41,953</bold></td></tr><tr><td>BIC</td><td>33,216</td><td>33,056</td><td><bold>32,858</bold></td><td>42,085</td><td>42,090</td><td><bold>42,000</bold></td></tr><tr><td>Δχ<sup>2</sup>(<italic>df</italic>)</td><td>–</td><td>168.18 (1)***</td><td><bold>214.88 (2)***</bold></td><td>–</td><td>3.45 (1)</td><td><bold>107.83 (2)***</bold></td></tr></tbody></table> </ephtml> </p> <p>3 ***p <.001. <emph>Notes</emph>. The best-fitting model is depicted in bold font.</p> <p>Table 3. Measurement invariance testing across test versions.</p> <p> <ephtml> <table><thead><tr><td /><td>Level of Invariance</td><td /><td /></tr></thead><tbody valign="top"><tr><td /><td>Configural</td><td>Strict</td><td>Strong</td></tr><tr><td>AIC</td><td><bold>74,766</bold></td><td>74,940</td><td>75,140</td></tr><tr><td>BIC</td><td><bold>74,868</bold></td><td>75,028</td><td>75,206</td></tr><tr><td>Δχ<sup>2</sup>(<italic>df</italic>)</td><td>–</td><td>178.37 (2)***</td><td>206.36 (3)***</td></tr></tbody></table> </ephtml> </p> <p>4 <emph>Notes</emph>. The best-fitting model is depicted in bold font.</p> <p>To further explore the practical consequences of the models demonstrating non-invariance we refitted the configural model in a slightly different parameterization with sum-to-zero contrasts for the items (e.g., word level was coded −1, sentence level 1, and text level 0 for the first contrast variable) and the interaction between related contrast variables and group (i.e., CB vs. PP). This model allowed to check for the overall difference between groups with respect to average item-easiness on performance. We found that average item-easiness for the PP version of ELFE II is expected to reduce by a factor of 0.85, 95% CI: [0.83, 0.88]. Hence and in accordance with previous research, there seems to be a slight performance advantage for the group who took the CB mode. All interaction terms (all <emph>p</emph>s <.004) and the tested contrast for the deviation of the word score (<emph>p</emph> <.001) were significant. As illustrated in Figure 2, the difference between both administration modes for sentence and text scores was less pronounced as compared to the overall difference in average item easiness. For the word level, however, the difference between administration modes was even more pronounced.</p> <p>Graph: Figure 2. For each ELFE II score, the difference in the average easiness in each of the groups is plotted (centered for both groups). 95% confidence ellipses are plotted around the item parameters: If they cover the reference line, parameters do not significantly differ across groups. Top-left: All items in one plot. Top-right: Plot for Word score only. Bottom-left: Plot for Sentence score only. Bottom-right: Plot for Text score only.</p> <p>Beyond item-easiness, also dispersion parameters had to be group-specific. All dispersion parameters were smaller for the CB mode (word: −4.00, 95%CI: [−4.45, −3.55]; sentence: −0.17, 95%CI: [−0.24, −0.10]; text: 0.11 [0.04, 0.18]) as compared to the PP mode (word: −0.47, 95%CI: [−0.67, −0.27]; sentence: −0.06, 95%CI: [−0.15, 0.03]; text: 0.37 [0.29, 0.45]). In line with the hypothesis, the strongest dispersion difference between CB and PP mode was found for the word score. The dispersion difference for the sentence score is further noteworthy because the shift in dispersion implied that the score was underdispersed for the CB mode but a Poisson item for the PP mode. Overall, the pattern implies a strong tendency toward underdispersion for the CB mode and a rather mixed pattern of dispersion parameters for the PP mode with underdispersion at the word level, equidispersion at the sentence level, and overdispersion at the text level. The dispersion differences between both administration modes further imply differences in empirical reliability estimates. Based on the dispersion findings of this analysis and in contrast to subtest reliability on manifest data level, one would expect the CB mode to be more reliable as compared to the PP mode. This was indeed found in the data (CB: 0.996; PP: 0.952), with excellent reliability estimates in both modes. This somehow contrasts with the results from retest reliability estimates: While in the PP version, the findings are in line with the reliability of the test reported in the manual, the CMPCM estimated reliability for the computer-based version is considerably higher and reaches unrealistically high levels.</p> <hd id="AN0187097618-13">Discussion</hd> <p>In the present study, we investigated the fit of CMPCMs with varying flexibility for modeling the raw score distributions of a reading test. The statistical approach was suited to successfully apply the Item Response Theory framework to model data, which is based on a test with a time cutoff per scale and which features mixed accuracy and speed components. The three subtests measure reading comprehension at the word, sentence, and text levels and differ in their respective speed component and difficulties. To this aim, we evaluated the relative model fit of the classical RPCM, a CMPCM with a constant dispersion parameter, and a CMPCM with an item-specific dispersion parameter. The CMPCM with item-specific dispersion parameter showed the best model fit and was thus most flexible in modeling the data.</p> <p>In addition, we examined mode effects on item easiness and dispersion parameters in the context of measurement invariance testing. As expected, underdispersion for the word level test was stronger in computer-based testing compared to PP testing, as children showed a higher working speed in this subtest on computers. In the following, we will discuss our main findings and provide some considerations about what we can learn from the model analyses regarding mode effects in reading before we give an outlook on future research concerning the modeling of speeded assessments.</p> <hd id="AN0187097618-14">Different speed components require subtest-specific dispersion parameters</hd> <p>Our comparison of the Poisson counts models with varying flexibility showed that the RPCM and the CMPCM with a constant dispersion parameter fitted the data worse compared to the CMPCM with subtest-specific dispersion parameters, which was in line with our assumptions. Item easiness was highest for the word subtest, followed by the sentence subtest, while item easiness was lowest for the text subtest. In turn, the dispersion parameter for the subtest at the word level was lowest, followed by the dispersion parameters at the sentence level, which in turn was lower than that at the text level, thus coherently reflecting the speed factor of each scale. Importantly, while the conditional distributions of the word and sentence scores were underdispersed with stronger underdispersion for the word scores, the conditional distributions of the text score were overdispersed. Thus, the CMPCM with item-specific dispersion parameters played out its advantage to flexibly model under- and overdispersion. At the same time, the dispersion patterns can be interpreted to reflect the relative differences between the subtests in speededness and can thus be used to characterize subscales within a larger test battery.</p> <p>Overall, the CMPCM with item-specific dispersion parameters corresponds well to the bimodal shape of the observed frequencies. The model-implied count frequencies for the first peak, however, were higher than the empirically observed frequencies, whereas the agreement between model and data for the second peak was very good. Given the different number of items for the three subtests (76 word items, 36 sentence items, and 26 text items), the second peak only includes item frequencies from the word subtest, which is characterized by high easiness and the lowest dispersion parameter. Additionally, although the CMPCM with item-specific dispersion parameters showed the best model fit of the three models, it might still be better suited to modeling classic speed tests with low average item difficulty than tests with a speed component and increasing item difficulties.</p> <p>Finally, it should be noted that we started model comparisons with the RPCM. One might as well start with a direct comprehensive analysis of the model with item-specific dispersion parameters. This model immediately allows all the inferences one wants to make about the dispersion parameters. Of course, this could also be the starting point for reducing the more complex CMPCM with item-specific dispersion to less complex models, if empirical findings allow. Here, in line with previous work using the CMPCM (Forthmann, Grotjahn, et al., [<reflink idref="bib18" id="ref81">18</reflink>]; Forthmann, Gühne, et al., [<reflink idref="bib19" id="ref82">19</reflink>]; Forthmann & Doebler, [<reflink idref="bib16" id="ref83">16</reflink>]), we have simply chosen the RPCM as a starting point for model comparison, as this approach compares a less established count data model (i.e., the CMPCM) with a more established one.</p> <hd id="AN0187097618-15">Mode effects in reading</hd> <p>The better model fit of the CMPCM was found for both presentation modes. In line with prior findings on the measurement invariance of the ELFE II, our analyses revealed that ELFE II displayed configural measurement invariance between the PP and the CB mode. Both item-easiness and item-dispersion parameters, however, varied significantly between both modes with slight performance advantages and smaller dispersion parameters for the CB mode. The overall pattern of these results might indicate that reading processes on paper and on the screen differ, which is in line with the analysis of manifest data and measurement invariance from multiple-group confirmatory analyses (Lenhard, Schroeders, et al., 2017). Thus, using CMPCM modeling, apart from the inconsistent reliability estimates, we could replicate the findings of other analysis strategies and add support for the validity of this approach, while adding results for the interpretation of under- and overdispersion on a content level of psychometric tests.</p> <hd id="AN0187097618-16">Limitations and future research</hd> <p>Since IRT models of tests with cutoffs are still not widely available, our analysis contributes to the development of approaches for this use case. Still, the drawback lies in the level of data analysis, which focuses on the results of subtest scales within a larger test battery. This means, the CMPCM is not suited for item revision and test construction on the subtest level and only comes into play once the subtests are finalized. Thus, at least in the pilot phase of the construction process, it is inevitable to collect both time-on-task and accuracy data and model them with appropriate procedures, such as log-normal IRT models (Fox et al., [<reflink idref="bib21" id="ref84">21</reflink>]), to identify poorly fitting items. The final version of the test can then be constructed based on reaction time and accuracy information, and the time cutoff can be set accordingly. In this way, it is also possible to avoid problems with ceiling effects when modeling subtests with the CMPCM (Forthmann, Grotjahn, et al., [<reflink idref="bib18" id="ref85">18</reflink>]), which complicate simple interpretations of item easiness parameters and prevent reliable assessment of the ability of individuals at the upper end of the ability distribution. Consequently, the CMPCM modeling is best combined with additional prior steps in the test development. Since the approach underlying this analysis is based on GLMM, there are many further possibilities, which could enrich the inventory of psychometric techniques. For example, we did not yet represent clustering in the data, though this is possible via the formula syntax of the underlying R packages—an information source, so far seldomly regarded in test construction.</p> <p>Regarding applied aspects of CMPCM, a downside is the complexity of the analysis when it comes to predicting person parameters for new data, as it is for example the case in individual diagnostics. It is of course possible to apply existing models to new results, but in this case, computer-based scoring is inevitable. In our opinion, the fitted total scores could be used for further steps like norming and the same is true for predicted scores of new data. So far, however, there are no established guidelines on fit indices for test construction based on CMPCM. Accordingly, practical recommendations still have to be developed.</p> <hd id="AN0187097618-17">Conclusion</hd> <p>Our findings highlight the call for a count data modeling approach that flexibly handles all possible variants of dispersion. The method demonstrated in this study helps to close a gap within IRT-based test construction, namely the modeling of speeded tests with time cutoffs.</p> <hd id="AN0187097618-18">Disclosure statement</hd> <p>Two of the authors were involved in the development of the ELFE II reading comprehension test and receive royalties. This analysis is not influenced by this since the raw data simply serve to demonstrate a new modeling approach unrelated to the original test. The scripts are available via https://t.ly/FS7os. The raw data are available on request to the second author.</p> <ref id="AN0187097618-19"> <title> References </title> <blist> <bibl id="bib1" idref="ref38" type="bt">1</bibl> <bibtext> Akaike, H. (1973). Information theory and an extension of the maximum likelihood principle. In B. N. Petrov & F. Csáki (Eds.), 2nd International Symposium on Information Theory (pp. 267 – 281). Akadémiai Kiadó.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref32" type="bt">2</bibl> <bibtext> Baghaei, P., Ravand, H., & Nadri, M. (2019). Is the d2 test of attention Rasch scalable? Analysis with the Rasch Poisson counts model. Perceptual and Motor Skills, 126 (1), 70 – 86. https://doi.org/10.1177/0031512518812183</bibtext> </blist> <blist> <bibl id="bib3" idref="ref51" type="bt">3</bibl> <bibtext> Brooks, M. E., Kristensen, K., Benthem, K. J. v., Magnusson, A., Berg, C. W., Nielsen, A., Skaug, H. J., Mächler, M., & Bolker, B. M. (2017). GlmmTMB balances speed and flexibility among packages for zero-inflated generalized linear mixed modeling. The R Journal, 9 (2), 378. https://doi.org/10.32614/RJ-2017-066</bibtext> </blist> <blist> <bibl id="bib4" idref="ref71" type="bt">4</bibl> <bibtext> Brown, A., & Croudace, T. J. (2015). Scoring and estimating score precision using multidimensional IRT models. In S. P. Reise & D. A. Revicki (Eds.), Multivariate applications series. Handbook of item response theory modeling: Applications to typical performance assessment (pp. 307 – 333). Routledge/Taylor & Francis Group.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref74" type="bt">5</bibl> <bibtext> Brown, A., & Maydeu-Olivares, A. (2011). Item response modeling of forced-choice questionnaires. Educational and Psychological Measurement, 71 (3), 460 – 502. https://doi.org/10.1177/0013164410375112</bibtext> </blist> <blist> <bibl id="bib6" idref="ref25" type="bt">6</bibl> <bibtext> Cain, K., Oakhill, J., & Bryant, P. (2004). Children's reading comprehension ability: Concurrent predicition by working memory, verbal ability, and component skills. Journal of Educational Psychology, 96 (1), 31 – 42. https://doi.org/10.1037/0022-0663.96.1.31</bibtext> </blist> <blist> <bibl id="bib7" idref="ref26" type="bt">7</bibl> <bibtext> Cain, K., Oakhill, J., & Lemmon, K. (2004). Individual differences in the inference of word meanings from context: The influence of reading comprehension, vocabulary knowledge, and memory capacity. Journal of Educational Psychology, 96 (4), 671 – 681. https://doi.org/10.1037/0022-0663.96.4.671</bibtext> </blist> <blist> <bibl id="bib8" idref="ref11" type="bt">8</bibl> <bibtext> Coltheart, M., Rastle, K., Perry, C., Langdon, R., & Ziegler, J. (2001). DRC: A dual route cascaded model of visual word recognition and reading aloud. Psychological Review, 108 (1), 204 – 256. https://doi.org/10.1037/0033-295X.108.1.204</bibtext> </blist> <blist> <bibl id="bib9" idref="ref44" type="bt">9</bibl> <bibtext> Consul, P. C., & Famoye, F. (1992). Generalized Poisson regression model. Communications in Statistics - Theory and Methods, 21 (1), 89 – 109. https://doi.org/10.1080/03610929208830766</bibtext> </blist> <blist> <bibtext> Czado, C., Erhardt, V., Min, A., & Wagner, S. (2007). Zero-inflated generalized Poisson models with regression effects on the mean, dispersion and zero-inflation level applied to patent outsourcing rates. Statistical Modelling, 7 (2), 125 – 153. https://doi.org/10.1177/1471082X0700700202</bibtext> </blist> <blist> <bibtext> De Ayala, R. J. (2009). The theory and practice of item response theory. Guilford Publications.</bibtext> </blist> <blist> <bibtext> De Boeck, P., Bakker, M., Zwitser, R., Nivard, M., Hofman, A., Tuerlinckx, F., & Partchev, I. (2011). The Estimation of Item Response Models with the lmer function from the lme4 Package in R. Journal of Statistical Software, 39 (12), 1 – 28. https://doi.org/10.18637/jss.v039.i12</bibtext> </blist> <blist> <bibtext> Doebler, A., Doebler, P., & Holling, H. (2014). A latent ability model for count data and application to processing speed. Applied Psychological Measurement, 38 (8), 587 – 598. https://doi.org/10.1177/0146621614543513</bibtext> </blist> <blist> <bibtext> Ehri, L. C. (1995). Stages of development in learning to read words by sight. Journal of Research in Reading, 18 (2), 116 – 125. https://doi.org/10.1111/j.1467-9817.1995.tb00077.x</bibtext> </blist> <blist> <bibtext> Faddy, M. J., & Bosch, R. J. (2001). Likelihood-based modeling and analysis of data underdispersed relative to the Poisson distribution. Biometrics, 57 (2), 620 – 624. https://doi.org/10.1111/j.0006-341X.2001.00620.x</bibtext> </blist> <blist> <bibtext> Forthmann, B., & Doebler, P. (2021). Reliability of researcher capacity estimates and count data dispersion: A comparison of Poisson, negative binomial, and Conway-Maxwell-Poisson models. Scientometrics, 126 (4), 3337 – 3354. https://doi.org/10.1007/s11192-021-03864-8</bibtext> </blist> <blist> <bibtext> Forthmann, B., Gerwig, A., Holling, H., Çelik, P., Storme, M., & Lubart, T. (2016). The be-creative effect in divergent thinking: The interplay of instruction and object frequency. Intelligence, 57, 25 – 32. https://doi.org/10.1016/j.intell.2016.03.00</bibtext> </blist> <blist> <bibtext> Forthmann, B., Grotjahn, R., Doebler, P., & Baghaei, P. (2020). A comparison of different item response theory models for scaling speeded C-tests. Journal of Psychoeducational Assessment, 38 (6), 692 – 705. https://doi.org/10.1177/0734282919889262</bibtext> </blist> <blist> <bibtext> Forthmann, B., Gühne, D., & Doebler, P. (2020). Revisiting dispersion in count data item response theory models: The Conway–Maxwell–Poisson counts model. British Journal of Mathematical and Statistical Psychology, 73 (S1), 32 – 50. https://doi.org/10.1111/bmsp.12184</bibtext> </blist> <blist> <bibtext> Fox, J.-P., Klein Entink, R., & van der Linden, W. (2007). Modeling of responses and response times with the Package cirt. Journal of Statistical Software, 20 (7), 1 – 14. https://doi.org/10.18637/jss.v020.i07</bibtext> </blist> <blist> <bibtext> Fox, J. P., Klotzke, K., & Simsek, A. S. (2023). R-package LNIRT for joint modeling of response accuracy and times. PeerJ Computer Science, 9, e1232. https://doi.org/10.7717/peerj-cs.1232</bibtext> </blist> <blist> <bibtext> Fuchs, L. S., Fuchs, D., Hosp, M. K., & Jenkins, J. R. (2001). Oral reading fluency as an indicator of reading competence: A theoretical, empirical, and historical analysis. Scientific Studies of Reading, 5 (3), 239 – 256. https://doi.org/10.1207/S1532799XSSR0503_3</bibtext> </blist> <blist> <bibtext> Goldhammer, F. (2015). Measuring ability, speed, or both? Challenges, psychometric solutions, and what can be gained from experimental control. Measurement: Interdisciplinary Research and Perspectives, 13 (3-4), 133 – 164. https://doi.org/10.1080/15366367.2015.1100020</bibtext> </blist> <blist> <bibtext> Green, B. F., Bock, R. D., Humphreys, L. G., Linn, R. L., & Reckase, M. D. (1984). Technical guidelines for assessing computerized adaptive tests. Journal of Educational Measurement, 21 (4), 347 – 360. https://doi.org/10.1111/j.1745-3984.1984.tb01039.x</bibtext> </blist> <blist> <bibtext> Guikema, S. D., & Coffelt, J. P. (2008). A flexible count data regression model for risk analysis. Risk Analysis: An Official Publication of the Society for Risk Analysis, 28 (1), 213 – 223. https://doi.org/10.1111/j.1539-6924.2008.01014.x</bibtext> </blist> <blist> <bibtext> Hilbe, J. M. (2011). Negative binomial regression. Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Holling, H., Böhning, W., & Böhning, D. (2015). The covariate-adjusted frequency plot for the Rasch Poisson counts model. Thailand Statistician, 13, 67 – 78.</bibtext> </blist> <blist> <bibtext> Holling, H., Böhning, W., Böhning, D., & Formann, A. K. (2016). The covariate-adjusted frequency plot. Statistical Methods in Medical Research, 25 (2), 902 – 916. https://doi.org/10.1177/0962280212473386</bibtext> </blist> <blist> <bibtext> Huang, A. (2017). Mean-parametrized Conway–Maxwell–Poisson regression models for dispersed counts. Statistical Modelling, 17 (6), 359 – 380. https://doi.org/10.1177/1471082X17697749</bibtext> </blist> <blist> <bibtext> Hung, L. F. (2012). A negative binomial regression model for accuracy tests. Applied Psychological Measurement, 36 (2), 88 – 103. https://doi.org/10.1177/0146621611429548</bibtext> </blist> <blist> <bibtext> Jendryczko, D., Berkemeyer, L., & Holling, H. (2020). Introducing a computerized figural memory test based on automatic item generation: An analysis with the Rasch Poisson counts model. Frontiers in Psychology, 11, 945. https://doi.org/10.3389/fpsyg.2020.00945</bibtext> </blist> <blist> <bibtext> Jenkins, J. R., Fuchs, L. S., van den Broek, P., Espin, C., & Deno, S. L. (2003). Sources of individual differences in reading comprehension and reading fluency. Journal of Educational Psychology, 95 (4), 719 – 729. https://doi.org/10.1037/0022-0663.95.4.719</bibtext> </blist> <blist> <bibtext> Kintsch, W. (1998). Comprehension: A paradigm for cognition. Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Kirschmann, N., Lenhard, W., & Suggate, S. (2021). Prediction of passage comprehension through word and sentence reading skills in a transparent language. Journal of Research in Reading, 44 (4), 817 – 836. https://doi.org/10.1111/1467-9817.12373</bibtext> </blist> <blist> <bibtext> Klauda, S. L., & Guthrie, J. T. (2008). Relationships of three components of reading fluency to reading comprehension. Journal of Educational Psychology, 100 (2), 310 – 321. https://doi.org/10.1037/0022-0663.100.2.310</bibtext> </blist> <blist> <bibtext> Klicpera, C., & Gasteiger-Klicpera, B. (1993). Psychologie der Lese- und Rechtschreibschwierigkeiten: Entwicklung, Ursachen, Förderung. Beltz.</bibtext> </blist> <blist> <bibtext> Landerl, K., Wimmer, H., & Frith, U. (1997). The impact of orthographic consistency on dyslexia: A German-English comparison. Cognition, 63, 315 – 334. https://doi.org/10.1016/S0010-0277(97)00005-X</bibtext> </blist> <blist> <bibtext> Lenhard, W., & Lenhard, A. (2020). Improvement of Norm Score Quality via regression-based continuous norming. Educational and Psychological Measurement, 81 (2), 229 – 261. https://doi.org/10.1177/0013164420928457</bibtext> </blist> <blist> <bibtext> Lenhard, W., Lenhard, A., & Schneider, W. (2017). ELFE II - Ein Leseverständnistest für Erst- bis Siebtklässer [ELFE II - A reading comprehension test for Grade one to seven]. Hogrefe.</bibtext> </blist> <blist> <bibtext> Lenhard, W., Schroeders, U., & Lenhard, A. (2017). Equivalence of screen versus print reading comprehension depends on task complexity and proficiency. Discourse Processes, 54 (5-6), 427 – 445. https://doi.org/10.1080/0163853X.2017.1319653</bibtext> </blist> <blist> <bibtext> Meredith, W. (1993). Measurement invariance, factor analysis and factorial invariance. Psychometrika, 58 (4), 525 – 543. https://doi.org/10.1007/BF02294825</bibtext> </blist> <blist> <bibtext> National Reading Panel. (2000). Teaching children to read: An evidence-based assessment of the scientific research literature on reading and its implications for reading instruction. National Institute of Child Health and Human Development.</bibtext> </blist> <blist> <bibtext> Ogasawara, H. (1996). Rasch's multiplicative Poisson model with covariates. Psychometrika, 61 (1), 73 – 92. https://doi.org/10.1007/BF02296959</bibtext> </blist> <blist> <bibtext> Pikulski, J. J., & Chard, D. J. (2005). Fluency: Bridge between decoding and reading comprehension. The Reading Teacher, 58 (6), 510 – 519. https://doi.org/10.1598/RT.58.6.2</bibtext> </blist> <blist> <bibtext> R Core Team. (2019). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. Retrieved from https://<ulink href="http://www.R-project.org/">www.R-project.org/</ulink></bibtext> </blist> <blist> <bibtext> Roskam, E. E. (1997). Models for speed and time-limited tests. In W. van der Linden & R. K. Hambleton (Eds.), Handbook of modern item response theory (pp. 187 – 208) Springer.</bibtext> </blist> <blist> <bibtext> Schindler, J., Richter, T., Isberner, M.-B., Neeb, Y., & Naumann, J. (2018). Construct validity of a process-oriented test assessing syntactic skills in German primary school children. Language Assessment Quarterly, 15 (2), 183 – 203. https://doi.org/10.1080/15434303.2018.1446142</bibtext> </blist> <blist> <bibtext> Schwarz, G. (1978). Estimating the dimension of a model. The Annals of Statistics, 6 (2), 461 – 464. https://doi.org/10.1214/aos/1176344136</bibtext> </blist> <blist> <bibtext> Sellers, K. F., & Shmueli, G. (2010). A flexible regression model for count data. The Annals of Applied Statistics, 4 (2), 943 – 961. https://doi.org/10.1214/09-AOAS306</bibtext> </blist> <blist> <bibtext> Stanovich, K. E. (1980). Toward an interactive-compensatory model of individual differences in the development of reading fluency. Reading Research Quarterly, 16 (1), 32. https://doi.org/10.2307/747348</bibtext> </blist> <blist> <bibtext> Thorndike, E. L., Bregman, E. O., Cobb, M. V., & Woodyard, E. (1926). The measurement of intelligence. Teachers College Bureau of Publications.</bibtext> </blist> <blist> <bibtext> van der Linden, W. J., Klein Entink, R. H., & Fox, J.-P. (2010). IRT parameter estimation with response times as collateral information. Applied Psychological Measurement, 34 (5), 327 – 347. https://doi.org/10.1177/0146621609349800</bibtext> </blist> <blist> <bibtext> Wimmer, H., & Goswami, U. (1994). The influence of orthographic consistency on reading development: Word recognition in English and German children. Cognition, 51, 91 – 103. https://doi.org/10.1016/0010-0277(94)90010-8</bibtext> </blist> <blist> <bibtext> Winkelmann, R. (1995). Duration dependence and dispersion in count-data models. Journal of Business & Economic Statistics, 13 (4), 467 – 474. https://doi.org/10.1080/07350015.1995.10524620</bibtext> </blist> </ref> <aug> <p>By Boris Forthmann; Wolfgang Lenhard; Alexandra Lenhard and Natalie Förster</p> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib11" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib38" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib23" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib51" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib20" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib52" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib46" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib37" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib53" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib14" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib44" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib22" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib50" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib36" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib42" firstref="ref18"></nolink> <nolink nlid="nl16" bibid="bib32" firstref="ref19"></nolink> <nolink nlid="nl17" bibid="bib21" firstref="ref21"></nolink> <nolink nlid="nl18" bibid="bib47" firstref="ref23"></nolink> <nolink nlid="nl19" bibid="bib33" firstref="ref24"></nolink> <nolink nlid="nl20" bibid="bib35" firstref="ref27"></nolink> <nolink nlid="nl21" bibid="bib34" firstref="ref28"></nolink> <nolink nlid="nl22" bibid="bib40" firstref="ref29"></nolink> <nolink nlid="nl23" bibid="bib43" firstref="ref31"></nolink> <nolink nlid="nl24" bibid="bib13" firstref="ref33"></nolink> <nolink nlid="nl25" bibid="bib27" firstref="ref34"></nolink> <nolink nlid="nl26" bibid="bib31" firstref="ref35"></nolink> <nolink nlid="nl27" bibid="bib19" firstref="ref36"></nolink> <nolink nlid="nl28" bibid="bib17" firstref="ref37"></nolink> <nolink nlid="nl29" bibid="bib26" firstref="ref39"></nolink> <nolink nlid="nl30" bibid="bib15" firstref="ref40"></nolink> <nolink nlid="nl31" bibid="bib30" firstref="ref42"></nolink> <nolink nlid="nl32" bibid="bib16" firstref="ref43"></nolink> <nolink nlid="nl33" bibid="bib25" firstref="ref45"></nolink> <nolink nlid="nl34" bibid="bib49" firstref="ref46"></nolink> <nolink nlid="nl35" bibid="bib10" firstref="ref47"></nolink> <nolink nlid="nl36" bibid="bib29" firstref="ref48"></nolink> <nolink nlid="nl37" bibid="bib18" firstref="ref56"></nolink> <nolink nlid="nl38" bibid="bib54" firstref="ref58"></nolink> <nolink nlid="nl39" bibid="bib39" firstref="ref61"></nolink> <nolink nlid="nl40" bibid="bib28" firstref="ref65"></nolink> <nolink nlid="nl41" bibid="bib48" firstref="ref69"></nolink> <nolink nlid="nl42" bibid="bib41" firstref="ref70"></nolink> <nolink nlid="nl43" bibid="bib24" firstref="ref72"></nolink> <nolink nlid="nl44" bibid="bib45" firstref="ref76"></nolink> <nolink nlid="nl45" bibid="bib12" firstref="ref77"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1501045
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Flexible Item Response Modeling for Timed Reading Comprehension Assessment
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Boris+Forthmann%22">Boris Forthmann</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9755-7304">0000-0001-9755-7304</externalLink>)<br /><searchLink fieldCode="AR" term="%22Wolfgang+Lenhard%22">Wolfgang Lenhard</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-8184-6889">0000-0002-8184-6889</externalLink>)<br /><searchLink fieldCode="AR" term="%22Alexandra+Lenhard%22">Alexandra Lenhard</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-8680-4381">0000-0001-8680-4381</externalLink>)<br /><searchLink fieldCode="AR" term="%22Natalie+Förster%22">Natalie Förster</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0634-5993">0000-0003-0634-5993</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Experimental+Education%22"><i>Journal of Experimental Education</i></searchLink>. 2025 93(4):770-786.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 17
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Item+Response+Theory%22">Item Response Theory</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Comprehension%22">Reading Comprehension</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Rate%22">Reading Rate</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Psychometrics%22">Psychometrics</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Tests%22">Reading Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Timed+Tests%22">Timed Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Raw+Scores%22">Raw Scores</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1080/00220973.2024.2367162
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0973<br />1940-0683
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: While a rich methodology for analyzing response patterns for accuracy and time-on-task is at hand via Item Response Theory (IRT), tests with time cutoffs are so far harder to handle. Given that this test mode is widely applied, especially in the context of paper-and-pencil testing, there is a lack of psychometric techniques for a relevant number of tests. In this context, the original work of Rasch and his Rasch Poisson Counts model indeed offers an approach for this scenario that is adequate to solve the problem but which leads to model violations in many cases. Recent developments in statistical modeling -- the so-called Conway Maxwell Poisson Counts Model (CMPCM) -- can solve the problem of under- and overdispersion. We apply this model to the norm data of the ELFE II reading comprehension test and analyze patterns of over- and underdispersion with regard to speededness and mode effects. CMPCM with subtest-specific dispersion was adequate to model the raw test data, with underdispersion occurring mainly in highly speeded subtests with low difficulty and overdispersion in less speeded subtests with high difficulty. Thus, the CMPCM could contribute to psychometric methodology to appropriately model tests with time cutoffs on the subtest level.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1501045
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1501045
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/00220973.2024.2367162
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 17
        StartPage: 770
    Subjects:
      – SubjectFull: Item Response Theory
        Type: general
      – SubjectFull: Reading Comprehension
        Type: general
      – SubjectFull: Reading Rate
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Psychometrics
        Type: general
      – SubjectFull: Reading Tests
        Type: general
      – SubjectFull: Timed Tests
        Type: general
      – SubjectFull: Raw Scores
        Type: general
    Titles:
      – TitleFull: Flexible Item Response Modeling for Timed Reading Comprehension Assessment
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Boris Forthmann
      – PersonEntity:
          Name:
            NameFull: Wolfgang Lenhard
      – PersonEntity:
          Name:
            NameFull: Alexandra Lenhard
      – PersonEntity:
          Name:
            NameFull: Natalie Förster
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0022-0973
            – Type: issn-electronic
              Value: 1940-0683
          Numbering:
            – Type: volume
              Value: 93
            – Type: issue
              Value: 4
          Titles:
            – TitleFull: Journal of Experimental Education
              Type: main
ResultId 1