Incorporating Test-Taking Engagement into Multistage Adaptive Testing Design for Large-Scale Assessments

Saved in:
Bibliographic Details
Title: Incorporating Test-Taking Engagement into Multistage Adaptive Testing Design for Large-Scale Assessments
Language: English
Authors: Okan Bulut (ORCID 0000-0001-5853-1267), Guher Gorgun, Hacer Karamese
Source: Journal of Educational Measurement. 2025 62(1):57-80.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 24
Publication Date: 2025
Document Type: Journal Articles
Reports - Research
Education Level: Secondary Education
Descriptors: Response Style (Tests), Testing Problems, Testing Accommodations, Measurement, Monte Carlo Methods, Learner Engagement, Reaction Time, Test Format, Achievement Tests, Foreign Countries, International Assessment, Sample Size
Assessment and Survey Identifiers: Program for International Student Assessment
DOI: 10.1111/jedm.12380
ISSN: 0022-0655
1745-3984
Abstract: The use of multistage adaptive testing (MST) has gradually increased in large-scale testing programs as MST achieves a balanced compromise between linear test design and item-level adaptive testing. MST works on the premise that each examinee gives their best effort when attempting the items, and their responses truly reflect what they know or can do. However, research shows that large-scale assessments may suffer from a lack of test-taking engagement, especially if they are low stakes. Examinees with low test-taking engagement are likely to show noneffortful responding (e.g., answering the items very rapidly without reading the item stem or response options). To alleviate the impact of noneffortful responses on the measurement accuracy of MST, test-taking engagement can be operationalized as a latent trait based on response times and incorporated into the on-the-fly module assembly procedure. To demonstrate the proposed approach, a Monte-Carlo simulation study was conducted based on item parameters from an international large-scale assessment. The results indicated that the on-the-fly module assembly considering both ability and test-taking engagement could minimize the impact of noneffortful responses, yielding more accurate ability estimates and classifications. Implications for practice and directions for future research were discussed.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1464057
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGMZnu2H8HK-REZ9U8-gG_nAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDMQrlPAb9igM-cI9OgIBEICBm4CDdoUKsYO9r8FpBiA-nqrBLaLJE7jtLOXcDBp5-vRIA8l89pJGgMvAu0iRsb7gVOcLEjv0JSTXXFeDVaVRvXq58iBfsWeMiIDf90a1PqQ6_otesQug3dOtKiVHwB5jO6O8FbuomQJIL76IblZZBmIHucLCQVu4XQ_4dpdcpe_FNWf9hdW13XFrpWa73EjYmxaWvjbKLObz784U
Text:
  Availability: 1
  Value: <anid>AN0183951184;mea01mar.25;2025Mar25.06:40;v2.2.500</anid> <title id="AN0183951184-1">Incorporating Test‐Taking Engagement into Multistage Adaptive Testing Design for Large‐Scale Assessments </title> <p>The use of multistage adaptive testing (MST) has gradually increased in large‐scale testing programs as MST achieves a balanced compromise between linear test design and item‐level adaptive testing. MST works on the premise that each examinee gives their best effort when attempting the items, and their responses truly reflect what they know or can do. However, research shows that large‐scale assessments may suffer from a lack of test‐taking engagement, especially if they are low stakes. Examinees with low test‐taking engagement are likely to show noneffortful responding (e.g., answering the items very rapidly without reading the item stem or response options). To alleviate the impact of noneffortful responses on the measurement accuracy of MST, test‐taking engagement can be operationalized as a latent trait based on response times and incorporated into the on‐the‐fly module assembly procedure. To demonstrate the proposed approach, a Monte‐Carlo simulation study was conducted based on item parameters from an international large‐scale assessment. The results indicated that the on‐the‐fly module assembly considering both ability and test‐taking engagement could minimize the impact of noneffortful responses, yielding more accurate ability estimates and classifications. Implications for practice and directions for future research were discussed.</p> <p>Large‐scale assessments play a significant role in measuring educational achievement and progress of students, schools, and entire education systems. National large‐scale assessments, such as the National Assessment of Educational Progress (NAEP) in the United States, the National Assessment Program–Literacy and Numeracy in Australia, and the Pan‐Canadian Assessment Program in Canada, are administered to a nationally representative sample of students to produce achievement data at the individual, school, and group levels (e.g., school districts, states, and provinces/territories). The results of these assessment programs inform educators, policymakers, and parents about students' learning experiences in various subject areas (e.g., reading, mathematics, and science) while enabling them to evaluate the effectiveness of their educational systems across the country (Kamens & McNeely, [<reflink idref="bib23" id="ref1">23</reflink>]; Verger et al., [<reflink idref="bib52" id="ref2">52</reflink>]).</p> <p>Similarly, international large‐scale assessments (ILSAs), such as the Organization for Economic Cooperation and Development's (OECD) Programme for International Student Assessment (PISA) and the International Association for the Evaluation of Educational Achievement's (IEA) Trends in Mathematics and Science Study (TIMSS), provide a way to compare the educational achievement and performance of students across different countries, languages, and educational systems. The results of ILSAs have been used to benchmark the performances of different education systems, allowing countries to identify areas where they need to improve and prioritize educational investments (Addey & Sellar, [<reflink idref="bib1" id="ref3">1</reflink>]; Kirsch et al., [<reflink idref="bib26" id="ref4">26</reflink>]). Also, ILSAs provide a data infrastructure for research into various topics and best practices in teaching and learning (Ertl et al., [<reflink idref="bib13" id="ref5">13</reflink>]; Hernández‐Torrano & Courtney, [<reflink idref="bib22" id="ref6">22</reflink>]).</p> <p>With the rapid evolution of computers over the last two decades, ILSAs have begun to transition from pencil‐and‐paper testing (PBT) to computer‐based assessment (CBA), enabling students to complete achievement tests and background questionnaires using a computer. For example, IEA offered a computer‐based version of the Progress in International Reading Literacy Study (PIRLS) in 2016, with the reading passages exclusively designed for this version called ePIRLS. Similarly, the OECD introduced a CBA version of PISA in the 2015 cycle, with nearly 90% of the participating countries choosing the CBA option (Shin et al., [<reflink idref="bib46" id="ref7">46</reflink>]). In addition to offering a wide range of benefits, such as increased test security and interactive item formats, the long‐awaited shift from PBT to CBA in ILSAs has also enabled the possibility of harnessing adaptive testing in large‐scale assessments. The 2012 cycle of PIAAC (Kirsch & Lennon, [<reflink idref="bib25" id="ref8">25</reflink>]), followed by the 2018 cycle of PISA (Yamamoto et al., [<reflink idref="bib62" id="ref9">62</reflink>]), implemented a multistage adaptive testing (MST) design to administer the most informative items over multiple stages based on students' provisional or interim ability levels estimated at each stage. These MST applications yielded a higher test efficiency and increased measurement precision compared to their nonadaptive counterparts (Yamamoto et al., [<reflink idref="bib61" id="ref10">61</reflink>]).</p> <p>Although integrating the MST design into ILSAs yielded more accurate and efficient assessments, especially for heterogeneous groups of students (e.g., Shin et al., [<reflink idref="bib46" id="ref11">46</reflink>]), a significant threat against the validity of ILSAs remains unresolved: test‐taking engagement. Research shows that large‐scale assessments, especially ILSAs, may suffer from a lack of test‐taking engagement (e.g., Kuang & Sahin, [<reflink idref="bib29" id="ref12">29</reflink>]; Lee & Jia, [<reflink idref="bib30" id="ref13">30</reflink>]; Wise et al., [<reflink idref="bib56" id="ref14">56</reflink>]). Students with low test‐taking engagement may show noneffortful (i.e., disengaged) responding—answering the items either very rapidly without reading the item stem or response options (Wise & Kuhfeld, [<reflink idref="bib58" id="ref15">58</reflink>]) or very slowly, thereby failing to complete the test (e.g., Gorgun & Bulut, [<reflink idref="bib19" id="ref16">19</reflink>]; Pools, [<reflink idref="bib40" id="ref17">40</reflink>]). Noneffortful responses jeopardize both provisional scoring and routing mechanisms in the MST design because MSTs rely on the premise that each student gives their best effort in responding to the items and, thus, their responses are expected to reflect what they know or can do.</p> <p>To alleviate the problem of test‐taking engagement in the MST design, we propose a new method that assembles test modules on the fly after each stage based on two latent traits: ability and test‐taking engagement. This method combines the on‐the‐fly MST (OMST; Zheng & Chang, [<reflink idref="bib65" id="ref18">65</reflink>]) design with an adjusted module assembly procedure based on a latent "test‐taking engagement" trait (from now on referred to as OMST‐LE). We utilize a modified version of Goldhammer et al.'s ([<reflink idref="bib17" id="ref19">17</reflink>]) model‐based approach to measure students' test‐taking engagement as a latent trait. Through response time thresholds, this approach transforms item responses into binary disengagement indicators (i.e., 1 = disengaged or 0 = engaged). Then, test‐taking disengagement as a latent trait is estimated based on these indicators. Our proposed method, OMST‐LE, considers estimates of both ability and test‐taking engagement after each stage to assemble subsequent modules with the most informative and engaging items for each student. In this study, we aim to compare the performance of the MST, OMST, and OMST‐LE designs based on the accuracy of ability estimates when students participating in the ILSA are heterogeneous regarding achievement and test‐taking engagement.</p> <hd id="AN0183951184-2">Background</hd> <p></p> <hd id="AN0183951184-3">Introduction of MST Design in ILSAs</hd> <p>A higher degree of interest in cross‐national comparisons has motivated many countries and economies around the world to participate in ILSAs such as PISA, TIMSS, and PIRLS over the last two decades. A steady increase in the number of countries and economies participating in ILSAs has also coincided with increased heterogeneity in student achievement (Rutkowski et al., [<reflink idref="bib42" id="ref20">42</reflink>]). Several psychometric challenges arise when administering an ILSA to heterogeneous educational systems, such as tailoring the assessment internationally, covering a wider achievement distribution, and dealing with floor effects for low‐performing educational systems (Rutkowski et al., [<reflink idref="bib43" id="ref21">43</reflink>]). Although testing organizations have adopted different remedies to deal with large achievement differences among the participating countries (e.g., using less difficult versions of the ILSA), a more promising and innovative solution has been the introduction of the MST design to ILSAs. Several ILSAs, such as the OECD's PISA and PIAAC, have already begun to harness the MST design to target a broad spectrum of student achievement accurately while balancing the content composition in each test form (Yamamoto et al., [<reflink idref="bib61" id="ref22">61</reflink>]).</p> <p>The MST design strikes a balanced compromise between linear test designs and item‐level adaptive testing by selecting a set of suitable items (i.e., modules) adaptively after each stage (Hendrickson, [<reflink idref="bib21" id="ref23">21</reflink>]; Zenisky et al., [<reflink idref="bib64" id="ref24">64</reflink>]). Researchers have discussed several advantages of the MST design over other testing formats in ILSAs. For example, MST shows greater measurement accuracy and efficiency when compared to linear fixed‐length tests. Yamamoto et al. ([<reflink idref="bib61" id="ref25">61</reflink>]) found that MST was 10% to 30% more efficient for the Literacy subtest and 4% to 31% for the Numeracy subtest in PIAAC. Also, through a simulation study, Yamamoto et al. ([<reflink idref="bib62" id="ref26">62</reflink>]) found that the measurement precision of PISA 2018 was improved by 4% to 5% with gains and up to 10% at very low and high achievement levels. In addition to greater measurement accuracy and efficiency, MST also facilitates the control of test constraints based on nonstatistical item characteristics (e.g., content, item type, and item position) while enabling examinees to review their responses within each module (Yamamoto et al., [<reflink idref="bib61" id="ref27">61</reflink>]; Yan et al., [<reflink idref="bib63" id="ref28">63</reflink>]).</p> <hd id="AN0183951184-4">Indicators of Test‐Taking Engagement</hd> <p>Numerous studies have attempted to explain the impact of test‐taking engagement on student performance and test characteristics in ILSAs (e.g., Kuang & Sahin, [<reflink idref="bib29" id="ref29">29</reflink>]). To investigate whether students have sufficient test‐taking engagement or motivation to participate in the test, researchers have used various self‐report measures that enable students to indicate their level of test‐taking engagement after completing the test (e.g., Wise & DeMars, [<reflink idref="bib55" id="ref30">55</reflink>]; Eklöf, [<reflink idref="bib12" id="ref31">12</reflink>]). The testing organizations have also chosen to embed various self‐report measures within the student background questionnaires of ILSAs to measure students' motivation. For example, the student questionnaire in the 2015 cycle of PISA included a section on several aspects of student well‐being, including motivation to achieve. The questionnaire results indicated large differences in motivation across all participating countries and economies (Mo, [<reflink idref="bib35" id="ref32">35</reflink>]). PISA also includes a three‐item measure called an effort thermometer to measure how much effort students invested in completing the PISA test (OECD, [<reflink idref="bib37" id="ref33">37</reflink>]). The results of this measure from the 2018 cycle of PISA indicated that more than 65% of the students across OECD countries put less effort into the PISA test than they would have done in a test that counted toward their school grades (OECD, [<reflink idref="bib38" id="ref34">38</reflink>]).</p> <p>Despite their ease of use, self‐report measures on test‐taking engagement have been widely criticized because they rely on the assumption that students can accurately evaluate their engagement and report it truthfully (Finn, [<reflink idref="bib15" id="ref35">15</reflink>]; Swerdzewski et al., [<reflink idref="bib47" id="ref36">47</reflink>]). Thus, researchers have shifted their attention to examining test‐taking engagement based on students' behaviors during the test. Computer‐based versions of the current ILSAs (e.g., PISA, PIAAC, and TIMSS) enable the collection of additional information about test‐taking behaviors, such as response times and process data (von Davier et al., [<reflink idref="bib53" id="ref37">53</reflink>]). Numerous studies have used response time as a relatively more objective indicator of test‐taking engagement, as response times can be collected without student awareness (Finn, [<reflink idref="bib15" id="ref38">15</reflink>]). These studies have typically focused on identifying a threshold that could separate solution behavior from rapid guessing behavior (e.g., Bulut et al., [<reflink idref="bib6" id="ref39">6</reflink>]; Kroehne et al., [<reflink idref="bib28" id="ref40">28</reflink>]; Soland et al., [<reflink idref="bib45" id="ref41">45</reflink>]). Researchers have proposed various methods to determine response time thresholds for rapid guessing, such as visual inspection (Setzer et al., [<reflink idref="bib44" id="ref42">44</reflink>]), the normative threshold method (Wise & Ma, [<reflink idref="bib60" id="ref43">60</reflink>]), and data‐driven threshold search with random search and genetic algorithm (Bulut et al., [<reflink idref="bib6" id="ref44">6</reflink>]).</p> <p>In addition to detecting rapid guessing as an indicator of test‐taking engagement (i.e., effortful responding or rapid guessing), some researchers have also utilized response times to conceptualize test‐taking engagement as a latent trait. For example, Gorgun and Bulut ([<reflink idref="bib19" id="ref45">19</reflink>]) proposed a polytomous scoring method based on the operationalization of speed and ability as a single, composite latent trait. This approach transforms dichotomous responses into polytomous responses based on response accuracy (correct or incorrect) and response time (rapid guessing, slow responding, or optimal time use). Goldhammer et al. ([<reflink idref="bib17" id="ref46">17</reflink>]) defined a latent dimension of test‐taking disengagement based on binary item disengagement indicators created through response time thresholds. First, the authors identified a response time threshold for each item and used these thresholds to transform response times into binary disengagement indicators (i.e., if response time < threshold, 1 = disengaged; otherwise, 0 = engaged). Next, they used these binary indicators to estimate test‐taking engagement as a continuous latent trait with the one‐parameter logistic item response theory (IRT) model. Goldhammer et al. evaluated their approach using empirical data from the literacy, numeracy, and problem‐solving subtests of PIAAC 2012. Their findings indicated that test‐takers with higher educational attainment showed lower test‐taking disengagement in PIAAC and that low‐achieving test‐takers responding to more difficult items were more likely to disengage than high‐achieving test‐takers during the test. These findings highlight the need for a test design procedure for ILSAs that considers the alignment among ability, item difficulty, and test‐taking engagement as a source of individual differences.</p> <hd id="AN0183951184-5">Utility of Response Time in Large‐Scale Assessments</hd> <p>Researchers have proposed different ways to incorporate response time as collateral information into ILSAs. For example, von Davier et al. ([<reflink idref="bib53" id="ref47">53</reflink>]) suggested using response time to monitor students' test‐taking engagement to detect and eliminate noneffortful responses. Weeks et al. ([<reflink idref="bib54" id="ref48">54</reflink>]) analyzed empirical data from the PIAAC literacy and numeracy tests to identify response time thresholds distinguishing omitted responses unrelated to the latent trait from those related to the latent trait. The authors used the identified thresholds to decide whether an omitted response should be treated as not administered or incorrect. Lee and Jia ([<reflink idref="bib30" id="ref49">30</reflink>]) used response time collected from the NAEP to investigate the impact of test‐taking behaviors (e.g., rapid guessing) on IRT calibration results.</p> <p>Researchers have also investigated the utility of response time in adaptive testing. For example, Fan et al. ([<reflink idref="bib14" id="ref50">14</reflink>]) and Choe et al. ([<reflink idref="bib8" id="ref51">8</reflink>]) incorporated response time into the item selection procedure to shorten the test duration without sacrificing estimation accuracy and test security. Du et al. ([<reflink idref="bib11" id="ref52">11</reflink>]) evaluated the performance of the same item selection methods within the OMST design. A common characteristic of these studies is the operationalization of response time through van der Linden's ([[<reflink idref="bib49" id="ref53">49</reflink>]]) lognormal modeling approach in which two time‐related parameters for each item (i.e., time discrimination and time intensity) and a latent speed parameter for each examinee are estimated. Wise and Kingsbury ([<reflink idref="bib57" id="ref54">57</reflink>]) proposed an effort‐guided adaptive testing approach to correct the score distortion due to the lack of test‐taking engagement. As students respond to the items, their responses are monitored in real time and categorized as either solution behavior or rapid guessing based on their response time. If a particular response is identified as rapid guessing, then it is not considered in the estimation of provisional ability. Instead, the item selection algorithm selects the next item based on the previous provisional ability estimate. The effort‐guided adaptive testing can be an effective solution for situations where the noneffortful response behavior occurs intermittently. However, if a gradual and continuous decline in test‐taking engagement throughout the test is present, then the effort‐guided adaptive testing may fail to capture the student's true ability.</p> <hd id="AN0183951184-6">Current Study</hd> <p>As more ILSAs have begun to employ MST, it is necessary to find the most suitable MST design based on the student populations for which the ILSAs are designed. In this study, we propose a new MST design for ILSAs that combines Zheng and Chang's ([<reflink idref="bib65" id="ref55">65</reflink>]) OMST design with Goldhammer et al.'s ([<reflink idref="bib17" id="ref56">17</reflink>]) approach to modeling test‐taking engagement as a latent trait. This design, referred to as OMST‐LE, assumes that there are individual differences among test‐takers based on their test‐taking engagement—a continuous latent trait that may influence whether test‐takers can answer the items correctly. Our goal is to strengthen the OMST design to measure achievement more accurately while enhancing the test‐taking experience for test‐takers. The two research questions underlying our study are as follows:</p> <p></p> <ulist> <item> Does OMST‐LE yield more accurate ability estimates than the conventional MST and OMST designs when students vary based on their test‐taking engagement?</item> <p></p> <item> How do various test characteristics affect the performance of OMST‐LE?</item> </ulist> <p>To address these research questions, we employed a Monte‐Carlo simulation approach with the item and engagement parameters obtained from empirical data from PISA. In the next section, we describe the methodological framework used in the current study, with an emphasis on the simulation design, simulation conditions, and evaluation criteria.</p> <hd id="AN0183951184-7">Methods</hd> <p>In this study, we used the item parameters based on the two‐parameter logistic (2PL) model obtained from the 2018 cycle of PISA (OECD, [<reflink idref="bib39" id="ref57">39</reflink>]). There were 253 items from four subtests in PISA 2018: global competence (46 items), math (61 items), reading (<reflink idref="bib34" id="ref58">34</reflink>), and science (112 items). To create a large item bank for this study, we assumed that the items from the four subtests measured a single (i.e., unidimensional) hypothetical construct. To estimate the test‐taking engagement parameters, we used the response time data from the Canadian sample of students (<emph>n</emph> = 22,653) who participated in the 2018 cycle of PISA. For the selected sample, the dataset consisted of individual response times for 244 items in the item bank, except for 9 items from the math subtests, which were dropped for the subsequent analyses due to the lack of item‐level response time data in the Canadian sample.[<reflink idref="bib1" id="ref59">1</reflink>] To have a balanced item bank regarding the number of items per content domain, we assigned the items to one of the three hypothetical content domains (see Table 1).</p> <p>1 Table The Number of Items per Content Domain in the Item Bank</p> <p> <ephtml> <table><thead><tr><th>PISA Subtest</th><th align="center">Number of Items</th><th align="center">Content Domain ID</th></tr></thead><tbody><tr><td>Global competence</td><td>46</td><td>1</td></tr><tr><td>Math</td><td>52</td><td>2</td></tr><tr><td>Reading</td><td>34</td><td>1</td></tr><tr><td>Science</td><td>112</td><td>2 (28 items)3 (84 items)</td></tr></tbody></table> </ephtml> </p> <p>In the Monte‐Carlo simulation study, several testing conditions were treated as fixed, including the sample size (<reflink idref="bib3" id="ref60">3</reflink>,000 examinees), the total number of items in the item bank (244 items), and the content constraints in MST, OMST, and OMST‐LE (approximately 12 items per content domain). Four simulation conditions were manipulated: (<reflink idref="bib1" id="ref61">1</reflink>) MST design; (<reflink idref="bib2" id="ref62">2</reflink>) the relationship between ability and test‐taking engagement; (<reflink idref="bib3" id="ref63">3</reflink>) the number of stages; and (<reflink idref="bib4" id="ref64">4</reflink>) the number of items per module. One hundred replications were performed for each simulation condition. In the following sections, we elaborate on the data generation process for test‐taking engagement indicators and item response data, describe the simulation conditions, and outline the data analysis process with the outcomes of our analyses and criteria for evaluating the results.</p> <hd id="AN0183951184-8">Engagement Indicators</hd> <p>We created test‐taking engagement indicators using a modified version of Goldhammer et al.'s ([<reflink idref="bib17" id="ref65">17</reflink>]) model‐based approach. First, we identified 20% (the lower threshold for rapid guessing) and 180% (the upper threshold for slow responding or idling) of the average response time for each item as normative thresholds. This approach acknowledges that some test‐takers may exhibit disengaged behaviors, such as rapid guessing with very short response times and idling without making any effort to answer the items (Bulut & Gorgun, [<reflink idref="bib5" id="ref66">5</reflink>]; Goldhammer et al., [<reflink idref="bib17" id="ref67">17</reflink>]; Gorgun & Bulut, [<reflink idref="bib19" id="ref68">19</reflink>]; Gorgun & Bulut, [<reflink idref="bib20" id="ref69">20</reflink>]). Numerous studies have utilized the 20% response time threshold to identify students who lack test‐taking engagement during low‐stakes large‐scale assessments (e.g., Svetina Valdivia et al., [<reflink idref="bib48" id="ref70">48</reflink>]; Wise & Kuhfeld, [<reflink idref="bib59" id="ref71">59</reflink>]).</p> <p>Once the lower and upper thresholds were identified for each item, we transformed the response times into binary engagement indicators based on the thresholds as follows: 1 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>B</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo linebreak="badbreak">=</mo><mfenced separators="" open="{" close="">1ifTi1≤RTij≤Ti20otherwise</mfenced><mspace width="0.33em" /><mo>,</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}S{B}_{ij} = \left\{ { \def\eqcellsep{&}\begin{array}{@{}*{1}{c}@{}} {1\ if\ {T}_{i1} \le R{T}_{ij} \le {T}_{i2}}\\ {0\ {\mathrm{otherwise}}\ } \end{array} } \right.\ ,\end{equation}$$</annotation></semantics></math> </ephtml> where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>B</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding="application/x-tex">$S{B}_{ij}$</annotation></semantics></math> </ephtml> is the solution behavior (i.e., effortful responding) indicator for examinee <emph>j</emph> on item <emph>i</emph>, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>R</mi><msub><mi>T</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub></mrow><annotation encoding="application/x-tex">$R{T}_{ij}$</annotation></semantics></math> </ephtml> is the response time of examinee <emph>j</emph> on item <emph>i</emph>, and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>T</mi><mrow><mi>i</mi><mn>1</mn></mrow></msub><annotation encoding="application/x-tex">${T}_{i1}$</annotation></semantics></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>T</mi><mrow><mi>i</mi><mn>2</mn></mrow></msub><annotation encoding="application/x-tex">${T}_{i2}$</annotation></semantics></math> </ephtml> are the lower and upper thresholds for determining rapid guessing and slow responding (i.e., disengaged responding) for item <emph>i</emph>, respectively. Figure 1 demonstrates an example of how binary engagement indicators are determined based on the response time distribution of an item.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar25/jedm12380-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12380-fig-0001.jpg" title="1 An example of how binary engagement indicators are determined based on the response time thresholds. Note: The dashed lines refer to the lower and upper thresholds for response time." /> </p> <p></p> <p>Next, we fit the Rasch model to the dataset with binary engagement indicators as item response variables. We reviewed the information‐weighted (Infit) and unweighted (Outfit) mean‐squared fit statistics to judge the item fit. Infit and Outfit values between.5 and 1.5 are considered acceptable (de Ayala, [<reflink idref="bib9" id="ref72">9</reflink>]). In the dataset with binary engagement indicators, the Infit values ranged from.984 to 1.067, and the Outfit values ranged from.871 to 1.076, suggesting that the items indicated an acceptable fit. The comparison of expected and observed item characteristic curves for the items also showed that the model fit the data quite well. Figure 2 shows the distribution of the item engagement parameters obtained from the Rasch model. Given the low rate of test‐taking disengagement in the empirical data, most parameters were below zero (<emph>M</emph> = −2.290, <emph>SD</emph> =.399), indicating that the items were generally easy to engage. The item engagement parameters indicated weak correlations (<emph>r</emph> <.10) with difficulty and discrimination parameters in the item bank.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar25/jedm12380-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12380-fig-0002.jpg" title="2 The distribution of the engagement parameters from the Rasch model for the dataset with binary engagement indicators." /> </p> <p></p> <hd id="AN0183951184-11">Simulation Conditions</hd> <p></p> <hd id="AN0183951184-12">MST design</hd> <p>In this study, we utilized three different MST designs: MST, OMST, and OMST‐LE. MST refers to administering a test panel with multiple preassembled modules across two or more stages. Each module can be constructed to meet statistical and nonstatistical objectives. The first stage of MST is nonadaptive and contains items of moderate difficulty, called the routing module, while the subsequent stages include multiple preassembled modules with varying difficulty levels. Test‐takers' overall performance on the administered modules is used to route them to the most appropriate module in the next stage. We used the bottom‐up assembly method (Luecht & Nungester, [<reflink idref="bib32" id="ref73">32</reflink>]) to consider both statistical (i.e., item difficulty and discrimination parameters) and nonstatistical constraints (i.e., content) in module assembly. The test information function (TIF) was maximized at <emph>θ</emph> = −1, <emph>θ</emph> = 0, and <emph>θ</emph> = 1 for easy, medium, and hard modules to have a moderate difficulty difference between the modules (e.g., Kim et al., [<reflink idref="bib24" id="ref74">24</reflink>]). The target TIF values at these theta points were determined using the average maximum information method (Luecht, [<reflink idref="bib31" id="ref75">31</reflink>]).</p> <p>A single test panel was built using an automated test assembly procedure for each combination of the number of stages and the module length. Each item was used only once across the modules within each panel. Content balancing was implemented using a constraint of <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>n</mi><mo>/</mo><mn>3</mn><mo>±</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$n/3 \pm 1$</annotation></semantics></math> </ephtml> , where <emph>n</emph> is the total number of items in each module, to ensure that each module contains an approximately equal number of items from all three content domains. For the panels with unequal module lengths, the distribution of items across the modules was determined to be proportional to the number of items in each module. The content distributions of OMST and OMST‐LE were mirrored to match that of MST. For the 1‐3 design, the two routing cut‐off points were set to <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mo>−</mo><mo>.</mo><mn>44</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \ -.44$</annotation></semantics></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mo>.</mo><mn>44</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \.44$</annotation></semantics></math> </ephtml> to partition the area under the standard normal distribution into three equal sections. Similarly, the routing cut‐off point was set at <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>0</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \ 0$</annotation></semantics></math> </ephtml> for the 1‐2‐2 panels to divide the area into two halves. The provisional and final ability estimates were obtained using the expected a posteriori (EAP) method.</p> <p>In OMST and OMST‐LE, each examinee started the test with a routing module built for the MST design. After completing Stage 1, responses to the routing module were used to estimate the provisional ability. Starting from Stage 2, OMST and OMST‐LE differed based on the inclusion of engagement in module assembly. In the OMST design, the next module (i.e., Stage 2) was assembled using the items that maximized the Fisher information based on the provisional ability estimate (Zheng & Chang, [<reflink idref="bib65" id="ref76">65</reflink>]). This process was repeated to create another module (Stage 3) on the fly in the three‐stage design. Provisional ability estimates after each stage and the final ability estimate after the last stage were obtained using the EAP method.</p> <p>The OMST‐LE design assumed that the testing environment stored both item responses and response times for each test‐taker and converted response times into binary engagement indicators using the procedure described in Equation 1. After completing Stage 1, test‐takers' provisional ability (based on responses) and latent engagement levels (based on binary engagement indicators) were estimated. The next module was assembled using the items that maximized the Fisher information based on ability and latent engagement. Two different methods were used to determine the best items for each examinee. In the first method, the Fisher information was computed based on ability and latent engagement for each item in the item pool, and the total Fisher information was calculated as follows: 2 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>I</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo linebreak="badbreak">=</mo><msub><mi>I</mi><mi>i</mi></msub><mspace width="0.33em" /><mfenced separators="" open="(" close=")"><msub><mi>θ</mi><mi>j</mi></msub></mfenced><mo linebreak="goodbreak">+</mo><msub><mi>I</mi><mi>i</mi></msub><mfenced separators="" open="(" close=")"><msub><mi>η</mi><mi>j</mi></msub></mfenced><mo>,</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{I}_{ij} = {I}_i\ \left({{\theta }_j} \right) + {I}_i\left({{\eta }_j} \right),\end{equation}$$</annotation></semantics></math> </ephtml></p> <p>where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>I</mi><mi>i</mi></msub><mrow><mo>(</mo><msub><mi>θ</mi><mi>j</mi></msub><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">${I}_i({{\theta }_j})$</annotation></semantics></math> </ephtml> is the Fisher information for examinee <emph>j</emph> and item <emph>i</emph> based on the ability parameter (i.e., <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mi>j</mi></msub><annotation encoding="application/x-tex">${\theta }_j$</annotation></semantics></math> </ephtml> ), <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>I</mi><mi>i</mi></msub><mrow><mo>(</mo><msub><mi>η</mi><mi>j</mi></msub><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">${I}_i({{\eta }_j})$</annotation></semantics></math> </ephtml> is the Fisher information for examinee <emph>j</emph> and item <emph>i</emph> based on the latent engagement parameter (i.e., <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mi>j</mi></msub><annotation encoding="application/x-tex">${\eta }_j$</annotation></semantics></math> </ephtml> ), and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>I</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><annotation encoding="application/x-tex">${I}_{ij}$</annotation></semantics></math> </ephtml> is the total Fisher information for examinee <emph>j</emph> and item <emph>i</emph>. Depending on the module length, a unique module was built for each examinee by sorting the items by their total information ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>I</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><annotation encoding="application/x-tex">${I}_{ij}$</annotation></semantics></math> </ephtml> ) and selecting the items with the highest total information.</p> <p>In the second method, the Fisher information calculated based on the ability parameter was multiplied by the probability of an engaged response: 3 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>I</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo linebreak="badbreak">=</mo><msub><mi>I</mi><mi>i</mi></msub><mfenced separators="" open="(" close=")"><msub><mi>θ</mi><mi>j</mi></msub></mfenced><mo linebreak="goodbreak">×</mo><mi>P</mi><mrow><mo>(</mo><mspace width="0.33em" /><msub><mi>Y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo linebreak="goodbreak">=</mo><mspace width="0.33em" /><mn>1</mn><mo>|</mo><msub><mi>η</mi><mi>j</mi></msub><mo>,</mo><msub><mi>b</mi><mi>i</mi></msub><mo>)</mo></mrow><mo>,</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{I}_{ij} = {I}_i\left({{\theta }_j} \right) \times P(\ {Y}_{ij} = \ 1|{\eta }_j,{b}_i),\end{equation}$$</annotation></semantics></math> </ephtml></p> <p>where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0018" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>I</mi><mi>i</mi></msub><mrow><mo>(</mo><msub><mi>θ</mi><mi>j</mi></msub><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">${I}_i({{\theta }_j})$</annotation></semantics></math> </ephtml> is the Fisher information for examinee <emph>j</emph> and item <emph>i</emph> based on the ability parameter, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0019" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo>(</mo><mspace width="0.33em" /><msub><mi>Y</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><mo>=</mo><mspace width="0.33em" /><mn>1</mn><mo>|</mo><msub><mi>η</mi><mi>j</mi></msub><mo>,</mo><msub><mi>b</mi><mi>i</mi></msub><mo>)</mo></mrow><annotation encoding="application/x-tex">$P(\ {Y}_{ij} = \ 1|{\eta }_j,{b}_i)$</annotation></semantics></math> </ephtml> is examinee <emph>j</emph>'s probability of being engaging in responding to item <emph>i</emph> given his or her latent engagement level ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0020" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mi>j</mi></msub><annotation encoding="application/x-tex">${\eta }_j$</annotation></semantics></math> </ephtml> ) and the engagement parameter of item <emph>i</emph> ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0021" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>b</mi><mi>i</mi></msub><annotation encoding="application/x-tex">${b}_i$</annotation></semantics></math> </ephtml> ). Equation 3 yields the adjusted Fisher information ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0022" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>I</mi><mrow><mi>i</mi><mi>j</mi></mrow></msub><annotation encoding="application/x-tex">${I}_{ij}$</annotation></semantics></math> </ephtml> ) based on the probability of an engaged response. As the probability of engagement decreases, the adjusted Fisher information also decreases, thereby reducing the possibility of selecting the item for the next stage of OMST‐LE. Like the first method, this method also leads to a unique module for each examinee with the most informative items based on the adjusted Fisher information values. For both types of OMST‐LE, provisional ability and latent engagement estimates after each stage and the final ability estimate after the last stage were obtained using the EAP method.</p> <hd id="AN0183951184-13">Relationship between ability and test‐taking engagement</hd> <p>Previous studies have revealed a positive correlation between students' test‐taking engagement and performance on ILSAs (e.g., Eklöf, [<reflink idref="bib12" id="ref77">12</reflink>]; Goldhammer et al., [<reflink idref="bib16" id="ref78">16</reflink>]), suggesting that lower‐achieving students may not be as motivated to perform well on the test as their higher‐achieving peers. These studies reported a moderate to high (Pearson) correlation between test‐taking engagement and test performance (e.g., Goldhammer et al., [<reflink idref="bib16" id="ref79">16</reflink>]), indicating a linear and monotonic relationship between test‐taking engagement and ability. However, some scholars argued that a linear association between test‐taking engagement and ability may not hold in every setting (e.g., Asseburg & Frey, [<reflink idref="bib2" id="ref80">2</reflink>]; Goldhammer et al., [<reflink idref="bib18" id="ref81">18</reflink>]). For example, Asseburg and Frey ([<reflink idref="bib2" id="ref82">2</reflink>]) examined the relationship between test‐taking engagement and ability using empirical data from a low‐stakes mathematics assessment administered to a large sample of German students in the course of PISA 2006. Their results showed that although a linear relationship existed between test‐taking engagement and ability, this trend did not hold for students with very low or very high ability levels. Also, the authors found that as the items on the test became more difficult, even higher‐achieving students reported lower levels of test‐taking engagement and higher levels of boredom and daydreaming.</p> <p>Several studies also demonstrated a nonlinear relationship between response accuracy and response time. For instance, Bolsinova and Molenaar ([<reflink idref="bib4" id="ref83">4</reflink>]) found that as test‐takers spent more time answering the items, their response accuracy initially increased; however, after a certain time, the response accuracy reached a plateau and started to drop. De Boeck and Jeon ([<reflink idref="bib10" id="ref84">10</reflink>]) described this nonlinear dependence using a dual‐processing theory. The authors argued that although the speed‐accuracy balance works in favor of the test‐taker up to a particular time point, longer response times eventually lead to a decreasing response accuracy. As the time required to solve an item correctly increases, the test‐taker begins to perceive a lower chance of finding the correct answer that may no longer compensate for the cost of their effort.</p> <p>For our simulation study, we took inspiration from research that explored the nonlinear correlation between ability and test‐taking engagement. We included a nonlinear relationship condition where low‐achievers and high‐achievers could experience low engagement in test‐taking due to spending too little or too much time on the items. Specifically, we hypothesized two conditions regarding the relationship between the latent variables of ability and test‐taking engagement. The linear condition assumes that as test‐takers' ability levels increase, their engagement levels also increase linearly, or vice versa. The nonmonotonic condition assumes a curvilinear relationship where the positive relationship between ability and test‐taking engagement tapers off and turns into a negative relationship eventually (see the data generation and analysis section for further details).</p> <hd id="AN0183951184-14">Number of stages</hd> <p>The MST, OMST, and OMST‐LE designs involved either two or three stages in the simulation study. In the conventional MST, the 1‐3 two‐stage and 1‐2‐2 three‐stage designs were used. Each design started with a routing module with moderate difficulty ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0023" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mi>b</mi></msub><annotation encoding="application/x-tex">${M}_b$</annotation></semantics></math> </ephtml> = 0, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0024" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$S{D}_b$</annotation></semantics></math> </ephtml> = 1) in the first stage. Then, we created easy ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0025" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mi>b</mi></msub><annotation encoding="application/x-tex">${M}_b$</annotation></semantics></math> </ephtml> = −1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0026" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$S{D}_b$</annotation></semantics></math> </ephtml> =.6) and hard ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0027" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mi>b</mi></msub><annotation encoding="application/x-tex">${M}_b$</annotation></semantics></math> </ephtml> = 1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0028" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$S{D}_b$</annotation></semantics></math> </ephtml> =.6) modules for the second and third stages of the 1‐2‐2 design and easy ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0029" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mi>b</mi></msub><annotation encoding="application/x-tex">${M}_b$</annotation></semantics></math> </ephtml> = −1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0030" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$S{D}_b$</annotation></semantics></math> </ephtml> =.6), moderate ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0031" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mi>b</mi></msub><annotation encoding="application/x-tex">${M}_b$</annotation></semantics></math> </ephtml> = 0, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0032" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$S{D}_b$</annotation></semantics></math> </ephtml> =.6), and hard ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0033" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mi>b</mi></msub><annotation encoding="application/x-tex">${M}_b$</annotation></semantics></math> </ephtml> = 1, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0034" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$S{D}_b$</annotation></semantics></math> </ephtml> =.6) modules for the second stage of the 1‐3 design. In the OMST and OMST‐LE designs, the test started with the same routing module from MST in the first stage and continued with on‐the‐fly generated modules for each examinee in the second and third stages.</p> <hd id="AN0183951184-15">Number of items per stage</hd> <p>We were interested in examining different module lengths as the impact of test‐taking engagement could change depending on the number of items within a module. Across all stages, either 9, 12, 18, or 24 items per module were selected while keeping the total test length fixed (i.e., 36 items) for all examinees. This resulted in two specific designs, with balanced and imbalanced module lengths across the stages: (<reflink idref="bib1" id="ref85">1</reflink>) equal design had either 18 items per module in the two‐stage design (i.e., 18‐18) and 12 items per module in the three‐stage design (i.e., 12‐12‐12), and (<reflink idref="bib2" id="ref86">2</reflink>) short‐to‐long design had either 12 items in the routing stage followed by 24 items in the second stage (i.e., 12‐24 in the 1‐3 design) or 9 items in the routing stage followed by 9 and 18 items in the second and third stages (i.e., 9‐9‐18 in the 1‐2‐2 design).</p> <hd id="AN0183951184-16">Data Generation and Analysis</hd> <p>Ability parameters were generated between −3 and 3 with equal intervals of.2, resulting in a vector of 31 unique ability (<emph>θ</emph>) values (<emph>θ</emph> = [−3, −2.8, ..., 2.8, 3]). Next, 100 examinees were drawn from a uniform distribution at each ability interval (e.g., the first interval was <emph>θ</emph> = [−3, −2.8]), yielding a total sample size of 3,000 examinees between <emph>θ</emph> = −3 and <emph>θ</emph> = 3 <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0035" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>(</mo><mi>M</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>0</mn><mo>,</mo><mspace width="0.33em" /><mi>S</mi><mi>D</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>1.7</mn></mrow><annotation encoding="application/x-tex">$(M\ = \ 0,\ SD\ = \ 1.7$</annotation></semantics></math> </ephtml> ). Next, two sets of latent engagement parameters (η) were generated. First, a linear equation ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0036" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mrow><mi>η</mi><mspace width="0.33em" /></mrow><mo>=</mo><mspace width="0.33em" /><mn>1.5</mn><mo>+</mo><mi>θ</mi><mo>+</mo><mi>ε</mi></mrow><annotation encoding="application/x-tex">${\mathrm{\eta \ }} = \ 1.5 + \theta + \varepsilon $</annotation></semantics></math> </ephtml> , where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0037" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ε</mi><mo>∼</mo><mi>N</mi><mo>(</mo><mn>0</mn><mo>,</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$\varepsilon \sim N(0,1$</annotation></semantics></math> </ephtml> ) and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0038" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msup><mi>R</mi><mn>2</mn></msup><mo>=</mo><mspace width="0.33em" /><mo>.</mo><mn>65</mn></mrow><annotation encoding="application/x-tex">${R}^2 = \.65$</annotation></semantics></math> </ephtml> ) with ability parameters was used to generate a vector of linear engagement values ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0039" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>M</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>1.5</mn><mo>,</mo><mspace width="0.33em" /><mi>S</mi><mi>D</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>2</mn></mrow><annotation encoding="application/x-tex">$M\ = \ 1.5,\ SD\ = \ 2$</annotation></semantics></math> </ephtml> ). Second, a quadratic function ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0040" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mrow><mi>η</mi><mspace width="0.33em" /></mrow><mo>=</mo><mspace width="0.33em" /><mn>3</mn><mo>+</mo><mo>.</mo><mn>5</mn><mrow><mo>(</mo><mi>θ</mi><mo>)</mo></mrow><mo>−</mo><mo>.</mo><mn>5</mn><msup><mrow><mo>(</mo><mi>θ</mi><mo>)</mo></mrow><mn>2</mn></msup><mo>+</mo><mi>ε</mi></mrow><annotation encoding="application/x-tex">${\mathrm{\eta \ }} = \ 3 +.5(\theta) -.5{(\theta)}^2 + \varepsilon $</annotation></semantics></math> </ephtml> ; where <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0041" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ε</mi><mo>∼</mo><mi>N</mi><mo>(</mo><mn>0</mn><mo>,</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$\varepsilon \sim N(0,1$</annotation></semantics></math> </ephtml> )) was used to generate a vector of nonmonotonic η values ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0042" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>M</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>1.5</mn><mo>,</mo><mspace width="0.33em" /><mi>S</mi><mi>D</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>1.9</mn></mrow><annotation encoding="application/x-tex">$M\ = \ 1.5,\ SD\ = \ 1.9$</annotation></semantics></math> </ephtml> ). Figure 3 shows the scatterplots of the ability and latent engagement parameters for each condition. Based on the generated latent engagement values, the proportion of disengaged responses per item ranged from 4% to 25%, similar to the disengaged response rates reported by Goldhammer et al. ([<reflink idref="bib17" id="ref87">17</reflink>]) based on the PIAAC dataset.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar25/jedm12380-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12380-fig-0003.jpg" title="3 Hypothesized relationships between ability and latent engagement parameters." /> </p> <p></p> <p>To generate dichotomous item responses, we followed a simple algorithm that considered the examinee's engagement status for each item. We used the probability of correctly answering the item, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0043" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo>(</mo><mi>X</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>1</mn><mo>|</mo><mi>θ</mi><mo>)</mo></mrow><annotation encoding="application/x-tex">$P(X\ = \ 1|\theta)$</annotation></semantics></math> </ephtml> , based on the 2PL model when the examinee's engagement indicator was equal to 1 (i.e., engaged). If, however, the engagement indicator for a given item was zero (i.e., disengaged), then we set the probability of correctly answering the item to.25 (i.e., the probability of guessing the correct answer in a multiple‐choice item with four response options). This assumed that if the examinee were disengaged, he or she would attempt to select one of the response options randomly. In the final step, a random uniform value between 0 and 1 ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0044" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>u</mi><annotation encoding="application/x-tex">$u$</annotation></semantics></math> </ephtml> ) was drawn to determine the dichotomous response (i.e., if <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0045" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo>(</mo><mi>X</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mn>1</mn><mo>|</mo><mi>θ</mi><mo>)</mo><mo>></mo><mi>u</mi></mrow><annotation encoding="application/x-tex">$P(X\ = \ 1|\theta) > u$</annotation></semantics></math> </ephtml> , 1; otherwise, 0).[<reflink idref="bib2" id="ref88">2</reflink>]</p> <p>The datasets with dichotomous item responses were analyzed using the conventional MST, OMST, and OMST‐LE designs. In the conventional MST and OMST designs, item responses were used to estimate the provisional ability parameters after each module and select which module (MST) or items (OMST) to administer in the next stage. In the OMST‐LE design, item responses and engagement indicators were used to estimate the provisional values of ability and latent engagement parameters after each stage. As described earlier, both ability and latent engagement were considered in selecting the most appropriate items for the next module. We used the mirt package (Chalmers, [<reflink idref="bib7" id="ref89">7</reflink>]) for item calibration for the response time dataset, the xxIRT package (Luo, [<reflink idref="bib33" id="ref90">33</reflink>]) for module assembly, the mstR package (Magis et al., [<reflink idref="bib34" id="ref91">34</reflink>]) for the MST simulation, and a custom function written using the R programming language (R Core Team, [<reflink idref="bib41" id="ref92">41</reflink>]) for the OMST‐LE simulation. The simulation results were also summarized and visualized using R.</p> <hd id="AN0183951184-18">Evaluation Criteria</hd> <p>The results of the simulation study were evaluated based on the recovery of the ability parameters. Our evaluation criteria included the correlation between estimated and true ability parameters, absolute bias, and root‐mean‐squared error (RMSE). Absolute bias and RMSE were calculated conditioning on each ability (<emph>θ</emph>) interval as follows: 3 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0046" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mrow><mi>Absolute</mi><mspace width="0.33em" /><mi>Bias</mi></mrow><mo linebreak="badbreak">=</mo><mfrac><mrow><msubsup><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>N</mi></msubsup><mfenced open="(" close=")"><mfenced separators="" open="|" close="|"><mrow><msub><mover accent="true"><mi>θ</mi><mo>̂</mo></mover><mi>j</mi></msub><mo>−</mo><msub><mi>θ</mi><mi>j</mi></msub></mrow></mfenced></mfenced></mrow><mi>N</mi></mfrac><mo>,</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{\mathrm{Absolute\ Bias}} = \frac{{\mathop \sum \nolimits_{j = 1}^N \left({\left| {{{\hat{\theta }}}_j - {\theta }_j} \right|} \right)}}{N},\end{equation}$$</annotation></semantics></math> </ephtml> and 4 <ephtml> <math display="block" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0047" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>RMSE</mi><mo linebreak="badbreak">=</mo><msqrt><mfrac><mrow><msubsup><mo>∑</mo><mrow><mi>j</mi><mo>=</mo><mn>1</mn></mrow><mi>N</mi></msubsup><msup><mfenced separators="" open="(" close=")"><mrow><msub><mover accent="true"><mi>θ</mi><mo>̂</mo></mover><mi>j</mi></msub><mo>−</mo><msub><mi>θ</mi><mi>j</mi></msub></mrow></mfenced><mn>2</mn></msup></mrow><mi>N</mi></mfrac></msqrt><mo>.</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation}{\mathrm{RMSE}} = \sqrt {\frac{{\mathop \sum \nolimits_{j = 1}^N {{\left({{{\hat{\theta }}}_j - {\theta }_j} \right)}}^2}}{N}}.\end{equation}$$</annotation></semantics></math> </ephtml></p> <p>In Equations 3 and 4 above, <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0048" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>θ</mi><mo>̂</mo></mover><mi>j</mi></msub><annotation encoding="application/x-tex">${\hat{\theta }}_j$</annotation></semantics></math> </ephtml> is examinee <emph>j</emph>'s final ability estimate obtained from the adaptive test administration (MST, OMST, or OMST‐LE), <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0049" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>θ</mi><mi>j</mi></msub><mspace width="0.33em" /></mrow><annotation encoding="application/x-tex">${\theta }_j\ $</annotation></semantics></math> </ephtml> is examinee <emph>j</emph>'s true ability level, and <emph>N</emph> is the number of examinees at each ability interval (<reflink idref="bib100" id="ref93">100</reflink>). Smaller values of RMSE and absolute bias and higher correlations between the estimated and true ability values indicated higher accuracy in recovering true ability parameters. In addition, we calculated classification accuracy (CA) based on a set of hypothetical cut‐off scores. CA indicates the extent to which the classifications based on the cut‐off scores are aligned with the true classifications of examinees. PISA uses a six‐point proficiency scale with Levels 1 through 6 (OECD, [<reflink idref="bib39" id="ref94">39</reflink>]). To set realistic cut‐off values, we used the average of cut‐off scores, <emph>θ</emph> = (−.876, −.178,.521, 1.220, 1.919), from the reading, mathematics, and science tests in PISA 2018 to divide the examinees into six proficiency levels.[<reflink idref="bib3" id="ref95">3</reflink>] CA was computed as the proportion of the correct classifications across the five cut‐off points. Absolute bias, RMSE, and CA values were averaged over 100 replications.</p> <hd id="AN0183951184-19">Results</hd> <p>In the following sections, we report the findings based on the relationship between ability and test‐taking engagement (i.e., linear or nonmonotonic).</p> <hd id="AN0183951184-20">Linear Condition</hd> <p>Table 2 presents the results for the linear condition across all simulation conditions. The results showed that MST was generally the least accurate among the four MST designs (MST, OMST, OMST‐LE based on total Fisher information, and OMST‐LE based on probability‐adjusted Fisher information). This was an expected finding because the OMST designs can better match each test‐taker's ability level through custom modules assembled on the fly (Zheng & Chang, [<reflink idref="bib65" id="ref96">65</reflink>]). Despite notable differences in conditional RMSE and absolute bias, MST, OMST, and OMST‐LE yielded highly similar results regarding correlations between true and estimated ability values. All of the MST designs produced high correlations (<emph>r</emph> >.97), with OMST‐LE being marginally better. The OMST‐LE designs yielded slightly higher CA values than MST and OMST across most simulation conditions.</p> <p>2 Table Simulation Results for the Linear Condition</p> <p> <ephtml> <table><thead><tr><th>Design</th><th align="center">Stages</th><th align="center">Items per Stage</th><th align="center">Absolute Bias</th><th align="center">RMSE</th><th align="center">Correlation</th><th align="center">CA</th></tr></thead><tbody><tr><td>MST</td><td>2</td><td>12‐24</td><td>.319</td><td>.401</td><td>.978</td><td>.756</td></tr><tr><td /><td>2</td><td>18‐18</td><td>.330</td><td>.414</td><td>.977</td><td>.741</td></tr><tr><td /><td>3</td><td>9‐9‐18</td><td>.317</td><td>.401</td><td>.978</td><td>.742</td></tr><tr><td /><td>3</td><td>12‐12‐12</td><td>.331</td><td>.414</td><td>.977</td><td>.746</td></tr><tr><td>OMST</td><td>2</td><td>12‐24</td><td>.298</td><td>.381</td><td>.980</td><td>.780</td></tr><tr><td /><td>2</td><td>18‐18</td><td>.321</td><td>.409</td><td>.978</td><td>.752</td></tr><tr><td /><td>3</td><td>9‐9‐18</td><td>.322</td><td>.412</td><td>.980</td><td>.756</td></tr><tr><td /><td>3</td><td>12‐12‐12</td><td>.338</td><td>.423</td><td>.977</td><td>.745</td></tr><tr><td>OMST‐LE(Total Information)</td><td>2</td><td>12‐24</td><td>.293</td><td>.376</td><td>.982</td><td>.781</td></tr><tr><td>2</td><td>18‐18</td><td>.315</td><td>.401</td><td>.980</td><td>.758</td></tr><tr><td>3</td><td>9‐9‐18</td><td>.327</td><td>.419</td><td>.979</td><td>.748</td></tr><tr><td>3</td><td>12‐12‐12</td><td>.333</td><td>.421</td><td>.978</td><td>.744</td></tr><tr><td>OMST‐LE(Probability Adjusted)</td><td>2</td><td>12‐24</td><td>.293</td><td>.375</td><td>.982</td><td>.781</td></tr><tr><td>2</td><td>18‐18</td><td>.320</td><td>.410</td><td>.980</td><td>.756</td></tr><tr><td>3</td><td>9‐9‐18</td><td>.322</td><td>.412</td><td>.978</td><td>.752</td></tr><tr><td>3</td><td>12‐12‐12</td><td>.335</td><td>.421</td><td>.977</td><td>.746</td></tr></tbody></table> </ephtml> </p> <p>Table 2 also shows that the short‐to‐long module design with two stages (i.e., 12‐24) produced the best results, whereas the equal‐length design with three stages (i.e., 12‐12‐12) yielded the least accurate results for the OMST and OMST‐LE designs. As the number of stages increased from two to three, the accuracy of both OMST and OMSTE‐LE diminished. However, the differences between the accuracy of OMST and OMST‐LE were mostly negligible. For OMST‐LE, the most accurate results were obtained when the routing module consisted of 12 items based on the short‐to‐long design (i.e., 12‐24). This finding emphasizes the impact of module length in Stage 1 on the recovery of ability levels when using the OMST‐LE designs. The difference between the two variants of OMST‐LE remained negligible across all simulation conditions.</p> <p>To conduct a more fine‐grained analysis of the simulation results, we also examined the RMSE and bias values across different intervals of true ability (see Figure 4). In all simulation conditions, the performance of OMST and OMST‐LE was superior to that of MST, which aligns with the results displayed in Table 2. The OMST‐LE designs yielded more accurate results for lower ability levels (from <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0050" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mo>−</mo><mn>3</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \ - 3$</annotation></semantics></math> </ephtml> and <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0051" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mo>−</mo><mo>.</mo><mn>6</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \ -.6$</annotation></semantics></math> </ephtml> ) than OMST, suggesting that taking the latent engagement into account improved the accuracy of the on‐the‐fly module assembly as a result of using more engaging items. The gap between the OMST and OMST‐LE designs gradually disappeared for higher ability levels ( <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0052" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mo>></mo><mo>.</mo><mn>6</mn></mrow><annotation encoding="application/x-tex">$\theta >.6$</annotation></semantics></math> </ephtml> ). This was an expected finding because, in the linear condition, the level of latent engagement was high among higher‐achieving examinees, reducing its influence on the module assembly process. That is, OMST and the two variants of OMST‐LE selected nearly the same items for higher‐achieving students and thus yielded similar ability estimates. Another important finding was that all MST designs, regardless of the module length and the number of stages, produced overestimated ability values for examinees with <emph>θ</emph> < 0 and underestimated ability values for examinees with <emph>θ</emph> > 0. This finding was mainly due to the shrinkage phenomenon in Bayesian estimators, such as the EAP method used in this study, leading to overestimating negative proficiency levels and underestimating positive ones (Kim et al., [<reflink idref="bib24" id="ref97">24</reflink>]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar25/jedm12380-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12380-fig-0004.jpg" title="4 Average RMSE and bias values across the ability continuum in the linear condition. Note. Bias was used to demonstrate the over‐ and underestimation of ability." /> </p> <p></p> <p>Overall, the findings presented in this section suggest that test‐taking engagement can be operationalized as a latent trait to optimize the on‐the‐fly module assembly process using the OMST‐LE design when a linear relationship exists between ability and test‐taking engagement.</p> <hd id="AN0183951184-22">Nonmonotonic Condition</hd> <p>Table 3 presents the simulation results for the nonmonotonic condition. Compared to the linear condition, the accuracy of all MST designs decreased in the nonmonotonic condition due to having more disengaged examinees at the lower and higher end of the ability continuum. Negligible differences in correlations between true and estimated values of ability and the superior performance of OMST‐LE over OMST and MST in terms of CA were common findings between the linear and nonmonotonic conditions. Similarly, the OMST‐LE designs produced the most accurate results when the routing module consisted of 12 items based on the short‐to‐long design (i.e., 12‐24). The three‐stage design was again less accurate than the two‐stage design based on conditional RMSE and absolute bias. In the nonmonotonic condition, OMST‐LE based on probability‐adjusted Fisher information appeared to perform slightly better than OMST‐LE based on total Fisher information. Surprisingly, the performance of the conventional MST design was less prone to changes in the number of stages and module length than that of the OMST‐based designs. Also, OMST still appeared to be a strong contender, especially for the two‐stage design (12–24), although it disregarded latent engagement in the on‐the‐fly module assembly process.</p> <p>3 Table Simulation Results for the Nonmonotonic Condition</p> <p> <ephtml> <table><thead><tr><th>Design</th><th align="center">Stages</th><th align="center">Items per Stage</th><th align="center">Absolute Bias</th><th align="center">RMSE</th><th align="center">Correlation</th><th align="center">CA</th></tr></thead><tbody><tr><td>MST</td><td>2</td><td>12‐24</td><td>.349</td><td>.471</td><td>.969</td><td>.722</td></tr><tr><td /><td>2</td><td>18‐18</td><td>.361</td><td>.479</td><td>.968</td><td>.713</td></tr><tr><td /><td>3</td><td>9‐9‐18</td><td>.341</td><td>.466</td><td>.969</td><td>.729</td></tr><tr><td /><td>3</td><td>12‐12‐12</td><td>.345</td><td>.457</td><td>.971</td><td>.743</td></tr><tr><td>OMST</td><td>2</td><td>12‐24</td><td>.330</td><td>.461</td><td>.972</td><td>.743</td></tr><tr><td /><td>2</td><td>18‐18</td><td>.354</td><td>.482</td><td>.969</td><td>.727</td></tr><tr><td /><td>3</td><td>9‐9‐18</td><td>.353</td><td>.488</td><td>.970</td><td>.731</td></tr><tr><td /><td>3</td><td>12‐12‐12</td><td>.363</td><td>.493</td><td>.969</td><td>.716</td></tr><tr><td>OMST‐LE(Total Information)</td><td>2</td><td>12‐24</td><td>.336</td><td>.471</td><td>.971</td><td>.747</td></tr><tr><td>2</td><td>18‐18</td><td>.350</td><td>.479</td><td>.969</td><td>.731</td></tr><tr><td>3</td><td>9‐9‐18</td><td>.361</td><td>.501</td><td>.969</td><td>.721</td></tr><tr><td>3</td><td>12‐12‐12</td><td>.365</td><td>.502</td><td>.967</td><td>.711</td></tr><tr><td>OMST‐LE(Probability Adjusted)</td><td>2</td><td>12‐24</td><td>.329</td><td>.460</td><td>.972</td><td>.749</td></tr><tr><td>2</td><td>18‐18</td><td>.354</td><td>.482</td><td>.969</td><td>.726</td></tr><tr><td>3</td><td>9‐9‐18</td><td>.353</td><td>.493</td><td>.969</td><td>.724</td></tr><tr><td>3</td><td>12‐12‐12</td><td>.363</td><td>.487</td><td>.971</td><td>.731</td></tr></tbody></table> </ephtml> </p> <p>Figure 5 shows the average RMSE and bias values across different ability levels when a nonmonotonic relationship exists between ability and test‐taking engagement. The results show that the positive impact of incorporating test‐taking engagement into on‐the‐fly module assembly in OMST‐LE was more prominent for examinees whose ability levels ranged from <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0053" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mo>−</mo><mn>3</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \ - 3$</annotation></semantics></math> </ephtml> to <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0054" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mspace width="0.33em" /><mo>=</mo><mspace width="0.33em" /><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$\theta \ = \ - 1$</annotation></semantics></math> </ephtml> . However, this effect was less pronounced for the higher‐achieving examinees (i.e., <ephtml> <math display="inline" altimg="urn:x-wiley:00220655:media:jedm12380:jedm12380-math-0055" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>θ</mi><mo>></mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$\theta > 1$</annotation></semantics></math> </ephtml> ). This finding was mainly because of the difference in the number of disengaged responses at the lower and higher ends of the ability continuum. In the simulations, low‐achieving examinees were likely to show disengaged responding due to having lower levels of test‐taking engagement. A similar trend regarding overestimated ability values for <emph>θ</emph> < 0 and underestimated ability values for <emph>θ</emph> > 0 was also present in the nonmonotonic condition.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar25/jedm12380-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12380-fig-0005.jpg" title="5 Average RMSE and bias values across the ability continuum in the nonmonotonic condition. Note. Bias was used to demonstrate the over‐ and underestimation of ability." /> </p> <p></p> <p>In conclusion, the OMST‐LE design appears to be a more suitable option than the conventional MST and OMST designs when test‐takers are heterogeneous in terms of both ability and test‐taking engagement. OMST‐LE can be particularly useful for large‐scale assessments where test‐takers' relative positions and performance levels are the main priority.</p> <hd id="AN0183951184-24">Discussion</hd> <p>During the past decade, the national and international large‐scale assessment programs have continued to switch from PBT and CBT to benefit from the advantages of computerized test administration (e.g., increased test security and efficient data collection). However, given the widening heterogeneity in student achievement across countries (Rutkowski et al., [<reflink idref="bib42" id="ref98">42</reflink>]), it has become essential to build adaptive testing applications capable of accommodating large achievement differences among student populations. Several ILSAs, such as the OECD's PISA and PIAAC, have adopted the MST design to accurately target a broader spectrum of achievement while maintaining the content distribution on the test (Yamamoto et al., [<reflink idref="bib61" id="ref99">61</reflink>]). Unlike item‐level adaptive testing applications, MST offers a relatively more straightforward design where the most suitable items are adaptively selected for each examinee at the module level after each stage of the test. This approach also facilitates the control of various test constraints such as content, item type, and item position.</p> <p>Regardless of the test format (e.g., CBT vs. MST), students' performance in low‐stakes large‐scale assessments such as ILSAs has been a serious concern because tests with low to no personal consequences are often prone to reduced test‐taking engagement and performance (e.g., Attali, [<reflink idref="bib3" id="ref100">3</reflink>]; Wise & DeMars, [<reflink idref="bib55" id="ref101">55</reflink>]). Especially for ILSAs, as the number of students lacking test‐taking engagement increases, the gap between students' achievement may also widen due to the score distortion (i.e., underestimated test scores), contaminating the meaning of achievement differences within and across the participating countries. To control the impact of test‐taking engagement on large‐scale test administrations, this study introduced an alternative MST design (OMST‐LE) that considers examinees' ability and test‐taking engagement levels jointly when assembling subsequent modules on the fly after each stage. Using Goldhammer et al.'s ([<reflink idref="bib17" id="ref102">17</reflink>]) model‐based approach, test‐taking engagement was operationalized as a continuous, latent trait based on binary indicators derived from item response times. For each examinee, provisional levels of ability and test‐taking engagement were estimated after each stage. Then, the Fisher information adjusted based on test‐taking engagement was used as an index for identifying the most informative items.</p> <p>Through a Monte‐Carlo simulation study using item parameters from PISA 2018, we compared the performance of OMST‐LE with that of conventional MST and OMST designs. During the simulations, we modified the relationship between ability and test‐taking engagement, the number of stages, and the number of items per module. OMST‐LE showed a better performance compared to the conventional MST and OMST designs under both linear and nonmonotonic conditions. The best results were achieved by OMST‐LE using a two‐stage design, where Stage 1 had fewer items than Stage 2. As the number of stages increased, the accuracy of OMST‐LE gradually decreased. Furthermore, the difference in performance between OMST‐LE and other MST designs was particularly noticeable in situations where there was a linear correlation between ability and test‐taking involvement. This emphasizes the need for further examination of the relationship between ability and test‐taking engagement to better address the issue of low test‐taking engagement during MST.</p> <p>In summary, this study provides an alternative MST method for national or international large‐scale assessment programs. Our findings show that defining test‐taking engagement as a continuous latent trait based on response times can inform the design and implementation of MSTs in practice. The practical value of OMST‐LE is that it can help avoid selecting items for which examinees are likely to show disengaged response behavior and fail to demonstrate their true ability. It should be noted that the implementation of OMST‐LE relies on the real‐time collection and processing of response time data to measure test‐takers' latent engagement levels. When accurate response time data at the item level is not available, OMST can be an alternative method for measuring test‐takers' ability.</p> <hd id="AN0183951184-25">Limitations and Future Research</hd> <p>This study has several limitations. First, we did not manipulate several testing conditions (e.g., test length) that may steer the impact of test‐taking engagement on examinees' test‐taking behaviors. For example, the lack of test‐taking engagement could be more detrimental in shorter tests, especially when examinees remain disengaged throughout the test. Second, we assumed that all examinees could complete the maximum number of items on the test. However, some students participating in ILSAs may not have an opportunity to answer all items on the test. For example, a large percentage of students could not reach the end of the problem‐solving and inquiry tasks in the 2019 cycle of TIMSS (Mullis et al., [<reflink idref="bib36" id="ref103">36</reflink>]). In future studies, the OMST‐LE design presented in this study could be developed further to alleviate the negative impact of not‐reached items on ability estimation. Lastly, we utilized Goldhammer et al.'s ([<reflink idref="bib17" id="ref104">17</reflink>]) model‐based approach to quantify test‐taking engagement as a latent trait and incorporated it into the OMST design. Other researchers have proposed alternative frameworks combining accuracy and speed within a scoring rule (e.g., Gorgun & Bulut, [<reflink idref="bib19" id="ref105">19</reflink>]; Klinkenberg et al., [<reflink idref="bib27" id="ref106">27</reflink>]; van Rijn & Ali, [<reflink idref="bib51" id="ref107">51</reflink>]). Future research is needed to evaluate and compare the utility of these scoring methods in improving the accuracy of ability estimation in MST applications for large‐scale assessments.</p> <ref id="AN0183951184-26"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref3" type="bt">1</bibl> <bibtext> These values were reported as <emph>TT</emph> or <emph>total time</emph> in the additional public user file for the cognitive items available on the PISA website. See <ulink href="http://www.oecd.org/pisa/data/2018database">http://www.oecd.org/pisa/data/2018database</ulink> for the response and response time datasets for PISA 2018.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref62" type="bt">2</bibl> <bibtext> Supplementary materials, including the simulation files and item bank usage results for the MST and OMST‐LE designs, can be found at https://osf.io/a3jwk/.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref60" type="bt">3</bibl> <bibtext> For the original PISA cut scores with a mean of 500 and a standard deviation of 100, see https://<ulink href="http://www.oecd.org/pisa/data/pisa2018technicalreport/PISA2018%20TecReport‐Ch‐15‐Proficiency‐Scales.pdf">www.oecd.org/pisa/data/pisa2018technicalreport/PISA2018%20TecReport‐Ch‐15‐Proficiency‐Scales.pdf</ulink>.</bibtext> </blist> </ref> <ref id="AN0183951184-27"> <title> References </title> <blist> <bibtext> Addey, C., & Sellar, S. (2018). Why do countries participate in PISA? Understanding the role of international large‐scale assessments in global education policy. In A. Verger, M. Novelli, & H. K. Altinyelken (Eds.), Global education policy and international development: New agendas, issues and policies (pp. 97 – 117). New York, NY : Bloomsbury Academic.</bibtext> </blist> <blist> <bibtext> Asseburg, R., & Frey, A. (2013). Too hard, too easy, or just right? The relationship between effort or boredom and ability‐difficulty fit. Psychological Test and Assessment Modeling, 55 (1), 92. Retrieved from https://www.psychologie‐aktuell.com/fileadmin/download/ptam/1‐2013_20130326/06_Asseburg.pdf</bibtext> </blist> <blist> <bibtext> Attali, Y. (2016). Effort in low‐stakes assessments: What does it take to perform as well as in a high‐stakes setting? Educational and Psychological Measurement, 76 (6), 1045 – 1058. https://doi.org/10.1177/0013164416634789</bibtext> </blist> <blist> <bibl id="bib4" idref="ref64" type="bt">4</bibl> <bibtext> Bolsinova, M., & Molenaar, D. (2018). Modeling nonlinear conditional dependence between response time and accuracy. Frontiers in Psychology, 9, 1525. https://doi.org/10.3389/fpsyg.2018.01525</bibtext> </blist> <blist> <bibl id="bib5" idref="ref66" type="bt">5</bibl> <bibtext> Bulut, O., & Gorgun, G. (2023, June). Utilizing response time for scoring the TIMSS 2019 problem solving and inquiry tasks. Paper presented at the 10th IEA International Research Conference (IEA IRC), Dublin, Ireland. https://doi.org/10.31234/osf.io/zc98s</bibtext> </blist> <blist> <bibl id="bib6" idref="ref39" type="bt">6</bibl> <bibtext> Bulut, O., Gorgun, G., Wongvorachan, T., & Tan, B. (2023). Rapid guessing in low‐stakes assessments: Finding the optimal response time threshold with random search and genetic algorithm. Algorithms, 16 (2), 89. https://doi.org/10.3390/a16020089</bibtext> </blist> <blist> <bibl id="bib7" idref="ref89" type="bt">7</bibl> <bibtext> Chalmers, R. P. (2012). mirt: A Multidimensional Item Response Theory Package for the R Environment. Journal of Statistical Software, 48 (6), 1 – 29. https://doi.org/10.18637/jss.v048.i06</bibtext> </blist> <blist> <bibl id="bib8" idref="ref51" type="bt">8</bibl> <bibtext> Choe, E. M., Kern, J. L., & Chang, H.‐H. (2018). Optimizing the use of response times for item selection in computerized adaptive testing. Journal of Educational and Behavioral Statistics, 43, 135 – 158. https://doi.org/10.3102/1076998617723642</bibtext> </blist> <blist> <bibl id="bib9" idref="ref72" type="bt">9</bibl> <bibtext> de Ayala, R. J. (2009). The theory and practice of item response theory. New York : Guilford Press.</bibtext> </blist> <blist> <bibtext> De Boeck, P., & Jeon, M. (2019). An overview of models for response times and processes in cognitive tests. Frontiers in Psychology, 10, 422756. https://doi.org/10.3389/fpsyg.2019.00102</bibtext> </blist> <blist> <bibtext> Du, Y., Li, A., & Chang, H. H. (2019). Utilizing response time in on‐the‐fly multistage adaptive testing. In Quantitative Psychology: 83rd Annual Meeting of the Psychometric Society, New York, NY 2018 (pp. 107 – 117). Springer International Publishing. https://doi.org/10.1007/978‐3‐030‐01310‐3_10</bibtext> </blist> <blist> <bibtext> Eklöf, H. (2007). Test‐taking motivation and mathematics performance in TIMSS 2003. International Journal of Testing, 7 (3), 311 – 326. https://doi.org/10.1080/15305050701438074</bibtext> </blist> <blist> <bibtext> Ertl, B., Hartmann, F. G., & Heine, J. H. (2020). Analyzing large‐scale studies: Benefits and challenges. Frontiers in Psychology, 11, 577410. https://doi.org/10.3389/fpsyg.2020.577410</bibtext> </blist> <blist> <bibtext> Fan, Z., Wang, C., Chang, H.‐H., & Douglas, J. (2012). Utilizing response time distributions for item selection in CAT. Journal of Educational and Behavioral Statistics, 37 (5), 655 – 670. https://doi.org/10.3102/1076998611422912</bibtext> </blist> <blist> <bibtext> Finn, B. (2015). Measuring motivation in low‐stakes assessments. ETS Research Report Series, 2015 (2), 1 – 17. https://doi.org/10.1002/ets2.12067</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Martens, T., Christoph, G., & Lüdtke, O. (2016). Test‐taking engagement in PIAAC. OECD Education Working Papers, No. 133, OECD Publishing, Paris. https://doi.org/10.1787/5jlzfl6fhxs2‐en</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Martens, T., & Lüdtke, O. (2017). Conditioning factors of test‐taking engagement in PIAAC: An exploratory IRT modelling approach considering person and item characteristics. Large‐Scale Assessments in Education, 5, 1 – 25. https://doi.org/10.1186/s40536‐017‐0051‐9</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Naumann, J., Stelter, A., Tóth, K., Rölke, H., & Klieme, E. (2014). The time on task effect in reading and problem solving is moderated by task difficulty and skill: Insights from a computer‐based large‐scale assessment. Journal of Educational Psychology, 106, 608 – 626. https://doi.org/10.1037/a0034716</bibtext> </blist> <blist> <bibtext> Gorgun, G., & Bulut, O. (2021). A polytomous scoring approach to handle not‐reached items in low‐stakes assessments. Educational and Psychological Measurement, 81 (5), 847 – 871. https://doi.org/10.1177/0013164421991211</bibtext> </blist> <blist> <bibtext> Gorgun, G., & Bulut, O. (2023). Incorporating test‐taking engagement into the item selection algorithm in low‐stakes computerized adaptive tests. Large‐Scale Assessments in Education, 11, Article 27. https://doi.org/10.1186/s40536‐023‐00177‐5</bibtext> </blist> <blist> <bibtext> Hendrickson, A. (2007). An NCME instructional module on multistage testing. Educational Measurement: Issues and Practice, 26 (2), 44 – 52. https://doi‐org/10.1111/j.1745‐3992.2007.00093.x</bibtext> </blist> <blist> <bibtext> Hernández‐Torrano, D., & Courtney, M. G. (2021). Modern international large‐scale assessment in education: An integrative review and mapping of the literature. Large‐scale Assessments in Education, 9 (1), 17. https://doi.org/10.1186/s40536‐021‐00109‐1</bibtext> </blist> <blist> <bibtext> Kamens, D. H., & McNeely, C. L. (2010). Globalization and the growth of international educational testing and national assessment. Comparative Education Review, 54 (1), 5 – 25. https://doi.org/10.1086/648471</bibtext> </blist> <blist> <bibtext> Kim, S., Moses, T., & Yoo, H. H. (2015). Effectiveness of item response theory (IRT) proficiency estimation methods under adaptive multistage testing. ETS Research Report Series, 2015 (1), 1 – 19. https://doi.org/10.1002/ets2.12057</bibtext> </blist> <blist> <bibtext> Kirsch, I., & Lennon, M. L. (2017). PIAAC: A new design for a new era. Large‐Scale Assessments in Education, 5 (1). https://doi.org/10.1186/s40536‐017‐0046‐6</bibtext> </blist> <blist> <bibtext> Kirsch, I., Lennon, M., von Davier, M., Gonzalez, E., & Yamamoto, K. (2013). On the growing importance of international large‐scale assessments. In M. von Davier, E. Gonzalez, I. Kirsch, & K. Yamamoto (Eds.), The role of international large‐scale assessments: Perspectives from technology, economy, and educational research (pp. 1 – 11). Dordrecht, The Netherlands : Springer.</bibtext> </blist> <blist> <bibtext> Klinkenberg, S., Straatemeier, M., & van der Maas, H. (2011). Computer adaptive practice of maths ability using a new item response model for on the fly ability and difficulty estimation. Computers & Education, 57, 1813 – 1824. https://doi.org/10.1016/j.compedu.2011.02.003</bibtext> </blist> <blist> <bibtext> Kroehne, U., Deribo, T., & Goldhammer, F. (2020). Rapid guessing rates across administration mode and test setting. Psychological Test and Assessment Modeling, 62 (2), 147 – 177. https://doi.org/10.25656/01:23630</bibtext> </blist> <blist> <bibtext> Kuang, H., & Sahin, F. (2023). Comparison of disengagement levels and the impact of disengagement on item parameters between PISA 2015 and PISA 2018 in the United States. Large‐scale Assessments in Education, 11 (1), 1 – 31. https://doi.org/10.1186/s40536‐023‐00152‐0</bibtext> </blist> <blist> <bibtext> Lee, Y. H., & Jia, Y. (2014). Using response time to investigate students' test‐taking behaviors in a NAEP computer‐based study. Large‐scale Assessments in Education, 2, 1 – 24. https://doi.org/10.1186/s40536‐014‐0008‐1</bibtext> </blist> <blist> <bibtext> Luecht, R. M. (2000, April). Implementing the computer‐adaptive sequential testing (CAST) framework to mass produce high‐quality computer‐adaptive and mastery tests. Paper presented at the annual meeting of the National Council on Measurement in Education, New Orleans, LA.</bibtext> </blist> <blist> <bibtext> Luecht, R. M., & Nungester, R. J. (1998). Some practical examples of computer‐adaptive sequential testing. Journal of Educational Measurement, 35 (3), 229 – 249. https://doi‐org/10.1111/j.1745‐3984.1998.tb00537.x</bibtext> </blist> <blist> <bibtext> Luo, X. (2019). xxIRT: Item response theory and computer‐based testing in R. Version 2.1.2. https://github.com/xluo11/xxIRT</bibtext> </blist> <blist> <bibtext> Magis, D., Yan, D., & von Davier, A. A. (2018). mstR: Procedures to generate patterns under multistage testing. Version 1.2. https://github.com/cran/mstR</bibtext> </blist> <blist> <bibtext> Mo, J. (2019). How is students' motivation related to their performance and anxiety? PISA in Focus. No. 92. OECD Publishing.</bibtext> </blist> <blist> <bibtext> Mullis, I. V., Martin, M. O., Fishbein, B., Foy, P., & Moncaleano, S. (2021). Findings from the TIMSS 2019 problem solving and inquiry tasks. TIMSS ve PIRLS International Study Center. Retrieved from https://timss2019.org/psi/wp‐content/uploads/exhibits/TIMSS‐2019‐Findings‐Problem‐Solving‐and‐Inquiry.pdf</bibtext> </blist> <blist> <bibtext> OECD. (2016). Low‐performing students: Why they fall behind and how to help them succeed. Paris : OECD Publishing. <ulink href="http://doi.org/10.1787/9789264250246‐en">http://doi.org/10.1787/9789264250246‐en</ulink></bibtext> </blist> <blist> <bibtext> OECD. (2019a). PISA 2018 results (volume I): What students know and can do. Paris : OECD Publishing. https://doi.org/10.1787/5f07c754‐en</bibtext> </blist> <blist> <bibtext> OECD. (2019b). PISA 2018 technical report. Paris, France : OECD Publishing. Retrieved from https://<ulink href="http://www.oecd.org/pisa/data/pisa2018technicalreport/">www.oecd.org/pisa/data/pisa2018technicalreport/</ulink></bibtext> </blist> <blist> <bibtext> Pools, E. (2022). Not‐reached items: An issue of time and of test‐taking disengagement? The case of PISA 2015 reading data. Applied Measurement in Education, 35 (3), 197 – 221. https://doi.org/10.1080/08957347.2022.2103136</bibtext> </blist> <blist> <bibtext> R Core Team. (2022). R: A language and environment for statistical computing. Vienna, Austria : R Foundation for Statistical Computing.</bibtext> </blist> <blist> <bibtext> Rutkowski, D., Rutkowski, L., & Liaw, Y.‐L. (2018). Measuring widening proficiency differences in international assessments: Are current approaches enough? Educational Measurement: Issues and Practice, 37 (4), 40 – 48. https://doi.org/10.1111/emip.12225</bibtext> </blist> <blist> <bibtext> Rutkowski, L., Rutkowski, D., & Svetina Valdivia, D. (2022). Multistage test design considerations in international large‐scale assessments of educational achievement. In T. Nilsen, A. Stancel‐Piątak, & J. E. Gustafsson (Eds.), International handbook of comparative large‐scale studies in education: Perspectives, methods and findings (pp. 1 – 19). Cham : Springer International Publishing.</bibtext> </blist> <blist> <bibtext> Setzer, J. C., Wise, S. L., van denHeuvel, J. R., & Ling, G. (2013). An investigation of examinee test‐taking effort on a large‐scale assessment. Applied Measurement in Education, 26 (1), 34 – 49. https://doi.org/10.1080/08957347.2013.739453</bibtext> </blist> <blist> <bibtext> Soland, J., Kuhfeld, M., & Rios, J. (2021). Comparing different response time threshold setting methods to detect low effort on a large‐scale assessment. Large‐Scale Assessments in Education, 9 (1), 8. https://doi.org/10.1186/s40536‐021‐00100‐w</bibtext> </blist> <blist> <bibtext> Shin, H. J., Yamamoto, K., Khorramdel, L., & Robin, F. (2021). Increasing measurement precision of PISA through multistage adaptive testing. In Quantitative Psychology: The 85th annual meeting of the Psychometric Society, Virtual (pp. 325 – 334). Springer International Publishing. https://doi.org/10.1007/978‐3‐030‐74772‐5_29</bibtext> </blist> <blist> <bibtext> Swerdzewski, P. J., Harmes, J. C., & Finney, S. J. (2011). Two approaches for identifying low‐motivated students in a low‐stakes assessment context. Applied Measurement in Education, 24, 162 – 188. https://doi.org/10.1080/08957347.2011.555217</bibtext> </blist> <blist> <bibtext> Svetina Valdivia, D., Rutkowski, L., Rutkowski, D., Canbolat, Y., & Underhill, S. (2023). Test engagement and rapid guessing: Evidence from a large‐scale state assessment. Frontiers in Education, 8, 1127644. https://doi.org/10.3389/feduc.2023.1127644</bibtext> </blist> <blist> <bibtext> van der Linden, W. J. (2006). A lognormal model for response times on test items. Journal of Educational and Behavioral Statistics, 31, 181 – 204. https://doi.org/10.3102/10769986031002</bibtext> </blist> <blist> <bibtext> van der Linden, W. J. (2007). A hierarchical framework for modeling speed and accuracy on test items. Psychometrika, 72, 287 – 308. https://doi.org/10.1007/s11336‐006‐1478‐z</bibtext> </blist> <blist> <bibtext> van Rijn, P. W., & Ali, U. S. (2017). A comparison of item response models for accuracy and speed of item responses with applications to adaptive testing. British Journal of Mathematical and Statistical Psychology, 70 (2), 317 – 345. https://doi.org/10.1111/bmsp.12101</bibtext> </blist> <blist> <bibtext> Verger, A., Parcerisa, L., & Fontdevila, C. (2019). The growth and spread of large‐scale assessments and test‐based accountabilities: A political sociology of global education reforms. Educational Review, 71 (1), 5 – 30. https://doi.org/10.1080/00131911.2019.1522045</bibtext> </blist> <blist> <bibtext> von Davier, M., Khorramdel, L., He, Q., Shin, H. J., & Chen, H. (2019). Developments in psychometric population models for technology‐based large‐scale assessments: An overview of challenges and opportunities. Journal of Educational and Behavioral Statistics, 44 (6), 671 – 705. https://doi.org/10.3102/1076998619881789</bibtext> </blist> <blist> <bibtext> Weeks, J. P., von Davier, M., & Yamamoto, K. (2016). Using response time data to inform the coding of omitted responses. Psychological Test and Assessment Modeling, 58 (4), 671.</bibtext> </blist> <blist> <bibtext> Wise, S. L., & DeMars, C. E. (2005). Low examinee effort in low‐stakes assessment: Problems and potential solutions. Educational Assessment, 10, 1 – 17. https://doi.org/10.1207/s15326977ea1001_1</bibtext> </blist> <blist> <bibtext> Wise, S. L., Im, S., & Lee, J. (2021). The impact of disengaged test taking on a state's accountability test results. Educational Assessment, 26 (3), 163 – 174. https://doi.org/10.1080/10627197.2021.1956897</bibtext> </blist> <blist> <bibtext> Wise, S. L., & Kingsbury, G. G. (2016). Modeling student test‐taking motivation in the context of an adaptive achievement test. Journal of Educational Measurement, 53 (1), 86 – 105. https://doi.org/10.1111/jedm.12102</bibtext> </blist> <blist> <bibtext> Wise, S. L., & Kuhfeld, M. R. (2020). A cessation of measurement: Identifying test taker disengagement using response time. In M. J. Margolis & R. A. Feinberg (Eds.), Integrating timing considerations to improve testing practices (1st ed., pp. 150 – 164). Routledge.</bibtext> </blist> <blist> <bibtext> Wise, S. L., & Kuhfeld, M. R. (2021). Using retest data to evaluate and improve effort‐moderated scoring. Journal of Educational Measurement, 58 (1), 130 – 149. https://doi.org/10.1111/jedm.12275</bibtext> </blist> <blist> <bibtext> Wise, S. L., & Ma, L. (2012). Setting response time thresholds for a CAT item pool: The normative threshold method. Paper presented at the annual meeting of the National Council on Measurement in Education, Vancouver, Canada.</bibtext> </blist> <blist> <bibtext> Yamamoto, K., Khorramdel, L., & Shin, H. J. (2018). Introducing multistage adaptive testing into international large‐scale assessment designs using the example of PIAAC. Psychological Test and Assessment Modeling, 60 (3), 347 – 368.</bibtext> </blist> <blist> <bibtext> Yamamoto, K., Shin, H., & Khorramdel, L. (2019). Introduction of multistage adaptive testing design in PISA 2018. OECD Education Working Papers, No. 209, OECD Publishing: Paris. https://doi.org/10.1787/b9435d4b‐en</bibtext> </blist> <blist> <bibtext> Yan, D., Lewis, C., & von Davier, A. A. (2014). Overview of computerized multistage tests. In D. Yan, A. A. von Davier, & C. Lewis (Eds), Computerized multistage testing: Theory and applications (pp. 3 – 20). CRC Press.</bibtext> </blist> <blist> <bibtext> Zenisky, A., Hambleton, R. K., & Luecht, R. (2010). Multistage testing: Issues, designs and research. In W. J. Van der Linden & C. A. W. Glas (Eds.), Elements of adaptive testing (pp. 355 – 372). Berlin, Germany : Springer. https://doi.org/10.1007/978‐0‐387‐85461‐8_18</bibtext> </blist> <blist> <bibtext> Zheng, Y., & Chang, H.‐H. (2015). On‐the‐fly assembled multistage adaptive testing. Applied Psychological Measurement, 39 (2), 104 – 118. https://doi.org/10.1177/0146621614544519</bibtext> </blist> </ref> <aug> <p>By Okan Bulut; Guher Gorgun and Hacer Karamese</p> <p>Reported by Author; Author; Author</p> <p></p> <p>OKAN BULUT is an Associate Professor of Measurement, Evaluation, and Data Science, Centre for Research in Applied Measurement and Evaluation, University of Alberta, 6‐110 Education Centre North, 11210 87 Ave NW, Edmonton, AB T6G 2G5 CANADA; bulut@ualberta.ca. His primary research interests include the use of data mining, machine learning, and artificial intelligence techniques to improve the quality and effectiveness of digital assessments.</p> <p>GUHER GORGUN is a Ph.D. candidate in Measurement, Evaluation, and Data Science, Faculty of Education, University of Alberta, 6‐110 Education Centre North, 11210 87 Ave NW, Edmonton, ABT6G 2G5 CANADA; gorgun@ualberta.ca. Her primary research interests include digital assessments, student modeling, educational data mining, human‐centered AI applications in education, and learning analytics.</p> <p>HACER KARAMESE is a psychometrician at WIDA at the University of Wisconsin‐Madison, 1025 W Johnson St, Madison, WI 53706; hacer.karamese@wisc.edu. Her primary research interests include the application of psychometric methods to large‐scale operational programs, with a specific focus on the development and implementation of multistage adaptive testing.</p> </aug> <nolink nlid="nl1" bibid="bib23" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib52" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib26" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib13" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib22" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib46" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib25" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib62" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib61" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib29" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib30" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib56" firstref="ref14"></nolink> <nolink nlid="nl13" bibid="bib58" firstref="ref15"></nolink> <nolink nlid="nl14" bibid="bib19" firstref="ref16"></nolink> <nolink nlid="nl15" bibid="bib40" firstref="ref17"></nolink> <nolink nlid="nl16" bibid="bib65" firstref="ref18"></nolink> <nolink nlid="nl17" bibid="bib17" firstref="ref19"></nolink> <nolink nlid="nl18" bibid="bib42" firstref="ref20"></nolink> <nolink nlid="nl19" bibid="bib43" firstref="ref21"></nolink> <nolink nlid="nl20" bibid="bib21" firstref="ref23"></nolink> <nolink nlid="nl21" bibid="bib64" firstref="ref24"></nolink> <nolink nlid="nl22" bibid="bib63" firstref="ref28"></nolink> <nolink nlid="nl23" bibid="bib55" firstref="ref30"></nolink> <nolink nlid="nl24" bibid="bib12" firstref="ref31"></nolink> <nolink nlid="nl25" bibid="bib35" firstref="ref32"></nolink> <nolink nlid="nl26" bibid="bib37" firstref="ref33"></nolink> <nolink nlid="nl27" bibid="bib38" firstref="ref34"></nolink> <nolink nlid="nl28" bibid="bib15" firstref="ref35"></nolink> <nolink nlid="nl29" bibid="bib47" firstref="ref36"></nolink> <nolink nlid="nl30" bibid="bib53" firstref="ref37"></nolink> <nolink nlid="nl31" bibid="bib28" firstref="ref40"></nolink> <nolink nlid="nl32" bibid="bib45" firstref="ref41"></nolink> <nolink nlid="nl33" bibid="bib44" firstref="ref42"></nolink> <nolink nlid="nl34" bibid="bib60" firstref="ref43"></nolink> <nolink nlid="nl35" bibid="bib54" firstref="ref48"></nolink> <nolink nlid="nl36" bibid="bib14" firstref="ref50"></nolink> <nolink nlid="nl37" bibid="bib11" firstref="ref52"></nolink> <nolink nlid="nl38" bibid="bib49" firstref="ref53"></nolink> <nolink nlid="nl39" bibid="bib57" firstref="ref54"></nolink> <nolink nlid="nl40" bibid="bib39" firstref="ref57"></nolink> <nolink nlid="nl41" bibid="bib34" firstref="ref58"></nolink> <nolink nlid="nl42" bibid="bib20" firstref="ref69"></nolink> <nolink nlid="nl43" bibid="bib48" firstref="ref70"></nolink> <nolink nlid="nl44" bibid="bib59" firstref="ref71"></nolink> <nolink nlid="nl45" bibid="bib32" firstref="ref73"></nolink> <nolink nlid="nl46" bibid="bib24" firstref="ref74"></nolink> <nolink nlid="nl47" bibid="bib31" firstref="ref75"></nolink> <nolink nlid="nl48" bibid="bib16" firstref="ref78"></nolink> <nolink nlid="nl49" bibid="bib18" firstref="ref81"></nolink> <nolink nlid="nl50" bibid="bib10" firstref="ref84"></nolink> <nolink nlid="nl51" bibid="bib33" firstref="ref90"></nolink> <nolink nlid="nl52" bibid="bib41" firstref="ref92"></nolink> <nolink nlid="nl53" bibid="bib100" firstref="ref93"></nolink> <nolink nlid="nl54" bibid="bib36" firstref="ref103"></nolink> <nolink nlid="nl55" bibid="bib27" firstref="ref106"></nolink> <nolink nlid="nl56" bibid="bib51" firstref="ref107"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1464057
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Incorporating Test-Taking Engagement into Multistage Adaptive Testing Design for Large-Scale Assessments
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Okan+Bulut%22">Okan Bulut</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-5853-1267">0000-0001-5853-1267</externalLink>)<br /><searchLink fieldCode="AR" term="%22Guher+Gorgun%22">Guher Gorgun</searchLink><br /><searchLink fieldCode="AR" term="%22Hacer+Karamese%22">Hacer Karamese</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2025 62(1):57-80.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 24
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Response+Style+%28Tests%29%22">Response Style (Tests)</searchLink><br /><searchLink fieldCode="DE" term="%22Testing+Problems%22">Testing Problems</searchLink><br /><searchLink fieldCode="DE" term="%22Testing+Accommodations%22">Testing Accommodations</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Monte+Carlo+Methods%22">Monte Carlo Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Learner+Engagement%22">Learner Engagement</searchLink><br /><searchLink fieldCode="DE" term="%22Reaction+Time%22">Reaction Time</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Format%22">Test Format</searchLink><br /><searchLink fieldCode="DE" term="%22Achievement+Tests%22">Achievement Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22International+Assessment%22">International Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Sample+Size%22">Sample Size</searchLink>
– Name: SubjectThesaurus
  Label: Assessment and Survey Identifiers
  Group: Su
  Data: <searchLink fieldCode="SU" term="%22Program+for+International+Student+Assessment%22">Program for International Student Assessment</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/jedm.12380
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0655<br />1745-3984
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The use of multistage adaptive testing (MST) has gradually increased in large-scale testing programs as MST achieves a balanced compromise between linear test design and item-level adaptive testing. MST works on the premise that each examinee gives their best effort when attempting the items, and their responses truly reflect what they know or can do. However, research shows that large-scale assessments may suffer from a lack of test-taking engagement, especially if they are low stakes. Examinees with low test-taking engagement are likely to show noneffortful responding (e.g., answering the items very rapidly without reading the item stem or response options). To alleviate the impact of noneffortful responses on the measurement accuracy of MST, test-taking engagement can be operationalized as a latent trait based on response times and incorporated into the on-the-fly module assembly procedure. To demonstrate the proposed approach, a Monte-Carlo simulation study was conducted based on item parameters from an international large-scale assessment. The results indicated that the on-the-fly module assembly considering both ability and test-taking engagement could minimize the impact of noneffortful responses, yielding more accurate ability estimates and classifications. Implications for practice and directions for future research were discussed.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1464057
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1464057
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/jedm.12380
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 24
        StartPage: 57
    Subjects:
      – SubjectFull: Response Style (Tests)
        Type: general
      – SubjectFull: Testing Problems
        Type: general
      – SubjectFull: Testing Accommodations
        Type: general
      – SubjectFull: Measurement
        Type: general
      – SubjectFull: Monte Carlo Methods
        Type: general
      – SubjectFull: Learner Engagement
        Type: general
      – SubjectFull: Reaction Time
        Type: general
      – SubjectFull: Test Format
        Type: general
      – SubjectFull: Achievement Tests
        Type: general
      – SubjectFull: Foreign Countries
        Type: general
      – SubjectFull: International Assessment
        Type: general
      – SubjectFull: Sample Size
        Type: general
      – SubjectFull: Program for International Student Assessment
        Type: general
    Titles:
      – TitleFull: Incorporating Test-Taking Engagement into Multistage Adaptive Testing Design for Large-Scale Assessments
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Okan Bulut
      – PersonEntity:
          Name:
            NameFull: Guher Gorgun
      – PersonEntity:
          Name:
            NameFull: Hacer Karamese
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 0022-0655
            – Type: issn-electronic
              Value: 1745-3984
          Numbering:
            – Type: volume
              Value: 62
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Journal of Educational Measurement
              Type: main
ResultId 1