Sensemaking of Process Data from Evaluation Studies of Educational Games: An Application of Cross-Classified Item Response Theory Modeling

Saved in:
Bibliographic Details
Title: Sensemaking of Process Data from Evaluation Studies of Educational Games: An Application of Cross-Classified Item Response Theory Modeling
Language: English
Authors: Tianying Feng (ORCID 0000-0003-2215-9234), Li Cai
Source: Journal of Educational Measurement. 2026 63(1).
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 37
Publication Date: 2026
Sponsoring Agency: Institute of Education Sciences (ED)
Contract Number: R305D210032
Document Type: Journal Articles
Reports - Research
Descriptors: Educational Games, Item Response Theory, Models, Misconceptions, Item Analysis, Game Based Learning, Pretests Posttests
DOI: 10.1111/jedm.12396
ISSN: 0022-0655
1745-3984
Abstract: Process information collected from educational games can illuminate how students approach interactive tasks, complementing assessment outcomes routinely examined in evaluation studies. However, the two sources of information are historically analyzed and interpreted separately, and diagnostic process information is often underused. To tackle these issues, we present a new application of cross-classified item response theory modeling, using indicators of knowledge misconceptions and item-level assessment data collected from a multisite game-based randomized controlled trial. This application addresses (a) the joint modeling of students' pretest and posttest item responses and game-based processes described by indicators of misconceptions; (b) integration of gameplay information when gauging the intervention effect of an educational game; (c) relationships among game-based misconception, pretest initial status, and pre-to-post change; and (d) nesting of students within schools, a common aspect in multisite research. We also demonstrate how to structure the data and set up the model to enable our proposed application, and how our application compares to three other approaches to analyzing gameplay and assessment data. Lastly, we note the implications for future evaluation studies and for using analytic results to inform learning and instruction.
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2026
Accession Number: EJ1501291
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHVG9ivFwz0ITGoKXdeEOGtAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNBI0Oj0JTWtS0lk0AIBEICBm58kjppfXHBl5Gx19qDkebphPPEH99EQuTZumybEqHME1VwAQXRppugOvj36z7Tao7qyUVjiJYEtjI_QL_zx5Z0JrJp2dzBRHCIm1NiQpQeziePA76mlweVou6YUlb7rMfatdz4xITfqkP8ErceYVmy5qzqSCmK_5qnyqe737hKD75kN2EC1Om2PYUS6hvFtF31Zej5yz-eIKMDl
Text:
  Availability: 1
  Value: <anid>AN0192629992;mea01mar.26;2026Apr01.06:22;v2.2.500</anid> <title id="AN0192629992-1">Sensemaking of Process Data from Evaluation Studies of Educational Games: An Application of Cross‐Classified Item Response Theory Modeling </title> <sbt id="AN0192629992-2">Sensemaking of Data From Game‐Based Evaluation Studies</sbt> <p>Process information collected from educational games can illuminate how students approach interactive tasks, complementing assessment outcomes routinely examined in evaluation studies. However, the two sources of information are historically analyzed and interpreted separately, and diagnostic process information is often underused. To tackle these issues, we present a new application of cross‐classified item response theory modeling, using indicators of knowledge misconceptions and item‐level assessment data collected from a multisite game‐based randomized controlled trial. This application addresses (a) the joint modeling of students' pretest and posttest item responses and game‐based processes described by indicators of misconceptions; (b) integration of gameplay information when gauging the intervention effect of an educational game; (c) relationships among game‐based misconception, pretest initial status, and pre‐to‐post change; and (d) nesting of students within schools, a common aspect in multisite research. We also demonstrate how to structure the data and set up the model to enable our proposed application, and how our application compares to three other approaches to analyzing gameplay and assessment data. Lastly, we note the implications for future evaluation studies and for using analytic results to inform learning and instruction.</p> <p>Understanding data from evaluation studies of educational games <emph>to answer questions about effectiveness and improvement</emph> involves, if not necessitates, sensemaking. Sensemaking is a deliberate effort to construct a plausible understanding of differences, complex situations, or ill‐structured problems (Dervin, [<reflink idref="bib23" id="ref1">23</reflink>]; Klein et al., [<reflink idref="bib47" id="ref2">47</reflink>]; Pirolli & Russell, [<reflink idref="bib63" id="ref3">63</reflink>]; Weick et al., [<reflink idref="bib76" id="ref4">76</reflink>]). Sensemaking starts with posing a frame. To frame is to initiate with a story, perspective, or framework that guides how one explores, defines, connects, and interprets data (Klein et al., [<reflink idref="bib47" id="ref5">47</reflink>]). One goal of sensemaking is to inform practices (Dervin, [<reflink idref="bib23" id="ref6">23</reflink>]) and actions (Klein et al., [<reflink idref="bib47" id="ref7">47</reflink>]; Weick et al., [<reflink idref="bib76" id="ref8">76</reflink>]). We use sensemaking as the conceptual underpinning of our paper and as a guide for prioritizing the preconditions upon which our paper is built.</p> <p>Sensemaking occurs when we grapple with information embedded in gameplay process data. Event‐based gameplay process data track individuals' moment‐to‐moment choices, interactions with in‐game elements, and other game‐related information, including the timing of each event and details about the elements interacted with (Chung, [<reflink idref="bib15" id="ref9">15</reflink>]). The resulting data can contain hundreds of rows associated with each player attempting one in‐game task or puzzle. Well‐instrumented (gameplay) process data contain information—construct‐relevant patterns and processes—that we can leverage to help infer how different individuals approach a task. The challenge lies in how we extract, analyze, and interpret such information to draw valid inferences about what individuals know or learn (Greiff et al., [<reflink idref="bib32" id="ref10">32</reflink>]; Lindner & Greiff, [<reflink idref="bib52" id="ref11">52</reflink>]).</p> <p>Sensemaking also occurs when we integrate gameplay <emph>process</emph> information with other <emph>outcome</emph> (Bergner & Davier, [<reflink idref="bib5" id="ref12">5</reflink>]) or <emph>product</emph> (Levy, [<reflink idref="bib51" id="ref13">51</reflink>]; Zumbo et al., [<reflink idref="bib84" id="ref14">84</reflink>]) information obtained from mediums like traditional assessments. This integration may serve several functions, such as holistically gauging the instructional effectiveness of a game‐based intervention or its absence thereof, exploring the relationship between the process and the outcome, and of equal importance, answering the question: "How and why did learning, growth, or change (not) occur?" Statistical modeling is one tool that accomplishes the integration. What is more, integration aimed at sensemaking demands an understanding of how factors beyond data analysis and model construction interact. Some of these factors are not exclusive to game‐based research, including (a) alignment in content and cognitive demand features between the game and external assessments (Baker et al., [<reflink idref="bib4" id="ref15">4</reflink>]; Mislevy et al., [<reflink idref="bib58" id="ref16">58</reflink>]); (b) game or task design features (Mislevy et al., [<reflink idref="bib59" id="ref17">59</reflink>], [<reflink idref="bib58" id="ref18">58</reflink>]; Plass et al., [<reflink idref="bib64" id="ref19">64</reflink>]), user‐interface design features (Chung & Baker, [<reflink idref="bib16" id="ref20">16</reflink>]), and their effects on gameplay and cognition; (c) data instrumentation (Bergner & Davier, [<reflink idref="bib5" id="ref21">5</reflink>]; Chung, [<reflink idref="bib15" id="ref22">15</reflink>]); and (d) use of theory‐informed or construct‐sensitive process information (Goldhammer et al., [<reflink idref="bib30" id="ref23">30</reflink>]; Lindner & Greiff, [<reflink idref="bib52" id="ref24">52</reflink>]). Overlooking these factors, such as poor data instrumentation and alignment, compromises data quality and validity of inferences made through a model that aims to integrate gameplay and assessment information for sensemaking.</p> <p>Of the many factors, we prioritize three as the preconditions for fruitful sensemaking and for the cross‐classified item response theory (IRT) modeling approach advocated in this paper. These preconditions are: (a) alignment in content and cognitive demand features between the game and the assessment, (b) use of diagnostic or theory‐informed process information, and (c) provision of a unified and flexible modeling framework. All three preconditions are crucial for creating a coherent analytical framework that enables us to use gameplay as a diagnostic tool, derive interpretable results, and provide feedback to inform game design, student learning, and instruction. In what follows, we discuss the existing shortfalls in meeting one or more of these preconditions.</p> <hd id="AN0192629992-3">Existing Shortfalls in Meeting Three Preconditions for Data Sensemaking</hd> <p></p> <hd id="AN0192629992-4">Shortfall 1: Lack of Alignment between Game Design and Assessment Design</hd> <p>The most critical shortfall that impedes sensemaking is the lack of alignment between what is measured and instructed by the game‐based intervention, and what is assessed by the assessment. The design of the learning system (e.g., educational games) is often distinct from that of the assessment system. This observation is articulated directly (Arieli‐Attali et al., [<reflink idref="bib3" id="ref25">3</reflink>]) or through calls, recommendations, and needs (e.g., Darling‐Hammond et al., [<reflink idref="bib19" id="ref26">19</reflink>]; Gane et al., [<reflink idref="bib27" id="ref27">27</reflink>]; Foster & Piacentini, [<reflink idref="bib26" id="ref28">26</reflink>]; National Research Council, [<reflink idref="bib60" id="ref29">60</reflink>]; Pellegrino & Quellmalz, [<reflink idref="bib61" id="ref30">61</reflink>]). This disconnection reduces the probability of leveraging rich diagnostic information gained from process data to provide feedback to the intended users, such as students and instructors. It can lead to unnecessarily burdensome summative assessments when information about progress and achievement is already available from process data, a perspective akin to what Pellegrino and Quellmalz ([<reflink idref="bib61" id="ref31">61</reflink>]) argued for technology‐enabled assessments. The lack of alignment in the content and cognitive demand features between the game and the assessment, as well as between the in‐game tasks and the assessment items, can also limit the range of questions and analyses available for exploration.</p> <hd id="AN0192629992-5">Shortfall 2: Underuse of Diagnostic or Theory‐Informed Process Information</hd> <p>The second shortfall stems from the first. By "underuse," we mean a tendency to prioritize the analysis of data from traditional assessments as the primary or exclusive source of evidence, whether in game‐based research (Garcia et al., [<reflink idref="bib28" id="ref32">28</reflink>]; Petri & Gresse von Wangenheim, [<reflink idref="bib62" id="ref33">62</reflink>]) or more broadly in studies with access to both traditional assessment and process data (e.g., computer‐based assessments; Greiff et al., [<reflink idref="bib32" id="ref34">32</reflink>]). Little attention is given to understanding learners' experiences within the digital medium, and in the case of game‐based evaluation research, within the intervention itself. This neglect renders conceptually meaningful information embedded in process data an "often‐noted but seldom used potential" (Greiff et al., [<reflink idref="bib32" id="ref35">32</reflink>], p. 93).</p> <p>By "underuse," we also mean that the use of diagnostic process information remains limited when compared to information on time, response accuracy, or generally, timing and counts of low‐level events (Greiff et al., [<reflink idref="bib32" id="ref36">32</reflink>]). Time spent on a task and response correctness are jointly modeled in cognitive diagnosis and psychometric modeling research (De Boeck & Jeon, [<reflink idref="bib21" id="ref37">21</reflink>]; Ercikan et al., [<reflink idref="bib24" id="ref38">24</reflink>]; Jiao et al., [<reflink idref="bib41" id="ref39">41</reflink>]; Lee & Jia, [<reflink idref="bib49" id="ref40">49</reflink>]; van der Linden, [<reflink idref="bib25" id="ref41">25</reflink>]). Indicators of time and event counts are also frequently featured in substantive research (e.g., Blanié et al., [<reflink idref="bib6" id="ref42">6</reflink>]; Chen et al., [<reflink idref="bib12" id="ref43">12</reflink>]; Cagiltay et al., [<reflink idref="bib7" id="ref44">7</reflink>]; Gauthier et al., [<reflink idref="bib29" id="ref45">29</reflink>]; Goldhammer et al., [<reflink idref="bib31" id="ref46">31</reflink>]; Hahnel et al., [<reflink idref="bib33" id="ref47">33</reflink>]; Hautala et al., [<reflink idref="bib36" id="ref48">36</reflink>]; Kiili et al., [<reflink idref="bib46" id="ref49">46</reflink>]; Shute & Rahimi, [<reflink idref="bib70" id="ref50">70</reflink>]; Tenorio Delgado et al., [<reflink idref="bib72" id="ref51">72</reflink>]).</p> <p>In comparison, diagnostic or theory‐informed process information is underreported and underused. This kind of information includes indicators of (mis)conceptions (Chung & Feng, [<reflink idref="bib18" id="ref52">18</reflink>]; Kerr, [<reflink idref="bib43" id="ref53">43</reflink>]), strategies (Chung & Baker, [<reflink idref="bib16" id="ref54">16</reflink>]; Greiff et al., [<reflink idref="bib32" id="ref55">32</reflink>]; Wüstenberg et al., [<reflink idref="bib80" id="ref56">80</reflink>]), expert‐defined rules (Hao et al., [<reflink idref="bib34" id="ref57">34</reflink>]), or cognitively meaningful patterns (Liu & Israel, [<reflink idref="bib53" id="ref58">53</reflink>]). The requirements for substantive knowledge and technical expertise affect how feasible it is to translate theories into gameplay and to extract diagnostic process information. Research on computer‐based assessments has also articulated similar challenges (Greiff et al., [<reflink idref="bib32" id="ref59">32</reflink>]; Lindner & Greiff, [<reflink idref="bib52" id="ref60">52</reflink>]).</p> <hd id="AN0192629992-6">Shortfall 3.1: Lack of a Unified Modeling Framework</hd> <p>We divide the third shortfall into two parts, Shortfalls 3.1 and 3.2, both of which concern modeling. Gameplay and assessment data are often not jointly analyzed to connect the assessment and the intervention. This also means that the estimation of the intervention effect is based on only a subset of the collected information or separate analyses, the findings of which are not integrated or cannot be integrated.</p> <p>Analyses used in game‐based research tend to employ a two‐stage procedure to examine the relationships among indicators derived from gameplay process data and assessment scores. In stage one, indicators are aggregated to the person level. In stage two, aggregated indicators are correlated with assessment scores via correlational analysis (Petri & Gresse von Wangenheim, [<reflink idref="bib62" id="ref61">62</reflink>]; Zhu et al., [<reflink idref="bib83" id="ref62">83</reflink>]), included as covariates in regression analysis (Hautala et al., [<reflink idref="bib36" id="ref63">36</reflink>]; Weiner & Sanchez, [<reflink idref="bib77" id="ref64">77</reflink>]), or used in a predictive framework with deep learning models (Min et al., [<reflink idref="bib56" id="ref65">56</reflink>]). Notable exceptions to the aforementioned analyses include (a) an analysis by Kerr and Chung ([<reflink idref="bib45" id="ref66">45</reflink>]), which investigated how individuals' overall game‐based performance scores mediated the relationship between pretest and posttest sum scores; (b) an analysis by Reese et al. ([<reflink idref="bib68" id="ref67">68</reflink>]), which used multilevel modeling to examine the velocity and acceleration in individuals' progresses toward the game goal; and (c) an analysis by Levy ([<reflink idref="bib50" id="ref68">50</reflink>]), which applied dynamic Bayesian network modeling to nonaggregated, longitudinal game‐based indicator data.</p> <p>It is also worth noting that outside the game‐based research context, more modern statistical and computational approaches have been developed to make use of process information (Jiao et al., [<reflink idref="bib40" id="ref69">40</reflink>]; Lindner & Greiff, [<reflink idref="bib52" id="ref70">52</reflink>]). For instance, in computerized assessments, process information has been used or integrated to refine assessment information and improve test reliability (Tang et al., [<reflink idref="bib71" id="ref71">71</reflink>]; Xiao et al., [<reflink idref="bib81" id="ref72">81</reflink>]; Zhang et al., [<reflink idref="bib82" id="ref73">82</reflink>]). Techniques used in categorical sequence analysis (Abbot & Tsay, [<reflink idref="bib1" id="ref74">1</reflink>]) are used to extract latent features from categorical process data. Along the line of embedding and dimensionality reduction, latent space modeling has been applied to process data (Chen et al., [<reflink idref="bib13" id="ref75">13</reflink>]).</p> <hd id="AN0192629992-7">Shortfall 3.2: Lack of a Flexible Modeling Framework</hd> <p>Commonly used statistical tools in game‐based research, such as correlational analysis and multiple regression modeling, often fail to address data complexities stemming from features of study design. Nor do they provide the flexibility for researchers to explore additional hypotheses. In multisite evaluation research, four types of dependency can occur in the collected assessment data (Cai et al., [<reflink idref="bib9" id="ref76">9</reflink>]), apart from complexities that arise in gameplay data. First, there is a dependency between the underlying constructs being measured over time. Second, when the same set of items is administered to students repeatedly, there is item‐level residual dependence. Third, the assumption of full exchangeability of individuals across experimental conditions often fails to hold, particularly in randomized controlled trials (RCTs) of learning games, where varied degrees of student learning are expected to occur as a result of the intervention. Fourth, there is a dependency among individuals nested in sites. In addition to the acknowledged dependencies, researchers may want to investigate how characteristics at the site, in‐game task, or person level influence individuals' performance and progress.</p> <hd id="AN0192629992-8">This Paper</hd> <p>We pose one frame for sensemaking of data collected in evaluation studies of educational games. This frame presupposes the first two preconditions are met: (a) alignment (Chung et al., [<reflink idref="bib17" id="ref77">17</reflink>]; Center for Advanced Technology in Schools, [<reflink idref="bib11" id="ref78">11</reflink>]; Vendlinski et al., [<reflink idref="bib75" id="ref79">75</reflink>]) and (b) use of diagnostic process information (Kerr, [<reflink idref="bib43" id="ref80">43</reflink>]; Kerr & Chung, [<reflink idref="bib44" id="ref81">44</reflink>]). With this frame, we address the third precondition about modeling.</p> <p>Our frame consists of two parts. First, among various methods for extracting information from process data, we use indicators of knowledge misconceptions derived from gameplay processes (Kerr, [<reflink idref="bib43" id="ref82">43</reflink>]). We use these indicators as a means to "notice and bracket" (Chia, [<reflink idref="bib14" id="ref83">14</reflink>]) diagnostic insights from gameplay, moving away from simpler indicators that often fall short in making "students' behaviours and thought processes visible" (Foster & Piacentini, [<reflink idref="bib26" id="ref84">26</reflink>], p. 37). Second, we present a new application of cross‐classified IRT modeling to integrate data from multiple sources, specifically data of gameplay and assessments.</p> <p>The new application addresses issues and data complexities mentioned above. It jointly models individuals' responses on game‐based diagnostic indicators and responses to assessment items, and incorporates gameplay information when gauging the intervention effect of an educational game. It also relates individuals' game‐based performance to changes in assessment outcomes and accounts for the nesting of individuals within sites in multisite studies.</p> <p>In the following sections, we present the notations and the data structure that combines gameplay indicator data and assessment item data. We introduce the general framework and components of cross‐classified IRT modeling (Huang & Cai, [<reflink idref="bib39" id="ref85">39</reflink>]; van den Noortgate et al., [<reflink idref="bib74" id="ref86">74</reflink>]). We then compare our proposed application with three other approaches applied to a data set collected from a multisite RCT of math games (Chung et al., [<reflink idref="bib17" id="ref87">17</reflink>]) to highlight advantages of our application. Lastly, we note the implications for future (game‐based) evaluation studies and for using analytic results to inform learning and instruction.</p> <hd id="AN0192629992-9">Data Structure: Game‐Based Indicator Data Combined with Item‐Level Assessment Data</hd> <p></p> <hd id="AN0192629992-10">Game‐Based Indicator Data</hd> <p></p> <hd id="AN0192629992-11">What are game‐based indicators?</hd> <p>Game‐based indicators are <emph>observable variables</emph> (Mislevy et al., [<reflink idref="bib59" id="ref88">59</reflink>]) derived from gameplay process data. These indicators capture different aspects of players' interactions with a game. Examples include indicators of time spent on tasks, low‐level event counts, patterns, strategies, and knowledge misconceptions. Often, we derive one or more game‐based indicators for each in‐game task and for each player.</p> <hd id="AN0192629992-12">What are in‐game tasks?</hd> <p>An in‐game task is a game level that is explicitly defined during the game's design and creation process. We use <emph>in‐game tasks</emph>, as opposed to <emph>game levels</emph>, to distinguish levels in a game from levels in the multilevel modeling framework. We use <emph>tasks</emph> to also emphasize that researchers can adopt our proposed application not only to analyze game‐based indicators but also data from other mediums such as simulations.</p> <hd id="AN0192629992-13">Data structure with one indicator</hd> <p>The simplest game‐based indicator data consists of one binary game‐based indicator derived per in‐game task and per player. Let us consider a scenario where three players have completed three in‐game tasks and provided responses on one binary indicator. The indicator shows whether a player used a particular strategy to complete a task. Table 1 presents the indicator data associated with this example.</p> <p>1 Table Example of Indicator Data with One Binary Indicator</p> <p> <ephtml> <table><thead><tr><th>Player</th><th>Task 1</th><th>Task 2</th><th>Task 3</th></tr></thead><tbody><tr><td>1</td><td>1</td><td>0</td><td>1</td></tr><tr><td>2</td><td>0</td><td>1</td><td>0</td></tr><tr><td>3</td><td>1</td><td>1</td><td>1</td></tr></tbody></table> </ephtml> </p> <p>In this example, each row contains one player's data, and each column corresponds to one in‐game task. A value of 1 indicates that the player used the strategy of interest to complete the task, while a value of 0 indicates that the player did not use the strategy.</p> <hd id="AN0192629992-14">Data structure with more than one indicator</hd> <p>Often, we derive more than one indicator to describe different aspects of individuals' gameplay. Suppose that for the same three tasks (Tasks 1‐3) and the same three players (Players 1‐3), we derive multiple indicators to measure different aspects of player performance. Without loss of generality, we assume that there are two binary indicators that capture the completion (<reflink idref="bib1" id="ref89">1</reflink>) or failure (0) of two specific objectives of a task. Table 2 shows the data in a semi‐long format, where each row corresponds to a unique player‐task combination and contains the player's responses on two game‐based indicators.</p> <p>2 Table Example of Indicator Data with Three Binary Indicators (Semi‐Long Format)</p> <p> <ephtml> <table><thead><tr><th>Player</th><th>Task</th><th>Indicator 1</th><th>Indicator 2</th></tr></thead><tbody><tr><td>1</td><td>1</td><td>1</td><td>0</td></tr><tr><td>1</td><td>2</td><td>1</td><td>0</td></tr><tr><td>1</td><td>3</td><td>0</td><td>1</td></tr><tr><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>⋯</mi><annotation encoding="application/x-tex">$\cdots$</annotation></semantics></math></p></td></tr><tr><td>3</td><td>1</td><td>0</td><td>1</td></tr><tr><td>3</td><td>2</td><td>1</td><td>1</td></tr><tr><td>3</td><td>3</td><td>1</td><td>0</td></tr></tbody></table> </ephtml> </p> <hd id="AN0192629992-15">Game‐Based Indicator Data as Cross‐Classified Data</hd> <p>Individuals' in‐game performance is affected by their knowledge and skills and the characteristics of the tasks they encounter. This observation is also evident in the data structure demonstrated in the earlier examples, where the responses on indicators are simultaneously organized or classified by both players (persons) and tasks. The relationship between persons and tasks does not exhibit a clear hierarchical structure. Instead, persons and tasks represent two distinct sources of variation. Therefore, it is appropriate to consider the responses on the derived indicators as cross‐classified, or simultaneously influenced, by both persons and tasks.</p> <p>Figure 1 shows game‐based indicator data and item response data as two examples of cross‐classified data. Responses on game‐based indicators are cross‐classified by persons and in‐game tasks, while responses to the assessment items are cross‐classified by persons and items. Moreover, if we represent an individual's assessment‐based competency and game‐based performance as two latent variables, we can correlate the two variables to connect in‐game performance with assessment outcomes, and vice versa.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar26/jedm12396-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12396-fig-0001.jpg" title="1 Cross‐classified structure of assessment item data and gameplay indicator data." /> </p> <p></p> <p>Modeling in‐game tasks as a random component offers several advantages, similar to modeling test items as random (e.g., De Boeck, [<reflink idref="bib20" id="ref90">20</reflink>]; De Boeck & Wilson, [<reflink idref="bib22" id="ref91">22</reflink>]; van den Noortgate et al., [<reflink idref="bib74" id="ref92">74</reflink>]). First, the random task approach is useful when a vast pool of in‐game tasks are generated by manipulating specific design elements, and the tasks of interest are a sample drawn from this pool. Second, modeling tasks as random instead of fixed reduces the computational burden by reducing the number of parameters to be estimated, while still considering the task effects on gameplay. Consequently, the random task approach may be more viable for studies with limited sample sizes. When adopting the random task approach, however, we are not interested in the individual tasks themselves but rather in describing their variability. Lastly, by regressing the random task variable on design variables, researchers can explore the extent to which different design variables explain the task effects.</p> <hd id="AN0192629992-17">Structure of the Combined Data</hd> <p>We now combine the game‐based indicator data with item‐level data from traditional assessments. Table 3 shows a general representation of the combined data in the wide format. In this table, two blocks of data are presented: the assessment block of items and the gameplay block of in‐game tasks. Here, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>P</mi><annotation encoding="application/x-tex">$P$</annotation></semantics></math> </ephtml> is the number of persons, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>I</mi><annotation encoding="application/x-tex">$I$</annotation></semantics></math> </ephtml> is the number of items, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>J</mi><annotation encoding="application/x-tex">$J$</annotation></semantics></math> </ephtml> is the number of in‐game tasks. The boldface <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold">y</mi><mrow><mi>j</mi><mi>p</mi></mrow></msub><annotation encoding="application/x-tex">$\mathbf {y}_{jp}$</annotation></semantics></math> </ephtml> indicates that one or more indicators can be derived for each Task <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>j</mi><annotation encoding="application/x-tex">$j$</annotation></semantics></math> </ephtml> . For example, we may be interested in whether a person uses Skill A and Skill B to complete an in‐game task, creating one indicator per skill. Similarly, the boldface <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold">y</mi><mrow><mi>i</mi><mi>p</mi></mrow></msub><annotation encoding="application/x-tex">$\mathbf {y}_{ip}$</annotation></semantics></math> </ephtml> indicates that one or more responses can be made for each Item <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>i</mi><annotation encoding="application/x-tex">$i$</annotation></semantics></math> </ephtml> . One example is a test‐retest setting, where individuals respond to a set of identical items at two different time points (e.g., pretest and posttest).</p> <p>3 Table Structure of the Combined Data (Wide Format)</p> <p> <ephtml> <table><thead><tr><th /><th /><th>Assessment Block</th><th>Gameplay Block</th></tr><tr><th /><th /><th>Item</th><th>Task</th></tr><tr><th /><th /><th>1 <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>⋯</mi><annotation encoding="application/x-tex">$\cdots$</annotation></semantics></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>i</mi><annotation encoding="application/x-tex">$i$</annotation></semantics></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>⋯</mi><annotation encoding="application/x-tex">$\cdots$</annotation></semantics></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>I</mi><annotation encoding="application/x-tex">$I$</annotation></semantics></math></p></th><th>1 <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>⋯</mi><annotation encoding="application/x-tex">$\cdots$</annotation></semantics></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>j</mi><annotation encoding="application/x-tex">$j$</annotation></semantics></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>⋯</mi><annotation encoding="application/x-tex">$\cdots$</annotation></semantics></math></p><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>J</mi><annotation encoding="application/x-tex">$J$</annotation></semantics></math></p></th></tr></thead><tbody><tr><td /><td>1</td><td /><td /></tr><tr><td /><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mo>⋮</mo><annotation encoding="application/x-tex">$\vdots$</annotation></semantics></math></p></td><td /><td /></tr><tr><td>Person</td><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math></p></td><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi mathvariant="bold">y</mi><mi>i</mi><mi>p</mi><annotation encoding="application/x-tex">$\mathbf {y}_{ip}$</annotation></semantics></math></p></td><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi mathvariant="bold">y</mi><mi>j</mi><mi>p</mi><annotation encoding="application/x-tex">$\mathbf {y}_{jp}$</annotation></semantics></math></p></td></tr><tr><td /><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mo>⋮</mo><annotation encoding="application/x-tex">$\vdots$</annotation></semantics></math></p></td><td /><td /></tr><tr><td /><td><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>P</mi><annotation encoding="application/x-tex">$P$</annotation></semantics></math></p></td><td /><td /></tr></tbody></table> </ephtml> </p> <hd id="AN0192629992-18">Example of the Combined Data</hd> <p>To bridge the previous discussions with a concrete example, let us consider a data set containing three individuals ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>P</mi><mo>=</mo><mn>3</mn></mrow><annotation encoding="application/x-tex">$P=3$</annotation></semantics></math> </ephtml> ) who attempted two pre‐ and post‐assessment items ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>I</mi><mo>=</mo><mn>2</mn></mrow><annotation encoding="application/x-tex">$I=2$</annotation></semantics></math> </ephtml> ) and two in‐game tasks ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>J</mi><mo>=</mo><mn>2</mn></mrow><annotation encoding="application/x-tex">$J=2$</annotation></semantics></math> </ephtml> ). For each of the tasks, we have also derived two indicators ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>K</mi><mo>=</mo><mn>2</mn></mrow><annotation encoding="application/x-tex">$K=2$</annotation></semantics></math> </ephtml> ) using the gameplay process data. All item and indicator responses are dichotomously scored.</p> <p>To connect this example with the data set discussed in the "Empirical Data Analysis" section, where we present the new application of cross‐classified IRT modeling, we assume that the individuals (Persons 1‐3) belong to different schools, where individuals within the same school share more similar experiences compared to individuals from different schools. Specifically, Person 1 and Person 2 are from School 1, and Person 3 from School 2. Table 4 presents the data in the wide format. Table 5 presents the data in the semi‐long format. For subsequent modeling and analyses, we use data in the semi‐long format.</p> <p>4 Table Example Combined Data (Wide Format)</p> <p> <ephtml> <table><thead><tr><th /><th /><th>Assessment (Block 1)</th><th>Gameplay (Block 2)</th></tr><tr><th>School</th><th>Person</th><th>Pre_I1</th><th>Pre_I2</th><th>Pst_I1</th><th>Pst_I2</th><th>J1_K1</th><th>J1_K2</th><th>J2_K1</th><th>J2_K2</th></tr></thead><tbody><tr><td>1</td><td>1</td><td>0</td><td>0</td><td>1</td><td>0</td><td>1</td><td>0</td><td>0</td><td>0</td></tr><tr><td>1</td><td>2</td><td>1</td><td>0</td><td>1</td><td>0</td><td>0</td><td>0</td><td>1</td><td>0</td></tr><tr><td>2</td><td>3</td><td>0</td><td>1</td><td>1</td><td>1</td><td>0</td><td>1</td><td>0</td><td>0</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note</emph>. Pre: pretest; Pst: posttest; J: in‐game task. The suffix K denotes the game‐based indicators, and the suffix I denotes the assessment items.</p> <p>5 Table Example Combined Data (Semi‐Long Format)</p> <p> <ephtml> <table><thead><tr><th /><th /><th /><th>Variables</th><th /><th /></tr><tr><th /><th /><th /><th>Item Responses</th><th>Indicator Responses</th><th>School Dummy Variables</th></tr><tr><th>Person</th><th>Block</th><th>Element</th><th>Pre</th><th>Pst</th><th>K1</th><th>K2</th><th>Sch1</th><th>Sch2</th></tr></thead><tbody><tr><td>1</td><td>1</td><td>I1</td><td>0</td><td>1</td><td /><td /><td>1</td><td>0</td></tr><tr><td>1</td><td>1</td><td>I2</td><td>0</td><td>0</td><td /><td /><td>1</td><td>0</td></tr><tr><td>1</td><td>2</td><td>J1</td><td /><td /><td>1</td><td>0</td><td>1</td><td>0</td></tr><tr><td>1</td><td>2</td><td>J2</td><td /><td /><td>0</td><td>0</td><td>1</td><td>0</td></tr><tr><td>2</td><td>1</td><td>I1</td><td>1</td><td>1</td><td /><td /><td>1</td><td>0</td></tr><tr><td>2</td><td>1</td><td>I2</td><td>0</td><td>0</td><td /><td /><td>1</td><td>0</td></tr><tr><td>2</td><td>2</td><td>J1</td><td /><td /><td>0</td><td>0</td><td>1</td><td>0</td></tr><tr><td>2</td><td>2</td><td>J2</td><td /><td /><td>1</td><td>0</td><td>1</td><td>0</td></tr><tr><td>3</td><td>1</td><td>I1</td><td>0</td><td>1</td><td /><td /><td>0</td><td>1</td></tr><tr><td>3</td><td>1</td><td>I2</td><td>1</td><td>1</td><td /><td /><td>0</td><td>1</td></tr><tr><td>3</td><td>2</td><td>J1</td><td /><td /><td>0</td><td>1</td><td>0</td><td>1</td></tr><tr><td>3</td><td>2</td><td>J2</td><td /><td /><td>0</td><td>0</td><td>0</td><td>1</td></tr></tbody></table> </ephtml> </p> <p>2 Pre: pretest; Pst: posttest; Sch: school; K: game‐based indicator.</p> <p>In Table 5, we use <emph>Variables</emph> to denote the columns of observed responses on game‐based indicators or assessment items. It is important to note that <emph>Variables</emph> (capitalized) refers specifically to observed responses (or observed variables) and should not be confused with any discussion of latent variables. We use the term <emph>Block</emph> to denote which block of data an observed variable belongs to, distinguishing between responses associated with items and with in‐game tasks. The <emph>Elements</emph> within each block are the items or tasks, labeled using the same suffixes shown in Table 4. For example, the first item in the assessment block is denoted as I1.</p> <p>Lastly, we add school dummy variables (Sch1 and Sch2) to indicate the respective school that each person belongs to. For example, Person 1 is from School 1, given that every cell associated with Person 1 under the Sch1 column has a value of one. This person has responses on two binary indicators (K1 and K2) from two in‐game tasks (J1 and J2) and responses to two items (I1 and I2) administered during both pretest and posttest.</p> <hd id="AN0192629992-19">A General Modeling Framework Using Cross‐Classified Item Response Theory Models</hd> <p></p> <hd id="AN0192629992-20">Overview</hd> <p>To jointly analyze data from different sources, such as game‐based indicator data and item‐level assessment data, we consider three sources of influence on individuals' observed responses: (a) the blocks, such as one block of items and another block of in‐game tasks; (b) the persons; and (c) the sites. Persons are nested in sites, such as students in different classrooms or schools.</p> <p>To account for these sources of influence in modeling, we model the blocks and the persons as latent variables, and we include school dummy variables to account for the nesting of persons in schools. When examining how students respond to assessment items and in‐game tasks, we assume that the person‐specific latent variables reflect the ability, knowledge, or skills of the students, and the block‐specific latent variables reflect the relative difficulties of in‐game tasks and assessment items.</p> <p>Here, we have a cross‐classified structure: the observed responses are cross‐classified by blocks and persons. The observed responses are situated on Level 1, while persons and blocks are situated on Level 2. We incorporate school dummy variables to address the nesting of persons within sites and estimate a separate intercept for each school, as with a fixed effects model. This approach differs from treating school effects as random, where the emphasis is on summarizing the distribution of school effects using a group mean and variance.</p> <hd id="AN0192629992-21">Notations</hd> <p>Let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> index sites ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>s</mi><mo>=</mo><mn>1</mn><mo>,</mo><mi>⋯</mi><mo>,</mo><mi>S</mi></mrow><annotation encoding="application/x-tex">$s = 1,\dots, S$</annotation></semantics></math> </ephtml> ), <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> index persons ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo>=</mo><mn>1</mn><mo>,</mo><mi>⋯</mi><mo>,</mo><mi>P</mi></mrow><annotation encoding="application/x-tex">$p=1,\dots,P$</annotation></semantics></math> </ephtml> ), <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>b</mi><annotation encoding="application/x-tex">$b$</annotation></semantics></math> </ephtml> index blocks ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>b</mi><mo>=</mo><mn>1</mn><mo>,</mo><mi>⋯</mi><mo>,</mo><mi>B</mi></mrow><annotation encoding="application/x-tex">$b = 1,\dots, B$</annotation></semantics></math> </ephtml> ), <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>e</mi><mi>b</mi></msub><annotation encoding="application/x-tex">$e_b$</annotation></semantics></math> </ephtml> index elements in Block <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>b</mi><annotation encoding="application/x-tex">$b$</annotation></semantics></math> </ephtml> (e.g., an item in the assessment block; <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>e</mi><mi>b</mi></msub><mo>=</mo><mn>1</mn><mo>,</mo><mi>⋯</mi><mo>,</mo><msub><mi>E</mi><mi>b</mi></msub></mrow><annotation encoding="application/x-tex">$e_b = 1,\dots, E_b$</annotation></semantics></math> </ephtml> ), and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>v</mi><annotation encoding="application/x-tex">$v$</annotation></semantics></math> </ephtml> index the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>v</mi><annotation encoding="application/x-tex">$v$</annotation></semantics></math> </ephtml> th Variable (e.g., one of the Variable columns in Table 5). With the data structure demonstrated in Table 5, each element is essentially associated with only one block—either the gameplay or the assessment block. Therefore, we can drop the Block subscript <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>b</mi><annotation encoding="application/x-tex">$b$</annotation></semantics></math> </ephtml> , and let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>y</mi><mrow><mi>v</mi><mi>e</mi><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$y_{veps}$</annotation></semantics></math> </ephtml> denote the response on Variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>v</mi><annotation encoding="application/x-tex">$v$</annotation></semantics></math> </ephtml> for Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>e</mi><annotation encoding="application/x-tex">$e$</annotation></semantics></math> </ephtml> by Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> in Site <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> . The person‐specific latent variables are referred to as the <emph>cluster‐side</emph> latent variables (e.g., <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mi>p</mi></msub><annotation encoding="application/x-tex">$\eta _p$</annotation></semantics></math> </ephtml> ). The item‐ and task‐specific latent variables are referred to as the <emph>block‐side</emph> latent variables (e.g., <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>ξ</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\xi _e$</annotation></semantics></math> </ephtml> ).</p> <hd id="AN0192629992-22">The Linear Predictor</hd> <p>We denote the conditional probability of a response given relevant parameters as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>π</mi><annotation encoding="application/x-tex">$\pi$</annotation></semantics></math> </ephtml> , omitting all subscripts. In the context of modeling binary responses, we consider an IRT model with the following basic form: 1 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>π</mi><mo linebreak="badbreak">=</mo><mfrac><mn>1</mn><mrow><mn>1</mn><mo>+</mo><mi>exp</mi><mrow><mo>(</mo><mo>−</mo><mi>z</mi><mo>)</mo></mrow></mrow></mfrac><mo>.</mo></mrow><annotation encoding="application/x-tex">$$\begin{equation} \pi = \frac{1}{1+\exp {(-z)}}. \end{equation}$$</annotation></semantics></math> </ephtml> In Equation 1, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>z</mi><annotation encoding="application/x-tex">$z$</annotation></semantics></math> </ephtml> is the linear predictor. We will gradually build the desired cross‐classified IRT model, beginning with the simplest form of the linear predictor and moving toward the full model specification.</p> <hd id="AN0192629992-23">The simplest form</hd> <p>We start by considering a simple scenario where there is no nesting of individuals within sites, and we do not include any explanatory predictors in the latent regression equations. We also assume a one‐parameter logistic (1PL) IRT model with one person‐specific latent variable and one block‐specific latent variable and the corresponding loadings, or slopes, are constrained to be 1. We use <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>z</mi><mrow><mi>v</mi><mi>e</mi><mi>p</mi></mrow></msub><annotation encoding="application/x-tex">$z_{vep}$</annotation></semantics></math> </ephtml> to denote the linear predictor in a cross‐classified 1PL model for Variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>v</mi><annotation encoding="application/x-tex">$v$</annotation></semantics></math> </ephtml> of Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>e</mi><annotation encoding="application/x-tex">$e$</annotation></semantics></math> </ephtml> (in Block <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>b</mi><annotation encoding="application/x-tex">$b$</annotation></semantics></math> </ephtml> ) responded by Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> . The block subscript <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>b</mi><annotation encoding="application/x-tex">$b$</annotation></semantics></math> </ephtml> is omitted when each element is only associated with one block. With the simplifications, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>z</mi><mrow><mi>v</mi><mi>e</mi><mi>p</mi></mrow></msub><annotation encoding="application/x-tex">$z_{vep}$</annotation></semantics></math> </ephtml> can be written as: 2 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>zvep=αv+ξe+ηp.<annotation encoding="application/x-tex">$$\begin{align} z_{vep} &= \alpha _{v} + \xi _{e} + \eta _p. \end{align}$$</annotation></semantics></math> </ephtml></p> <p>Two sources of variation are present: one arising from the persons and the other arising from the blocks. The person‐specific latent variable is <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>η</mi><mi>p</mi></msub><mo>∼</mo><mi>N</mi><mrow><mo>(</mo><mn>0</mn><mo>,</mo><msubsup><mi>σ</mi><mrow><mi>P</mi><mi>e</mi><mi>r</mi><mi>s</mi><mi>o</mi><mi>n</mi></mrow><mn>2</mn></msubsup><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">$\eta _p \sim N(0, \sigma _{Person}^2)$</annotation></semantics></math> </ephtml> , which varies over the persons. The block‐specific latent variable is <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>ξ</mi><mi>e</mi></msub><mo>∼</mo><mi>N</mi><mrow><mo>(</mo><mn>0</mn><mo>,</mo><msubsup><mi>σ</mi><mrow><mi>B</mi><mi>l</mi><mi>o</mi><mi>c</mi><mi>k</mi></mrow><mn>2</mn></msubsup><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">$\xi _{e} \sim N(0, \sigma _{Block}^2)$</annotation></semantics></math> </ephtml> , which varies over elements in a given block. Again, the block‐specific subscript <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>b</mi><annotation encoding="application/x-tex">$b$</annotation></semantics></math> </ephtml> is omitted, given that all elements belong to only one block. Although we will not consider this in detail, the modeling framework could accommodate a generalizability theory (G‐theory) like interaction term between persons and blocks. It is crucial, however, to have enough replications, determined by design, within each of the cross‐classified cells to help disentangle this interaction from other within‐cell unidentified sources of variation ("error"); without replications, they would be confounded (e.g., Shavelson & Webb, [<reflink idref="bib69" id="ref93">69</reflink>], pp. 20‐23; Raudenbush & Bryk, [<reflink idref="bib65" id="ref94">65</reflink>], pp. 377‐378).</p> <p>Substantively speaking, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>η</mi><mo>̂</mo></mover><mi>p</mi></msub><annotation encoding="application/x-tex">$\hat{\eta }_p$</annotation></semantics></math> </ephtml> is the estimated score of Person <emph>p</emph>'s latent factor of interest. This latent factor could represent an individual's underlying ability, skill, or knowledge. For a block of elements, such as a block of assessment items or in‐game tasks, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>ξ</mi><mo>̂</mo></mover><mi>e</mi></msub><annotation encoding="application/x-tex">$\hat{\xi }_{e}$</annotation></semantics></math> </ephtml> is the estimated relative intercept of Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>e</mi><annotation encoding="application/x-tex">$e$</annotation></semantics></math> </ephtml> in a block. <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>α</mi><mo>̂</mo></mover><mi>v</mi></msub><annotation encoding="application/x-tex">$\hat{\alpha }_v$</annotation></semantics></math> </ephtml> is the intercept of the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>v</mi><annotation encoding="application/x-tex">$v$</annotation></semantics></math> </ephtml> th Variable. Using content of Table 5 as an example, for Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>I</mi><mn>1</mn></mrow><annotation encoding="application/x-tex">$I1$</annotation></semantics></math> </ephtml> , which is a test item administered at both pretest and posttest, Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>I</mi><mn>1</mn></mrow><annotation encoding="application/x-tex">$I1$</annotation></semantics></math> </ephtml> gets the same predicted random effect <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>ξ</mi><mo>̂</mo></mover><mi>e</mi></msub><annotation encoding="application/x-tex">$\hat{\xi }_{e}$</annotation></semantics></math> </ephtml> , irrespective of whether the testing occasion is pretest or posttest. <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>α</mi><mo>̂</mo></mover><mrow><mi>P</mi><mi>r</mi><mi>e</mi></mrow></msub><annotation encoding="application/x-tex">$\hat{\alpha }_{Pre}$</annotation></semantics></math> </ephtml> is the estimated intercept for Variable <emph>Pre</emph> (the average intercept for all pretest items), <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>α</mi><mo>̂</mo></mover><mrow><mi>P</mi><mi>s</mi><mi>t</mi></mrow></msub><annotation encoding="application/x-tex">$\hat{\alpha }_{Pst}$</annotation></semantics></math> </ephtml> is the estimated intercept for Variable <emph>Pst</emph> (the average intercept for all posttest items), and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mover accent="true"><mi>α</mi><mo>̂</mo></mover><mrow><mi>K</mi><mn>1</mn></mrow></msub><annotation encoding="application/x-tex">$\hat{\alpha }_{K1}$</annotation></semantics></math> </ephtml> is the estimated intercept for a game‐based indicator named <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>K</mi><mn>1</mn></mrow><annotation encoding="application/x-tex">$K1$</annotation></semantics></math> </ephtml> . Each intercept can be understood as inversely related to a corresponding difficulty. We refer interested readers to Reckase ([<reflink idref="bib67" id="ref95">67</reflink>]) for detailed discussions on multidimensional intercept and difficulty terms.</p> <hd id="AN0192629992-24">Adding explanatory predictors</hd> <p>We can regress the random terms, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mi>p</mi></msub><annotation encoding="application/x-tex">$\eta _{p}$</annotation></semantics></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>ξ</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\xi _{e}$</annotation></semantics></math> </ephtml> , on explanatory predictors, such as design variables and covariates, to answer substantive research questions. Considering the case of a single predictor <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>x</mi><mi>p</mi></msub><annotation encoding="application/x-tex">$x_p$</annotation></semantics></math> </ephtml> for Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> and another predictor <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>z</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$z_{e}$</annotation></semantics></math> </ephtml> for Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>e</mi><annotation encoding="application/x-tex">$e$</annotation></semantics></math> </ephtml> , we have the following latent regression equations: 3 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ηp=β0+β1xp+εp,<annotation encoding="application/x-tex">$$\begin{align} \eta _p &= \beta _{0} + \beta _{1}x_p + \epsilon _p, \end{align}$$</annotation></semantics></math> </ephtml> 4 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ξe=γ0+γ1ze+εe,<annotation encoding="application/x-tex">$$\begin{align} \xi _{e} &= \gamma _{0} + \gamma _{1}z_{e} + \epsilon _{e}, \end{align}$$</annotation></semantics></math> </ephtml> where the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>β</mi><mi mathvariant="normal">s</mi></mrow><annotation encoding="application/x-tex">$\beta{\rm s}$</annotation></semantics></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>γ</mi><mi mathvariant="normal">s</mi></mrow><annotation encoding="application/x-tex">$\gamma{\rm s}$</annotation></semantics></math> </ephtml> are the latent regression coefficients, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>ε</mi><mi>p</mi></msub><annotation encoding="application/x-tex">$\epsilon _p$</annotation></semantics></math> </ephtml> and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>ε</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\epsilon _{e}$</annotation></semantics></math> </ephtml> are the random effects. Substituting Equations 3 and 4 into Equation 2, we obtain 5 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>zvep=αv+(γ0+γ1ze+εe)+(β0+β1xp+εp)<annotation encoding="application/x-tex">$$\begin{align} z_{vep} &= \alpha _{v} + (\gamma _{0} + \gamma _{1}z_{e} + \epsilon _{e}) + (\beta _{0} + \beta _{1}x_p + \epsilon _p) \end{align}$$</annotation></semantics></math> </ephtml> or more compactly, 6 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>zvep=αv+(γ′ze+εe)+(β′xp+εp).<annotation encoding="application/x-tex">$$\begin{align} z_{vep} &= \alpha _{v} + (\bm{\gamma }^{\prime }\bm{z}_{e} + \epsilon _{e}) + (\bm{\beta }^{\prime }\bm{x}_p + \epsilon _p). \end{align}$$</annotation></semantics></math> </ephtml></p> <hd id="AN0192629992-25">Adding site dummy variables</hd> <p>Note that Equation 2 does not reflect the nesting of individuals within sites. Suppose there are <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>S</mi><annotation encoding="application/x-tex">$S$</annotation></semantics></math> </ephtml> sites, one way to account for such nesting is to regress the person‐specific random term on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mo>−</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$S-1$</annotation></semantics></math> </ephtml> school dummy variables (e.g., on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>d</mi><mn>2</mn></msub><mo>,</mo><mi>⋯</mi><mo>,</mo><msub><mi>d</mi><mi>S</mi></msub></mrow><annotation encoding="application/x-tex">$d_{2}, \dots, d_{S}$</annotation></semantics></math> </ephtml> assuming School 1 is the reference school). Now, the person latent variable for Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> in Site <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> is written as 7 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ηps=μsds+εps=μs+εps,<annotation encoding="application/x-tex">$$\begin{align} \eta _{ps} = \mu _s d_s + \epsilon _{ps} = \mu _s + \epsilon _{ps}, \end{align}$$</annotation></semantics></math> </ephtml> where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>ε</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\epsilon _{ps}$</annotation></semantics></math> </ephtml> becomes Person <emph>p</emph>'s deviation from Site <emph>s</emph>' intercept <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>μ</mi><mi>s</mi></msub><annotation encoding="application/x-tex">$\mu _s$</annotation></semantics></math> </ephtml> .</p> <hd id="AN0192629992-26">The Full Model</hd> <p>We have now discussed key components of a cross‐classified IRT model that can jointly model game‐based indicator data and assessment data. Following Huang and Cai ([<reflink idref="bib39" id="ref96">39</reflink>])'s descriptions, the full model has two main parts: (a) the latent structural model, which examines the relationships between different predictors and the latent variables, and (b) the measurement model, which examines the relationships between the latent variables and the observed responses.</p> <hd id="AN0192629992-27">The latent structural model</hd> <p></p> <hd id="AN0192629992-28">Cluster Side</hd> <p>On the cluster side, we formulate the latent regression equations as follows. 8 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ηps=θps+μs,<annotation encoding="application/x-tex">$$\begin{align} \bm{\eta }_{ps} &= \bm{\theta }_{ps} + \mu _s, \end{align}$$</annotation></semantics></math> </ephtml> 9 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>θps=Bxps+εps.<annotation encoding="application/x-tex">$$\begin{align} \bm{\theta }_{ps} &= \bm{B}\bm{x}_{ps} + \bm{\epsilon }_{ps}. \end{align}$$</annotation></semantics></math> </ephtml></p> <p>In Equation 8, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">η</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\eta }_{ps}$</annotation></semantics></math> </ephtml> decomposes into a vector of person‐specific latent variables <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">θ</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\theta }_{ps}$</annotation></semantics></math> </ephtml> for Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> in Site <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> and an intercept <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>μ</mi><mi>s</mi></msub><annotation encoding="application/x-tex">$\mu _s$</annotation></semantics></math> </ephtml> associated with Site <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> . Equation 9 shows that we can regress the latent variables on external predictors. Specifically, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">x</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{x}_{ps}$</annotation></semantics></math> </ephtml> contains predictors specific to Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> in Site <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> , and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi mathvariant="bold-italic">B</mi><annotation encoding="application/x-tex">$\bm{B}$</annotation></semantics></math> </ephtml> contains the regression coefficients associated with the person‐level predictors. <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">ε</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\epsilon }_{ps}$</annotation></semantics></math> </ephtml> contains the random terms, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi mathvariant="bold-italic">ε</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><mo>∼</mo><mi>M</mi><mi>V</mi><mi>N</mi><mrow><mo>(</mo><mrow><mn mathvariant="bold">0</mn></mrow><mo>,</mo><msub><mi mathvariant="bold">Σ</mi><mrow><mi>P</mi><mi>e</mi><mi>r</mi><mi>s</mi><mi>o</mi><mi>n</mi></mrow></msub><mo>)</mo></mrow></mrow><annotation encoding="application/x-tex">$\bm{\epsilon }_{ps} \sim MVN(\bm{0}, \bm{\Sigma }_{Person})$</annotation></semantics></math> </ephtml> .</p> <p>If all game‐based indicators are designed to measure a general game‐based performance, and if all assessment items are administered at two time points (e.g., pretest and posttest), <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">θ</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\theta }_{ps}$</annotation></semantics></math> </ephtml> becomes a three‐dimensional vector. The first element of this vector represents Person <emph>p</emph>'s game‐based performance. The second element represents Person <emph>p</emph>'s knowledge measured at the pretest. The third element represents Person <emph>p</emph>'s knowledge measured at the posttest. Equation 8 can be expanded as follows: 10 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ηps=μs+θps,gameplayμs+θps,pretestμs+θps,posttest.<annotation encoding="application/x-tex">$$\begin{align} \bm{\eta }_{ps} = \def\eqcellsep{&}\begin{bmatrix} \mu _s + \theta _{ps, gameplay} \\ \mu _s + \theta _{ps, pretest} \\ \mu _s + \theta _{ps, posttest} \end{bmatrix}. \end{align}$$</annotation></semantics></math> </ephtml></p> <p>With pretest and posttest administrations, we might also want to measure the extent of change that has occurred between the two assessments as an direct indicator of learning. To do so, we can introduce a latent change parameterization (see Cai & Houts, [<reflink idref="bib10" id="ref97">10</reflink>]; McArdle, [<reflink idref="bib55" id="ref98">55</reflink>]).</p> <p>The core concept underlying a latent change parameterization is as follows. With two time points, we include an additional latent variable, denoted as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{p\Delta }$</annotation></semantics></math> </ephtml> , at Time 2. This latent variable represents a difference or change, which is added to the overall status measured at Time 1. A more general version is discussed in Cai and Houts ([<reflink idref="bib10" id="ref99">10</reflink>]).</p> <p>If we let <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mrow><mi>p</mi><mn>1</mn></mrow></msub><annotation encoding="application/x-tex">$\eta _{p1}$</annotation></semantics></math> </ephtml> denote Person <emph>p</emph>'s measured knowledge at Time 1 (pretest) and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mrow><mi>p</mi><mn>2</mn></mrow></msub><annotation encoding="application/x-tex">$\eta _{p2}$</annotation></semantics></math> </ephtml> denote this person's measured knowledge at Time 2 (posttest), we can re‐express <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mrow><mi>p</mi><mn>2</mn></mrow></msub><annotation encoding="application/x-tex">$\eta _{p2}$</annotation></semantics></math> </ephtml> as <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>η</mi><mrow><mi>p</mi><mn>2</mn></mrow></msub><mo>=</mo><msub><mi>η</mi><mrow><mi>p</mi><mn>1</mn></mrow></msub><mo>+</mo><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub></mrow><annotation encoding="application/x-tex">$\eta _{p2} = \eta _{p1} + \theta _{p\Delta }$</annotation></semantics></math> </ephtml> . When we rearrange this equation, we can see that <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><mo>=</mo><msub><mi>η</mi><mrow><mi>p</mi><mn>2</mn></mrow></msub><mo>−</mo><msub><mi>η</mi><mrow><mi>p</mi><mn>1</mn></mrow></msub></mrow><annotation encoding="application/x-tex">$\theta _{p\Delta } = \eta _{p2} - \eta _{p1}$</annotation></semantics></math> </ephtml> indeed represents a difference in Person <emph>p</emph>'s knowledge measured at two time points. Based on variance algebra, the variance of the latent change variable <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{p\Delta }$</annotation></semantics></math> </ephtml> is typically smaller than the variance of a nonchange latent variable like <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>η</mi><mrow><mi>p</mi><mn>1</mn></mrow></msub><annotation encoding="application/x-tex">$\eta _{p1}$</annotation></semantics></math> </ephtml> .</p> <p>To connect our discussion on latent changes to Equation 8, we reexpress the third element of <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">θ</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\theta }_{ps}$</annotation></semantics></math> </ephtml> to include a time specific effect <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{p\Delta }$</annotation></semantics></math> </ephtml> . As discussed, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{p\Delta }$</annotation></semantics></math> </ephtml> represents a latent change from the baseline measured at the pretest, with estimable mean and variance. The expanded Equation 8 becomes 11 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ηps=μs+θps,gameplayμs+θps,baselineμs+θps,baseline+θpΔ.<annotation encoding="application/x-tex">$$\begin{align} \bm{\eta }_{ps} = \def\eqcellsep{&}\begin{bmatrix} \mu _s + \theta _{ps, gameplay} \\ \mu _s + \theta _{ps, baseline} \\ \mu _s + \theta _{ps, baseline} + \theta _{p\Delta } \end{bmatrix}. \end{align}$$</annotation></semantics></math> </ephtml></p> <p>In Equation 11, the time‐specific effect <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{p\Delta }$</annotation></semantics></math> </ephtml> is a latent change that captures the part of individual knowledge or performance measured at Time 2 that is not identical to the same individual's knowledge or performance measured at Time 1. When Time 1 corresponds to the pretest administration and Time 2 to the posttest administration, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi>s</mi><mo>,</mo><mi>b</mi><mi>a</mi><mi>s</mi><mi>e</mi><mi>l</mi><mi>i</mi><mi>n</mi><mi>e</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{ps, baseline}$</annotation></semantics></math> </ephtml> reflects the baseline status or performance measured at the pretest, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>θ</mi><mrow><mi>p</mi><mi mathvariant="normal">Δ</mi></mrow></msub><annotation encoding="application/x-tex">$\theta _{p\Delta }$</annotation></semantics></math> </ephtml> reflects the change from pretest to posttest.</p> <hd id="AN0192629992-29">Block Side</hd> <p>The blocks are uncorrelated with the clusters. The latent regression equation associating Element <emph>e</emph>'s predictors and block‐side latent variables is 12 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>ξe=Γze+εe.<annotation encoding="application/x-tex">$$\begin{align} \bm{\xi }_{e} &= \bm{\Gamma }\bm{z}_{e} + \bm{\epsilon }_{e}. \end{align}$$</annotation></semantics></math> </ephtml></p> <p>In Equation 12, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi mathvariant="bold">Γ</mi></mrow><annotation encoding="application/x-tex">$\bm{\Gamma }$</annotation></semantics></math> </ephtml> contains the regression coefficients. <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">z</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\bm{z}_{e}$</annotation></semantics></math> </ephtml> is a vector containing the predictors. <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">ε</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\bm{\epsilon }_{e}$</annotation></semantics></math> </ephtml> is a vector of random effects following <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>M</mi><mi>V</mi><mi>N</mi><mo>(</mo><mrow><mn mathvariant="bold">0</mn></mrow><mo>,</mo><msub><mi mathvariant="bold">Σ</mi><mrow><mi>B</mi><mi>l</mi><mi>o</mi><mi>c</mi><mi>k</mi></mrow></msub><mo>)</mo></mrow><annotation encoding="application/x-tex">$MVN(\bm{0}, \bm{\Sigma }_{Block})$</annotation></semantics></math> </ephtml> . With the data structure demonstrated in Table 5, each element is essentially associated with only one block—either the gameplay or the assessment block.</p> <hd id="AN0192629992-30">The measurement model</hd> <p>We assume a 1PL measurement model, with binary observed responses denoted by <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>y</mi><mrow><mi>v</mi><mi>e</mi><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$y_{veps}$</annotation></semantics></math> </ephtml> . One could explore the potential uses of more complex models if the data support such exploration. With a 1PL model, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>y</mi><mrow><mi>v</mi><mi>e</mi><mi>p</mi><mi>s</mi></mrow></msub><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$y_{veps} = 1$</annotation></semantics></math> </ephtml> indicates a correct response or the presence of a game‐based pattern, whereas <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>y</mi><mrow><mi>v</mi><mi>e</mi><mi>p</mi><mi>s</mi></mrow></msub><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">$y_{veps} = 0$</annotation></semantics></math> </ephtml> indicates an incorrect response or the absence of a game‐based pattern. The conditional response probability is 13 <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><semantics>P(yveps=1∣ηps,ξe)=11+exp[−(αv+λ1′ηps+λ2′ξe)].<annotation encoding="application/x-tex">$$\begin{align} P(y_{veps} = 1 \mid \bm{\eta }_{ps}, \bm{\xi }_{e}) &= \frac{1}{1+\exp {[-(\alpha _v + \bm{\lambda }_{1}^{\prime }\bm{\eta }_{ps} + \bm{\lambda }_{2}^{\prime }\bm{\xi }_{e})]}}. \end{align}$$</annotation></semantics></math> </ephtml></p> <p>In Equation 13, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>α</mi><mi>v</mi></msub><annotation encoding="application/x-tex">$\alpha _v$</annotation></semantics></math> </ephtml> is Variable <emph>v</emph>'s intercept. <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">λ</mi><mn>1</mn></msub><annotation encoding="application/x-tex">$\bm{\lambda }_{1}$</annotation></semantics></math> </ephtml> is a vector that contains Variable <emph>v</emph>'s slopes on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">η</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\eta }_{ps}$</annotation></semantics></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">η</mi><mrow><mi>p</mi><mi>s</mi></mrow></msub><annotation encoding="application/x-tex">$\bm{\eta }_{ps}$</annotation></semantics></math> </ephtml> is a vector of latent variables for Person <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>p</mi><annotation encoding="application/x-tex">$p$</annotation></semantics></math> </ephtml> in Site <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>s</mi><annotation encoding="application/x-tex">$s$</annotation></semantics></math> </ephtml> . <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">λ</mi><mn>2</mn></msub><annotation encoding="application/x-tex">$\bm{\lambda }_{2}$</annotation></semantics></math> </ephtml> is a vector that contains Variable <emph>v</emph>'s slopes on <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">ξ</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\bm{\xi }_{e}$</annotation></semantics></math> </ephtml> , where <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi mathvariant="bold-italic">ξ</mi><mi>e</mi></msub><annotation encoding="application/x-tex">$\bm{\xi }_{e}$</annotation></semantics></math> </ephtml> is a vector of latent variables for Element <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>e</mi><annotation encoding="application/x-tex">$e$</annotation></semantics></math> </ephtml> . With a 1PL model, the slopes across observed variables loading on the same latent variable are constrained to be equal.</p> <hd id="AN0192629992-31">Empirical Data Analysis</hd> <p>Now that we have discussed the data structure and the general modeling framework, we present our proposed application using cross‐classified IRT modeling and highlight its advantages over three alternative approaches. We apply each of the approaches to a data set collected from a large‐scale RCT of math games.</p> <p>We include the alternative approaches for illustrative purposes and for responding to a potential inquiry from general researchers. If simpler methods like correlational analysis, multiple linear regression, or more advanced ones like SEM using aggregated outcomes, produce similar patterns of findings regarding the intervention effect and the relationship between gameplay and assessments, why use cross‐classified IRT modeling?</p> <hd id="AN0192629992-32">Data Source</hd> <p>The data set used by this paper consists of gameplay and pre‐post item response data collected from 1,711 students in 24 schools participating in a multisite RCT of math learning games (Chung et al., [<reflink idref="bib17" id="ref100">17</reflink>]). Students were randomly assigned to the treatment condition ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>n</mi><mrow><mi>t</mi><mi>r</mi><mi>e</mi><mi>a</mi><mi>t</mi><mi>m</mi><mi>e</mi><mi>n</mi><mi>t</mi></mrow></msub><mo>=</mo><mn>873</mn></mrow><annotation encoding="application/x-tex">$n_{treatment} = 873$</annotation></semantics></math> </ephtml> ) and the control condition ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mi>n</mi><mrow><mi>c</mi><mi>o</mi><mi>n</mi><mi>t</mi><mi>r</mi><mi>o</mi><mi>l</mi></mrow></msub><mo>=</mo><mn>838</mn></mrow><annotation encoding="application/x-tex">$n_{control} = 838$</annotation></semantics></math> </ephtml> ) at the classroom level within each school. The treatment group played four games about rational numbers. The control group played four games about solving equations. Both groups received the same amount of instructional time. All eight games were developed with the same development process and personnel (i.e., the same team developing the knowledge specifications and the same team developing the games). The RCT met the group design standards of What Works Clearinghouse ([<reflink idref="bib78" id="ref101">78</reflink>]). Its study protocol was reviewed and approved by the University of California, Los Angeles (UCLA) Institutional Review Board (IRB).</p> <hd id="AN0192629992-33">Game of Interest</hd> <p>The analysis focuses on <emph>Save Patch</emph>, an educational game designed to target concepts of fraction addition. Figure 2 shows annotated screenshots of two tasks in the game. The top screenshot shows an earlier task that targets the addition of whole units. The bottom screenshot shows a later task that targets the addition of fractions.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar26/jedm12396-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12396-fig-0002.jpg" title="2 Two tasks in Save Patch. [Colour figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p>As shown in Figure 2, in each task, students are presented with a grid and rope pieces that have the length of one unit or fractions of a unit (e.g., halves). Students attempt each task by dragging various rope pieces to the signposts on the grid to form a path. The path is then followed by the game character. The objective is to use ropes with the correct lengths and guide the character from the starting location to the target location. The addition of the rope pieces follows the rules of fraction addition. For example, only rope pieces with the same unit or part of an unit (same denominator) can be added.</p> <hd id="AN0192629992-35">Assessment of Fractions Knowledge</hd> <p>The assessment data contain students' responses to 10 dichotomously scored items used for the pre‐ and post‐assessments. The items target the meanings of the unit, denominator, numerator, and addition of fractions. Figure 3 shows three sample items. Past research has shown that the assessment has high technical qualities based on traditional metrics, including classical item statistics. For more information on its development and validation, please refer to Vendlinski et al. ([<reflink idref="bib75" id="ref102">75</reflink>]) and Chung et al. ([<reflink idref="bib17" id="ref103">17</reflink>]).</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/MEA/01mar26/jedm12396-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jedm12396-fig-0003.jpg" title="3 Example assessment items." /> </p> <p></p> <hd id="AN0192629992-37">Game‐Based Indicators of Misconceptions about Fractions</hd> <p>The gameplay indicator data contain students' responses to nine binary indicators of misconceptions, derived using process data from 27 in‐game tasks. The choice of these indicators was driven by the set of knowledge specifications central to learning rational numbers (Vendlinski et al., [<reflink idref="bib75" id="ref104">75</reflink>]). The same set of knowledge specifications also underlay the development of the game and the pre‐ and post‐assessment items. The development of these indicators was influenced by specifications of the conceptual assessment framework within the evidence‐centered design framework (Kerr and Chung, [<reflink idref="bib44" id="ref105">44</reflink>]; Mislevy et al., [<reflink idref="bib57" id="ref106">57</reflink>]). The development also heavily relied on (a) the moment‐to‐moment information in the gameplay process data, (b) how well the in‐game interactions addressed cognitive demands related to fractions knowledge, and (c) the degree to which the gameplay data captured cognitively meaningful interactions (Kerr, [<reflink idref="bib43" id="ref107">43</reflink>], p. 90). For more information on development and validation, please refer to Kerr ([<reflink idref="bib43" id="ref108">43</reflink>]) and Kerr and Chung ([<reflink idref="bib44" id="ref109">44</reflink>], [<reflink idref="bib45" id="ref110">45</reflink>]).</p> <p>Table 6 presents the nine indicators and their definitions. For this paper, only data based on students' first submissions are used. The first submission window starts when each task begins and ends when a student clicks on the submit button and observes whether or not the submitted solution leads the game character to the target location.</p> <p>6 Table List of Game‐Based Indicators of Misconceptions about Fractions</p> <p> <ephtml> <table><thead><tr><th>No.</th><th align="center">Misconception</th><th align="center">Indicator</th><th align="center">Definition</th></tr></thead><tbody><tr><td>1</td><td>Avoiding math</td><td>Placed everything in order</td><td>Used all resources in the order that they were presented in a task.</td></tr><tr><td>2</td><td>Unitizing error</td><td>Saw as one unit</td><td>Saw the entire grid as one unit.</td></tr><tr><td>3</td><td /><td>Saw as wholes</td><td>Saw each fractional piece as a whole unit.</td></tr><tr><td>4</td><td>Partitioning error</td><td>Counted hash marks</td><td>Appeared to count the hash marks on the grid to determine the denominator.</td></tr><tr><td>5</td><td /><td>Counted hash marks and posts</td><td>Appeared to count the hash marks and posts on the grid to determine the denominator.</td></tr><tr><td>6</td><td>Unitizing and partitioning error</td><td>Saw as one unit and counted hash marks</td><td>Saw the entire grid as one unit but also appeared to count the hash marks on the grid to determine the denominator.</td></tr><tr><td>7</td><td /><td>Saw as one unit and counted hash marks and posts</td><td>Saw the entire grid as one unit but also appeared to count the hash marks and posts on the grid to determine the denominator.</td></tr><tr><td>8</td><td>Iterating error</td><td>Wrong numerator</td><td>Added the wrong number of fractional units.</td></tr><tr><td>9</td><td>Converting to wholes error</td><td>Saw as mixed number</td><td>Saw the solution as a mixed number and tried to add a whole unit and a fractional unit without converting everything to have the same denominator.</td></tr></tbody></table> </ephtml> </p> <hd id="AN0192629992-38">Core Research Questions</hd> <p>We address two research questions that are often asked in RCTs of learning games. The first question concerns evaluating the effectiveness of the intervention in promoting understanding of fraction concepts. The second question concerns how students' performance in the game relates to their performance on traditional assessments.</p> <p></p> <ulist> <item> <bold> RQ1. _B_Treatment Effect</bold> : To what extent do students in the treatment group, who are assigned to play <emph>Save Patch</emph> , differ from students in the control group in their knowledge of fractions as measured by the pre‐ and post‐assessments?</item> <p></p> <item> <bold> RQ2. _B_Relationship between Gameplay and Assessments</bold> : To what extent is students' game‐based performance (e.g., misconception) related to their knowledge of fractions?</item> </ulist> <hd id="AN0192629992-39">Data Analyses</hd> <p>We compare our proposed approach, cross‐classified IRT modeling, to three alternative approaches used for summarizing and analyzing gameplay and assessment data. By juxtaposing these approaches, we discuss the extent to which each approach addresses the two research questions, discuss key model estimates, and argue for the application of cross‐classified IRT modeling.</p> <p>The four approaches proceed from simpler techniques to more advanced ones, including (a) basic descriptive statistics and correlations, (b) multiple linear regression, (c) structural equation modeling (SEM), and our proposed approach (d) cross‐classified IRT modeling. For (a)‐(c), we used assessment sum scores and aggregated game‐based indicators. For (b)‐(d), we created school dummy variables (school‐level fixed effects) to account for the nesting of students within schools and the potential impact of school‐level differences on individuals' outcomes.</p> <hd id="AN0192629992-40">Why We Used School Fixed Effects Rather than Random Effects (Variability)?</hd> <p>To account for the nesting of individuals within schools and school‐level differences, there are two common approaches: (a) random effects or multilevel models and (b) fixed effects models that include a dummy variable for each school. The choice between fixed effects and random effects depends on the underlying assumptions about the nature of these effects and aims of the study. With the random effects approach, we can use multilevel modeling to allow for varying school‐specific intercepts (random intercepts) and further examine varying treatment effects across schools (random coefficients).</p> <p>In our case, however, we are not interested in modeling variability, whether it is school‐level variability or variability (heterogeneity) in treatment effects across schools, an analysis that would be useful in substantive research. Instead, our objective is to enable a gradual transition from building a simpler model, such as a multiple linear regression model, to building more complex (latent variable) models covered in later sections. Modeling school fixed effects achieves this objective while providing one way to control for nesting and school‐level differences. Therefore, we include school‐specific dummy variables in examples using multiple linear regression, SEM, or cross‐classified IRT modeling.</p> <hd id="AN0192629992-41">Summary of Results</hd> <p>Table 7 summarizes results from all four analyses. Please note that the effect sizes, along with other results, are provided to show a consistent overall trend (e.g., positive intervention effect). Because values displayed in Table 7 were obtained using different analytic approaches, and each approach used different scaling and parameters, values generated by different analyses, including the effect sizes, are not comparable.</p> <p>7 Table Summary of Results</p> <p> <ephtml> <table><thead><tr><th>RQ1. Treatment Effect</th></tr></thead><tbody><tr><td>Approach</td><td>Effect size</td><td align="center">Effect size is relative to</td></tr><tr><td>Descriptive statistics</td><td>.18</td><td>Pooled variance of observed posttest sum scores</td></tr><tr><td>Multiple regression</td><td>.19</td><td>Unit variance of observed posttest sum scores</td></tr><tr><td>SEM<sup>a</sup></td><td>.29</td><td>Unit variance of posttest latent variable</td></tr><tr><td>Cross‐classified IRT</td><td>.26</td><td>Estimated variance of pre‐to‐post latent change variable</td></tr></tbody></table> </ephtml> </p> <p></p> <p> <ephtml> <table><thead><tr><th>RQ2. Relationship between Gameplay and Assessments</th></tr></thead><tbody><tr><td>With observed variables only</td></tr><tr><td>Correlational analysis</td><td>1.</td><td>Proportion of in‐game tasks missed negatively correlated with pretest sum scores (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo>27<mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo>001<annotation encoding="application/x-tex">$\rho = -.27, p <.001$</annotation></semantics></math></p>) and posttest sum scores (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo>33<mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo>001<annotation encoding="application/x-tex">$\rho = -.33, p <.001$</annotation></semantics></math></p>).</td></tr><tr><td /><td>2.</td><td>Average number of misconceptions negatively correlated with pretest (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo>46<mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo>001<annotation encoding="application/x-tex">$\rho = -.46, p <.001$</annotation></semantics></math></p>) and posttest sum scores (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo>47<mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo>001<annotation encoding="application/x-tex">$\rho = -.47, p <.001$</annotation></semantics></math></p>).</td></tr><tr><td>Multiple linear regression</td><td>1.</td><td>N/A. The treatment indicator and game‐based indicators could not be included in the same model. A separate model with data from only the treatment group is needed to gauge the relationship.</td></tr><tr><td>With observed and latent variables</td></tr><tr><td>SEM</td><td>1.</td><td>The gameplay latent variable, measured through average misconception and proportion of in‐game tasks missed, had a negative impact on the posttest (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>β</mi><mo>̂</mo><annotation encoding="application/x-tex">$\hat{\beta }$</annotation></semantics></math></p> = –3.33, <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>S</mi><mi>E</mi><annotation encoding="application/x-tex">$SE$</annotation></semantics></math></p> =.51, <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>p</mi><mo><</mo><mo>.</mo>001<annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math></p>).</td></tr><tr><td /><td>2.</td><td>The gameplay and pretest latent variables were negatively correlated (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>ϕ</mi><mo>̂</mo><annotation encoding="application/x-tex">$\hat{\phi }$</annotation></semantics></math></p> = –.20, <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>S</mi><mi>E</mi><annotation encoding="application/x-tex">$SE$</annotation></semantics></math></p> =.03, <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>p</mi><mo><</mo><mo>.</mo>001<annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math></p>).</td></tr><tr><td /><td>3.</td><td>With the SEM framework, alternative analyse (e.g., mediation) may also be performed.</td></tr><tr><td>Cross‐classified IRT</td><td>1.</td><td>The game‐based misconception latent variable was negatively correlated with the baseline latent variable (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>σ</mi><mo>̂</mo><annotation encoding="application/x-tex">$\hat{\sigma }$</annotation></semantics></math></p> = –.64, <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>S</mi><mi>E</mi><annotation encoding="application/x-tex">$SE$</annotation></semantics></math></p> =.07).</td></tr><tr><td /><td>2.</td><td>The game‐based misconception latent variable was negatively correlated with the pre‐ to post‐latent change variable (<p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>σ</mi><mo>̂</mo><annotation encoding="application/x-tex">$\hat{\sigma }$</annotation></semantics></math></p> = –.15).</td></tr></tbody></table> </ephtml> </p> <ulist> <item>3 <emph>Note</emph>. Each treatment effect is a form of standardized estimate. Each approach uses different scaling and parameters. Values listed in this table are meant to show a general pattern or trend in results (e.g., positive treatment effects). These values are not comparable.</item> <item>4 SEM: structural equation modeling.</item> </ulist> <p>We also note that our proposed approach using cross‐classified IRT modeling was able to directly examine the effect of the game‐based intervention on the pretest‐to‐posttest change while making use of item‐level assessment data and nonaggregated game‐based indicator data.</p> <hd id="AN0192629992-42">Descriptive statistics and correlational analysis</hd> <p>Descriptive statistics is a typical starting point for understanding both assessment and gameplay data. These statistics include numerical summaries such as the mean and standard deviation for each variable.</p> <hd id="AN0192629992-43">RQ1. Treatment Effect</hd> <p>Table 8 presents the sample size excluding missing data (<emph>n</emph>), mean (<emph>M</emph>), and standard deviation (<emph>SD</emph>) of the pretest and posttest sum scores for the treatment and control groups.</p> <p>8 Table Pretest and Posttest Assessment Sum Scores by Experimental Conditions</p> <p> <ephtml> <table><thead><tr><th>Condition</th><th align="center">Time</th><th><italic>n</italic></th><th><italic>M</italic></th><th><italic>SD</italic></th></tr></thead><tbody><tr><td>Control</td><td>Pretest</td><td>801</td><td>4.57</td><td>2.64</td></tr><tr><td /><td>Posttest</td><td>774</td><td>4.80</td><td>2.79</td></tr><tr><td>Treatment</td><td>Pretest</td><td>851</td><td>4.52</td><td>2.68</td></tr><tr><td /><td>Posttest</td><td>801</td><td>5.33</td><td>2.95</td></tr><tr><td>Cohen's <p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><mi>d</mi><annotation encoding="application/x-tex">$d$</annotation></semantics></math></p> for posttest</td><td align="center">.18</td></tr></tbody></table> </ephtml> </p> <p>5 <emph>Note</emph>. The total sample size is 1,711 across 24 schools, with 873 students assigned to the treatment condition and 838 assigned to the control condition.</p> <p>We make two observations from results presented in this table. First, the treatment and control groups had similar mean pretest scores and the standard deviations of the scores ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mrow><mi>c</mi><mi>o</mi><mi>n</mi><mi>t</mi><mi>r</mi><mi>o</mi><mi>l</mi></mrow></msub><annotation encoding="application/x-tex">$M_{control}$</annotation></semantics></math> </ephtml> = 4.57, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mrow><mi>c</mi><mi>o</mi><mi>n</mi><mi>t</mi><mi>r</mi><mi>o</mi><mi>l</mi></mrow></msub></mrow><annotation encoding="application/x-tex">$SD_{control}$</annotation></semantics></math> </ephtml> = 2.64; <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msub><mi>M</mi><mrow><mi>t</mi><mi>r</mi><mi>e</mi><mi>a</mi><mi>t</mi><mi>m</mi><mi>e</mi><mi>n</mi><mi>t</mi></mrow></msub><annotation encoding="application/x-tex">$M_{treatment}$</annotation></semantics></math> </ephtml> = 4.52, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><msub><mi>D</mi><mrow><mi>t</mi><mi>r</mi><mi>e</mi><mi>a</mi><mi>t</mi><mi>m</mi><mi>e</mi><mi>n</mi><mi>t</mi></mrow></msub></mrow><annotation encoding="application/x-tex">$SD_{treatment}$</annotation></semantics></math> </ephtml> = 2.68). The similarity in pretest scores suggests that the randomization process succeeded in creating two groups with similar levels of prior knowledge as measured by the pretest. Second, following the game‐based intervention, the mean posttest score of the treatment group was.53 points higher than that of the control group, indicating a difference of.18 pooled posttest standard deviation. The second observation suggests that the learning game helped enhance students' knowledge of fractions. However, a difference derived from the posttest data alone offers limited insights, as it does not account for students' pretest performance, their performance in the game‐based intervention, and the complexities introduced by the multisite design of the RCT. We can address these limitations by using a more sophisticated approach, as discussed in later sections.</p> <p>Here, we also include two aggregated game‐based indicators: the proportion of unique in‐game tasks missed and the average frequency of misconceptions. These indicators are used in subsequent analyses, except for cross‐classified IRT modeling where nonaggregated indicators can be used. Table 9 shows the descriptive statistics for the two aggregated indicators.</p> <p>9 Table Descriptive Statistics of Two Aggregated Game‐Based Indicators</p> <p> <ephtml> <table><thead><tr><th>Variable</th><th><italic>M</italic></th><th><italic>SD</italic></th><th>Min.</th><th>Max.</th></tr></thead><tbody><tr><td>Average frequency of misconceptions</td><td>.29</td><td>.16</td><td>.00</td><td>.71</td></tr><tr><td>Proportion of unique tasks missed</td><td>.06</td><td>.16</td><td>.00</td><td>.93</td></tr></tbody></table> </ephtml> </p> <p>6 <emph>Note</emph>. Results are based on data from 826 students with analyzable gameplay data. The average frequency of misconceptions is computed based on values summed across nine indicators and averaged across in‐game tasks played by each individual.</p> <p>The rationale for choosing the two indicators is twofold. First, the proportion of in‐game tasks missed (or conversely, tasks played) represents count‐ or proportion‐based indicators commonly used in existing literature to gauge in‐game performance or progress. The preference for this type of indicators may be attributed to their straightforward definition and ease of computation. In our case, the more unique tasks players attempted, the further they progressed in the game, and likely the more they knew or learned. To compute the proportion, we counted the number of in‐game tasks not played by an individual and divided this count by 27, which was the total number of unique tasks. Of note, the tasks were organized into stages, and each stage had its gameplay flow and fraction knowledge specifications[<reflink idref="bib1" id="ref111">1</reflink>] (Center for Advanced Technology in Schools, [<reflink idref="bib11" id="ref112">11</reflink>]).</p> <p>Second, the average frequency of misconceptions was calculated based on the nine specific misconception indicators shown in Table 6. When computing the average frequency of misconceptions for each student, we noted that each of the nine binary indicators denoted the presence or absence of a particular fraction‐related misconception. We aggregated these indicator values up to the individual level by summing the number of 1s across the nine indicators and then averaging this sum by the number of tasks played by an individual. The aggregated indicator provides an overall measure of the frequency of misconceptions related to adding fractions (in students' initial submissions).</p> <hd id="AN0192629992-44">RQ2. Relationship between Gameplay and Assessments</hd> <p>We can display gameplay and assessment data graphically in more than one dimension to inspect potential relationships among the variables. On the other hand, creating graphs becomes more complex and cumbersome as the number of variables grows. In such cases, correlational analysis provides one way to summarize the pairwise relationships, including their directions and strengths, observed through visualization.</p> <p>Table 10 shows the pairwise nonparametric correlations (Spearman's rho) between the aggregated game‐based indicators and assessment scores. The magnitudes of the correlations between indicators and assessment scores were weak to moderate. Specifically, students who missed more in‐game tasks tended to exhibit more misconceptions in their first submissions ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ρ</mi><mo>=</mo><mo>.</mo><mn>29</mn><mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$\rho =.29, p <.001$</annotation></semantics></math> </ephtml> ) while having lower pretest ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo><mn>27</mn><mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$\rho = -.27, p <.001$</annotation></semantics></math> </ephtml> ) and posttest sum scores ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo><mn>33</mn><mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$\rho = -.33, p <.001$</annotation></semantics></math> </ephtml> ). Students exhibiting more misconceptions tended to have lower pretest ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo><mn>46</mn><mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$\rho = -.46, p <.001$</annotation></semantics></math> </ephtml> ) and posttest sum scores ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>ρ</mi><mo>=</mo><mo>−</mo><mo>.</mo><mn>47</mn><mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$\rho = -.47, p <.001$</annotation></semantics></math> </ephtml> ).</p> <p>10 Table Correlations between Game‐Based and Assessment Measures (n$n$ = 786‐826)</p> <p> <ephtml> <table><thead><tr><th /><th align="center">1</th><th>2</th><th>3</th></tr></thead><tbody><tr><td>1. Prop. unique tasks missed</td><td align="center">‐</td><td /><td /></tr><tr><td>2. Avg. frequency of misconceptions</td><td>.29<ext-link href="***" /></td><td>‐</td><td /></tr><tr><td>3. Pretest sum scores</td><td>−.27<ext-link href="***" /></td><td>−.46<ext-link href="***" /></td><td>‐</td></tr><tr><td>4. Posttest sum scores</td><td>−.33<ext-link href="***" /></td><td>−.47<ext-link href="***" /></td><td>.74<ext-link href="***" /></td></tr></tbody></table> </ephtml> </p> <p>7 <emph>Note</emph>. *** <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> .</p> <p>The findings presented above are consistent with our expected relationships between misconceptions detected during gameplay and assessment scores. By inspecting descriptive statistics and pairwise correlations, we have two findings relevant for answering the core research questions. First, we found a difference of.18 pooled standard deviations based on the posttest sum scores, which falls within the typical range observed in studies focusing on game‐based (math) learning.[<reflink idref="bib2" id="ref113">2</reflink>] While this magnitude may appear small when compared to the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mo>></mo><mo>.</mo><mn>40</mn></mrow><annotation encoding="application/x-tex">$>.40$</annotation></semantics></math> </ephtml> hinge point (Hattie, [<reflink idref="bib35" id="ref114">35</reflink>]; Mayer, [<reflink idref="bib54" id="ref115">54</reflink>]), we caution against overreliance on using a fixed threshold to evaluate instructional effectiveness universally, as argued by Hattie ([<reflink idref="bib35" id="ref116">35</reflink>], p. 30) and Kraft ([<reflink idref="bib48" id="ref117">48</reflink>]). Second, we observed negative relationships between game‐based nonpositive behaviors (tasks missed and misconceptions) and higher assessment scores. These two findings offer preliminary insights into whether the game‐based intervention positively impacted students' knowledge of fractions and how students' in‐game performance related to their assessment performance.</p> <p>However, insights gleaned from descriptive statistics and pairwise correlations are limited and fragmented. First, the assertion about the treatment effect relied on comparing the experimental groups' posttest sum scores without considering the nesting of individuals within schools. Second, correlations were only about pairwise associations. These concerns can be addressed by using a more advanced method that accommodates multiple variables or predictors, and models either school fixed effects or random effects to account for nesting. Lastly, findings were derived from separate analyses and were not integrated, as they would be in a unified modeling framework. For example, findings from the correlational analysis were not considered in the calculation of the noted.18 difference. In the next section, we use a multiple linear regression model with added school dummy variables to address some of the concerns discussed thus far.</p> <hd id="AN0192629992-45">Multiple linear regression</hd> <p>With a multiple linear regression model, we can examine the relationships between students' posttest scores and other variables, including pretest sum scores and the treatment assignment indicator.</p> <hd id="AN0192629992-46">Model Specification</hd> <p>We specified the multiple regression model with school fixed effects as follows. We regressed the dependent variable, posttest sum scores, on pretest sum scores, treatment condition indicator, and 23 school dummy variables with School 1 as the reference school [<emph>F</emph>(<reflink idref="bib25" id="ref118">25</reflink>, 1504) = 90.74, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> , <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><msup><mi>R</mi><mn>2</mn></msup><annotation encoding="application/x-tex">$R^2$</annotation></semantics></math> </ephtml> =.60]. The inclusion of school dummy variables allowed each school to have its own intercept. Each coefficient associated with a dummy variable denoted a shift in the intercept for a specific school in comparison to the reference school.</p> <hd id="AN0192629992-47">RQ1. Treatment Effect</hd> <p>Results presented in Table 11 showed a positive estimated coefficient on the treatment assignment indicator ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mover accent="true"><mi>β</mi><mo>̂</mo></mover><annotation encoding="application/x-tex">$\hat{\beta }$</annotation></semantics></math> </ephtml> =.19, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.04, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>t</mi><annotation encoding="application/x-tex">$t$</annotation></semantics></math> </ephtml> = 4.74, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> ) after controlling for pretest sum scores and school‐level differences. In simpler terms, we observed a positive estimated effect of the game‐based intervention, which was about.19 <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>D</mi></mrow><annotation encoding="application/x-tex">$SD$</annotation></semantics></math> </ephtml> increase in posttest sum scores for the treatment group compared to the control group. Note that the effect of.19 was obtained after standardizing the pretest and posttest sum scores. Each sum score variable was standardized by subtracting from it its sample mean and dividing it by its standard deviation.</p> <p>11 Table Multiple Regression Results Using Posttest Sum Scores as the Dependent Variable</p> <p> <ephtml> <table><thead><tr><th /><th align="center">Estimate</th><th align="center"><italic>SE</italic></th><th align="center"><italic>t</italic></th></tr></thead><tbody><tr><td>Intercept</td><td>.23<ext-link href="*" /></td><td>.09</td><td>2.46</td></tr><tr><td>Pretest sum scores</td><td>.73<ext-link href="***" /></td><td>.02</td><td>41.91</td></tr><tr><td>Treatment</td><td>.19<ext-link href="***" /></td><td>.04</td><td>4.74</td></tr></tbody></table> </ephtml> </p> <ulist> <item>8 <emph>Note</emph>. Pretest and posttest sum scores are mean‐centered and scaled by one <emph>SD</emph>. The estimated coefficients on the school dummy variables are not shown in the table; these estimates range between −.68 ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.12) and −.07 ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.11).</item> <item>9 * <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>05</mn></mrow><annotation encoding="application/x-tex">$p <.05$</annotation></semantics></math> </ephtml> ; *** <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> .</item> </ulist> <p>Given a 10‐item assessment that was dichotomously scored, it may also be of interest to examine the unstandardized coefficient on the treatment assignment indicator. With the same set of predictors, the estimated coefficient became <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>B</mi><mo>̂</mo></mover><mo>=</mo><mo>.</mo><mn>55</mn></mrow><annotation encoding="application/x-tex">$\hat{B} =.55$</annotation></semantics></math> </ephtml> ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.12, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>t</mi><annotation encoding="application/x-tex">$t$</annotation></semantics></math> </ephtml> = 4.74, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> ). In other words, on average, students who played the game about adding fractions ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mi>r</mi><mi>e</mi><mi>a</mi><mi>t</mi><mi>m</mi><mi>e</mi><mi>n</mi><mi>t</mi><mo>=</mo><mn>1</mn></mrow><annotation encoding="application/x-tex">$Treatment = 1$</annotation></semantics></math> </ephtml> ) scored.55 points higher in their posttest sum scores or answered.55 more items correctly in their posttest, compared to students who did not play the game about adding fractions ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>T</mi><mi>r</mi><mi>e</mi><mi>a</mi><mi>t</mi><mi>m</mi><mi>e</mi><mi>n</mi><mi>t</mi><mo>=</mo><mn>0</mn></mrow><annotation encoding="application/x-tex">$Treatment = 0$</annotation></semantics></math> </ephtml> ).</p> <hd id="AN0192629992-48">RQ2. Relationship between Gameplay and Assessments</hd> <p>We did not include any game‐based indicator into the model, and as a result, we did not have results pertaining to the second research question.</p> <p>Why didn't we include any game‐based indicators? Recall that roughly half of the sample was assigned to the treatment condition and asked to play the game <emph>Save Patch</emph> about adding fractions, and there were no <emph>Save Patch</emph> data for the other half of the sample. If we were to include any game‐based indicator into the multiple regression model, we would lose half of our data—data of students in the control condition, making it impossible to estimate the coefficient on the treatment assignment indicator. Said differently, it would be impossible to estimate the treatment effect given the exclusion of the control group's data. An alternative is to create two separate models: one considering the treatment assignment indicator and another considering the game‐based indicators.</p> <p>Multiple linear regression has several advantages over descriptive statistics and pairwise correlations. One advantage is its ability to control for multiple predictors, such as students' pretest sum scores, and obtain a more precise estimate of the treatment effect. Another advantage is its ability to account for school‐level differences by adding school dummy variables (school‐level fixed effects).</p> <p>However, with the aforementioned multiple linear regression model, we cannot simultaneously examine the effectiveness of the game‐based intervention and the relationship between gameplay and assessments. The model also limits our ability to pose additional questions or conduct further analyses. For example, we cannot investigate how student‐specific factors affect assessment performance by regressing a variable representing the overall assessment performance on background variables. Meanwhile, we cannot explore how elements of game design affect in‐game performance by regressing a variable representing the overall in‐game performance on indicators of game‐specific cognitive or mechanic features. To overcome these limitations, we turn to latent variable models with added school dummy variables. We begin with the use of structural equation modeling.</p> <hd id="AN0192629992-49">Structural equation modeling (SEM)</hd> <p>Researchers may use structural equation models to jointly analyze assessment sum scores and aggregated game‐based indicators, moving from the modeling of observed variables to that of both observed and latent variables. We include this SEM example for illustrative purposes only, showing a latent variable modeling approach different from cross‐classified IRT modeling.</p> <p>Before we specify the model, we note two potential modeling limitations associated with using SEM to jointly analyze aggregated gameplay and assessment outcomes. First, joint modeling presented in this SEM example relies on treating the two game‐based indicators as outcome variables instead of covariates, and on assuming all outcome variables to be conditionally normal. Second, in this SEM example, the pretest and posttest latent variables are single‐indicator latent variables. This setup assumes that the a single indicator (e.g., pretest sum score) fully reflects the measured phenomenon (e.g., pretest performance) without measurement error (Raykov & Marcoulides, [<reflink idref="bib66" id="ref119">66</reflink>]). Both the assumption of normality and the assumption of no measurement error likely do not hold in this specific context. Therefore, we caution readers who still want to apply SEM to the analysis of aggregated gameplay and assessment data to stay vigilant about the single‐indicator situation and to conduct thorough checks for possible violations of model assumptions.</p> <hd id="AN0192629992-50">Model Specification</hd> <p>The specified model had three latent variables. The first latent variable, labeled PRETEST, represented pretest performance. The second, labeled POSTTEST, represented posttest performance. The third latent variable, labeled GAMEPLAY, reflected nonpositive gameplay performance. PRETEST and POSTTEST were single‐indicator latent variables. For example, the PRETEST latent variable predicted only one observed outcome variable, the pretest sum score. Similarly, the POSTTEST latent variable predicted only one observed outcome variable, the posttest sum score. When connecting an observed outcome variable with a single‐indicator latent variable, we fixed the observed variable's factor loading to 1 and its error variance to 0. We also fixed the constant intercept term of the observed variable to be the observed variable's mean.</p> <p>The third latent variable, GAMEPLAY, predicted two observed outcome variables: the proportion of in‐game tasks missed and the average frequency of misconceptions. GAMEPLAY reflected nonpositive performance because the two game‐based indicators were conceptually and empirically shown, for example, through pairwise correlations, to be negatively related to higher test scores. In the model, we fixed the factor loading of the misconception indicator to 1, leaving the other indicator's factor loading freely estimated. We also freely estimated the error variances associated with both game‐based indicators.</p> <p>Among the latent variables, we regressed POSTTEST on PRETEST and GAMEPLAY, respectively. We estimated the variance associated with each latent variable or its disturbance, and we also estimated the covariance between PRETEST and GAMEPLAY. We assumed that the means of the latent variables were 0.</p> <p>To estimate the treatment effect, we regressed POSTTEST on the treatment assignment indicator. We estimated both unstandardized and standardized treatment effects. The standardized solution was obtained by invoking the <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>S</mi></mrow><annotation encoding="application/x-tex">$SS$</annotation></semantics></math> </ephtml> command in LISREL 12 (Jöreskog & Sörbom, [<reflink idref="bib42" id="ref120">42</reflink>]), which standardized all latent variables.</p> <p>To account for school‐level differences, we also regressed POSTTEST on 23 school dummy variables, using School 1 as the reference school. As explained at the start of the "Data Analyses" section, we are not interested in exploring the variability of treatment effects across schools. Instead, we include school dummy variables as a method to address the nesting of students in schools. Our focus remains on the joint analysis of gameplay and assessment data.</p> <p>Note that the specified SEM model is one of the possible models for examining the relationship between gameplay and assessments. An alternative model could involve GAMEPLAY acting as a mediator, in which case the interpretation of the treatment effect, as indicated by the coefficient on the treatment assignment indicator, can get more complicated. The model discussed in this section is used to build on the multiple regression analysis, where we regressed posttest sum scores on pretest sum scores, the treatment assignment indicator, and school dummy variables.</p> <p>With the above specifications, we estimated the model with full information maximum likelihood in LISREL 12 (Jöreskog & Sörbom, [<reflink idref="bib42" id="ref121">42</reflink>]).</p> <hd id="AN0192629992-51">RQ1. Treatment Effect</hd> <p>We examined the coefficient on the treatment assignment indicator to gauge the treatment effect. The standardized estimated coefficient was.29 ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.01, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>t</mi><annotation encoding="application/x-tex">$t$</annotation></semantics></math> </ephtml> = 24.57, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> ), suggesting a positive impact of the game‐based intervention on the latent posttest performance, after considering the relationships between POSTTEST and other variables of interest (e.g., PRETEST).</p> <p>Compared with the effect obtained in the multiple regression analysis, this effect of.29 was estimated while considering the following information: (a) students' pretest performance; (b) information or process evidence from the intervention (game) itself; and (c) the interplay between game‐ and assessment‐based outcomes. With this joint analysis, our claim about the effectiveness of the game‐based intervention on student knowledge outcomes becomes grounded in multiple sources of information. Such a joint analysis differs from the previous analyses that solely relied on assessment data.</p> <hd id="AN0192629992-52">RQ2. Relationship between Gameplay and Assessments</hd> <p>There were two findings on the relationship between gameplay and assessments. First, GAMEPLAY, a latent variable measured through the average frequency of misconception and the proportion of in‐game tasks missed, had a negative impact on POSTTEST ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mover accent="true"><mi>β</mi><mo>̂</mo></mover><annotation encoding="application/x-tex">$\hat{\beta }$</annotation></semantics></math> </ephtml> = −3.33, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.51, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>t</mi><annotation encoding="application/x-tex">$t$</annotation></semantics></math> </ephtml> = 6.50, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> ). This result suggests that students with better performance in the game, such as exhibiting fewer misconceptions, tended to score higher on the posttest, after accounting for influences of other variables of interest on POSTTEST. Second, GAMEPLAY and PRETEST were negatively correlated ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mover accent="true"><mi>ϕ</mi><mo>̂</mo></mover><annotation encoding="application/x-tex">$\hat{\phi }$</annotation></semantics></math> </ephtml> = −.20, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.03, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mi>t</mi><annotation encoding="application/x-tex">$t$</annotation></semantics></math> </ephtml> = 7.13, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>p</mi><mo><</mo><mo>.</mo><mn>001</mn></mrow><annotation encoding="application/x-tex">$p <.001$</annotation></semantics></math> </ephtml> ), suggesting that students who scored higher on the pretest or started with greater prior knowledge of fractions tended to perform better in the game, and hence having fewer misconceptions and fewer missed tasks.</p> <p>The patterns we observed with the SEM approach, including the sign and magnitude of the treatment effect and how gameplay related to assessments, were consistent with patterns seen in previous analyses. First, we observed positive treatment effects across all analyses discussed thus far. Second, in both correlational and SEM analyses, we observed an overall negative association between nonpositive gameplay, such as showing misconceptions, and better assessment performance.</p> <p>We also note the methodological differences in what model parameters were involved and how the effects were estimated. For instance, the SEM‐based effect size was obtained after standardizing the latent variables and thus was relative to the unit variance of a latent variable, not an observed variable. In comparison, the effect size computed using descriptive statistics was based on observed posttest sum scores exclusively. The purpose of discussing these effect sizes is not to imply their comparability. Rather, we discuss them to discern the overall trend in how the game‐based intervention influenced student knowledge outcomes.</p> <p>By using a structural equation model, we can address a key drawback of the multiple linear regression approach: the inability to jointly include gameplay and assessment data. With multiple regression, we cannot include gameplay data with the treatment assignment indicator in a single regression model without losing roughly half of the sample. Consequently, the multiple regression model that includes the treatment assignment indicator lacks any variables related to gameplay. What sets a more advanced approach like SEM apart from earlier, simpler methods is its ability to simultaneously analyze gameplay and assessment data while also having the flexibility to handle certain data complexities and enable researchers to investigate secondary hypotheses.</p> <p>In addition to the two modeling limitations mentioned earlier, another limitation of the SEM approach lies in the use of aggregated data. Item‐level assessment data were aggregated into sum scores, and task‐specific game‐based indicators were aggregated into individual‐specific values. This aggregation procedure, particularly concerning gameplay data, overlooks the varying characteristics of the in‐game tasks. In <emph>Save Patch</emph>, the tasks are designed to gradually introduce key concepts and skills about adding fractions as players progress through the game (Center for Advanced Technology in Schools, [<reflink idref="bib11" id="ref122">11</reflink>]). For example, players first tackle the addition of whole units before delving into the addition of fractions with the same denominator but different numerators, and ultimately, fractions with varying denominators and numerators. The varying characteristics of the in‐game tasks can affect how players engage with and perform in the game, encouraging us to explore an alternative that leverages nonaggregated assessment and gameplay data.</p> <hd id="AN0192629992-53">Cross‐classified item response theory modeling</hd> <p>We have analyzed assessment and gameplay data using three approaches: descriptive statistics and correlational analysis, multiple linear regression, and SEM. For each approach, we discussed the findings and limitations, highlighting the additional insights gained from using a more flexible and complex approach.</p> <p>With descriptive statistics and correlational analysis, we obtained preliminary insights into the extent to which the game‐based intervention improved student knowledge outcomes (e.g., posttest sum scores). However, this approach had several limitations, including (a) the inability to adjust for students' pretest or baseline performance when comparing posttest performance between the treatment and control groups; (b) the inability to account for data complexities, such as the nesting of students within schools; and (c) the absence of integrated findings due to the lack of a unified and flexible modeling framework that jointly analyzes data from different sources, such as gameplay and assessments.</p> <p>With multiple linear regression, we addressed two limitations associated with the use of descriptive statistics and correlational analysis. First, we compared students' posttest sum scores while also considering their pretest sum scores. Second, we included school dummy variables into the regression model (school fixed effects) as a means to account for the nesting of students within schools and the impact of school‐level differences on student outcomes. However, one serious drawback remained: the inability to jointly analyze gameplay and assessment data. Including any gameplay data into this regression model would result in the loss of approximately half of the data, specifically data of students in the control condition who did not play the game, and the loss of the ability to estimate the treatment effect.</p> <p>With SEM, we included both aggregated assessment sum scores and game‐based indicators in the same model while retaining the ability to address complexities such as the nesting of students within schools. While we appreciated the flexibility and the potential for alternative model specifications offered by the SEM framework, we also noted the limitations of a SEM‐based approach that relied on aggregated outcomes. Data aggregation, in the context of game‐based research, overlooked the varying characteristics of in‐game tasks introduced by design.</p> <p>Furthermore, all three preceding approaches did not directly model changes in students' performance. We would be one step closer to making claims about student learning if we could directly examine changes in performance, such as changes from pre‐ to post‐assessments. The earlier methods, including descriptive statistics, correlational analysis, and multiple linear regression, used posttest sum scores as the outcome, controlling for pretest sum scores when applicable. In the SEM application using aggregated outcomes, we regressed a latent variable representing posttest performance on the treatment assignment indicator. Although SEM, particularly longitudinal SEM using latent‐change concepts (e.g., McArdle, [<reflink idref="bib55" id="ref123">55</reflink>]), can model latent changes, we did not explore additional structural equation models because we wanted to transition away from using aggregated data, especially aggregated game‐based indicator data.</p> <p>Given the aforementioned considerations, we now present our proposed approach that uses a cross‐classified IRT model to jointly analyze nonaggregated gameplay and assessment data.</p> <hd id="AN0192629992-54">Structure of Input Data</hd> <p>The general structure of the input data mirrors the structure presented in Table 5 in the "Example of the Combined Data" section. The term "Block" refers to either the assessment block of items or the gameplay block of in‐game tasks. Person‐specific latent variables are referred to as cluster side latent variables. Item‐ or game task‐specific latent variables are referred to as block side latent variables.</p> <hd id="AN0192629992-55">Model Specification</hd> <p>The model is a specific case within the general cross‐classified IRT modeling framework described earlier.</p> <p>The latent structural component of the model had a total of five latent variables. Among these, three were on the cluster side, representing individual game‐based misconception, baseline performance, and pretest to posttest change, respectively. We assumed a consistent misconception latent variable because the gameplay duration and exposure for one single game was relatively short. The other two latent variables were on the block side. Specifically, the gameplay block captured varying relative intercepts of in‐game tasks, while the assessment block captured varying relative intercepts of pre‐ and post‐assessment items.</p> <p>We imposed or relaxed constraints on the means, variances, and covariances associated with the latent variables as follows. For the misconception and baseline latent variables, as well as the two block side latent variables, we constrained their means to 0 and variances to 1. We freely estimated (a) the mean and the variance of the latent change variable, (b) the covariance between the misconception and baseline latent variables, and (c) the covariance between the misconception and latent change variables.</p> <p>To estimate the treatment effect and account for school effects, we regressed the baseline and change latent variables on covariates. Specifically, we regressed both the baseline and change latent variables on the treatment assignment indicator. Also, we regressed the baseline latent variable on 23 school dummy variables, making School 1 the reference school.</p> <p>We specified the relationships between observed and latent variables as follows, starting with the cluster side latent variables and moving to the block side latent variables. First, the nine game‐based indicators loaded on the first latent variable (misconception), and their slopes were constrained to be equal.[<reflink idref="bib3" id="ref124">3</reflink>] Second, the two observed variables containing assessment item responses, denoted as pretest and posttest, loaded on the second latent variable (baseline), while the posttest variable also loaded on the third latent variable (latent change). This configuration enabled us to use the third latent variable to capture students' latent change scores. Third, we constrained slopes of the pretest and posttest variables loading on the second latent variable, as well as the slope of the posttest variable loading on the third latent variable, to be equal. Fourth, we constrained the intercepts of the pretest and posttest variables to be equal. Fifth, we constrained slopes of the game‐based indicators loading on the fourth latent variable (the gameplay block) to be equal. Lastly, we constrained slopes of the pretest and posttest variables loading on the fifth latent variable (the assessment block) to be equal.</p> <p>We estimated the specified model using the Metropolis‐Hastings Robbins‐Monro algorithm implemented in flexMIRT 3.65 (Cai, [<reflink idref="bib8" id="ref125">8</reflink>]). The chosen random seed was 1010. The number of imputations from the MH step per RM cycle was 1. The dispersion values of the Metropolis proposal densities for the first and second levels of the specified model were 3.0 and 2.0, respectively. The number of Stage I (constant gain) cycles was 10,000, and the number of Stage II (Stochastic EM, constant gain) cycles was 1,000. The chosen scoring method was Expected A Posteriori (EAP). In addition, we enabled options to save the iteration history, estimated model parameters, and individual IRT scale scores. All other estimation settings were the defaults described in Houts and Cai ([<reflink idref="bib38" id="ref126">38</reflink>]).</p> <p>Table 12 presents the estimated group parameters and regression coefficients on the treatment assignment indicator. The estimated coefficients on the school dummy variables ranged between −.44 ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.14) and.65 ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> = 0.11).</p> <p>12 Table Estimates and SEs of Group Parameters and Latent Regression Coefficients</p> <p> <ephtml> <table><thead><tr valign="bottom"><th>Parameter</th><th align="center">Game‐Based Misconception</th><th align="center">Baseline</th><th align="center">Latent Change</th><th align="center">Block (Tasks)</th><th align="center">Block (Items)</th></tr></thead><tbody><tr><td>Latent mean (SE)</td><td>.00 (—)</td><td>.00 (—)</td><td>.11 (.01)</td><td>.00 (—)</td><td>.00 (—)</td></tr><tr><td>Regression coefficient (SE)</td><td /><td>.08 (.06)</td><td>.26 (.03)</td><td /><td /></tr><tr><td>Covariance matrix (SE)</td><td>1.00 (—)</td><td /><td /><td /><td /></tr><tr><td /><td>−.64 (.07)</td><td>1.00 (—)</td><td /><td /><td /></tr><tr><td /><td>−.06 (.02)</td><td>.00 (—)</td><td>.17 (.03)</td><td /><td /></tr><tr><td /><td>.00 (—)</td><td>.00 (—)</td><td>.00 (—)</td><td>1.00 (—)</td><td /></tr><tr><td /><td>.00 (—)</td><td>.00 (—)</td><td>.00 (—)</td><td>.00 (—)</td><td>1.00 (—)</td></tr></tbody></table> </ephtml> </p> <p>Table 13 presents the estimated parameters associated with the assessment items and the game‐based indicators. A game‐based indicator can be likened to an item, with its estimated intercept inversely related to a corresponding difficulty. Each indicator's estimated intercept also reflects the degree of prevalence of a specific misconception about fraction addition.</p> <p>13 Table Item Parameter Estimates and SEs</p> <p> <ephtml> <table><thead><tr><th>Variable</th><th align="center">Label</th><th><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>λ</mi><mrow><mi>m</mi><mi>i</mi><mi>s</mi><mi>c</mi></mrow></msub><annotation encoding="application/x-tex">$\lambda _{misc}$</annotation></semantics></math></p></th><th><italic>SE</italic></th><th><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>λ</mi><mrow><mi>b</mi><mi>a</mi><mi>s</mi><mi>e</mi></mrow></msub><annotation encoding="application/x-tex">$\lambda _{base}$</annotation></semantics></math></p></th><th><italic>SE</italic></th><th><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>λ</mi><mrow><mi>c</mi><mi>h</mi><mi>g</mi></mrow></msub><annotation encoding="application/x-tex">$\lambda _{chg}$</annotation></semantics></math></p></th><th><italic>SE</italic></th><th><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>λ</mi><mrow><mi>b</mi><mn>1</mn></mrow></msub><annotation encoding="application/x-tex">$\lambda _{b1}$</annotation></semantics></math></p></th><th><italic>SE</italic></th><th><p><math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics xmlns=""><msub><mi>λ</mi><mrow><mi>b</mi><mn>2</mn></mrow></msub><annotation encoding="application/x-tex">$\lambda _{b2}$</annotation></semantics></math></p></th><th><italic>SE</italic></th><th align="center">Intercept</th><th><italic>SE</italic></th></tr></thead><tbody><tr><td>1</td><td>EvIO</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−3.80</td><td>.05</td></tr><tr><td>2</td><td>SAOU</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−3.65</td><td>.05</td></tr><tr><td>3</td><td>SAOUaCHM</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−4.71</td><td>.08</td></tr><tr><td>4</td><td>SAOUaCHMaP</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−5.95</td><td>.15</td></tr><tr><td>5</td><td>CHMaP</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−2.59</td><td>.03</td></tr><tr><td>6</td><td>SwAW</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−4.01</td><td>.06</td></tr><tr><td>7</td><td>CnHM</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−2.33</td><td>.03</td></tr><tr><td>8</td><td>WrnN</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−2.27</td><td>.03</td></tr><tr><td>9</td><td>SAMN</td><td>.47</td><td>.02</td><td /><td /><td /><td /><td /><td /><td>.87</td><td>.05</td><td>−5.04</td><td>.09</td></tr><tr><td>10</td><td>Pretest Items</td><td /><td /><td>1.49</td><td>.02</td><td /><td /><td>1.36</td><td>.03</td><td /><td /><td>−0.30</td><td>.02</td></tr><tr><td>11</td><td>Posttest Items</td><td /><td /><td>1.49</td><td>.02</td><td>1.49</td><td>.02</td><td>1.36</td><td>.03</td><td /><td /><td>−.30</td><td>.02</td></tr></tbody></table> </ephtml> </p> <p>10 <emph>Note</emph>. The nine game‐based indicators are abbreviated: EvIO = Everything in order, SAOU = Saw as one unit, SAOUaCHM = Saw as one unit and counted hash marks, SAOUaCHMaP = Saw as one unit and counted hash marks and posts, CHMaP = Counted hash marks and posts, SwAW = Saw as wholes, CnHM = Counted hash marks, WrnN = Wrong numerator, and SAMN = Saw as mixed number. Other abbreviations are <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>m</mi><mi>i</mi><mi>s</mi><mi>c</mi></mrow><annotation encoding="application/x-tex">$misc$</annotation></semantics></math> </ephtml> for game‐based misconception, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>b</mi><mi>a</mi><mi>s</mi><mi>e</mi></mrow><annotation encoding="application/x-tex">$base$</annotation></semantics></math> </ephtml> for baseline performance, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>c</mi><mi>h</mi><mi>g</mi></mrow><annotation encoding="application/x-tex">$chg$</annotation></semantics></math> </ephtml> for latent change, <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>b</mi><mn>1</mn></mrow><annotation encoding="application/x-tex">$b1$</annotation></semantics></math> </ephtml> for Block 1, the assessment block, and <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>b</mi><mn>2</mn></mrow><annotation encoding="application/x-tex">$b2$</annotation></semantics></math> </ephtml> for Block 2, the gameplay block.</p> <hd id="AN0192629992-56">RQ1. Treatment Effect</hd> <p>Based on the estimated regression coefficients presented in Table 12, the control and treatment groups did not differ significantly in their baseline performance measured at pretest ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>β</mi><mo>̂</mo></mover><mo>=</mo><mo>.</mo><mn>08</mn><mo>,</mo><mi>S</mi><mi>E</mi><mo>=</mo><mo>.</mo><mn>06</mn><mo>,</mo><mi>p</mi><mo>=</mo><mi>n</mi><mi>s</mi></mrow><annotation encoding="application/x-tex">$\hat{\beta } =.08, SE =.06, p = ns$</annotation></semantics></math> </ephtml> ). The two groups, in comparison, did differ significantly in their latent changes, and there was a positive treatment effect on the latent pretest‐to‐posttest change ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mover accent="true"><mi>β</mi><mo>̂</mo></mover><mo>=</mo><mo>.</mo><mn>26</mn><mo>,</mo><mi>S</mi><mi>E</mi><mo>=</mo><mo>.</mo><mn>03</mn><mo>,</mo><mi>p</mi><mo><</mo><mo>.</mo><mn>05</mn></mrow><annotation encoding="application/x-tex">$\hat{\beta } =.26, SE =.03, p <.05$</annotation></semantics></math> </ephtml> ). Note that this effect size of.26 is relative to the variance of a latent change variable. As discussed in "The Full Model" section, a latent change variable tends to have a variance that is smaller than the variance of a nonchange latent variable.</p> <hd id="AN0192629992-57">RQ2. Relationship between Gameplay and Assessments</hd> <p>The covariance matrix presented in Table 12 provides information on the relationship between gameplay and assessments. Game‐based misconception was negatively correlated with the baseline performance ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mover accent="true"><mi>σ</mi><mo>̂</mo></mover><mrow><mi>b</mi><mi>a</mi><mi>s</mi><mi>e</mi><mo>,</mo><mi>m</mi><mi>i</mi><mi>s</mi><mi>c</mi></mrow></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>64</mn></mrow><annotation encoding="application/x-tex">$\hat{\sigma }_{base,misc} = -.64$</annotation></semantics></math> </ephtml> , <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><mi>S</mi><mi>E</mi></mrow><annotation encoding="application/x-tex">$SE$</annotation></semantics></math> </ephtml> =.07). Because the game‐based indicators targeted misconceptions about adding fractions, the direction of this correlation aligned with the expectation: students with lower baseline performance tended to exhibit more misconceptions during gameplay. Game‐based misconception was also negatively correlated with the latent change ( <ephtml> <math display="inline" xmlns="http://www.w3.org/1998/Math/MathML"><semantics><mrow><msub><mover accent="true"><mi>σ</mi><mo>̂</mo></mover><mrow><mi>c</mi><mi>h</mi><mi>g</mi><mo>,</mo><mi>m</mi><mi>i</mi><mi>s</mi><mi>c</mi></mrow></msub><mo>=</mo><mo>−</mo><mo>.</mo><mn>06</mn><mo>/</mo><msqrt><mrow><mo>.</mo><mn>17</mn></mrow></msqrt><mo>=</mo><mo>−</mo><mo>.</mo><mn>15</mn></mrow><annotation encoding="application/x-tex">$\hat{\sigma }_{chg,misc} = -.06/\sqrt {.17} = -.15$</annotation></semantics></math> </ephtml> ). This result suggests that students who exhibited more misconceptions in their gameplay tended to experience a lesser degree of change from pretest to posttest.</p> <p>These findings have three important implications for learning and instruction. First, information extracted from gameplay process data, such as indicators of misconceptions, can help surface students' prior knowledge and common conceptions concerning the addition of fractions. Second, the presence of misconceptions during gameplay may hinder or limit the extent to which students can make progress in their understanding of adding fractions. Third, these findings can serve as valuable inputs for instructional planning. For example, teachers may adjust subsequent instructions to address common conceptions and misconceptions identified during students' gameplay, with the goal of promoting conceptual change and improving student learning outcomes.</p> <hd id="AN0192629992-58">Summary and Discussion</hd> <p>We began this paper by framing the understanding of data from game‐based evaluation studies as a sensemaking process. Recognizing the breadth and complexity of this process, we focused on presenting a new application of cross‐classified IRT modeling. The application showed how we could use a cross‐classified IRT model to jointly analyze item‐level data from traditional assessments and nonaggregated gameplay process information summarized by diagnostic indicators, while addressing key aspects of (multisite) evaluation research of educational games. With the latent change parameterization discussed in "The Full Model" section, the proposed application could connect individuals' performance in a digital game with changes in assessment outcomes, providing one way to directly evaluate the impact of a game‐based intervention on learning.</p> <p>One notable limitation of our paper is the lack of discussions on model fit, especially for parameterized latent variable models. While a major goal of this paper is to demonstrate the utility of a modern and flexible modeling framework, we recognize that it is still important to examine how well a model fits the data before interpreting any estimated parameters, given that inferences drawn from a model are contingent on the quality of the model itself.</p> <p>Another aspect that we did not discuss is player or learner engagement in games. Engagement in games, or defining, operationalizing, and measuring engagement in educational games, is a complex topic in itself (e.g., Abdul Jabbar & Felicia, [<reflink idref="bib2" id="ref127">2</reflink>]; Hookham & Nesbitt, [<reflink idref="bib37" id="ref128">37</reflink>]). It remains important to investigate the interplay of engagement in educational games, performance, and learning. We thus invite future research to measure more of the extent to which learners engage in educational games.</p> <p>Moreover, two preconditions need to be met in order to fully realize the benefits of the modeling application proposed in our paper. The application of cross‐classified IRT modeling to integrate game and assessment data places great value on (a) having a set of content and cognitive demand features that guide the development of both learning (e.g., educational games) and assessment systems, and (b) having diagnostic or theory‐informed indicators developed using process data. Meeting the first precondition greatly facilitates the second. We believe that these two preconditions, along with having a unified and flexible modeling framework like cross‐classified IRT modeling, are crucial for creating a coherent analytical framework that enables us to harness diagnostic information from process data, reduces our reliance on summative assessments, and enhances the instructional process by offering targeted feedback to students and instructors.</p> <p>Compared to fully data‐driven procedures, the shortcoming of the kind of approach advocated in our paper—emphasizing sensemaking and meeting preconditions—is also evident. Sensemaking is a volitional process, where we actively come to understand the differences as well as the interconnectivity embedded in multiple sources of information, amid other complexities. This process is time and labor intensive, requiring us to navigate massive volumes and diverse types of data and distill them into diagnostic, actionable information about effectiveness and areas for improvement. The challenges and requirements posed by this process can be daunting, especially when juxtaposed with high‐potential machine learning or artificial intelligence applications, wherein tasks such as indicator development can be relegated.</p> <p>Nevertheless, we believe that there is value in sensemaking, whether it is for the evaluation of educational games or for the analysis of game‐based evaluation data. Developing and executing a well‐defined theory of action, as well as meeting preconditions, is an investment that is likely to result in "more efficient analysis and more valid interpretation of the data" (Goldhammer et al., [<reflink idref="bib30" id="ref129">30</reflink>]; Lindner & Greiff, [<reflink idref="bib52" id="ref130">52</reflink>]; Zumbo et al., [<reflink idref="bib84" id="ref131">84</reflink>], pp. 245‐246). Returning to the <emph>frame</emph> argument in sensemaking, we also want to underscore the centrality of the researchers in the scientific discovery process. While data‐driven methods are valuable for identifying unseen trends, researchers play a crucial role in shaping, guiding, and engaging in conversations about the discovery process.</p> <hd id="AN0192629992-59">Acknowledgments</hd> <p>Dr. Li Cai's research is partially supported by a grant from the Institute of Education Sciences (R305D210032). The views expressed in this paper belong to the co‐authors and do not represent those of the funding agency.</p> <ref id="AN0192629992-60"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref74" type="bt">1</bibl> <bibtext> Of the 826 students with analyzable gameplay data, 704 students (85%) attempted all 27 in‐game tasks, and 752 students (91%) attempted at least 20 of the 27 in‐game tasks, where the gameplay flow in the first 20 tasks overlapped with the flow in the last 7 tasks. Starting Task 16, the numbers and types of fraction knowledge specifications targeted by each task were also similar. A total of 777 of 826 students (94%) attempted Task 16 or beyond.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref113" type="bt">2</bibl> <bibtext> Based on a meta‐analysis of 24 studies examining the impact of game‐based learning on student math achievement, the overall weighted effect size was .13 with an associated 95% confidence interval of [.02, .24] (Tokac et al., [73]). Based on another meta‐analysis of 48 studies that examined the impact of educational games on learning outcomes and used games as the only instructional method, the overall weighted effect size was .20 with an associated 95% confidence interval of [.03, .37] (Wouters et al., [79]).</bibtext> </blist> <blist> <bibl id="bib3" idref="ref25" type="bt">3</bibl> <bibtext> We imposed this constraint considering the small sample size relative to the complexity of the model, especially if all of these slopes were to be freely estimated. If warranted by theoretical and practical considerations, such as having a larger sample size and data conditions that would support the estimation of a more complex model, it is possible to relax this constraint.</bibtext> </blist> </ref> <ref id="AN0192629992-61"> <title> References </title> <blist> <bibtext> Abbot, A., & Tsay, A. (2000). Sequence analysis and optimal matching methods in sociology: Review and prospect. Sociological Methods & Research, 29(1), 3–33.</bibtext> </blist> <blist> <bibtext> Abdul Jabbar, A. I., & Felicia, P. (2015). Gameplay engagement and learning in game‐based learning: A systematic review. Review of Educational Research, 85(4), 740–779.</bibtext> </blist> <blist> <bibtext> Arieli‐Attali, M., Ward, S., Thomas, J., Deonovic, B., & von Davier, A. A. (2019). The expanded evidence‐centered design (e‐ECD) for learning and assessment systems: A framework for incorporating learning goals and processes within assessment design. Frontiers in Psychology, 10.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref15" type="bt">4</bibl> <bibtext> Baker, E., Chung, G., & Delacruz, G. (2008). Design and validation of technology‐based performance assessments. In J. M. Spector, M. D. Merrill, J. van Merriënboer, & M. P. Driscoll (Eds.), Handbook of research on educational communications and technology (3 edn., pp. 595–604). Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref12" type="bt">5</bibl> <bibtext> Bergner, Y., & von Davier, A. A. (2019). Process data in NAEP: Past, present, and future. Journal of Educational and Behavioral Statistics, 44(6), 706–732.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref42" type="bt">6</bibl> <bibtext> Blanié, A., Amorim, M.‐A., Meffert, A., Perrot, C., Dondelli, L., & Benhamou, D. (2020). Assessing validity evidence for a serious game dedicated to patient clinical deterioration and communication. Advances in Simulation, 5(4), 1–12.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref44" type="bt">7</bibl> <bibtext> Cagiltay, N. E., Ozcelik, E., & Ozcelik, N. S. (2015). The effect of competition on learning in games. Computers & Education, 87, 35–41.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref125" type="bt">8</bibl> <bibtext> Cai, L. (2022). flexMIRT<sups>®</sups>: Flexible multilevel multidimensional item analysis and test scoring. Computer software.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref76" type="bt">9</bibl> <bibtext> Cai, L., Choi, K., & Kuhfeld, M. (2016). On the role of multilevel item response models in multisite evaluation studies for serious games. In H. O'Neil, E. Baker, & R. Perez (Eds.), Using games and simulations for teaching and assessment (pp. 280–301). Routledge.</bibtext> </blist> <blist> <bibtext> Cai, L., & Houts, C. R. (2021). Longitudinal analysis of patient‐reported outcomes in clinical trials: Applications of multilevel and multidimensional item response theory. Psychometrika, 86(3), 754–777.</bibtext> </blist> <blist> <bibtext> Center for Advanced Technology in Schools (2012). CATS‐developed games. (CRESST Resource Paper No. 15). University of California, Los Angeles, National Center for Research on Evaluation, Standards, and Student Testing (CRESST). https://cresst.org/publications/cresst‐publication‐3255/</bibtext> </blist> <blist> <bibtext> Chen, F., Cui, Y., & Chu, M.‐W. (2020). Utilizing game analytics to inform and validate digital game‐based assessment with evidence‐centered game design: A case study. International Journal of Artificial Intelligence in Education, 30(3), 481–503.</bibtext> </blist> <blist> <bibtext> Chen, Y., Zhang, J., Yang, Y., & Lee, Y.‐S. (2022). Latent space model for process data. Journal of Educational Measurement, 59(4), 517–535.</bibtext> </blist> <blist> <bibtext> Chia, R. (2000). Discourse analysis organizational analysis. Organization, 7(3), 513–518.</bibtext> </blist> <blist> <bibtext> Chung, G. K. W. K. (2015). Guidelines for the design and implementation of game telemetry for serious games analytics. In C. S. Loh, Y. Sheng, & D. Ifenthaler (Eds.), Serious games analytics (pp. 59–79). Springer.</bibtext> </blist> <blist> <bibtext> Chung, G. K. W. K., & Baker, E. (2003). An exploratory study to examine the feasibility of measuring problem‐solving processes using a click‐through interface. Journal of Technology, Learning and Assessment, 2(2).</bibtext> </blist> <blist> <bibtext> Chung, G. K. W. K., Choi, K., Baker, E. L., & Cai, L. (2014). The effects of math video games on learning: A randomized evaluation study with innovative impact estimation techniques. CRESST.</bibtext> </blist> <blist> <bibtext> Chung, G. K. W. K., & Feng, T. (2024). From clicks to constructs: An examination of validity evidence of game‐based indicators derived from theory. In M. Sahin & D. Ifenthaler (Eds.), Assessment analytics in education—Designs, methods and solutions. Springer.</bibtext> </blist> <blist> <bibtext> Darling‐Hammond, L., Herman, J., Pellegrino, J., Abedi, J., Aber, J. L., Baker, E., Bennett, R., Gordon, E., Haertel, E., Hakuta, K., Ho, A., Linn, R. L., Pearson, P. D., Popham, J., Resnick, L., Schoenfeld, A. H., Shavelson, R., Shepard, A., Shulman, L., & Steele, C. M. (2013). Criteria for high‐quality assessment. Stanford Center for Opportunity Policy in Education.</bibtext> </blist> <blist> <bibtext> De Boeck, P. (2008). Random item IRT models. Psychometrika, 73(4), 533–559.</bibtext> </blist> <blist> <bibtext> De Boeck, P., & Jeon, M. (2019). An overview of models for response times and processes in cognitive tests. Frontiers in Psychology, 10(102), 1–11.</bibtext> </blist> <blist> <bibtext> De Boeck, P., & Wilson, M. (Eds.). (2004). Explanatory item response models: A generalized linear and nonlinear approach. Springer.</bibtext> </blist> <blist> <bibtext> Dervin, B. (2003). Sense‐making methodology reader: Selected writings of Brenda Dervin. Hampton Press.</bibtext> </blist> <blist> <bibtext> Ercikan, K., Guo, H., & He, Q. (2020). Use of response process data to inform group comparisons and fairness research. Educational Assessment, 25(3), 179–197.</bibtext> </blist> <blist> <bibtext> van der Linden, W. J. (2007). A hierarchical framework for modeling speed and accuracy on test items. Psychometrika, 72(3), 287–308.</bibtext> </blist> <blist> <bibtext> Foster, N., & Piacentini, M. (Eds.). (2023). Innovating assessments to measure and support complex skills. OECD Publishing.</bibtext> </blist> <blist> <bibtext> Gane, B. D., Zaidi, S. Z., & Pellegrino, J. W. (2018). Measuring what matters: Using technology to assess multidimensional learning. European Journal of Education, 53(2), 176–187.</bibtext> </blist> <blist> <bibtext> Garcia, I., Pacheco, C., Méndez, F., & Calvo‐Manzano, J. A. (2020). The effects of game‐based learning in the acquisition of "soft skills" on undergraduate software engineering courses: A systematic literature review. Computer Applications in Engineering Education, 28(5), 1327–1354.</bibtext> </blist> <blist> <bibtext> Gauthier, A., Corrin, M., & Jenkinson, J. (2015). Exploring the influence of game design on learning and voluntary use in an online vascular anatomy study aid. Computers & Education, 87, 24–34.</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Hahnel, C., Kroehne, U., & Zehner, F. (2021). From byproduct to design factor: On validating the interpretation of process indicators based on log data. Large‐scale Assessments in Education, 9(1), 20.</bibtext> </blist> <blist> <bibtext> Goldhammer, F., Naumann, J., Stelter, A., Tóth, K., Rölke, H., & Klieme, E. (2014). The time on task effect in reading and problem solving is moderated by task difficulty and skill: Insights from a computer‐based large‐scale assessment. Journal of Educational Psychology, 106(3), 608–626.</bibtext> </blist> <blist> <bibtext> Greiff, S., Wüstenberg, S., & Avvisati, F. (2015). Computer‐generated log‐file analyses as a window into students' minds? A showcase study based on the PISA 2012 assessment of problem solving. Computers & Education, 91, 92–105.</bibtext> </blist> <blist> <bibtext> Hahnel, C., Jung, A. J., & Goldhammer, F. (2023). Theory matters: An example of deriving process indicators from log data to assess decision‐making processes in web search tasks. European Journal of Psychological Assessment, 39(4), 271–279.</bibtext> </blist> <blist> <bibtext> Hao, J., Shu, Z., & Davier, A. Von. (2015). Analyzing process data from game/scenario‐based tasks: An edit distance approach. Journal of Educational Data Mining, 7(1), 33–50.</bibtext> </blist> <blist> <bibtext> Hattie, J. (Ed.). (2023). Visible learning: A synthesis of over 2,100 meta‐analyses relating to achievement. Routledge.</bibtext> </blist> <blist> <bibtext> Hautala, J., Heikkilä, R., Nieminen, L., Rantanen, V., Latvala, J.‐M., & Richardson, U. (2020). Identification of reading difficulties by a digital game‐based assessment technology. Journal of Educational Computing Research, 58(5), 1003–1028.</bibtext> </blist> <blist> <bibtext> Hookham, G., & Nesbitt, K. (2019). A systematic review of the definition and measurement of engagement in serious games. In Proceedings of the Australasian Computer Science Week Multiconference, ACSW'19 (pp. 1–10). Association for Computing Machinery.</bibtext> </blist> <blist> <bibtext> Houts, C. R., & Cai, L. (2020). flexMIRT<sups>®</sups> user's manual version 3.6: Flexible multilevel multidimensional item analysis and test scoring. Vector Psychometric Group. Software manual.</bibtext> </blist> <blist> <bibtext> Huang, S., & Cai, L. (2024). Cross‐classified item response theory modeling with an application to student evaluation of teaching. Journal of Educational and Behavioral Statistics, 49(3), 311–341. https://doi.org/10.3102/10769986231193351</bibtext> </blist> <blist> <bibtext> Jiao, H., He, Q., & Veldkamp, B. P. (2021). Editorial: Process data in educational and psychological measurement. Frontiers in Psychology, 12, 793399.</bibtext> </blist> <blist> <bibtext> Jiao, H., Liao, D., & Zhan, P. (2019). Utilizing process data for cognitive diagnosis. In: M. von Davier & Y. S. Lee (Eds.), Handbook of diagnostic classification models: Models and model extensions, applications, software packages. Methodology of Educational Measurement and Assessment (pp. 421–436). Springer.</bibtext> </blist> <blist> <bibtext> Jöreskog, K., & Sörbom, D. (2023). LISREL 12. Computer software.</bibtext> </blist> <blist> <bibtext> Kerr, D. (2014). Into the black box: Using data mining of in‐game actions to draw inferences from educational technology about students' math knowledge.</bibtext> </blist> <blist> <bibtext> Kerr, D., & Chung, G. K. W. K. (2012a). Identifying key features of student performance in educational video games and simulations through cluster analysis. Journal of Educational Data Mining, 4(1), 144–182.</bibtext> </blist> <blist> <bibtext> Kerr, D., & Chung, G. K. W. K. (2012b). The mediation effect of in‐game performance between prior knowledge and posttest score.</bibtext> </blist> <blist> <bibtext> Kiili, K., Moeller, K., & Ninaus, M. (2018). Evaluating the effectiveness of a game‐based rational number training—in‐game metrics as learning indicators. Computers & Education, 120, 13–28.</bibtext> </blist> <blist> <bibtext> Klein, G., Phillips, J. K., Rall, E. L., & Peluso, D. A. (2007). A data‐frame theory of sensemaking. In R. R. Hoffman (Ed.), Expertise out of Context: Proceedings of the Sixth International Conference on Naturalistic Decision Making (pp. 113–155). Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibtext> Kraft, M. A. (2020). Interpreting effect sizes of education interventions. Educational Researcher, 49(4), 241–253.</bibtext> </blist> <blist> <bibtext> Lee, Y.‐H., & Jia, Y. (2014). Using response time to investigate students' test‐taking behaviors in a NAEP computer‐based study. Large‐Scale Assessments in Education, 2(8), 1–24.</bibtext> </blist> <blist> <bibtext> Levy, R. (2019). Dynamic Bayesian network modeling of game‐based diagnostic assessments. Multivariate Behavioral Research, 54(6), 771–794.</bibtext> </blist> <blist> <bibtext> Levy, R. (2020). Implications of considering response process data for greater and lesser psychometrics. Educational Assessment, 25(3), 218–235.</bibtext> </blist> <blist> <bibtext> Lindner, M. A., & Greiff, S. (2023). Process data in computer‐based assessment: Challenges and opportunities in opening the black box. European Journal of Psychological Assessment, 39(4), 241–251.</bibtext> </blist> <blist> <bibtext> Liu, T., & Israel, M. (2022). Uncovering students' problem‐solving processes in game‐based learning environments. Computers & Education, 182, 104462.</bibtext> </blist> <blist> <bibtext> Mayer, R. E. (2019). Computer games in education. Annual Review of Psychology, 70(1), 531–549.</bibtext> </blist> <blist> <bibtext> McArdle, J. J. (2009). Latent variable modeling of differences and changes with longitudinal data. Annual Review of Psychology, 60(1), 577–605.</bibtext> </blist> <blist> <bibtext> Min, W., Frankosky, M. H., Mott, B. W., Rowe, J. P., Smith, A., Wiebe, E., Boyer, K. E., & Lester, J. C. (2020). Deepstealth: Game‐based learning stealth assessment with deep neural networks. IEEE Transactions on Learning Technologies, 13(2), 312–325.</bibtext> </blist> <blist> <bibtext> Mislevy, R., Behrens, J. T., Dicerbo, K. E., & Levy, R. (2012). Design and discovery in educational assessment: Evidence‐centered design, psychometrics, and educational data mining. Journal of Educational Data Mining, 4(1), 11–48.</bibtext> </blist> <blist> <bibtext> Mislevy, R., Corrigan, S., Oranje, A., DiCerbo, K., Bauer, M. I., von Davier, A., & John, M. (2015). Psychometrics and game‐based assessment. In F. Drasgow (Ed.), Technology and testing (pp. 23–48). Routledge.</bibtext> </blist> <blist> <bibtext> Mislevy, R., Oranje, A., Bauer, M. I., von Davier, A. A., Hao, J., Corrigan, S., Hoffman, E., DiCerbo, K., & John, M. (2014). Psychometric considerations in game‐based assessment. White paper, GlassLab Research, Institute of Play.</bibtext> </blist> <blist> <bibtext> National Research Council (2001). Knowing what students know: The science and design of educational assessment. National Academies Press.</bibtext> </blist> <blist> <bibtext> Pellegrino, J. W., & Quellmalz, E. S. (2010). Perspectives on the integration of technology and assessment. Journal of Research on Technology in Education, 43(2), 119–134.</bibtext> </blist> <blist> <bibtext> Petri, G., & Gresse von Wangenheim, C. (2017). How games for computing education are evaluated? A systematic literature review. Computers & Education, 107, 68–90.</bibtext> </blist> <blist> <bibtext> Pirolli, P., & Russell, D. (2011). Introduction to this special issue on sensemaking. Human‐Computer Interaction, 26(1), 1–8.</bibtext> </blist> <blist> <bibtext> Plass, J. L., Homer, B. D., & Kinzer, C. K. (2015). Foundations of game‐based learning. Educational Psychologist, 50(4), 258–283.</bibtext> </blist> <blist> <bibtext> Raudenbush, S. W., & Bryk, A. S. (2002). Hierarchical linear models: Applications and data analysis methods. Sage.</bibtext> </blist> <blist> <bibtext> Raykov, T., & Marcoulides, G. A. (Eds.). (2006). A first course in structural equation modeling. Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibtext> Reckase, M. D. (2009). Multidimensional item response theory. Springer.</bibtext> </blist> <blist> <bibtext> Reese, D. D., Tabachnick, B. G., & Kosko, R. E. (2015). Video game learning dynamics: Actionable measures of multidimensional learning trajectories. British Journal of Educational Technology, 46(1), 98–122.</bibtext> </blist> <blist> <bibtext> Shavelson, R., & Webb, N. (1991). Generalizability theory: A primer. Sage.</bibtext> </blist> <blist> <bibtext> Shute, V. J., & Rahimi, S. (2021). Stealth assessment of creativity in a physics video game. Computers in Human Behavior, 116, 106647.</bibtext> </blist> <blist> <bibtext> Tang, X., Wang, Z., He, Q., Liu, J., & Ying, Z. (2020). Latent feature extraction for process data via multidimensional scaling. Psychometrika, 85(2), 378–397.</bibtext> </blist> <blist> <bibtext> Tenorio Delgado, M., Arango Uribe, P., Aparicio Alonso, A., & Rosas Díaz, R. (2016). TENI: A comprehensive battery for cognitive assessment based on games and technology. Child Neuropsychology, 22(3), 276–291.</bibtext> </blist> <blist> <bibtext> Tokac, U., Novak, E., & Thompson, C. G. (2019). Effects of game‐based learning on students' mathematics achievement: A meta‐analysis. Journal of Computer Assisted Learning, 35(3), 407–420.</bibtext> </blist> <blist> <bibtext> van den Noortgate, W., De Boeck, P., & Meulders, M. (2003). Cross‐classification multilevel logistic models in psychometrics. Journal of Educational and Behavioral Statistics, 28, 369–386.</bibtext> </blist> <blist> <bibtext> Vendlinski, T. P., Delacruz, G. C., Buschang, R. E., Chung, G. K. W. K., & Baker, E. L. (2010). Developing high‐quality assessments that align with instructional video games. CRESST Report 774, University of California, Los Angeles, National Center for Research on Evaluation, Standards, and Student Testing (CRESST).</bibtext> </blist> <blist> <bibtext> Weick, K. E., Sutcliffe, K. M., & Obstfeld, D. (2005). Organizing and the process of sensemaking. Organization Science, 16(4), 409–421.</bibtext> </blist> <blist> <bibtext> Weiner, E. J., & Sanchez, D. R. (2020). Cognitive ability in virtual reality: Validity evidence for VR game‐based assessments. International Journal of Selection and Assessment, 28(3), 215–235.</bibtext> </blist> <blist> <bibtext> What Works Clearinghouse. (2015). WWC review of the report "The Effects of Math Video Games on Learning."</bibtext> </blist> <blist> <bibtext> Wouters, P., Nimwegen, C., Oostendorp, H., & Spek, E. (2013). A meta‐analysis of the cognitive and motivational effects of serious games. Journal of Educational Psychology, 105, 249.</bibtext> </blist> <blist> <bibtext> Wüstenberg, S., Greiff, S., & Funke, J. (2012). Complex problem solving—more than reasoning?Intelligence, 40(1), 1–14.</bibtext> </blist> <blist> <bibtext> Xiao, Y., Veldkamp, B., & Liu, H. (2022). Combining process information and item response modeling to estimate problem‐solving ability. Educational Measurement: Issues and Practice, 41(2), 36–54.</bibtext> </blist> <blist> <bibtext> Zhang, S., Wang, Z., Qi, J., Liu, J., & Ying, Z. (2023). Accurate assessment via process data. Psychometrika, 88(1), 76–97.</bibtext> </blist> <blist> <bibtext> Zhu, S., Guo, Q., & Yang, H. H. (2023). Beyond the traditional: A systematic review of digital game‐based assessment for students' knowledge, skills, and affections. Sustainability, 15(5), 4693.</bibtext> </blist> <blist> <bibtext> Zumbo, B. D., Maddox, B., & Care, N. M. (2023). Process and product in computer‐based assessments. European Journal of Psychological Assessment, 39(4), 252–262.</bibtext> </blist> </ref> <aug> <p>By Tianying Feng and Li Cai</p> <p>Reported by Author; Author</p> <p></p> <p>TIANYING FENG is a doctoral student in the Education ‐ Advanced Quantitative Methodology program at UCLA and a research assistant at the National Center for Research on Evaluation, Standards, and Student Testing (CRESST), SEIS Building, Los Angeles, CA 90095‐1522; tfeng0315@ucla.edu. Her primary research interests include technology‐based measurement and learning research and statistical computing.</p> <p>LI CAI is a Professor of Education in the Advanced Quantitative Methodology program at UCLA and Director of the National Center for Research on Evaluation, Standards, and Student Testing (CRESST), 315 SEIS Building, Los Angeles, CA 90095‐1522; cai@cresst.org. His primary research interests include psychometrics and statistical computing.</p> </aug> <nolink nlid="nl1" bibid="bib23" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib47" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib63" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib76" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib15" firstref="ref9"></nolink> <nolink nlid="nl6" bibid="bib32" firstref="ref10"></nolink> <nolink nlid="nl7" bibid="bib52" firstref="ref11"></nolink> <nolink nlid="nl8" bibid="bib51" firstref="ref13"></nolink> <nolink nlid="nl9" bibid="bib84" firstref="ref14"></nolink> <nolink nlid="nl10" bibid="bib58" firstref="ref16"></nolink> <nolink nlid="nl11" bibid="bib59" firstref="ref17"></nolink> <nolink nlid="nl12" bibid="bib64" firstref="ref19"></nolink> <nolink nlid="nl13" bibid="bib16" firstref="ref20"></nolink> <nolink nlid="nl14" bibid="bib30" firstref="ref23"></nolink> <nolink nlid="nl15" bibid="bib19" firstref="ref26"></nolink> <nolink nlid="nl16" bibid="bib27" firstref="ref27"></nolink> <nolink nlid="nl17" bibid="bib26" firstref="ref28"></nolink> <nolink nlid="nl18" bibid="bib60" firstref="ref29"></nolink> <nolink nlid="nl19" bibid="bib61" firstref="ref30"></nolink> <nolink nlid="nl20" bibid="bib28" firstref="ref32"></nolink> <nolink nlid="nl21" bibid="bib62" firstref="ref33"></nolink> <nolink nlid="nl22" bibid="bib21" firstref="ref37"></nolink> <nolink nlid="nl23" bibid="bib24" firstref="ref38"></nolink> <nolink nlid="nl24" bibid="bib41" firstref="ref39"></nolink> <nolink nlid="nl25" bibid="bib49" firstref="ref40"></nolink> <nolink nlid="nl26" bibid="bib25" firstref="ref41"></nolink> <nolink nlid="nl27" bibid="bib12" firstref="ref43"></nolink> <nolink nlid="nl28" bibid="bib29" firstref="ref45"></nolink> <nolink nlid="nl29" bibid="bib31" firstref="ref46"></nolink> <nolink nlid="nl30" bibid="bib33" firstref="ref47"></nolink> <nolink nlid="nl31" bibid="bib36" firstref="ref48"></nolink> <nolink nlid="nl32" bibid="bib46" firstref="ref49"></nolink> <nolink nlid="nl33" bibid="bib70" firstref="ref50"></nolink> <nolink nlid="nl34" bibid="bib72" firstref="ref51"></nolink> <nolink nlid="nl35" bibid="bib18" firstref="ref52"></nolink> <nolink nlid="nl36" bibid="bib43" firstref="ref53"></nolink> <nolink nlid="nl37" bibid="bib80" firstref="ref56"></nolink> <nolink nlid="nl38" bibid="bib34" firstref="ref57"></nolink> <nolink nlid="nl39" bibid="bib53" firstref="ref58"></nolink> <nolink nlid="nl40" bibid="bib83" firstref="ref62"></nolink> <nolink nlid="nl41" bibid="bib77" firstref="ref64"></nolink> <nolink nlid="nl42" bibid="bib56" firstref="ref65"></nolink> <nolink nlid="nl43" bibid="bib45" firstref="ref66"></nolink> <nolink nlid="nl44" bibid="bib68" firstref="ref67"></nolink> <nolink nlid="nl45" bibid="bib50" firstref="ref68"></nolink> <nolink nlid="nl46" bibid="bib40" firstref="ref69"></nolink> <nolink nlid="nl47" bibid="bib71" firstref="ref71"></nolink> <nolink nlid="nl48" bibid="bib81" firstref="ref72"></nolink> <nolink nlid="nl49" bibid="bib82" firstref="ref73"></nolink> <nolink nlid="nl50" bibid="bib13" firstref="ref75"></nolink> <nolink nlid="nl51" bibid="bib17" firstref="ref77"></nolink> <nolink nlid="nl52" bibid="bib11" firstref="ref78"></nolink> <nolink nlid="nl53" bibid="bib75" firstref="ref79"></nolink> <nolink nlid="nl54" bibid="bib44" firstref="ref81"></nolink> <nolink nlid="nl55" bibid="bib14" firstref="ref83"></nolink> <nolink nlid="nl56" bibid="bib39" firstref="ref85"></nolink> <nolink nlid="nl57" bibid="bib74" firstref="ref86"></nolink> <nolink nlid="nl58" bibid="bib20" firstref="ref90"></nolink> <nolink nlid="nl59" bibid="bib22" firstref="ref91"></nolink> <nolink nlid="nl60" bibid="bib69" firstref="ref93"></nolink> <nolink nlid="nl61" bibid="bib65" firstref="ref94"></nolink> <nolink nlid="nl62" bibid="bib67" firstref="ref95"></nolink> <nolink nlid="nl63" bibid="bib10" firstref="ref97"></nolink> <nolink nlid="nl64" bibid="bib55" firstref="ref98"></nolink> <nolink nlid="nl65" bibid="bib78" firstref="ref101"></nolink> <nolink nlid="nl66" bibid="bib57" firstref="ref106"></nolink> <nolink nlid="nl67" bibid="bib35" firstref="ref114"></nolink> <nolink nlid="nl68" bibid="bib54" firstref="ref115"></nolink> <nolink nlid="nl69" bibid="bib48" firstref="ref117"></nolink> <nolink nlid="nl70" bibid="bib66" firstref="ref119"></nolink> <nolink nlid="nl71" bibid="bib42" firstref="ref120"></nolink> <nolink nlid="nl72" bibid="bib38" firstref="ref126"></nolink> <nolink nlid="nl73" bibid="bib37" firstref="ref128"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1501291
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Sensemaking of Process Data from Evaluation Studies of Educational Games: An Application of Cross-Classified Item Response Theory Modeling
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Tianying+Feng%22">Tianying Feng</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-2215-9234">0000-0003-2215-9234</externalLink>)<br /><searchLink fieldCode="AR" term="%22Li+Cai%22">Li Cai</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Educational+Measurement%22"><i>Journal of Educational Measurement</i></searchLink>. 2026 63(1).
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 37
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305D210032
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Educational+Games%22">Educational Games</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Response+Theory%22">Item Response Theory</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Misconceptions%22">Misconceptions</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Analysis%22">Item Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Game+Based+Learning%22">Game Based Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Pretests+Posttests%22">Pretests Posttests</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/jedm.12396
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0655<br />1745-3984
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Process information collected from educational games can illuminate how students approach interactive tasks, complementing assessment outcomes routinely examined in evaluation studies. However, the two sources of information are historically analyzed and interpreted separately, and diagnostic process information is often underused. To tackle these issues, we present a new application of cross-classified item response theory modeling, using indicators of knowledge misconceptions and item-level assessment data collected from a multisite game-based randomized controlled trial. This application addresses (a) the joint modeling of students' pretest and posttest item responses and game-based processes described by indicators of misconceptions; (b) integration of gameplay information when gauging the intervention effect of an educational game; (c) relationships among game-based misconception, pretest initial status, and pre-to-post change; and (d) nesting of students within schools, a common aspect in multisite research. We also demonstrate how to structure the data and set up the model to enable our proposed application, and how our application compares to three other approaches to analyzing gameplay and assessment data. Lastly, we note the implications for future evaluation studies and for using analytic results to inform learning and instruction.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1501291
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1501291
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/jedm.12396
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 37
    Subjects:
      – SubjectFull: Educational Games
        Type: general
      – SubjectFull: Item Response Theory
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Misconceptions
        Type: general
      – SubjectFull: Item Analysis
        Type: general
      – SubjectFull: Game Based Learning
        Type: general
      – SubjectFull: Pretests Posttests
        Type: general
    Titles:
      – TitleFull: Sensemaking of Process Data from Evaluation Studies of Educational Games: An Application of Cross-Classified Item Response Theory Modeling
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Tianying Feng
      – PersonEntity:
          Name:
            NameFull: Li Cai
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0022-0655
            – Type: issn-electronic
              Value: 1745-3984
          Numbering:
            – Type: volume
              Value: 63
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Journal of Educational Measurement
              Type: main
ResultId 1