Self-Regulated Learning in the Digitally Enhanced Science Classroom: Toward an Early Warning System
Saved in:
| Title: | Self-Regulated Learning in the Digitally Enhanced Science Classroom: Toward an Early Warning System |
|---|---|
| Language: | English |
| Authors: | Marcus Kubsch (ORCID |
| Source: | Educational Psychology Review. 2025 37(2). |
| Availability: | Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ |
| Peer Reviewed: | Y |
| Page Count: | 38 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Education Level: | Junior High Schools Middle Schools Secondary Education Elementary Education Grade 7 Grade 8 |
| Descriptors: | Inquiry, Science Instruction, Electronic Books, Workbooks, Physics, Artificial Intelligence, Prediction, Data Collection, Cognitive Processes, Metacognition, Affective Behavior, Productivity, Self Management, Foreign Countries, At Risk Students, Middle School Students, Grade 7, Grade 8, Technology Uses in Education |
| Geographic Terms: | Germany |
| DOI: | 10.1007/s10648-025-10011-9 |
| ISSN: | 1040-726X 1573-336X |
| Abstract: | Recent research underscores the importance of inquiry learning for effective science education. Inquiry learning involves self-regulated learning (SRL), for example when students conduct investigations. Teachers face challenges in orchestrating and tracking student learning in such instruction; making it hard to adequately support students. Using AI methods such as machine learning (ML), the data that is generated when students interact in technology-enhanced classrooms can be used to track their learning and subsequently to inform teachers so that they can better support student learning. This study implemented digital workbooks in an inquiry-based physics unit, collecting cognitive, metacognitive, and affective data from 214 students. Using ML methods, an early warning system was developed to predict students' learning outcomes. Explainable ML methods were used to unpack these predictions and analyses were conducted for potential biases. Results indicate that an integration of cognitive, metacognitive, and affective data can predict students' productivity with an accuracy ranging from 60 to 100% as the unit progresses. Initially, affective and metacognitive variables dominate predictions, with cognitive variables becoming more significant later. Using only affective and metacognitive data, predictive accuracies ranged from 60 to 80% throughout. Bias was found to be highly dependent on the ML methods being used. The study highlights the potential of digital student workbooks to support SRL in inquiry-based science education, guiding future research and development to enhance instructional feedback and teacher insights into student engagement. Further, the study sheds new light on the data needed and the methodological challenges when using ML methods to investigate SRL processes in classrooms. |
| Abstractor: | As Provided |
| Notes: | https://osf.io/uv8tn |
| Entry Date: | 2025 |
| Accession Number: | EJ1466628 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHeKDcqbIa8R28L01YJ9uaOAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDClSI7WHIAvDFq6GgwIBEICBm7C_s0VYrmn79G7snEA6mMMGbvsI4KUS_IOM10OkdjBq0pCIVG3FL2yMh4iyeKCcP0bBkkHRyDsxHxvCwwU1QLs_a-SKgIV6_HYtjXUkrFy--fXG2ylaF5IY9Q9P3Jg4I6ReCx6_BFp5C7igqUkjsZ3ieZdbOu2vyPVt9YE2RS9q5Kit5rxIYbzZzfUYv_ygkA0FaapO0ILnHkDy Text: Availability: 1 Value: <anid>AN0184416951;epv01jun.25;2025Jul06.13:03;v2.2.500</anid> <title id="AN0184416951-1">Self-regulated Learning in the Digitally Enhanced Science Classroom: Toward an Early Warning System </title> <p>Recent research underscores the importance of inquiry learning for effective science education. Inquiry learning involves self-regulated learning (SRL), for example when students conduct investigations. Teachers face challenges in orchestrating and tracking student learning in such instruction; making it hard to adequately support students. Using AI methods such as machine learning (ML), the data that is generated when students interact in technology-enhanced classrooms can be used to track their learning and subsequently to inform teachers so that they can better support student learning. This study implemented digital workbooks in an inquiry-based physics unit, collecting cognitive, metacognitive, and affective data from 214 students. Using ML methods, an early warning system was developed to predict students' learning outcomes. Explainable ML methods were used to unpack these predictions and analyses were conducted for potential biases. Results indicate that an integration of cognitive, metacognitive, and affective data can predict students' productivity with an accuracy ranging from 60 to 100% as the unit progresses. Initially, affective and metacognitive variables dominate predictions, with cognitive variables becoming more significant later. Using only affective and metacognitive data, predictive accuracies ranged from 60 to 80% throughout. Bias was found to be highly dependent on the ML methods being used. The study highlights the potential of digital student workbooks to support SRL in inquiry-based science education, guiding future research and development to enhance instructional feedback and teacher insights into student engagement. Further, the study sheds new light on the data needed and the methodological challenges when using ML methods to investigate SRL processes in classrooms.</p> <p>Supplementary Information The online version contains supplementary material available at https://doi.org/10.1007/s10648-025-10011-9.</p> <p>Recent approaches to science teaching involve inquiry-based instruction where students take the role of a scientist to explore a phenomenon by engaging in scientific practices such as developing and testing hypotheses (Pedaste et al., [<reflink idref="bib86" id="ref1">86</reflink>]). Inquiry instruction affords learners with many degrees of freedom and allows them to investigate a problem in their own way. Thus, the inquiry process requires students to become active and regulate their learning which is usually challenging for them (Schraw et al., [<reflink idref="bib98" id="ref2">98</reflink>]). It is therefore important that teachers notice learners who struggle with self-regulation and provide them with adequate support early in the inquiry process. Technology-enhanced learning environments have the potential to support teachers in monitoring their students' inquiry and understanding by collecting data about cognitive, affective, motivational, and metacognitive processes of self-regulation during learning (CAMM processes; Molenaar et al., [<reflink idref="bib80" id="ref3">80</reflink>]). Machine learning (ML) techniques can leverage such data and use them to predict students' learning performance, which holds the potential to develop early warning systems for teachers (Macfadyen &amp; Dawson, [<reflink idref="bib67" id="ref4">67</reflink>]). Such early warning systems, in turn, can support teachers in making decisions about providing adaptive support to students.</p> <p>Ample research has provided insights into how CAMM processes during self-regulation of learning relate to learning outcomes (e.g., Azevedo et al., [<reflink idref="bib7" id="ref5">7</reflink>]; Saint et al., [<reflink idref="bib94" id="ref6">94</reflink>]; Xu et al., [<reflink idref="bib116" id="ref7">116</reflink>]). However, the relevance of the different processes and to what extent the insights from this body of work can guide the development of early warning systems for real-life classrooms[<reflink idref="bib1" id="ref8">1</reflink>] remains an open question because previous studies typically drew on data from laboratory, or otherwise well-controlled data collections over relatively short periods of time. While collecting data in such settings has its merits, for example, they allow collecting multimodal data streams—often including sensor data (e.g., Azevedo et al., [<reflink idref="bib7" id="ref9">7</reflink>]; Molenaar et al., [<reflink idref="bib80" id="ref10">80</reflink>]; Winne, [<reflink idref="bib112" id="ref11">112</reflink>])—it also has its challenges. Laboratory and laboratory-like settings do not reflect the affordances of data collections in real-life classrooms, such as having to find a balance between collecting data unobtrusively and collecting data that represents objective, reliable, and valid operationalizations of the constructs of interest, as well as collecting data in a time and effort efficient manner (i.e., economical), and considering legal and ethical bounds that are in place to protect students' privacy. Furthermore, it is challenging to collect data on SRL processes and learning outcomes that unfold during inquiry learning over weeks or even months of instruction. Against this background, there is a need to better understand how different CAMM processes can be modelled and used to predict learning outcomes, when represented by data that is collected under the affordances of the criteria for data collection in classrooms outlined above.</p> <p>The present study aims to address this need with a secondary analysis of data that was collected in middle school physics classrooms as part of a research project. This data set encompassed data on the cognitive, affective, and metacognitive processes (CAM) of learning in the classroom, while data on students' motivation was not available. With the present analysis, we address the methodological question whether and how different data streams that capture these processes can be combined to predict, as early as possible during instruction, which students might encounter difficulties. More specifically, we address this question by comparing machine learning models in terms of their predictive accuracy, while simultaneously investigating the predictive power of different combinations of CAM processes. In addition, we consider bias regarding gender and parental educational background as methodological challenges for predictions based on machine-learning.</p> <hd id="AN0184416951-2">Background</hd> <p></p> <hd id="AN0184416951-3">Inquiry Learning in Science Classrooms and Self-regulated Learning</hd> <p></p> <hd id="AN0184416951-4">Self-regulation of Learning During Inquiry-Instruction</hd> <p>Recent research underscores the benefits of inquiry learning for effective science education (e.g., Schneider et al., [<reflink idref="bib95" id="ref12">95</reflink>]). This approach of science instruction focuses on engaging the learners in an active investigation of natural phenomena through (in varying degrees) self-directed discovery (De Jong &amp; Lazonder, [<reflink idref="bib23" id="ref13">23</reflink>]). During inquiry-instruction the students actively construct new knowledge by carrying out inquiry processes such as learning about relationships between concepts through setting goals, formulating hypotheses, observing phenomena, conducting experiments, and testing hypotheses in the classroom (Pedaste et al., [<reflink idref="bib86" id="ref14">86</reflink>]). These activities are grouped into five general phases that form the inquiry cycle: Orientation, conceptualization (e.g., formulating hypotheses), investigation, conclusion, and discussion (i.e., reflection) (Pedaste et al., [<reflink idref="bib86" id="ref15">86</reflink>]). Due to the iterative nature of the inquiry process, self-regulation of the learning process is a necessary condition for successful learning during inquiry (De Jong &amp; Njoo, [<reflink idref="bib24" id="ref16">24</reflink>]; Lai et al., [<reflink idref="bib59" id="ref17">59</reflink>]; Pedaste et al., [<reflink idref="bib86" id="ref18">86</reflink>]). Several models exist that describe self-regulation of learning and while these models diverge in many regards, they all share the notion that self-regulation of learning can be understood as an active, constructive process in which learners set goals and continuously monitor their progress towards these goals. Based on the results of this monitoring, learners take control over their learning by regulating their cognitions, affects, metacognition, and motivation (Dent &amp; Koenka, [<reflink idref="bib25" id="ref19">25</reflink>]; Schraw et al., [<reflink idref="bib98" id="ref20">98</reflink>]). This collection of four components of self-regulation is sometimes referred to as "CAMM" processes (cognition, affect, metacognition, and motivation; e.g., Azevedo et al., [<reflink idref="bib6" id="ref21">6</reflink>]; Molenaar et al., [<reflink idref="bib80" id="ref22">80</reflink>]).</p> <hd id="AN0184416951-5">Cognitive, Affective, Motivational, and Metacognitive Processes During Self-regulation</hd> <p>Self-regulation of learning is assumed to be a cyclical process (e.g., Winne &amp; Perry, [<reflink idref="bib113" id="ref23">113</reflink>]; Zimmerman, [<reflink idref="bib120" id="ref24">120</reflink>]). According to the COPES model (conditions, operations products, evaluations, standards; Winne &amp; Perry, [<reflink idref="bib113" id="ref25">113</reflink>]), learners first generate a perception of the task, followed by setting goals and generating plans on how to reach these goals. In order to reach their goals, learners apply learning strategies and continuously monitor whether their current strategy is suitable to achieve the goal. Depending on the result of this monitoring, learners may choose to adapt their current strategy or select a different strategy. The metacognitive components of planning and monitoring can cause learners to revisit earlier steps in the regulation cycle. Thus, the cycle of self-regulation is not strictly linear, but recursive (Azevedo et al., [<reflink idref="bib7" id="ref26">7</reflink>]; Winne, [<reflink idref="bib111" id="ref27">111</reflink>]).</p> <p>In the classroom, instruction ultimately aims at promoting knowledge integration which encompasses students acquiring ideas, integrating them, and thus constructing increasingly complex and integrated networks of ideas (e.g., Linn et al., [<reflink idref="bib64" id="ref28">64</reflink>], [<reflink idref="bib65" id="ref29">65</reflink>]). This process includes phases of self-regulation, where cognitive, affective, metacognitive, and motivational processes play a role in explaining how self-regulation of learning benefits learning outcomes. In general, there is evidence suggesting that learners who are proficient in self-regulation exhibit higher learning performance (see meta-analysis by Dent and Koenka ([<reflink idref="bib25" id="ref30">25</reflink>]), as well as Lai et al. ([<reflink idref="bib59" id="ref31">59</reflink>]) and Sinatra and Taasoobshirazi ([<reflink idref="bib100" id="ref32">100</reflink>]) for respective findings for science learning). For inquiry in particular, Lai et al. ([<reflink idref="bib59" id="ref33">59</reflink>]) showed that guiding students' self-regulation during inquiry not only facilitates beneficial inquiry practices such as time management and self-evaluation, but also led to higher learning achievement as compared to learners who engaged in the same inquiry instruction, but without additional regulation support.</p> <p>Previous studies investigated the role of the different cognitive, affective, metacognitive, and motivational processes, and their interplay (for overviews see for example Azevedo et al., [<reflink idref="bib4" id="ref34">4</reflink>]; Molenaar et al., [<reflink idref="bib80" id="ref35">80</reflink>]) during learning. Cognitive processes encompass applying strategies that target processing learning material with the goal of acquiring and integrating new knowledge, whereas metacognitive processes encompass goal setting, planning, self-monitoring (e.g., judgements of learning), self-evaluation, and self-control (e.g., Azevedo et al., [<reflink idref="bib7" id="ref36">7</reflink>]; Dent &amp; Koenka, [<reflink idref="bib25" id="ref37">25</reflink>]). These processes impact how effectively learners are able to plan and adapt their learning process and thus explain how self-regulation occurs (Dent &amp; Koenka, [<reflink idref="bib25" id="ref38">25</reflink>]; Winne &amp; Perry, [<reflink idref="bib113" id="ref39">113</reflink>]). In their meta-analysis, Dent and Koenka ([<reflink idref="bib25" id="ref40">25</reflink>]) found that both cognitive and metacognitive strategies were positively associated with academic achievement, but that the association between metacognitive strategies, especially planning, and achievement was stronger than the association between cognitive strategies and achievement. Other studies highlight the interplay between students' prior knowledge and their use of learning strategies. For instance, in the context of learning with the intelligent tutoring system "MetaTutor" (Azevedo et al., [<reflink idref="bib7" id="ref41">7</reflink>]), students with low prior knowledge relied on ineffective cognitive strategies such as taking verbatim notes, while students with higher prior knowledge showed sequences of note taking and summarizing (thus organizing the learning content) (Azevedo et al., [<reflink idref="bib7" id="ref42">7</reflink>]; Taub &amp; Azevedo, [<reflink idref="bib102" id="ref43">102</reflink>]). Such superficial notes predicted lower learning outcomes (Azevedo et al., [<reflink idref="bib7" id="ref44">7</reflink>]; Trevors et al., [<reflink idref="bib105" id="ref45">105</reflink>]).</p> <p>These cognitive and metacognitive processes do not occur in a vacuum. Motivation initiates and maintains the self-regulated learning process (Dent &amp; Koenka, [<reflink idref="bib25" id="ref46">25</reflink>]) and lower learning gains are to be expected if learners possess high cognitive and metacognitive skills, but lack motivation (McDowell, [<reflink idref="bib72" id="ref47">72</reflink>]; Zusho et al., [<reflink idref="bib121" id="ref48">121</reflink>]). Another significant factor that should be considered are students' affects during learning. While Efklides et al. ([<reflink idref="bib33" id="ref49">33</reflink>]), Efklides and Petkaki ([<reflink idref="bib32" id="ref50">32</reflink>]), Zheng et al. ([<reflink idref="bib119" id="ref51">119</reflink>]), and Mega et al. ([<reflink idref="bib73" id="ref52">73</reflink>]) stress that affective processes are tightly interwoven with cognitive and metacognitive processes, Ainley ([<reflink idref="bib1" id="ref53">1</reflink>]) adds that the classic perspectives on motivation has to be expanded with an affective perspective as well, since the psychological processes underlying motivation and affect are closely linked. For example, emotions influence how learners allocate effort. Experiences such as disfluency, that is, a feeling that the task is more difficult than expected, inform metacognition and lead to emotions such as frustration or curiosity (Efklides et al., [<reflink idref="bib33" id="ref54">33</reflink>]). Further, positive effects were found by Mega et al. ([<reflink idref="bib73" id="ref55">73</reflink>]) who reported that positive emotions have a positive impact on learning, when effective self-regulatory actions and motivation during learning are present as well. While especially the role of different emotions during self-regulation of learning and for learning performance in the classroom is not understood as thoroughly as the role of cognition and metacognition (see SMA grid, Molenaar et al., [<reflink idref="bib80" id="ref56">80</reflink>], or the review by Azevedo et al., [<reflink idref="bib7" id="ref57">7</reflink>]), evidence underlines the central role of affect during learning. For instance, meta-analytical evidence yielded robust positive associations between activating, positive emotions (e.g., enjoyment) and learning outcomes, whereas activating negative emotions such as anger, but also deactivating negative emotions such as boredom were associated with low performance (Camacho-Morles et al., [<reflink idref="bib14" id="ref58">14</reflink>]; also see Pardos et al., [<reflink idref="bib84" id="ref59">84</reflink>]; Vilhunen et al., [<reflink idref="bib109" id="ref60">109</reflink>]; Zheng et al., [<reflink idref="bib119" id="ref61">119</reflink>]). For frustration (i.e., a deactivating, negative emotion) the meta-analysis by Camacho-Morles et al. ([<reflink idref="bib14" id="ref62">14</reflink>]) did not yield significant associations with achievement. As one reason for this finding the authors hypothesize that the effects of frustration on effort regulation differ between learners.</p> <p>In summary, research has accumulated evidence that explains how self-regulation as a whole, affords and affects learning, while acknowledging the crucial role of cognitive, metacognitive, and affective processes. At the same time, the reviewed studies show that there is an intricate interplay between the different processes, especially CAM processes. Presently, this interplay is gaining increasing attention (Molenaar et al., [<reflink idref="bib80" id="ref63">80</reflink>]). With respect to helping teachers in supporting their students during learning, being sensitive to the different cognitive, metacognitive, and affective processes appears promising.</p> <p>Previous studies primarily analyzed SRL on a micro level (e.g., specific learning strategies), for instance in studies that focus on patterns in the regulation processes (e.g., Bernacki, [<reflink idref="bib10" id="ref64">10</reflink>]) and its real-time aspects (e.g., Azevedo et al., [<reflink idref="bib6" id="ref65">6</reflink>]; Saint et al., [<reflink idref="bib93" id="ref66">93</reflink>]). The majority of these studies has been conducted in laboratory, or otherwise well-controlled settings that do not reflect the affordances of inquiry learning in classrooms, that is, the majority of studies has limited ecological validity with respect to real-life inquiry learning. The affordances of inquiry learning in classrooms are manifold. For example, many factors limit the available data-streams, e.g., the pace of changes between rooms, different teachers, and limited breaks between lessons may prohibit the collection of sensor data, already strained time resources may limit the amount of instruments that can be administered, and unplanned events such as illness, conflicts between students or sudden, unplanned room changes may disrupt data-collection protocols.</p> <p>Despite these challenges, we argue that monitoring indicators for the different CAM processes during instruction in the classroom can help us identify learners who struggle and thus may benefit from support, for example from their teacher. As described above, inquiry instruction is challenging for learners because it requires apt self-regulation skills. The steps of the inquiry process build on each other and struggling at one step may hamper students' ability to proceed or succeed at a subsequent step. Optimally, teachers notice this early on in the process. If the inquiry instruction is embedded in a digital learning environment, it is possible to collect data about different processes involved in SRL and leverage these insights for an early warning system.</p> <hd id="AN0184416951-6">Early Warning Systems</hd> <p>Early warning systems play a crucial role in identifying students who are at risk of academic failure, allowing for timely and adaptive interventions. By detecting early signs of struggle these systems enable educators to provide personalized support, either through direct teacher involvement or automated, adaptive mechanisms. The benefits of such systems are substantial, as they can mitigate potential dropouts, enhance student retention, and improve overall academic performance (e.g., Arnold et al., [<reflink idref="bib4" id="ref67">4</reflink>]; Jayaprakash et al., [<reflink idref="bib50" id="ref68">50</reflink>]; Macfadyen &amp; Dawson, [<reflink idref="bib67" id="ref69">67</reflink>]).</p> <p>The title "Mining LMS data to develop an "early warning system" for educators: A proof of concept" (Macfadyen &amp; Dawson, [<reflink idref="bib67" id="ref70">67</reflink>]) indicates the notion of early warning systems has been popularized in an educational data mining tradition, that is, a tradition of data analysis that emphasizes an inductive, data-driven approach leveraging data science methods such as machine learning to better understand and support learning using data that is available. Further, respective work has primarily focused on higher education settings and typically considers learning across a semester. Research on early warning systems has utilized various predictors to identify at-risk students. For instance, Arnold and Pistilli ([<reflink idref="bib3" id="ref71">3</reflink>]) identified course engagement metrics, such as frequency of login, assignment submission patterns, and participation in online discussions, as key indicators of academic success. Macfadyen and Dawson ([<reflink idref="bib67" id="ref72">67</reflink>]) found that learning management system data, including the number of discussion posts, quiz attempts, and resource views, were significant predictors of student performance. However, data related to self-regulated learning and the respective CAM processes has rarely been used in early warning systems, highlighting a gap in current research.</p> <p>A different approach to predicting student performance with the goal of supporting students has been followed in the research tradition around intelligent tutoring systems (ITS) (Corbett et al., [<reflink idref="bib20" id="ref73">20</reflink>]). In contrast to the early warning system tradition, this line of work has emphasized careful, theory guided modelling (e.g., using Bayesian Networks) of students' performance drawing on data deliberately collected using careful designs. Further, rather than predicting performance over the course of a semester, intelligent tutoring systems typically focus on predicting performance on a task or set of tasks with the goal of providing adaptivity to students (e.g., in the form of feedback, scaffolding, or in selecting what tasks the student should work on next). Recently, researchers have begun to incorporate CAM related dimensions such as affect into intelligent tutoring systems (e.g., Grawemeyer et al., [<reflink idref="bib42" id="ref74">42</reflink>]).</p> <p>Given the essential role of the different cognitive, metacognitive, and affective processes (see Camacho-Morles et al., [<reflink idref="bib14" id="ref75">14</reflink>]; Dent &amp; Koenka, [<reflink idref="bib25" id="ref76">25</reflink>]; Molenaar et al., [<reflink idref="bib80" id="ref77">80</reflink>]) for learning outcomes, we argue that modeling these factors in an early warning system should be considered, as they represent the self-regulated learning processes more immediately when compared to earlier attempts that often relied on straight-forward behavioral metrics such as dwell time or number of posts or quiz attempts. This lends itself to the question: which facets of self-regulated learning predict students' success most reliably and early enough?</p> <p>Evaluating the potential of an early-warning system further depends on the affordances of the learning environment in which self-regulated learning takes place. When a learning environment is designed with the possibility of providing insights into self-regulated learning in mind, the data-streams available can be more deliberate and theory informed than in research from a data mining tradition but probably remain less fine grained and rich—given the affordances of collecting data in schools—than in research from an intelligent tutoring system tradition or recent work that investigated self-regulated learning in controlled settings (e.g., Bernacki, [<reflink idref="bib10" id="ref78">10</reflink>]; Saint et al., [<reflink idref="bib94" id="ref79">94</reflink>]). Thus, it is not only relevant which facets of self-regulated learning predict students' success most reliably and early enough but also how readily respective data can be collected and analyzed.</p> <p>A last consideration concerns the methods that are used to model CAM related data. In the context of early warning systems, the interpretability of the models—typically machine learning models—underlying early warning systems is paramount. For these systems to be actionable, the insights generated must be understandable and accessible to educators, enabling them to make informed decisions. For these systems to add to our understanding of self-regulated learning, researchers must be able to relate the results from these systems to substantive theory. Lastly, the interpretability ensures that the interventions are not only data-driven but also contextually appropriate and feasible for implementation in diverse educational settings (Baker &amp; Siemens, [<reflink idref="bib9" id="ref80">9</reflink>]).</p> <p>As noted above, early-warning systems often leverage machine-learning approaches to identify patterns in data. Machine learning approaches, despite their clear benefits, are also associated with challenges that must be considered when employed in learning settings.</p> <hd id="AN0184416951-7">Machine Learning—Potentials and Challenges</hd> <p>In the past few years, machine learning has been positioned as a powerful framework to investigate the complexities of educational processes (Hilbert et al., [<reflink idref="bib45" id="ref81">45</reflink>]) and support learning (U.S. Department of Education, Office of Educational Technology, [<reflink idref="bib107" id="ref82">107</reflink>]). In the context of self-regulated learning, Molenaar et al. ([<reflink idref="bib80" id="ref83">80</reflink>]) emphasize how different machine learning techniques have been used to measure regulation processes or make predictions about students' self-regulated learning activity with the goal of supporting learning, e.g., through adaptive feedback or scaffolds. For example, Taub et al. ([<reflink idref="bib103" id="ref84">103</reflink>]) used commercial (and effectively black-box) software to automatically detect emotions in facial expression video data to investigate the relationship between affective processes in an educational game, and Lim et al. ([<reflink idref="bib63" id="ref85">63</reflink>]) used rule-based AI, that is, an AI system that relies on a set of predefined IF... THEN... statements to make decisions, to provide personalized scaffolds to learners, e.g., when they found that students had engaged in a certain learning activity based on trace data such as mouse-clicks, a predefined scaffold was displayed.</p> <p>While these examples show the potential of using machine learning and other AI methods to support (research into) self-regulated learning, they do not discuss how the threat of bias challenges these potentials. As Baker and Hawn ([<reflink idref="bib8" id="ref86">8</reflink>]) notice in a recent review, numerous definitions of bias[<reflink idref="bib2" id="ref87">2</reflink>] are used in the literature. A helpful distinction is to consider these definitions along a continuum along the extent to which these definitions focus on the technical versus the societal aspect of bias. On the technical end, bias can be considered as "a model's predictive performance (however defined) unjustifiably differs across disadvantaged groups along social axes such as race, gender, and class" (Mitchell et al., [<reflink idref="bib77" id="ref88">77</reflink>]). This represents the more technical end as the emphasis is one the machine learning model. In contrast, bias can also be considered as "unwanted or societally unfavorable outcome[s]" (Suresh &amp; Guttag, [<reflink idref="bib101" id="ref89">101</reflink>]). The latter definition of bias represents the societal end of the spectrum as the emphasis is on the (societal) outcome. This distinction is important because a machine learning model cannot exhibit technical bias but still produce societally unfavorable outcomes (and vice versa). In the context of this study, we focus on the technical aspect of bias as we are not (yet) concerned with the development of a system that is actually implemented in society (i.e., a classroom) but rather investigate the preliminary conditions for the development of such a system. More specifically, we consider the aspect of separation, i.e., the extent to which a machine learning model makes correct and incorrect predictions at similar rates for different groups (Kizilcec &amp; Lee, [<reflink idref="bib54" id="ref90">54</reflink>]).</p> <p>The extent to which bias has been found in machine learning models varies widely across and within contexts. Prominent examples of bias in educational contexts are repeated cases of online-exam supervision software failing to detect students with darker skin colors (e.g., Tiku, [<reflink idref="bib104" id="ref91">104</reflink>]). In the context of early warning systems, a recent study by Gándara et al. ([<reflink idref="bib38" id="ref92">38</reflink>]) found that across various machine learning models that used common predictors of student success in college were less accurate for students from racially minoritized students. In contrast, recent work in the context of self-regulated learning, has found little evidence of bias, concluding that the developed machine learning models are robust against bias (e.g., Zhang et al., [<reflink idref="bib118" id="ref93">118</reflink>]). Overall, the evidence suggests that bias-free systems can be developed but the already mentioned variability in the literature does not yet provide a clear picture regarding the conditions and factors, for instance, the choice of machine learning model or data-stream, affecting the presence of bias in machine learning models.</p> <p>In consequence, recent work in education (Krist &amp; Kubsch, [<reflink idref="bib56" id="ref94">56</reflink>]), especially in the learning analytics community (Cerratto Pargman et al., [<reflink idref="bib15" id="ref95">15</reflink>]; Kitto &amp; Knight, [<reflink idref="bib53" id="ref96">53</reflink>]), stresses that machine learning systems need to be critically evaluated with regard to bias. Recently, this sentiment was taken up in government reports (U.S. Department of Education, Office of Educational Technology, [<reflink idref="bib107" id="ref97">107</reflink>]) and in legislature in the European Union (European Commission, [<reflink idref="bib19" id="ref98">19</reflink>]). Thus, when leveraging data that is produced during learning in digital environments with the goal to make predictions based on machine learning, potential bias in ML models needs to be considered.</p> <p>When investigations into the bias in ML models reveal bias, this bias may be addressed or mitigated. In a recent review Idowu ([<reflink idref="bib48" id="ref99">48</reflink>]) identified several successful strategies to mitigate bias such as class balancing techniques or slicing analysis. However, these strategies represent fixes to underlying issues in the data, e.g., as a result of sampling bias. Thus Idowu ([<reflink idref="bib48" id="ref100">48</reflink>]) recommends prioritizing fairness in the data whenever possible.</p> <p>Another issue with machine learning that is tied to the issue of fairness is that models with increased predictive accuracy often are less scrutable (Kay et al., [<reflink idref="bib52" id="ref101">52</reflink>]) or interpretable (Breiman, [<reflink idref="bib12" id="ref102">12</reflink>], [<reflink idref="bib13" id="ref103">13</reflink>]). Highly accurate models, such as (deep) neural networks and ensemble methods (e.g., random forests), rely on complex, non-linear interactions among variables, making their decision-making processes opaque and difficult to trace. This opacity challenges scrutability, the ability (of people affected by a system, e.g., learners or teachers) to examine and understand how a model comes to its conclusions. Scrutable models clarify how inputs influence outputs, helping identify biases and errors. When models lack scrutability, assessing their fairness and reliability becomes more difficult, as their internal logic remains inaccessible. Beyond scrutability, reduced interpretability limits what we can learn about underlying theoretical relationships. In research, interpretable models provide not only predictions but also insights into causal mechanisms. When reasoning processes are obscured, deriving or testing theoretical knowledge and trusting ML-driven decisions becomes challenging. To address these concerns, explainable AI (XAI) (Minh et al., [<reflink idref="bib74" id="ref104">74</reflink>]) develops methods to improve both scrutability and interpretability. Techniques such as feature importance analysis, attention mechanisms, local approximations (e.g., LIME and SHAP), and inherently interpretable models aim to enhance transparency (Molnar, [<reflink idref="bib81" id="ref105">81</reflink>]).</p> <hd id="AN0184416951-8">Research Questions</hd> <p>Research shows that (guided) inquiry plays a key role in effective science education (Schneider et al., [<reflink idref="bib95" id="ref106">95</reflink>]). However, engaging in inquiry requires self-regulated learning from students. For teachers, it is often challenging to track their students' progress in such learning environments, making it hard to adequately support all students. In schools, the advent of technology-enhanced inquiry learning environments has enabled the gathering of data reflective of students' self-regulated learning processes across different data streams—although limited by the affordances of collecting data in classrooms. Machine learning techniques are promising in their ability to integrate these data in order to analyze students' learning, thus providing the basis for the generation of actionable insights. This can take the form of dashboards (Molenaar &amp; Knoop-van Campen, [<reflink idref="bib79" id="ref107">79</reflink>]) or early warning systems (Macfadyen &amp; Dawson, [<reflink idref="bib67" id="ref108">67</reflink>]) that support teachers in providing adequate support to their students. Research on SRL has demonstrated how CAM processes affect successful learning. Specifically, the interplay between different processes has received increasing attention. While previous work has worked towards modeling self-regulated learning on the level of individual processes by exploring individual processes as well as their interplay, it is currently not clear how data from multiple data streams that reflect different cognitive, metacognitive, and affective dimensions can be analyzed—using ML techniques—to make valid statements about the productiveness of students' learning. Thus, the present study seeks to explore different combinations of CAM aspects of learning and how they relate to students' learning success. Given the susceptibility of ML techniques to bias (Baker &amp; Hawn, [<reflink idref="bib8" id="ref109">8</reflink>]), it is further important to investigate the influence of analytical choices such as the used ML techniques. To address these research gaps, we ask the following research questions:</p> <p></p> <ulist> <item> Research question 1: To what extent can different data streams, CAM (cognition, affect, metacognition) dimensions, and ML techniques be used to predict the productivity of students' learning in ecologically valid learning environments such as schools?</item> <p></p> <item> Research question 2: To what extent are predictions of the productivity of students' learning trajectories susceptible to bias?</item> </ulist> <hd id="AN0184416951-9">Methods</hd> <p>To answer our research questions, we reanalyzed data from a larger research project that investigated students' learning trajectories about energy in middle school using a technology enhanced workbook. In the study, physics teachers were provided with an instructional unit about energy that followed inquiry-learning which was enriched with a technology-enhanced workbook that recorded student artifacts. This allowed us to collect data that reflect SRL processes. The dataset included indicators for CAM aspects of self-regulation; notably, the dataset did not include data on students' motivation during learning. The data and analysis scripts are available here: https://osf.io/uv8tn/.</p> <hd id="AN0184416951-10">Setting and Sample</hd> <p>The study was conducted in middle schools in northern Germany with <emph>N</emph> = 214 students from grade 7 and 8 from eight voluntarily participating teachers across two types of schools: the <emph>Gymnasium</emph> and the <emph>Gemeinschaftsschule.</emph> In the German school system, a <emph>Gymnasium</emph> is a type of secondary school that prepares students for higher education. It typically covers grades 5 to 12 or 13 and ends with the <emph>Abitur—</emph>a qualification for university entrance. The curriculum is academically rigorous, focusing on a broad range of subjects including languages, sciences, and humanities. A <emph>Gemeinschaftsschule</emph> is a more inclusive type of secondary school that caters to students of all abilities. It allows students to stay together longer before deciding on their educational path and often provides various qualifications, from vocational certificates to the <emph>Abitur</emph>, depending on the students' capabilities and goals. Table 1 shows the composition of the sample across school types and grades.</p> <p>Table 1 Distribution of students across grades and school type</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Grade&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="2"&gt;&lt;p&gt;School type&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;&amp;#931;&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;Gemeinschaftsschule&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Gymnasium&lt;/p&gt;&lt;/th&gt;&lt;th align="left" /&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;7&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;43&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;54&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;97&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;8&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;31&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;86&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;117&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;&amp;#931;&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;74&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;130&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;214&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0184416951-11">Instructional Unit</hd> <p>The instructional unit aligned to the local state curriculum and focused on the energy concept, a core idea in physics. The unit was developed[<reflink idref="bib3" id="ref110">3</reflink>] in an iterative co-design process with teachers in a previous project and emphasized inquiry learning. More specifically, the design followed the principles of project-based learning pedagogy (Krajcik &amp; Shin, [<reflink idref="bib55" id="ref111">55</reflink>]) which is a widely used framework for the development of inquiry learning in science education. This means that students engaged in a series of investigations—using a range of scientific practices such as asking questions, planning and conducting investigations, analyzing data and constructing explanations—to answer the driving question "What is the best way to mount solar cells on a house?". The instructional unit consisted of five parts (Fig. 1). The first part motivated the driving question, during parts two to four, students answered aspects of that driving question by investigating three sub-driving questions, and finally synthesized their learning to answer the unit-level driving question in the fifth and last part.</p> <p>Graph: Fig. 1 Structure and contents of the instructional unit</p> <p>The technology enhanced workbook was implemented by ways of a moodle (Dougiamas &amp; Taylor, [<reflink idref="bib30" id="ref112">30</reflink>]) course. Moodle is an open-source learning management system that is widely used in schools in Germany. The workbook featured a wide variety of task formats ranging from multiple-choice to open-ended answers (see Fig. 2 for examples) and replaced paper-based worksheets and notebooks. Each student had access to a laptop or tablet of their own to enter their answers.</p> <p>Graph: Fig. 2 Example tasks from the technology enhanced workbook. a) Fill-in-the-gap task. b) Task that asks students to conduct an experiment and record their results</p> <p>All teachers participated in professional learning opportunities to familiarize them with the principles of project-based pedagogy, the general structure of the unit, central activities, and the technology enhanced workbook.</p> <hd id="AN0184416951-12">Data Sources for CAM Dimensions</hd> <p></p> <hd id="AN0184416951-13">Student Artifacts—Cognition</hd> <p>The technology enhanced workbook allowed us to record students' answers to the 36 tasks in the workbook. These tasks prompted students to construct explanations based on evidence, plan and conduct investigations, or engage in arguments about the interpretations of evidence and thus required cognitive processes such as elaborating and organizing ideas, goal setting and planning, or reasoning and problem solving. Using a procedure grounded in evidence-centered design (Mislevy &amp; Haertel, [<reflink idref="bib75" id="ref113">75</reflink>]; Pellegrino et al., [<reflink idref="bib89" id="ref114">89</reflink>]), we scored every answer with respect to the demonstrated knowledge elements, engagement in scientific practices, and learning performances, leading to a total number of 107 scores per student. For example, when a task required students to integrate two knowledge elements and a scientific practice in learning performance, this resulted in four scores for this task (two scores for the knowledge elements, one score for the engagement in the scientific practice, and one score for the learning performance). A learning performance is an integration of knowledge and scientific practice, such as constructing an explanation to answer the question "Why does a solar cell provide less electric energy when it is cloudy outside?". Knowledge elements and engagement in scientific practices were scored dichotomously (0 = knowledge element/scientific practice not demonstrated, 1 = knowledge element/scientific practice demonstrated) and learning performances were scored with a partial credit system (0 = learning performance not demonstrated, 1 = learning performance partly demonstrated, 2 = learning performance fully demonstrated).</p> <p>All open-ended tasks were scored by instructed student workers with drift checks being implemented in the form of regular reviews of coded data. Intercoder reliability was assessed on the basis of approximately 10% of the data which was randomly selected and independently coded by two coders. Spearman's rho was chosen as a measure of intercoder reliability to account for the ordinal character of our codes. An average Spearman's rho of 0.97 with a minimum of 0.68 indicated substantial intercoder reliability. Across the different scores we found an average Spearman's rho of 0.94 for knowledge elements, an average Spearman's rho of 0.91 for engagement in scientific practices, and an average Spearman's rho of 0.94 for learning performances.</p> <p>Many students had not answered every of the 36 tasks due to various reasons, for example, a student might not know the answer to a task, the teacher may have instructed students to leave out a task, or the student might have missed a day of instruction due to illness. In this light, imputation of missing variables is challenging.[<reflink idref="bib4" id="ref115">4</reflink>] Further, as we have no information for why a student may not have answered a task given the plethora of potential reasons mentioned above, we did not assign a "0" if a student had not attempted a task as this would—potentially incorrectly— assume that the student would not hold the respective disposition. Rather, we decided to focus on the evidence presented in the tasks attempted. Thus, we decided to calculate sum scores of the knowledge elements, scientific practices, and learning performances for each task in the five parts of the units that was attempted, i.e., where students entered any kind of answer, to represent the amount of demonstrated evidence of proficiency with the respective knowledge element, scientific practice, or learning performance. In contrast to calculating mean scores, this has the benefit that confidence in the evidence is better reflected. Thus consider for example two students A and B: both have scored a 1 for a specific knowledge element—A in two tasks and B in three tasks. Both students would receive a mean score of 1. In contrast, in case of a sum score, student A has a sum score of 2 and student B a sum score of 3, reflecting that student B has demonstrated more evidence for holding the respective knowledge element.</p> <p>In addition, we also calculated the total sum score across all sub-scores (knowledge elements, scientific practices, and learning performances) for each part of the unit as a measure of the evidence that a student had met the overall goal of the respective part of the unit. The ideas behind this is that those total sum scores more reliably reflect the overall proficiency of a student as they rely on more data (and also lead to fewer predictors in the models) whereas the aforementioned sum scores per knowledge element, scientific practice, and learning performance provide more diagnostic information but are probably less reliable as they rely on less data and lead to more predictors in the model.</p> <hd id="AN0184416951-14">Self-report Measures</hd> <p></p> <hd id="AN0184416951-15">Affect and Metacognition</hd> <p>After each of the five parts of the instructional unit, students' epistemic emotions—including the critical control ("At the moment, I'm doing well.") and value ("The contents of the current activities are important to me") appraisals (e.g., Pekrun, [<reflink idref="bib87" id="ref116">87</reflink>])—and metacognition were assessed using a five-point Likert scale. For epistemic emotions, we used the seven-item short version of the epistemically related emotion scale that asked students to what extent they felt surprised, curious, confused, anxious, frustrated, joy, and bored during learning (Pekrun et al., [<reflink idref="bib88" id="ref117">88</reflink>]). The scale was validated with a sample of 438 students from three countries (Pekrun et al., [<reflink idref="bib88" id="ref118">88</reflink>]) and is assumed to be a cost-effective, minimally invasive instrument to assess epistemic emotions during learning. Based on experiences with the instrument during previous data collections and feedback from participants during piloting, we added "interested" as an eighth item to the scale.</p> <p>To assess metacognition, we developed three items that measured students' metacognitive evaluation of their understanding (judgement of learning), "I understand the contents well" and "I learn a lot during the lesson," as well as the perceived difficulty of the tasks, "The tasks are difficult."</p> <hd id="AN0184416951-16">Diversity Criteria</hd> <p>We assessed socio-demographic information, namely Gender and parental educational background, at the end of the unit. These students were asked to identify their gender: 74 students identified as female, 64 as male, and the remaining 76 students did not provide a gender identity. Further, we asked students whether any of their parents had attended an academic degree (Wößmann et al., [<reflink idref="bib115" id="ref119">115</reflink>]): 80 students reported that any of their parents had an academic degree, 30 students reported that this was not the case, and the remaining 110 students either reported that they did not know whether any of their parents had an academic degree or choose not to answer the question.</p> <hd id="AN0184416951-17">Analyses</hd> <p></p> <hd id="AN0184416951-18">Research Question 1: Predicting the Productivity of Students' Learning</hd> <p>To predict the productivity of students' learning, productivity needs to be defined in the first place. Based on students' sum score across all tasks of the unit—reflecting to what extent students met the learning goal of the unit, i.e., developing a better understanding of the energy concept and thus being able to answer the driving question—the simplest distinction of productivity would be to define scores below a certain threshold, e.g., the sample mean, as unproductive and scores above that threshold as productive. However, in the context of an early warning system, a more nuanced differentiation can be helpful. Therefore, we decided to not just differentiate between productive and unproductive but also included a moderately productive level that captures students that have room for improvement but are not in dire need of support. The mapping of scores to productivity levels is shown in Table 2.</p> <p>Table 2 Definition of productivity of learning</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Score&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Productivity of learning&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;In upper quartile&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Productive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;In middle quartiles&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Moderately productive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;In lower quartile&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Unproductive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Next, we fit different supervised machine learning models using the tidymodels (Kuhn &amp; Wickham, [<reflink idref="bib58" id="ref120">58</reflink>]) framework in R (R Development Core Team, [<reflink idref="bib91" id="ref121">91</reflink>]) to predict the productivity of students' learning based on combinations of the CAMM dimensions and for different time-points of the instructional unit which we will detail in the following.</p> <hd id="AN0184416951-19">Machine Learning Models</hd> <p>To predict students' productivity, we compared four increasingly complex and at the same time increasingly uninterpretable machine learning models: (<reflink idref="bib1" id="ref122">1</reflink>) (multinomial) logistic regression is straightforward to interpret but only accounts for linear relationships, (<reflink idref="bib2" id="ref123">2</reflink>) decision trees are also straightforward to interpret and can also handle non-linear relationships, and (<reflink idref="bib3" id="ref124">3</reflink>) boosting trees (specifically XGBoost (Chen &amp; Guestrin, [<reflink idref="bib17" id="ref125">17</reflink>])) are an extension of decision trees where multiple trees are combined and averaged to increase predictive accuracy. However, interpretability decreases as there is no single tree to interpret anymore. For tabular data—as in this study—these models routinely provide the best performance compared to other model classes (Grinsztajn et al., [<reflink idref="bib43" id="ref126">43</reflink>]). (<reflink idref="bib4" id="ref127">4</reflink>) Neural networks are a highly flexible class of models that can fit complex and non-linear relationships. Specifically, we used a feed-forward neural network with two layers consisting of ten nodes each between the input and output layer. These numbers were chosen based on modelling experiences and to avoid overfitting (Hastie et al., [<reflink idref="bib44" id="ref128">44</reflink>]). Neural networks with multiple layers between the input and output layers may be considered as instances of deep learning. Deep learning (and neural networks more generally) architectures have proven to be highly effective on complex data, for example in the context of computer vision and natural language processing (LeCun et al., [<reflink idref="bib61" id="ref129">61</reflink>]), but effectively are black-box models.</p> <hd id="AN0184416951-20">Combinations of CAM Dimensions</hd> <p>Our dataset encompassed data for three CAM dimensions: cognitive, affective, and metacognitive (see Table 3 for an overview of all variables and variable names). The affective and metacognitive data came from a self-report data stream while the cognitive data came from a student artifact data stream. For the cognitive data, we used the sum scores per sub-score (knowledge elements, engagement in scientific practices, and learning performances) in a part of the unit as well as the sum score across all scores in a part of the unit. Using sum scores per sub-score has the theoretical advantage of potentially providing finer grained diagnostic information (one sub-score, e.g., knowledge elements, may be more predictive than another). At the same time, the number of parameters that need to be fitted increases which will reduce precision of parameter estimates. In contrast, using the sum score across all scores has the benefit of having just one cognitive variable per part of the unit, allowing the respective parameters to be estimated with more precision, but sacrifices diagnostic specificity.</p> <p>Table 3 Overview of all variables, their names during analyses, and CAM dimension</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Variable&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Name&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;CAM dimension&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Sum scores of the knowledge elements, scientific practices, and learning performances for each of the five learning sets&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;[knowledge element]&lt;/p&gt;&lt;p&gt;P[x]&amp;#95;[practice]&lt;/p&gt;&lt;p&gt;P[x]&amp;#95;[learning performance]&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Cognitive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Total sum score across all sub-scores (knowledge elements, scientific practices, and learning performances)&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;total&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Cognitive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Epistemic emotion&amp;#8212;surprised&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;surprise&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Epistemic emotion&amp;#8212;curious&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;curiousity&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Epistemic emotion&amp;#8212;anxious&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;anxiety&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Epistemic emotion&amp;#8212;frustrated&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;frustration&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Epistemic emotion&amp;#8212;bored&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;boredom&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Epistemic emotion&amp;#8212;interested&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;interested&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Emotion&amp;#8212;control appraisal&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;control&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;Emotion&amp;#8212;value appraisal&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;emo&amp;#95;value&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;"I understand the contents well"&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;meta1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Metacognitive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;"I learn a lot during class"&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;meta2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Metacognitive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;"The tasks are difficult."&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;P[x]&amp;#95;meta3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Metacognitive&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>[x] denotes the timepoint during the instructional unit at which the variable was measured. The names in the "Name" column are later used in variable importance plots</p> <p>Overall, with these data sources we arrived at twelve different combinations of data sources that reflect unimodal, horizontal, and integrated analytical approaches (see Molenaar et al., [<reflink idref="bib80" id="ref130">80</reflink>]) (Table 4).</p> <p>Table 4 Combination of data sources and analytical approaches</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Cognitive (aggregated)&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Cognitive (all sub-scores)&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Affective&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Metacognitive&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Analytical approach&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;Unimodal&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;Unimodal&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;Unimodal&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Unimodal&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;Integrated&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Integrated&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Integrated&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;Integrated&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Integrated&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Integrated&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;Horizontal&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0184416951-21">Timepoints in the Instructional Unit</hd> <p>The instructional unit consisted of five parts. For an early warning system (and other interventions), it is important to learn about the productivity of a students' learning as early as possible in the unit. Thus, we systematically varied timepoints included in the prediction. Table 5 shows the relation between the timepoints and inclusion of data from the different parts of the unit. As the table shows, increasingly more data is used to predict the outcome as the unit progresses. This way, we can track how early in the unit precise predictions of the outcome become feasible and to what extent the importance of variables for predicting the outcome changes over the course of the unit.</p> <p>Table 5 Relation between timepoints and inclusion of data for the different machine learning models</p> <p> <ephtml> &lt;table frame="hsides" rules="groups"&gt;&lt;thead&gt;&lt;tr&gt;&lt;th align="left"&gt;&lt;p&gt;Timepoint&lt;/p&gt;&lt;/th&gt;&lt;th align="left" colspan="5"&gt;&lt;p&gt;Data included from part of the instructional unit&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;Outcome&lt;/p&gt;&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="left"&gt;&lt;p&gt;P1&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;P2&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;P3&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;P4&lt;/p&gt;&lt;/th&gt;&lt;th align="left"&gt;&lt;p&gt;P5&lt;/p&gt;&lt;/th&gt;&lt;th align="left" /&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;1&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" rowspan="5"&gt;&lt;p&gt;Productivity of learning&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;2&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;3&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;4&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left" /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;&lt;p&gt;5&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;td align="left"&gt;&lt;p&gt;x&lt;/p&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0184416951-22">Machine Learning Procedure</hd> <p>Machine learning starts with splitting the data into a training (75% of the data) and a test set (25% of the data). During this splitting, we specified that the gender balance was to remain the same between test and training set and all data from individual students is exclusively part of either the test or training set to avoid data leakage. Generally, the goal is to find an optimal model using the train set and validating it using the test set. However, finding the optimal model with the training set is not trivial because the different models have hyperparameters that determine model performance. Beyond rules of thumb and default settings for hyperparameters, which often do not provide the best performance, one can determine the hyperparameters based on the data. This involves testing various combinations of hyperparameters to identify the best set that yields the highest model accuracy or other performance metrics.</p> <p>Extensive hyperparameter search across a semi-random grid (Dupuy et al., [<reflink idref="bib31" id="ref131">31</reflink>]) of 75 combinations of hyperparameters was conducted for each combination of machine learning model, data source, and time point. When necessary, missing values for affective and metacognitive variables were imputed using bagged-trees (Kuhn &amp; Johnson, [<reflink idref="bib57" id="ref132">57</reflink>]). Altogether, this resulted in a total number of 4 (types of models) × 11 (combinations of data sources) × 5 (number of timepoints) = 220 models that were fit and evaluated using tenfold cross-validation (James et al., [<reflink idref="bib49" id="ref133">49</reflink>]).</p> <p>Next, using the combination of CAM dimensions that yielded the highest accuracies during hyper-parameter search, we evaluated the model predictions against the test set to assess how the models generalized.</p> <p>For the model that performed best against the test set, we investigated variable importance using the model specific variable importance measure gain (Chen &amp; Guestrin, [<reflink idref="bib17" id="ref134">17</reflink>]) which reflects what variables hold the greatest predictive power. Gain is a metric used in tree-based models like XGBoost to measure the improvement in the model's objective function (e.g., log loss, mean squared error) resulting from a split on a particular feature. It provides an intuitive understanding of the relative importance of features in the model. Gain-based importance is computed during training and is detailed in the XGBoost paper (Chen &amp; Guestrin, [<reflink idref="bib17" id="ref135">17</reflink>]).</p> <hd id="AN0184416951-23">Research Question 2: Investigating Bias</hd> <p>To investigate to what extent the predictions of students' learning outcomes exhibited bias with respect to gender or parental educational background, we focused on the best performing models from research question 1. For those models, we evaluated the predictive accuracy for each Gender and parental educational background separately. If the accuracies differ between students of different Gender or parental educational background, this reflects bias with bigger differences in accuracy indicating more pronounced bias. We focused on gender because there is a well-known gender bias in science (e.g., Miyake et al., [<reflink idref="bib78" id="ref136">78</reflink>]) in general and physics in particular (e.g., Avraamidou, [<reflink idref="bib5" id="ref137">5</reflink>]).</p> <p>Besides Gender, we investigated the role of parental educational background. Betancour et al. ([<reflink idref="bib11" id="ref138">11</reflink>]), using nationally representative data from the Early Childhood Longitudinal Study, have demonstrated that "higher levels of parental education (i.e., some college, high school, bachelor's, graduate school) predicted significant increases in science achievement." These effects remain after controlling for factors such as income or ethnicity. Further, based on PISA data these findings hold for school achievement more generally (OECD, [<reflink idref="bib83" id="ref139">83</reflink>]). Thus, these two dimensions, Gender and parental educational background, seemed particularly important to consider when investigating bias in a science context.</p> <hd id="AN0184416951-24">Results</hd> <p></p> <hd id="AN0184416951-25">Research Question 1</hd> <p></p> <hd id="AN0184416951-26">Predicting the Productivity of Learning</hd> <p>Figure 3 shows the cross-validation accuracy of the best performing models from the hyper-parameter search for the different combinations of CAM dimensions across all five timepoints. The highest predictive accuracy (as measured by the F1 scores; other metrics (see online supplement S1) provide equivalent results) was observed in most combinations of CAM dimensions for boosting, while logistic regression exhibited the lowest predictive accuracy in nearly all cases.</p> <p>Graph: Fig. 3 Hyperparameter search results: cross-validation F1 of ML models predicting the productivity of students' learning productivity based on unit outcome for different data streams and CAM dimensions</p> <p>The precision of decision trees tended to rank between boosting and logistic regression—sometimes nearly matching the performance of boosting, sometimes performing distinctly worse than boosting. Interestingly, neural networks had the highest predictive accuracy when only affect and metacognition were considered. In combinations of cognitive, affective and metacognitive dimensions, neural networks often initially outperformed the other models while falling back behind decision trees and boosting when more data was used (i.e., timepoints P4 and P5). With regard to the different combinations of CAM dimensions, it was notable that the models that included the affective and metacognitive dimensions only, barely improved in predictive accuracy over time. Instead, these models showed rather flat time-trends. In contrast, models that additionally included cognitive data yielded increasing predictive accuracy over time.</p> <p>Interestingly, the models that used only cognitive data had a lower accuracy at timepoints P1 and P2 compared to the models that use only affective and/or metacognitive data. At time point P3, models that use only cognitive and models that used only affective, and/or metacognitive data, were similar in accuracy. Finally, for timepoints P4 and P5, the models that relied only on cognitive data yielded higher predictive accuracy than the models that used only affective and/or metacognitive data.</p> <p>Overall, combining cognitive, affective, and metacognitive data yielded the highest accuracies (Meta &amp; Emo &amp; all knowledge / Meta &amp; Emo &amp; agg. knowledge), with the models that used the aggregated cognitive data (Meta &amp; Emo &amp; agg. knowledge) yielded higher accuracies than those that used all individual cognitive scores (Meta &amp; Emo &amp; all knowledge). Therefore, we selected the models that used the aggregated cognitive data, affective data, and metacognitive data for the evaluation against the test set. This evaluation allowed us to determine how well the models generalize to new data.</p> <p>Figure 4 shows the results of evaluating the models that used the aggregated cognitive data, affective data, and metacognitive data against the test set. Overall, we found a similar pattern compared to the respective panel in Fig. 3 but with a slightly decreased accuracy. This was to be expected as accuracies are typically lower when evaluating models against the test set as compared to cross-validation (Ghojogh &amp; Crowley, [<reflink idref="bib40" id="ref140">40</reflink>]). At P3, the best performing model (boosting) reached an accuracy of 65% which indicates that after 3 of the 5 parts of the unit (which includes the first 17 out of a total number of 36 tasks) the productivity of students' learning could be predicted with adequate accuracy.</p> <p>Graph: Fig. 4 Test-set F1 of ML models predicting the productivity of students' learning based on self-report and student artifact data streams for cognitive, affective, and metacognitive data (Meta &amp; Emo &amp; agg. knowledge from Fig. 3)</p> <hd id="AN0184416951-27">Variable Importance</hd> <p>Figures 5, 6, and 7 show variable importance plots for the best performing model—boosting using aggregated cognitive, affective, and meta-cognitive data—for time points P1, P2, and P3. For P4 and P5, variable importance was entirely dominated by the respective cognitive variable, i.e., the sum of all cognitive scores up to that time point. Still, Figs. 5, 6, and 7 provide valuable insights into which dimensions are predictive of the different learning trajectories earlier in the instructional unit.</p> <p>Graph: Fig. 5 Variable importance for boosting model predicting the productivity of students' learning based on self-report and student artifact data streams for cognitive, affective, and metacognitive data for timepoint 1</p> <p>Graph: Fig. 6 Variable importance for boosting model predicting the productivity of students' learning based on self-report and student artifact data streams for cognitive, affective, and metacognitive data for timepoint 2</p> <p>Graph: Fig. 7 Variable importance for boosting model predicting the productivity of students' learning based on self-report and student artifact data streams for cognitive, affective, and metacognitive data for timepoint 3</p> <p>Figure 5 shows that at time point P1, surprise and meta_1 ("I understand the content well."), and frustration are most important. Interestingly, meta_1 measured at time point P1 was still the most important variable when predicting the productivity of students' learning at time point P2 (Fig. 6). Only with some distance follow meta_2 ("I learn a lot during class.") and surprise as measured at time point P2. While the cognitive variable P1_total had the second to last importance in Fig. 1, we can see in Fig. 2 that the respective variable from time point P2 is the fifth most important variable.</p> <p>Finally, the trend of the cognitive variable becoming increasingly more important is visible in Fig. 7 where P3_total nearly dominated the importance. Interestingly, students' initial self-evaluation of their learning (P1_meta_1) remains somewhat important here with the third highest importance after interest as measured at P3.</p> <p>In sum, our results indicate that an integrating CAM dimensions across data streams (self-report and scored artifacts) leads to the highest predictive accuracy for the productivity of students' learning. Initially, students' metacognition and epistemic emotions are most important during the first two parts of the unit (P1, P2) and fade in importance compared to the cognitive dimension towards the end of the unit (parts P3, P4, and P5).</p> <hd id="AN0184416951-28">Research Question 2: Bias</hd> <p>Figure 8 shows the predictive accuracy by Gender for the best performing combination of CAM dimensions (Meta &amp; Emo &amp; agg. knowledge) across the four machine learning models. Overall, boosting exhibited the highest and most consistent accuracy across all time points, while logistic regression, neural network, and decision tree models yielded less consistent accuracies for both Genders. This indicates that boosting showed the least bias.</p> <p>Graph: Fig. 8 Test-set F1 of ML models predicting the productivity of learning based on self-report and student artifact data streams for cognitive, affective, and metacognitive data by Gender</p> <p>In contrast, in Fig. 9, the decision tree model showed the most consistent accuracies for both dimensions of parental educational background whereas boosting exhibited a pronounced difference between the academic and non-academic backgrounds at P1. Similar to Fig. 8, the consistency of the neural network and logistic regression was lower compared to boosting and decision tree.</p> <p>Graph: Fig. 9 Test-set F1 of ML models predicting the productivity of students' learning based on self-report and student artifact data streams for cognitive, affective, and metacognitive data by parental educational background</p> <p>Overall, these findings indicate that all modeling approaches were susceptible to bias—although to a different degree. Further, bias varied across categories such as parental educational background or Gender, while one model may exhibit little bias in one case, it can exhibit distinctly more in another.</p> <hd id="AN0184416951-29">Discussion</hd> <p>In this study, we explored to what extent different combinations of CAM aspects of SRL can be used to predict the productivity of learning in an authentic school setting (i.e., an instructional unit using a guided-inquiry approach). The goal of this work was to understand what data is needed and how it needs to be analyzed to provide valid, unbiased insights that allow supporting self-regulated learning processes, for example in the form of an early warning system (Macfadyen &amp; Dawson, [<reflink idref="bib67" id="ref141">67</reflink>]). Our results indicate that data streams about CAM dimensions that can be gathered relatively economically can be used to successfully predict the productivity of students' learning trajectories. Further, at least with respect to gender and parental educational background, bias is highly variable ranging from negligible to being very much present. We discuss these findings in more detail below.</p> <hd id="AN0184416951-30">Research Question 1—Predicting Learning Trajectories</hd> <p>Using a combination of short questionnaires and the analysis of students' artifacts in the digital learning environment, we found that the best predictions regarding the productivity of self-regulated learning were possible by combining data across the affective, metacognitive and cognitive dimension using AI methods such as machine learning. As such, our study can be classified as using an integrated approach where single instruments were used to assess each CAM dimension (see SMA grid, Molenaar et al., [<reflink idref="bib80" id="ref142">80</reflink>]). The cognitive dimension of SRL was measured using students' responses to tasks in the digital learning environment (i.e., behavioral). The respective tasks were carefully designed using evidence-centered design which afforded a reliable and valid measurement of students' knowledge and its development over time. The metacognitive, and affective dimensions of SRL were measured using short questionnaires (i.e., self-reports) which allowed not only an economical measurement at multiple times during the instruction, but also objective, reliable and valid measurements. Unlike studies that examine specific CAMM processes, our research adopted a broader perspective, considering the context (metacognitive, affective) and products (cognitive) of regulation during inquiry-instruction (see Winne &amp; Perry, [<reflink idref="bib113" id="ref143">113</reflink>]). While this combination of processes is not unique (see the SMA grid in Molenaar et al., [<reflink idref="bib80" id="ref144">80</reflink>]), our study stands out in two ways. First, our study was conducted in an ecological valid setting in which learners followed different instructional approaches such as working on tasks individually, conducting experiments or engaging in classroom discussions. Thus, our study represents a learning setting that afforded more degrees of freedom for teachers and learners. These activities were performed with the use of laptops and tablet computers to collect self-report and student artifact data. As such, our study stand in contrast to other studies that took an integrated approach to assess regulation of learning, employed sophisticated instruments to capture different modalities of the CAMM processes, such as bio-physiological data (e.g., EDA: Dindar et al., [<reflink idref="bib27" id="ref145">27</reflink>]; EEG: Mangaroska et al., [<reflink idref="bib69" id="ref146">69</reflink>]) or eye-tracking (e.g., Emerson et al., [<reflink idref="bib34" id="ref147">34</reflink>]). As a consequence, many studies that investigate self-regulation processes are conducted in laboratory or otherwise controlled settings (e.g., Emerson et al., [<reflink idref="bib34" id="ref148">34</reflink>]; Engelmann &amp; Bannert, [<reflink idref="bib35" id="ref149">35</reflink>]; Malmberg et al., [<reflink idref="bib68" id="ref150">68</reflink>]; Mangaroska et al., [<reflink idref="bib69" id="ref151">69</reflink>]; Rakovic et al., [<reflink idref="bib92" id="ref152">92</reflink>]; Taub et al., [<reflink idref="bib103" id="ref153">103</reflink>]; cf. e.g., Dindar et al., [<reflink idref="bib27" id="ref154">27</reflink>]; Gašević et al., [<reflink idref="bib39" id="ref155">39</reflink>] for studies that were conducted in classrooms or higher education settings), or used learning environments that limited possible learner behavior, such as intelligent tutoring systems (e.g., Azevedo et al., [<reflink idref="bib7" id="ref156">7</reflink>]; Saint et al., [<reflink idref="bib94" id="ref157">94</reflink>]). Against this background, our findings underscore the assumptions from research on the regulation of learning (e.g., Dent &amp; Koenka, [<reflink idref="bib25" id="ref158">25</reflink>]; Efklides et al., [<reflink idref="bib33" id="ref159">33</reflink>]; Winne &amp; Perry, [<reflink idref="bib113" id="ref160">113</reflink>]) that the different CAMM processes (in our case CAM) have individual contributions to learning achievement in the classroom. On the level of theory, modeling the interplay between the different processes is challenging; however, on an empirical level, we saw that an integrated perspective is effective in predicting learning achievement.</p> <p>One noteworthy finding of the present study is that in early phases of the instructional unit, metacognition (i.e., monitoring understanding) and epistemic emotions (i.e., interest and overall value appraisal) were the most important features in predicting the learning achievement at the end of the unit, whereas the cognitive dimension increased in importance towards the end of the unit. Finding that positive affect is predictive of learning is in line with findings from previous studies (e.g., Camacho-Morles et al., [<reflink idref="bib14" id="ref161">14</reflink>]; Mega et al., [<reflink idref="bib73" id="ref162">73</reflink>]; Pardos et al., [<reflink idref="bib84" id="ref163">84</reflink>]). What our study adds is the finding that the role of affect may change throughout the course of an instructional unit. We conjecture that early in the instruction, affective processes, such as interest and value, are likely more predictive of learning outcomes because they drive students to invest effort in the inquiry, motivated by the engaging nature of the driving questions. This engagement may also foster continuous self-monitoring of understanding, a metacognitive process that remains relevant throughout the cycle. However, as instruction progresses, the accumulation and integration of knowledge becomes increasingly predictive of final learning achievements. Prior knowledge, consistently a strong predictor as highlighted in meta-analyses like Simonsmeier et al. ([<reflink idref="bib99" id="ref164">99</reflink>]) becomes more informative, allowing for more precise estimations of learning achievement. It needs to be noted, however, that a portion of the predictive power of the cognitive component of SRL in our analyses stemmed from the high prediction power of prior knowledge (Simonsmeier et al., [<reflink idref="bib99" id="ref165">99</reflink>]). While the artifacts that were coded represent the result of cognitive processes (or combinations thereof) such as elaborating and organizing ideas, goal setting and planning, or reasoning and problem solving, they did not represent individual cognitive processes such as different learning strategies (Matcha et al., [<reflink idref="bib70" id="ref166">70</reflink>]) or selecting and processing different learning materials (Saint et al., [<reflink idref="bib94" id="ref167">94</reflink>]). It also needs to be noted that the initially relatively low predictive power of the cognitive scores may be a result of how we summarized the scores: the sum scores we used represent the amount of demonstrated evidence with respect to the cognitive dimension and thus higher scores compared to lower scores do not only reflect more evidence but are also more reliable as lower scores can also be the result of missing data. Thus, especially in the beginning of the unit when only data from a few tasks is considered, the data may be more noisy than towards the end of the unit. To address this issue in future research, it would be important to find ways, potentially using data from multiple modalities or process data, to reliably differentiate various reasons for missing data (i.e., absence, lack of knowledge) and to verify that the entered replies were the product of a students' own cognitive processes and not, for example, based on something they copied from their neighbor. Further, our findings need to be interpreted against several limitations. First, the dataset used for our analyses did not include measures for students' motivation during learning. Thus, our analyses covered only CAM, but not CAMM. While previous has shown that motivation and emotions are closely linked (e.g., Ainley, [<reflink idref="bib1" id="ref168">1</reflink>]; Mega et al., [<reflink idref="bib73" id="ref169">73</reflink>]) and while classical measures for motivation such as the Intrinsic Motivation Inventory include certain emotions (e.g., joy) (see e.g. McAuley et al. [<reflink idref="bib71" id="ref170">71</reflink>]), it may be worthwhile to empirically tease apart the effects of motivation and different emotions during learning. This may be especially relevant with respect to designing personalized support for students. Second, the instrument used to assess the metacognitive aspects of learning was self-developed and thus may have covered metacognition on the surface. For future studies, assessing the nuances of metacognition, such as metacognitive awareness (e.g., Tuononen et al., [<reflink idref="bib106" id="ref171">106</reflink>]) or confidence judgements (e.g., Dougherty et al., [<reflink idref="bib29" id="ref172">29</reflink>]) is a direction to explore.</p> <p>Another noteworthy finding is that—while we explored a range of ML techniques in our analysis—the model that performed best (i.e., boosting) was relatively simple and transparent compared to the complex and opaque techniques which are increasingly used (e.g., Ahmad Uzir et al., [<reflink idref="bib108" id="ref173">108</reflink>], see also Molenaar et al., [<reflink idref="bib80" id="ref174">80</reflink>]). In this way, the present study provided important support for the feasibility of measuring and analyzing CAM processes in real classrooms to support learning—without the need for substantial investments into digital technologies to collect, synchronize, and analyze process data, or the data privacy issues arising from collecting, e.g., video, sensor, or interview data. Further, even when data of different granularity and data streams are integrated, simple and transparent ML techniques can be successful. This contrasts with other studies such as Chan et al. ([<reflink idref="bib16" id="ref175">16</reflink>]) that use intransparent deep learning methods to arrive at similar results. Overall, the finding of relatively simple and transparent methods perform better than more complex and intransparent methods mirrors similar findings from the field of student modeling (Wilson et al., [<reflink idref="bib110" id="ref176">110</reflink>]). One prerequisite for achieving accurate and actionable predictions, is employing reliable and valid instruments to measure the different SRL processes. In the present study, we used established short-questionnaires (i.e., for epistemic emotions; Pekrun et al., [<reflink idref="bib88" id="ref177">88</reflink>]) as well as a knowledge test that covered different aspects of science-competence, based on evidence-centered design (Mislevy &amp; Haertel, [<reflink idref="bib76" id="ref178">76</reflink>]). However, instruments such as questionnaires may be subject to inaccuracies that stem from self-report, such as retention and generalizability problems, or social desirability (Dörrenbacher-Ulrich et al., [<reflink idref="bib28" id="ref179">28</reflink>]; Winne &amp; Perry, [<reflink idref="bib113" id="ref180">113</reflink>]). An alternative to using self-reports is using multimodal detectors, for instance to capture students' emotions during learning (e.g., D'Mello &amp; Kory, [<reflink idref="bib22" id="ref181">22</reflink>]). The benefits of such approaches are their high degree of unobtrusiveness and the opportunity to collect more fine-grained information about different emotions and how they change over time. At the same time, collecting behavioral traces from the learning environment requires a careful mapping of theoretically relevant constructs and patterns in the trace-data (Fan et al., [<reflink idref="bib36" id="ref182">36</reflink>]; Saint et al., [<reflink idref="bib94" id="ref183">94</reflink>]), especially when the goal is to achieve psychometrically valid measurements (for an example see Landers et al., [<reflink idref="bib60" id="ref184">60</reflink>]). This is currently still a challenge for multimodal learning analytics. For the context of learning in collaborative settings, Schneider et al. ([<reflink idref="bib96" id="ref185">96</reflink>]) identified inconsistent results in studies that used sensors to capture learning and collaboration behavior and whether the measurements were significantly related with outcomes. For example, while head-based metrics such as gate were associated with joint attention and learning gains, metrics that were based on physiological sensors were more often than not significant predictors of relevant outcomes. Based on their review, the authors make several suggestions, such as utilizing theory more deliberately to design MMLA measurements (also see Wise &amp; Shaffer, [<reflink idref="bib114" id="ref186">114</reflink>] for an in-depth discussion of this issue).</p> <p>Further, a specific challenge when deploying machine learning approaches that allow leverage trace that were collected unobtrusively is obtaining valid data that represent a "ground truth" that can be used to train a machine learning model (see, e.g., Saint et al., 20202015). This process includes rigorous preparation. However, it needs to be noted that, thus far, research has shown that models do not always generalize to populations or contexts outside their original training data (for the context of affect detection: D'Mello &amp; Kory, [<reflink idref="bib22" id="ref187">22</reflink>]). For example, Plumley et al., ([<reflink idref="bib90" id="ref188">90</reflink>]) observed that the accuracy of their at-risk detection model decreased when applying the model to classify at-risk students in a dataset from a subsequent semester. Thus, acknowledging that different methodological approaches are necessary to capture different components is an important feature of SRL research (see Dörrenbächer-Ulrich et al., [<reflink idref="bib28" id="ref189">28</reflink>]; Molenaar et al., [<reflink idref="bib80" id="ref190">80</reflink>]).</p> <p>The transparency of boosting allowed us to investigate the importance of the different CAM dimensions at different points in time during the instructional unit. We found that throughout the unit there was a shift from affective and metacognitive processes towards cognitive processes being most important for the predictions. While this shift from affective and metacognitive processes towards cognitive processes may be driven by the learning environment, it shows the importance of explainable ML techniques in researching self-regulated learning processes because this information is critical for informing potential instructional interventions. Had we not analyzed variable importance in the boosting models, it would have been challenging to determine which variables are driving predictive accuracy. For future research, this suggests that researchers need to consider a potential trade-off between accuracy and interpretability/explainability. Complex models such as a (deep) neural network may result in improved performance yet limited explainability may reduce the value of such models in informing potential instructional interventions and gains in substantive understanding. At the same time, advances in explainable AI methods and model architectures (e.g., Hollmann et al., [<reflink idref="bib46" id="ref191">46</reflink>]) require us to regularly revisit where the threshold for the accuracy/interpretability trade-off lies. In addition, depending on the use case, one also needs to consider how easy (a few parameters without interactions as in regression models or a shallow decisions tree) or how difficult (many parameters with interactions or requiring explainable AI methods) it is to interpret a model (see also Kay et al., [<reflink idref="bib52" id="ref192">52</reflink>]).</p> <hd id="AN0184416951-31">Research Question 2—Bias</hd> <p>Our findings indicate that the extent of bias exhibited by the machine learning models differed depending on the specific model utilized and was not consistent across the two dimensions under investigation (parental educational background and Gender). For example, for Gender, boosting exhibited the least bias and for parental educational background decision trees exhibited the least bias. Further, even within one model, we found large variations of bias across time, e.g., boosting exhibited very little bias with respect to parental educational background for all timepoints but P3 where bias was huge. In this way, our findings add to positions that warn against the potential of ML—and by extension AI methods more generally—to perpetuate inequalities (Cheuk, [<reflink idref="bib18" id="ref193">18</reflink>]; Crawford, [<reflink idref="bib21" id="ref194">21</reflink>]). At the same time, the potentials that these methods offer to support (research on) self-regulated learning are manifold as (Molenaar et al., [<reflink idref="bib80" id="ref195">80</reflink>]) point out. In this way, the field faces a dilemma that requires skillful navigation. In this light, one may be tempted to argue that the bias we found could be some kind of artifact of the training process. However, given that we specifically made sure to balance our test, training and validation set by gender, the gender bias cannot be dismissed easily. Given the structure of missing data, we were not able to do the same for the parental educational background dimension, meaning that the parental educational background bias could in fact be an artifact.</p> <p>Overall, these findings add to the growing literature on bias in the context of SRL (e.g., Zhang et al., [<reflink idref="bib118" id="ref196">118</reflink>]; Zambrano et al., [<reflink idref="bib117" id="ref197">117</reflink>]). While Zhang et al. ([<reflink idref="bib118" id="ref198">118</reflink>]), for example, found no bias across multiple dimensions, our mixed results suggest that the question to what extent bias is present or not may be highly context dependent and findings from one context or ML model may not easily generalize to another. Further, our study from a German context provides a data-point with respect to the presence of bias that extends beyond the US-centric evidence base that Zambrano et al. ([<reflink idref="bib117" id="ref199">117</reflink>]) lament.</p> <p>From a practical perspective, our study faced the challenge that empirically assessing bias is challenging because it (a) requires to collect very personal data that (b) participants may not be willing to disclose, and (c) can require very large samples as there may only be few instances in the population in the first place. We actually faced especially the latter issue because there were only a handful of students in our sample who indicated that German was not their first language—too few for any meaningful quantitative comparison (see also Zambrano et al., [<reflink idref="bib117" id="ref200">117</reflink>]). With regard to quantitative comparisons, one might be inclined to posit that the observed discrepancies, for instance in the context of the decision tree model as illustrated in Fig. 9, while objectively present, may be practically inconsequential due to their relatively modest magnitude. Here, Narayanan ([<reflink idref="bib82" id="ref201">82</reflink>]) convincingly contends that small biases can accumulate to big effects over time. The fields of learning analytics and artificial intelligence in education have acknowledged bias regarding diversity criteria as challenges and thus appeal to researchers to take into account bias in their studies (Alwahaby et al., [<reflink idref="bib2" id="ref202">2</reflink>]; Baker &amp; Hawn, [<reflink idref="bib8" id="ref203">8</reflink>]). For example, Idowu ([<reflink idref="bib48" id="ref204">48</reflink>]) describes suggestions for how to handle bias (i.e., de-bias) algorithms that are applied to educational data.</p> <p>What does the presence of bias mean for using ML to further study and support self-regulated learning? We argue that future research should ask how bias can be mitigated when measuring and analyzing CAMM processes. When we aim to actually support self-regulated learning using ML, bias will need to become a front and center research topic—not least for legal reasons (European Commission, [<reflink idref="bib19" id="ref205">19</reflink>])—but for ethical ones.</p> <hd id="AN0184416951-32">Implications for Practice and Research</hd> <p>What do our findings mean for practice and research? To answer this question we draw on the two paradigms model by Breiman ([<reflink idref="bib12" id="ref206">12</reflink>]). Breimann distinguishes between two modeling paradigms: prediction and explanation. Predictive models prioritize accuracy, often leveraging machine learning techniques that optimize performance but lack transparency. Explanatory models, in contrast, emphasize interpretability and alignment with theoretical frameworks but may sacrifice predictive power. From a <emph>practical perspective</emph>, our study demonstrates that integrating CAM data enables accurate predictions of students' learning productivity. In the context of an early warning system, where the primary goal is timely intervention rather than deep theoretical insight, high predictive accuracy is certainly useful. Knowing which students are in danger of not meeting learning goals can help a teacher to allocate resources (e.g., Holstein et al., [<reflink idref="bib47" id="ref207">47</reflink>]), provide feedback (e.g., Karademir et al., [<reflink idref="bib51" id="ref208">51</reflink>]), or implement interventions to support students' SRL (e.g. Dever et al., [<reflink idref="bib26" id="ref209">26</reflink>]). In a first step, knowledge about the interplay between different processes during learning, combined with person-level variable importance information, allows teachers to draw conclusions why a student may be in danger of missing learning goals such as lacking motivation due to low interest (although they would have to treat this information with a grain of salt as the underlying model is not designed and thus does not allow for causal inference (see, e.g., Schölkopf, [<reflink idref="bib97" id="ref210">97</reflink>]; Pearl &amp; Mackenzie, [<reflink idref="bib85" id="ref211">85</reflink>]). In a second step, a teacher can derive effective pedagogical interventions that target the students' current trajectory and needs. Thus, a well-supported theory of how different processes affect learning is still required in order to derive effective pedagogical interventions, even if predictive systems function with a high accuracy (Giannakos &amp; Cukurova, [<reflink idref="bib41" id="ref212">41</reflink>]; Wise &amp; Shaffer, [<reflink idref="bib114" id="ref213">114</reflink>]). In the context of an early warning system, the teacher functions as an intermediary between the AI-driven early warning system and the student would naturally serves as the human-in-the-loop (e.g., Li, [<reflink idref="bib62" id="ref214">62</reflink>]), mitigating risks of wrongful predictions and unreliable explanations by the early warning system (e.g., due to bias as was the case across a number of models in this study or remaining uncertainty in the model). In other contexts, however, the value and appropriateness of high accuracy from opaque models may be severely limited. Consider for example a teacher that wants to understand <emph>why</emph> their students are not meeting their learning goals or high stakes automated assessment where a final score is presented but the underlying reasons remain unclear. Here, a theory aligned, transparent model is paramount.</p> <p>Similarly, transparency and theory alignment are also paramount from a <emph>research perspective</emph>. Predictive models help identify (combinations of) variables that are associated with a given criterion (e.g., learning performance), while leaving open how these variables affect each other and the outcome. Thus, while predictive models can highlight potentially relevant variables and provide insights into the comparative strength of different variables, they offer only limited insights into the mechanisms underlying learning and self-regulation. Against the background of previous studies in this area our findings are not surprising in the way that the role of affect for learning has already been established (Camacho-Morles et al., [<reflink idref="bib14" id="ref215">14</reflink>]; Efklides et al., [<reflink idref="bib33" id="ref216">33</reflink>]; Mega et al., [<reflink idref="bib73" id="ref217">73</reflink>]). However, our results revealed a temporal shift in predictive relevance: Affective and metacognitive factors dominate early predictions, suggesting that students' affect and their metacognitive judgements allow to predict their learning performance with high accuracy, while cognitive indicators gain importance throughout the course of the instructional unit. Future research should explore whether this pattern holds for different students and across different learning contexts, and which affects during learning may prepare or afford sustainable knowledge integration. In this case, more explainable models (e.g., explainable boosting machines, Lou et al., [<reflink idref="bib66" id="ref218">66</reflink>]) may shed light on the question which degrees of the different emotions are particularly conducive to learning. This, in turn, can refine theoretical accounts that model the relationship between cognition, metacognition, and affect (e.g., Mega et al., [<reflink idref="bib73" id="ref219">73</reflink>]).</p> <p>Ultimately, our study underscores the balance between predictive accuracy and interpretability. While machine learning enhances practical applications like early warning systems, further research is needed to integrate these models with existing SRL theories and ensure their fairness and transparency. Taken together, we argue that (<reflink idref="bib1" id="ref220">1</reflink>) the usefulness and appropriateness of predictive vs. explanatory modelling is highly context dependent and (<reflink idref="bib2" id="ref221">2</reflink>) that predictive and explanatory modelling can play complementary roles where findings from predictive modelling can drive new research that uses explanatory approaches to test hypothesis derived from predictive approaches.</p> <hd id="AN0184416951-33">Conclusion</hd> <p>This study highlights the potential of integrating cognitive, affective, and metacognitive data to support self-regulated learning in digitally enhanced, inquiry-based science classrooms. By comparing the predictions of different machine learning models across different combinations of cognitive, affective, and metacognitive variables, we approached the goal of developing an early warning system. We demonstrated the feasibility of predicting students' learning trajectories with significant accuracy, enabling timely and personalized instructional support. Our findings show that combining data from cognitive, affective, and metacognitive domains enhances predictive accuracy. Early in the instructional unit, affective and metacognitive variables are more predictive, with cognitive variables gaining importance as the unit progresses. This dynamic underscores the need for integrated analysis approaches (Molenaar et al., [<reflink idref="bib80" id="ref222">80</reflink>]) to capture the full spectrum of learning processes. Additionally, we found that relatively simple and transparent ML techniques, like boosting, can effectively predict learning trajectories in real classroom settings, contrasting with the work that tends towards collecting multi-modal data through extensive technology use in controlled settings and black-box models. What remains an open question, is the extent to which integrating data on motivation that captures facets distinct from what we covered with the affective variables may further improve predictive accuracy.</p> <p>Despite the success of our ML models, we observed that bias varied across different models and demographic groups, highlighting the need for continuous scrutiny and improvement to ensure equitable educational support. Addressing bias is crucial for both ethical and practical reasons, as it maximizes the effectiveness of ML-driven interventions in diverse classroom settings. Lastly, we call for careful consideration of whether predictive or explanatory modelling (or a combination of both) is suitable for the task at hand and not fall prey to hypes around methodologies.</p> <hd id="AN0184416951-34">Funding</hd> <p>Open Access funding enabled and organized by Projekt DEAL.</p> <hd id="AN0184416951-35">Data Availability</hd> <p>Data is available here: https://doi.org/10.17605/OSF.IO/UV8TN https://osf.io/uv8tn/.</p> <hd id="AN0184416951-36">Declarations</hd> <p></p> <hd id="AN0184416951-37">Competing Interests</hd> <p>The authors declare no competing interests.</p> <hd id="AN0184416951-38">Supplementary Information</hd> <p>Below is the link to the electronic supplementary material.</p> <p>Graph: Supplementary file1 (ZIP 89.7 KB)</p> <hd id="AN0184416951-39">Publisher's Note</hd> <p>Springer Nature remains neutral with regard to jurisdictional claims in published maps and institutional affiliations.</p> <ref id="AN0184416951-40"> <title> References </title> <blist> <bibl id="bib1" idref="ref8" type="bt">1</bibl> <bibtext> Ainley M. Connecting with learning: Motivation, affect and cognition in interest processes. Educational Psychology Review. 2006; 18: 391-405. 10.1007/s10648-006-9033-0</bibtext> </blist> <blist> <bibl id="bib2" idref="ref87" type="bt">2</bibl> <bibtext> Alwahaby, H, Cukurova, M, Papamitsiou, Z, spsampsps Giannakos, M. (2022). The evidence of impact and ethical considerations of multimodal learning analytics: a systematic literature review. In M. Giannakos, D. Spikol, D. Di Mitri, K. Sharma, X. Ochoa, spsampsps R. Hammad (Eds.), The Multimodal Learning Analytics Handbook (pp. 289–325). Springer International Publishing. https://doi.org/10.1007/978-3-031-08076-0_12</bibtext> </blist> <blist> <bibl id="bib3" idref="ref71" type="bt">3</bibl> <bibtext> Arnold, K. E, &amp; Pistilli, M. D. (2012). Course signals at Purdue: Using learning analytics to increase student success. Association for Computing Machinery. https://doi.org/10.1145/2330601.2330666</bibtext> </blist> <blist> <bibl id="bib4" idref="ref34" type="bt">4</bibl> <bibtext> Arnold J, Kremer K, Mayer J. Scaffolding beim Forschenden Lernen: Eine empirische Untersuchung zur Wirkung von Lernunterstützungen. Zeitschrift Für Didaktik der Naturwissenschaften. 2017; 23; 1: 21-37. 10.1007/s40573-016-0053-0</bibtext> </blist> <blist> <bibl id="bib5" idref="ref137" type="bt">5</bibl> <bibtext> Avraamidou L. Identities in/out of physics and the politics of recognition. Journal of Research in Science Teaching. 2022; 59; 1: 58-94. 10.1002/tea.21721</bibtext> </blist> <blist> <bibl id="bib6" idref="ref21" type="bt">6</bibl> <bibtext> Azevedo R, Taub M, Mudrick NVSchunk DH, Greene JA. Understanding and reasoning about real-time cognitive, affective, and metacognitive processes to foster self-regulation with advanced learning technologies. Handbook of self-regulation of learning and performance. 2017; Routledge: 254-270. 10.4324/9781315697048-17</bibtext> </blist> <blist> <bibl id="bib7" idref="ref5" type="bt">7</bibl> <bibtext> Azevedo R, Bouchet F, Duffy M, Harley J, Taub M, Trevors G, Cloude E, Dever D, Wiedbusch M, Wortha F, Cerezo R. Lessons learned and future directions of MetaTutor: Leveraging multichannel data to scaffold self-regulated learning with an intelligent tutoring system. Frontiers in Psychology. 2022; 13: 813632. 10.3389/fpsyg.2022.813632</bibtext> </blist> <blist> <bibl id="bib8" idref="ref86" type="bt">8</bibl> <bibtext> Baker RS, Hawn A. Algorithmic bias in education. International Journal of Artificial Intelligence in Education. 2022; 32; 4: 1052-1092. 10.1007/s40593-021-00285-9</bibtext> </blist> <blist> <bibl id="bib9" idref="ref80" type="bt">9</bibl> <bibtext> Baker, R, spsampsps Siemens, G. (2014). Educational data mining and learning analytics. In R. K. Sawyer (Ed.), The Cambridge Handbook of the Learning Sciences (2nd ed, pp. 253–272). Cambridge University Press. https://doi.org/10.1017/CBO9781139519526.016</bibtext> </blist> <blist> <bibtext> Bernacki, M. L. (2017). Examining the cyclical, loosely sequenced, and contingent features of self-regulated learning: Trace data and their analysis. In D. H. Schunk &amp; J. A. Greene (Eds.), Handbook of Self-Regulation of Learning and Performance (pp. 370–387). Routledge.</bibtext> </blist> <blist> <bibtext> Betancur L, Votruba-Drzal E, Schunn C. Socioeconomic gaps in science achievement. International Journal of STEM Education. 2018; 5; 1: 38. 10.1186/s40594-018-0132-5</bibtext> </blist> <blist> <bibtext> Breiman L. Random Forests. Machine Learning. 2001; 45; 1: 5-32. 10.1023/A:1010933404324</bibtext> </blist> <blist> <bibtext> Breiman L. Statistical modeling: The two cultures (with comments and a rejoinder by the author). Statistical Science. 2001; 16; 3: 199-231. 10.1214/ss/1009213726</bibtext> </blist> <blist> <bibtext> Camacho-Morles J, Slemp GR, Pekrun R, Loderer K, Hou H, Oades LG. Activity achievement emotions and academic performance: A meta-analysis. Educational Psychology Review. 2021; 33; 3: 1051-1095. 10.1007/s10648-020-09585-3</bibtext> </blist> <blist> <bibtext> Cerratto Pargman, T, McGrath, C, Viberg, O, Kitto, K, Knight, S, &amp; Ferguson, R. (2021). Responsible learning analytics: creating just, ethical, and caring LA systems. Companion Proceedings. LAK21.</bibtext> </blist> <blist> <bibtext> Chan RY-Y, Wong CMV, Yum YN. Predicting behavior change in students with special education needs using multimodal learning analytics. IEEE Access. 2023; 11: 63238-63251. 10.1109/ACCESS.2023.3288695</bibtext> </blist> <blist> <bibtext> Chen, T, &amp; Guestrin, C. (2016). XGBoost: a scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 785–794. https://doi.org/10.1145/2939672.2939785</bibtext> </blist> <blist> <bibtext> Cheuk, T. (2021). Can AI be racist? Color‐evasiveness in the application of machine learning to science assessments. Science Education, sce.21671. https://doi.org/10.1002/sce.21671</bibtext> </blist> <blist> <bibtext> European Commission. (2024). Proposal for a regulation laying down harmonised rules on artificial intelligence (Artificial Intelligence Act) and amending certain union legislative acts. https://<ulink href="http://www.eur-lex.europa.eu">www.eur-lex.europa.eu</ulink></bibtext> </blist> <blist> <bibtext> Corbett, A. T, Koedinger, K. R, spsampsps Anderson, J. R. (1997). Intelligent tutoring systems. In Handbook of human-computer interaction (pp. 849–874). Elsevier.</bibtext> </blist> <blist> <bibtext> Crawford K. Atlas of AI: Power, politics, and the planetary costs of artificial intelligence. 2021; Yale University Press. 10.2307/j.ctv1ghv45t</bibtext> </blist> <blist> <bibtext> D'Mello, S. K, &amp; Kory, J. (2015). A review and meta-analysis of multimodal affect detection systems. ACM Computing Surveys, 47(3), 1–36. https://doi.org/10.1145/2682899</bibtext> </blist> <blist> <bibtext> De Jong, T, spsampsps Lazonder, A. W. (2014). The guided discovery learning principle in multimedia learning. In R. E. Mayer (Ed.), The Cambridge Handbook of Multimedia Learning (2nd ed, pp. 371–390). Cambridge University Press. https://doi.org/10.1017/CBO9781139547369.019</bibtext> </blist> <blist> <bibtext> De Jong, T, spsampsps Njoo, M. (1992). Learning and instruction with computer simulations: learning processes involved. In E. De Corte, M. C. Linn, H. Mandl, spsampsps L. Verschaffel (Eds.), Computer-Based Learning Environments and Problem Solving (pp. 411–427). Springer Berlin Heidelberg. https://doi.org/10.1007/978-3-642-77228-3_19</bibtext> </blist> <blist> <bibtext> Dent AL, Koenka AC. The relation between self-regulated learning and academic achievement across childhood and adolescence: A meta-analysis. Educational Psychology Review. 2016; 28; 3: 425-474. 10.1007/s10648-015-9320-8</bibtext> </blist> <blist> <bibtext> Dever, D.A, Sonnenfeld, N.A, Wiedbusch, M.D, Azevedo, R. (2022). Pedagogical agent support and its relationship to learners' self-regulated learning strategy use with an intelligent tutoring system. In: Rodrigo, M.M, Matsuda, N, Cristea, A.I, Dimitrova, V. (eds) Artificial Intelligence in Education. AIED 2022. Lecture Notes in Computer Science, vol 13355. Springer, Cham. https://doi.org/10.1007/978-3-031-11644-5_27</bibtext> </blist> <blist> <bibtext> Dindar M, Malmberg J, Järvelä S, Haataja E, Kirschner PA. Matching self-reports with electrodermal activity data: Investigating temporal changes in self-regulated learning. Education and Information Technologies. 2020; 25; 3: 1785-1802. 10.1007/s10639-019-10059-5</bibtext> </blist> <blist> <bibtext> Dörrenbächer-Ulrich L, Weißenfels M, Russer L, Perels F. Multimethod assessment of self-regulated learning in college students: Different methods for different components?. Instructional Science. 2021; 49; 1: 137-163. 10.1007/s11251-020-09533-2</bibtext> </blist> <blist> <bibtext> Dougherty MR, Robey AM, Buttaccio D. Do metacognitive judgments alter memory performance beyond the benefits of retrieval practice? A comment on and replication attempt of Dougherty, Scheck, Nelson, and Narens (2005). Memory &amp; Cognition. 2018; 46; 4: 558-565. 10.3758/s13421-018-0791-y</bibtext> </blist> <blist> <bibtext> Dougiamas, M, &amp; Taylor, P. (2003). Moodle: using learning communities to create an open source course management system. In D. Lassner &amp; C. McNaught (Eds.), Proceedings of EdMedia + Innovate Learning 2003 (pp. 171–178). Association for the Advancement of Computing in Education (AACE). https://<ulink href="http://www.learntechlib.org/p/13739">www.learntechlib.org/p/13739</ulink></bibtext> </blist> <blist> <bibtext> Dupuy, D, Helbert, C, &amp; Franco, J. (2015). DiceDesign and DiceEval: Two R Packages for Design and Analysis of Computer Experiments. Journal of Statistical Software, 65(11). https://doi.org/10.18637/jss.v065.i11</bibtext> </blist> <blist> <bibtext> Efklides A, Petkaki C. Effects of mood on students' metacognitive experiences. Learning and Instruction. 2005; 15; 5: 415-431. 10.1016/j.learninstruc.2005.07.010</bibtext> </blist> <blist> <bibtext> Efklides, A, Schwartz, B. L, spsampsps Brown, V. (2017). Motivation and affect in self-regulated learning. In D. H. Schunk spsampsps J. A. Greene (Eds.), Handbook of Self-Regulation of Learning and Performance (2nd ed, pp. 64–82). Routledge. https://doi.org/10.4324/9781315697048-5</bibtext> </blist> <blist> <bibtext> Emerson A, Cloude EB, Azevedo R, Lester J. Multimodal learning analytics for game-based learning. British Journal of Educational Technology. 2020; 51; 5: 1505-1526. 10.1111/bjet.12992</bibtext> </blist> <blist> <bibtext> Engelmann K, Bannert M. Analyzing temporal data for understanding the learning process induced by metacognitive prompts. Learning and Instruction. 2021; 72: 101205. 10.1016/j.learninstruc.2019.05.002</bibtext> </blist> <blist> <bibtext> Fan Y, Van Der Graaf J, Lim L, Raković M, Singh S, Kilgour J, Moore J, Molenaar I, Bannert M, Gašević D. Towards investigating the validity of measurement of self-regulated learning based on trace data. Metacognition and Learning. 2022; 17; 3: 949-987. 10.1007/s11409-022-09291-1</bibtext> </blist> <blist> <bibtext> Fischer, J. (2022). Basiskonzepte konkretisieren – Entwicklung und Evaluation einer Interventionsmaßnahme zur Förderung kumulativen Lernens durch den Einsatz von Basiskonzepten. Christian-Albrechts-Universität zu Kiel.</bibtext> </blist> <blist> <bibtext> Gándara D, Anahideh H, Ison MP, Picchiarini L. Inside the Black Box: Detecting and Mitigating Algorithmic Bias Across Racialized Groups in College Student-Success Prediction. AERA Open. 2024; 10: 23328584241258741. 10.1177/23328584241258741</bibtext> </blist> <blist> <bibtext> Gasevic, D, Jovanovic, J, Pardo, A, &amp; Dawson, S. (2017). Detecting learning strategies with analytics: links with self-reported measures and academic performance. Journal of Learning Analytics, 4(2). https://doi.org/10.18608/jla.2017.42.10</bibtext> </blist> <blist> <bibtext> Ghojogh, B, &amp; Crowley, M. (2023). The theory behind overfitting, cross validation, regularization, bagging, and boosting: tutorial (arXiv:1905.12787). arXiv. <ulink href="http://arxiv.org/abs/1905.12787">http://arxiv.org/abs/1905.12787</ulink></bibtext> </blist> <blist> <bibtext> Giannakos M, Cukurova M. The role of learning theory in multimodal learning analytics. British Journal of Educational Technology. 2023; 54; 5: 1246-1267. 10.1111/bjet.13320</bibtext> </blist> <blist> <bibtext> Grawemeyer B, Mavrikis M, Holmes W, Gutiérrez-Santos S, Wiedmann M, Rummel N. Affective learning: Improving engagement and enhancing learning with affect-aware feedback. User Modeling and User-Adapted Interaction. 2017; 27; 1: 119-158. 10.1007/s11257-017-9188-z</bibtext> </blist> <blist> <bibtext> Grinsztajn, L, Oyallon, E, &amp; Varoquaux, G. (2022). Why do tree-based models still outperform deep learning on tabular data? (arXiv:2207.08815). arXiv. <ulink href="http://arxiv.org/abs/2207.08815">http://arxiv.org/abs/2207.08815</ulink></bibtext> </blist> <blist> <bibtext> Hastie, T, Tibshirani, R, &amp; Friedman, J. H. (2009). The elements of statistical learning: Data mining, inference, and prediction (2nd ed). Springer.</bibtext> </blist> <blist> <bibtext> Hilbert, S, Coors, S, Kraus, E, Bischl, B, Lindl, A, Frei, M, Wild, J, Krauss, S, Goretzko, D, &amp; Stachl, C. (2021). Machine learning for the educational sciences. Review of Education, 9(3). https://doi.org/10.1002/rev3.3310</bibtext> </blist> <blist> <bibtext> Hollmann N, Müller S, Purucker L, Krishnakumar A, Körfer M, Hoo SB, Schirrmeister RT, Hutter F. Accurate predictions on small data with a tabular foundation model. Nature. 2025; 637; 8045: 319-326. 10.1038/s41586-024-08328-6</bibtext> </blist> <blist> <bibtext> Holstein, K, McLaren, B. M, &amp; Aleven, V. (2019). Co-designing a real-time classroom orchestration tool to support teacher–AI complementarity. Journal of Learning Analytics, 6(2), 27–52. https://doi.org/10.18608/jla.2019.62.3</bibtext> </blist> <blist> <bibtext> Idowu JA. Debiasing Education Algorithms. International Journal of Artificial Intelligence in Education. 2024. 10.1007/s40593-023-00389-4</bibtext> </blist> <blist> <bibtext> James, G, Witten, D, Hastie, T, spsampsps Tibshirani, R. (2013). An introduction to statistical learning (Vol. 103). Springer New York. https://doi.org/10.1007/978-1-4614-7138-7</bibtext> </blist> <blist> <bibtext> Jayaprakash, S. M, Moody, E. W, Lauría, E. J. M, Regan, J. R, &amp; Baron, J. D. (2014). Early alert of academically at-risk students: an open source analytics initiative. Journal of Learning Analytics, 1(1), 6–47. https://doi.org/10.18608/jla.2014.11.3</bibtext> </blist> <blist> <bibtext> Karademir, O, Borgards, L, Di Mitri, D, Strauß, S, Kubsch, M, Brobeil, M, Grimm, A, Gombert, S, Rummel, N, Neumann, K, &amp; Drachsler, H. (2024). Following the impact chain of the LA cockpit: an intervention study investigating a teacher dashboard's effect on student learning. Journal of Learning Analytics, 11(2), 215–228. https://doi.org/10.18608/jla.2024.8399</bibtext> </blist> <blist> <bibtext> Kay, J, Kummerfeld, B, Conati, C, Porayska-Pomsta, K, spsampsps Holstein, K. (2023). Scrutable AIED. In B. Du Boulay, A. Mitrovic, spsampsps K. Yacef (Eds.), Handbook of Artificial Intelligence in Education (pp. 101–125). Edward Elgar Publishing. https://doi.org/10.4337/9781800375413.00015</bibtext> </blist> <blist> <bibtext> Kitto K, Knight S. Practical ethics for building learning analytics. British Journal of Educational Technology. 2019; 50; 6: 2855-2870. 10.1111/bjet.12868</bibtext> </blist> <blist> <bibtext> Kizilcec, R. F, spsampsps Lee, H. (2022). Algorithmic fairness in education. In W. Holmes spsampsps K. Porayska-Pomsta, The Ethics of Artificial Intelligence in Education (1st ed, pp. 174–202). Routledge. https://doi.org/10.4324/9780429329067-10</bibtext> </blist> <blist> <bibtext> Krajcik, J. S, spsampsps Shin, N. (2014). Project-Based Learning. In R. K. Sawyer (Ed.), The Cambridge handbook of the learning sciences (Second edition, pp. 275–297). Cambridge University Press. https://doi.org/10.1017/CBO9781139519526.018</bibtext> </blist> <blist> <bibtext> Krist, C, &amp; Kubsch, M. (2023). Bias, bias everywhere: a response to Li et al. and Zhai and Nehm. Journal of Research in Science Teaching, tea.21913. https://doi.org/10.1002/tea.21913</bibtext> </blist> <blist> <bibtext> Kuhn M, Johnson K. Applied predictive modeling. Springer, New York. 2013. 10.1007/978-1-4614-6849-3</bibtext> </blist> <blist> <bibtext> Kuhn, M, &amp; Wickham, H. (2020). Tidymodels: A collection of packages for modeling and machine learning using tidyverse principles.https://<ulink href="http://www.tidymodels.org">www.tidymodels.org</ulink></bibtext> </blist> <blist> <bibtext> Lai C-L, Hwang G-J, Tu Y-H. The effects of computer-supported self-regulation in science inquiry on learning outcomes, learning processes, and self-efficacy. Educational Technology Research and Development. 2018; 66; 4: 863-892. 10.1007/s11423-018-9585-y</bibtext> </blist> <blist> <bibtext> Landers RN, Auer EM, Mersy G, Marin S, Blaik J. You are what you click: Using machine learning to model trace data for psychometric measurement. International Journal of Testing. 2022; 22; 3–4: 243-263. 10.1080/15305058.2022.2134394</bibtext> </blist> <blist> <bibtext> LeCun Y, Bengio Y, Hinton G. Deep learning. Nature. 2015; 521; 7553: 436-444. 10.1038/nature14539</bibtext> </blist> <blist> <bibtext> Li, T. (2024). Cognitive synergy: exploring the transformative intersection of human intelligence and artificial intelligence in designing equitable next generation science assessments (Doctoral dissertation, Michigan State University).</bibtext> </blist> <blist> <bibtext> Lim L, Bannert M, Van Der Graaf J, Singh S, Fan Y, Surendrannair S, Rakovic M, Molenaar I, Moore J, Gašević D. Effects of real-time analytics-based personalized scaffolds on students' self-regulated learning. Computers in Human Behavior. 2023; 139: 107547. 10.1016/j.chb.2022.107547</bibtext> </blist> <blist> <bibtext> Linn MC, Clark D, Slotta JD. WISE design for knowledge integration. Science Education. 2003; 87; 4: 517-538. 10.1002/sce.10086</bibtext> </blist> <blist> <bibtext> Linn, M. C, Eylon, B.-S, Rafferty, A, &amp; Vitale, J. M. (2017). Designing instruction to improve lifelong inquiry learning. EURASIA Journal of Mathematics, Science and Technology Education, 11(2). https://doi.org/10.12973/eurasia.2015.1317a</bibtext> </blist> <blist> <bibtext> Lou, Y, Caruana, R, Gehrke, J, &amp; Hooker, G. (2013). Accurate intelligible models with pairwise interactions. Proceedings of the 19th ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, 623–631. https://doi.org/10.1145/2487575.2487579</bibtext> </blist> <blist> <bibtext> Macfadyen LP, Dawson S. Mining LMS data to develop an "early warning system" for educators: A proof of concept. Computers &amp; Education. 2010; 54; 2: 588-599. 10.1016/j.compedu.2009.09.008</bibtext> </blist> <blist> <bibtext> Malmberg J, Fincham O, Pijeira-Díaz HJ, Järvelä S, Gašević D. Revealing the hidden structure of physiological states during metacognitive monitoring in collaborative learning. Journal of Computer Assisted Learning. 2021; 37; 3: 861-874. 10.1111/jcal.12529</bibtext> </blist> <blist> <bibtext> Mangaroska K, Sharma K, Gašević D, Giannakos M. Exploring students' cognitive and affective states during problem solving through multimodal data: Lessons learned from a programming activity. Journal of Computer Assisted Learning. 2022; 38; 1: 40-59. 10.1111/jcal.12590</bibtext> </blist> <blist> <bibtext> Matcha, W, Gašević, D, Ahmad Uzir, N, Jovanović, J, Pardo, A, Lim, L, Maldonado-Mahauad, J, Gentili, S, Pérez-Sanagustín, M, &amp; Tsai, Y.-S. (2020). Analytics of learning strategies: role of course design and delivery modality. Journal of Learning Analytics, 7(2), 45–71. https://doi.org/10.18608/jla.2020.72.3</bibtext> </blist> <blist> <bibtext> McAuley E, Duncan T, Tammen VV. Psychometric properties of the intrinsic motivation inventory in a competitive sport setting: A confirmatory factor analysis. Research Quarterly for Exercise and Sport. 1989; 60; 1: 48-58. 10.1080/02701367.1989.10607413</bibtext> </blist> <blist> <bibtext> McDowell LD. The roles of motivation and metacognition in producing self-regulated learners of college physical science: A review of empirical studies. International Journal of Science Education. 2019; 41; 17: 2524-2541. 10.1080/09500693.2019.1689584</bibtext> </blist> <blist> <bibtext> Mega C, Ronconi L, De Beni R. What makes a good student? How emotions, self-regulated learning, and motivation contribute to academic achievement. Journal of Educational Psychology. 2014; 106; 1: 121-131. 10.1037/a0033546</bibtext> </blist> <blist> <bibtext> Minh D, Wang HX, Li YF, Nguyen TN. Explainable artificial intelligence: A comprehensive review. Artificial Intelligence Review. 2022; 55; 5: 3503-3568. 10.1007/s10462-021-10088-y</bibtext> </blist> <blist> <bibtext> Mislevy RJ, Almond RG, Lukas JF. A brief introduction to evidence-centered design. ETS Research Report Series. 2003; 2003; 1: i-29. 10.1002/j.2333-8504.2003.tb01908.x</bibtext> </blist> <blist> <bibtext> Mislevy, R. J, &amp; Haertel, G. D. (2007). Implications of evidence-centered design for educational testing. Educational Measurement: Issues and Practice, 25(4), 6–20. https://doi.org/10.1111/j.1745-3992.2006.00075.x</bibtext> </blist> <blist> <bibtext> Mitchell S, Potash E, Barocas S, D'Amour A, Lum K. Algorithmic fairness: Choices, assumptions, and definitions. Annual Review of Statistics and Its Application. 2021; 8; 1: 141-163. 10.1146/annurev-statistics-042720-125902</bibtext> </blist> <blist> <bibtext> Miyake A, Kost-Smith LE, Finkelstein ND, Pollock SJ, Cohen GL, Ito TA. Reducing the gender achievement gap in college science: A classroom study of values affirmation. Science. 2010; 330; 6008: 1234-1237. 10.1126/science.1195996</bibtext> </blist> <blist> <bibtext> Molenaar, I, spsampsps Knoop-van Campen, C. (2017). Teacher Dashboards in Practice: Usage and Impact. In É. Lavoué, H. Drachsler, K. Verbert, J. Broisin, spsampsps M. Pérez-Sanagustín (Eds.), Data Driven Approaches in Digital Education (Vol. 10474, pp. 125–138). Springer International Publishing. https://doi.org/10.1007/978-3-319-66610-5_10</bibtext> </blist> <blist> <bibtext> Molenaar I, Mooij SD, Azevedo R, Bannert M, Järvelä S, Gašević D. Measuring self-regulated learning and the role of AI: Five years of research using multimodal multichannel data. Computers in Human Behavior. 2023; 139: 107540. 10.1016/j.chb.2022.107540</bibtext> </blist> <blist> <bibtext> Molnar, C. (2022). Interpretable machine learning: a guide for making black box models explainable (Second edition). Christoph Molnar.</bibtext> </blist> <blist> <bibtext> Narayanan, A. (2022). The limits of the quantitative approach to discrimination.</bibtext> </blist> <blist> <bibtext> OECD. PISA 2022 results (volume I): The state of learning and equity in education. OECD. 2023. 10.1787/53f23881-en</bibtext> </blist> <blist> <bibtext> Pardos, Z. A, Baker, R. S. J. D, San Pedro, M. O. C. Z, Gowda, S. M, &amp; Gowda, S. M. (2013). Affective states and state tests: investigating how affect throughout the school year predicts end of year learning outcomes. Proceedings of the Third International Conference on Learning Analytics and Knowledge, 117–124. https://doi.org/10.1145/2460296.2460320</bibtext> </blist> <blist> <bibtext> Pearl, J, &amp; Mackenzie, D. (2018). The book of why: The new science of cause and effect (First edition). Basic Books.</bibtext> </blist> <blist> <bibtext> Pedaste M, Mäeots M, Siiman LA, de Jong T, van Riesen SAN, Kamp ET, Manoli CC, Zacharia ZC, Tsourlidaki E. Phases of inquiry-based learning: Definitions and the inquiry cycle. Educational Research Review. 2015; 14: 47-61. 10.1016/j.edurev.2015.02.003</bibtext> </blist> <blist> <bibtext> Pekrun R. The control-value theory of achievement emotions: Assumptions, corollaries, and implications for educational research and practice. Educational Psychology Review. 2006; 18; 4: 315-341. 10.1007/s10648-006-9029-9</bibtext> </blist> <blist> <bibtext> Pekrun R, Vogl E, Muis KR, Sinatra GM. Measuring emotions during epistemic activities: The Epistemically-Related Emotion Scales. Cognition and Emotion. 2017; 31; 6: 1268-1276. 10.1080/02699931.2016.1204989</bibtext> </blist> <blist> <bibtext> Pellegrino JW, DiBello LV, Goldman SR. A framework for conceptualizing and evaluating the validity of instructionally relevant assessments. Educational Psychologist. 2015; 51; 1: 59-81. 10.1080/00461520.2016.1145550</bibtext> </blist> <blist> <bibtext> Plumley RD, Bernacki ML, Greene JA, Kuhlmann S, Raković M, Urban CJ, Hogan KA, Lee C, Panter AT, Gates KM. Co-designing enduring learning analytics prediction and support tools in undergraduate biology courses. British Journal of Educational Technology. 2024; 55; 5: 1860-1883. 10.1111/bjet.13472</bibtext> </blist> <blist> <bibtext> R Development Core Team. (2008). R: A Language and Environment for Statistical Computing. R Foundation for Statistical Computing. <ulink href="http://www.R-project.org">http://www.R-project.org</ulink></bibtext> </blist> <blist> <bibtext> Rakovic, M, Fan, Y, Van Der Graaf, J, Singh, S, Kilgour, J, Lim, L, Moore, J, Bannert, M, Molenaar, I, &amp; Gasevic, D. (2022). Using Learner Trace Data to Understand Metacognitive Processes in Writing from Multiple Sources. LAK22: 12th International Learning Analytics and Knowledge Conference, 130–141. https://doi.org/10.1145/3506860.3506876</bibtext> </blist> <blist> <bibtext> Saint J, Fan Y, Gašević D, Pardo A. Temporally-focused analytics of self-regulated learning: A systematic review of literature. Computers and Education: Artificial Intelligence. 2022; 3: 100060. 10.1016/j.caeai.2022.100060</bibtext> </blist> <blist> <bibtext> Saint J, Whitelock-Wainwright A, Gasevic D, Pardo A. Trace-SRL: A Framework for Analysis of Microlevel Processes of Self-Regulated Learning From Trace Data. IEEE Transactions on Learning Technologies. 2020; 13; 4: 861-877. 10.1109/TLT.2020.3027496</bibtext> </blist> <blist> <bibtext> Schneider, B, Krajcik, J. S, Lavonen, J, &amp; Salmela-Aro, K. (2020). Learning science: The value of crafting engagement in science environments. Yale University Press.</bibtext> </blist> <blist> <bibtext> Schneider B, Sung G, Chng E, Yang S. How can high-frequency sensors capture collaboration? A review of the empirical links between multimodal metrics and collaborative constructs. Sensors. 2021; 21; 24: 8185. 10.3390/s21248185</bibtext> </blist> <blist> <bibtext> Schölkopf, B. (2019). Causality for machine learning. arXiv:1911.10500 [Cs, Stat]. <ulink href="http://arxiv.org/abs/1911.10500">http://arxiv.org/abs/1911.10500</ulink></bibtext> </blist> <blist> <bibtext> Schraw G, Crippen KJ, Hartley K. Promoting self-regulation in science education: Metacognition as part of a broader perspective on learning. Research in Science Education. 2006; 36; 1–2: 111-139. 10.1007/s11165-005-3917-8</bibtext> </blist> <blist> <bibtext> Simonsmeier BA, Flaig M, Deiglmayr A, Schalk L, Schneider M. Domain-specific prior knowledge and learning: A meta-analysis. Educational Psychologist. 2022; 57; 1: 31-54. 10.1080/00461520.2021.1939700</bibtext> </blist> <blist> <bibtext> Sinatra, G. M, spsampsps Taasoobshirazi, G. (2017). The self-regulation of learning and conceptual change in science. In D. H. Schunk spsampsps J. A. Greene (Eds.), Handbook of Self-Regulation of Learning and Performance (2nd ed, pp. 153–165). Routledge. https://doi.org/10.4324/9781315697048-10</bibtext> </blist> <blist> <bibtext> Suresh, H, &amp; Guttag, J. V. (2021). A framework for understanding sources of harm throughout the machine learning life cycle. Equity and Access in Algorithms, Mechanisms, and Optimization, 1–9. https://doi.org/10.1145/3465416.3483305</bibtext> </blist> <blist> <bibtext> Taub M, Azevedo R. How does prior knowledge influence eye fixations and sequences of cognitive and metacognitive SRL processes during learning with an intelligent tutoring system?. International Journal of Artificial Intelligence in Education. 2019; 29; 1: 1-28. 10.1007/s40593-018-0165-4</bibtext> </blist> <blist> <bibtext> Taub M, Sawyer R, Lester J, Azevedo R. The impact of contextualized emotions on self-regulated learning and scientific reasoning during learning with a game-based learning environment. International Journal of Artificial Intelligence in Education. 2020; 30; 1: 97-120. 10.1007/s40593-019-00191-1</bibtext> </blist> <blist> <bibtext> Tiku, N. (2020, October 9). The dark side of "AI-powered" proctoring. Wired. https://<ulink href="http://www.wired.com/story/student-exam-software-bias-proctorio/">www.wired.com/story/student-exam-software-bias-proctorio/</ulink></bibtext> </blist> <blist> <bibtext> Trevors G, Duffy M, Azevedo R. Note-taking within MetaTutor: Interactions between an intelligent tutoring system and prior knowledge on note-taking and learning. Educational Technology Research and Development. 2014; 62; 5: 507-528. 10.1007/s11423-014-9343-8</bibtext> </blist> <blist> <bibtext> Tuononen T, Hyytinen H, Räisänen M, Hailikari T, Parpala A. Metacognitive awareness in relation to university students' learning profiles. Metacognition and Learning. 2023; 18; 1: 37-54. 10.1007/s11409-022-09314-x</bibtext> </blist> <blist> <bibtext> U.S. Department of Education, Office of Educational Technology. (2023). Artificial Intelligence and Future of Teaching and Learning: Insights and Recommendations.</bibtext> </blist> <blist> <bibtext> Uzir, N, Gašević, D, Matcha, W, Jovanović, J, Pardo, A, Lim, L.-A, spsampsps Gentili, S. (2019). Discovering Time Management Strategies in Learning Processes Using Process Mining Techniques. In M. Scheffel, J. Broisin, V. Pammer-Schindler, A. Ioannou, spsampsps J. Schneider (Eds.), Transforming Learning with Meaningful Technologies (Vol. 11722, pp. 555–569). Springer International Publishing. https://doi.org/10.1007/978-3-030-29736-7_41</bibtext> </blist> <blist> <bibtext> Vilhunen E, Turkkila M, Lavonen J, Salmela-Aro K, Juuti K. Clarifying the relation between epistemic emotions and learning by using experience sampling method and pre-posttest design. Frontiers in Education. 2022; 7: 826852. 10.3389/feduc.2022.826852</bibtext> </blist> <blist> <bibtext> Wilson, K. H, Karklin, Y, Han, B, &amp; Ekanadham, C. (2016). Back to the basics: Bayesian extensions of IRT outperform neural networks for proficiency estimation (arXiv:1604.02336). arXiv. <ulink href="http://arxiv.org/abs/1604.02336">http://arxiv.org/abs/1604.02336</ulink></bibtext> </blist> <blist> <bibtext> Winne PSchunk DH, Greene JA. Cognition and metacognition within self-regulated learning. Handbook of Self-Regulation of Learning and Performance. 2017; Routledge: 36-48. 10.4324/9781315697048-3</bibtext> </blist> <blist> <bibtext> Winne, P. H. (2017b). Learning analytics for self-regulated learning. In Handbook of Learning Analytics (pp. 241–249). https://doi.org/10.18608/hla17.021</bibtext> </blist> <blist> <bibtext> Winne, P. H, spsampsps Perry, N. E. (2000). Measuring self-regulated learning. In Handbook of Self-Regulation (pp. 531–566). Elsevier. https://doi.org/10.1016/B978-012109890-2/50045-7</bibtext> </blist> <blist> <bibtext> Wise, A. F, &amp; Shaffer, D. W. (2015). Why theory matters more than ever in the age of big data. Journal of Learning Analytics, 2(2), 5–13. https://doi.org/10.18608/jla.2015.22.2</bibtext> </blist> <blist> <bibtext> Wößmann L, Schoner F, Freundl V, Pfaehler F. Der ifo-" Ein Herz für Kinder"-Chancenmonitor: Wie (un-) gerecht sind die Bildungschancen von Kindern aus verschiedenen Familien in Deutschland verteilt?. Ifo Schnelldienst. 2023; 76; 04: 29-47</bibtext> </blist> <blist> <bibtext> Xu Z, Zhao Y, Liew J, Zhou X, Kogut A. Synthesizing research evidence on self-regulated learning and academic achievement in online and blended learning environments: A scoping review. Educational Research Review. 2023; 39: 100510. 10.1016/j.edurev.2023.100510</bibtext> </blist> <blist> <bibtext> Zambrano, A. F, Zhang, J, &amp; Baker, R. S. (2024). Investigating algorithmic bias on Bayesian knowledge tracing and carelessness detectors. Proceedings of the 14th Learning Analytics and Knowledge Conference, 349–359. https://doi.org/10.1145/3636555.3636890</bibtext> </blist> <blist> <bibtext> Zhang, J, J. Ma, Andres, R. L, Hutt, S, Baker, R. S, Ocumpaugh, J, Mills, C, Jamiella Brooks, Sethuraman, S, &amp; Young, T. (2022). Detecting SMART Model Cognitive Operations in Mathematical Problem-Solving Process. https://doi.org/10.5281/ZENODO.6853161</bibtext> </blist> <blist> <bibtext> Zheng J, Lajoie S, Li S. Emotions in self-regulated learning: A critical literature review and meta-analysis. Frontiers in Psychology. 2023; 14: 1137010. 10.3389/fpsyg.2023.1137010</bibtext> </blist> <blist> <bibtext> Zimmerman, B. J. (2002). Becoming a self-regulated learner: an overview. Theory Into Practice, 41(2), 64–70. JSTOR.</bibtext> </blist> <blist> <bibtext> Zusho A, Pintrich PR, Coppola B. Skill and will: The role of motivation and cognition in the learning of college chemistry. International Journal of Science Education. 2003; 25; 9: 1081-1094. 10.1080/0950069032000052207</bibtext> </blist> </ref> <ref id="AN0184416951-41"> <title> Footnotes </title> <blist> <bibtext> With "real-life classroom" we refer to classrooms that are typical for instruction in schools, i.e., classrooms where traditional face-to-face settings are the norm.</bibtext> </blist> <blist> <bibtext> Note that while some researchers use the terms bias and fairness interchangeably others have started to use bias with respect to the technical aspect and reserved fairness for the societal aspect. Given that within the context of our study, we do not use any machine learning model in practice, i.e., society, but focus on the research side, we use the term bias and focus on the technical aspect.</bibtext> </blist> <blist> <bibtext> For more details on the development and validation of the units, see Fischer ([37]).</bibtext> </blist> <blist> <bibtext> Note though that preliminary analyses using imputation led to overall very similar results compared to using the approach we opted for in this study.</bibtext> </blist> </ref> <aug> <p>Reported by Author; Author; Author; Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib86" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib98" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib80" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib67" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib94" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib116" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib112" firstref="ref11"></nolink> <nolink nlid="nl8" bibid="bib95" firstref="ref12"></nolink> <nolink nlid="nl9" bibid="bib23" firstref="ref13"></nolink> <nolink nlid="nl10" bibid="bib24" firstref="ref16"></nolink> <nolink nlid="nl11" bibid="bib59" firstref="ref17"></nolink> <nolink nlid="nl12" bibid="bib25" firstref="ref19"></nolink> <nolink nlid="nl13" bibid="bib113" firstref="ref23"></nolink> <nolink nlid="nl14" bibid="bib120" firstref="ref24"></nolink> <nolink nlid="nl15" bibid="bib111" firstref="ref27"></nolink> <nolink nlid="nl16" bibid="bib64" firstref="ref28"></nolink> <nolink nlid="nl17" bibid="bib65" firstref="ref29"></nolink> <nolink nlid="nl18" bibid="bib100" firstref="ref32"></nolink> <nolink nlid="nl19" bibid="bib102" firstref="ref43"></nolink> <nolink nlid="nl20" bibid="bib105" firstref="ref45"></nolink> <nolink nlid="nl21" bibid="bib72" firstref="ref47"></nolink> <nolink nlid="nl22" bibid="bib121" firstref="ref48"></nolink> <nolink nlid="nl23" bibid="bib33" firstref="ref49"></nolink> <nolink nlid="nl24" bibid="bib32" firstref="ref50"></nolink> <nolink nlid="nl25" bibid="bib119" firstref="ref51"></nolink> <nolink nlid="nl26" bibid="bib73" firstref="ref52"></nolink> <nolink nlid="nl27" bibid="bib14" firstref="ref58"></nolink> <nolink nlid="nl28" bibid="bib84" firstref="ref59"></nolink> <nolink nlid="nl29" bibid="bib109" firstref="ref60"></nolink> <nolink nlid="nl30" bibid="bib10" firstref="ref64"></nolink> <nolink nlid="nl31" bibid="bib93" firstref="ref66"></nolink> <nolink nlid="nl32" bibid="bib50" firstref="ref68"></nolink> <nolink nlid="nl33" bibid="bib20" firstref="ref73"></nolink> <nolink nlid="nl34" bibid="bib42" firstref="ref74"></nolink> <nolink nlid="nl35" bibid="bib45" firstref="ref81"></nolink> <nolink nlid="nl36" bibid="bib107" firstref="ref82"></nolink> <nolink nlid="nl37" bibid="bib103" firstref="ref84"></nolink> <nolink nlid="nl38" bibid="bib63" firstref="ref85"></nolink> <nolink nlid="nl39" bibid="bib77" firstref="ref88"></nolink> <nolink nlid="nl40" bibid="bib101" firstref="ref89"></nolink> <nolink nlid="nl41" bibid="bib54" firstref="ref90"></nolink> <nolink nlid="nl42" bibid="bib104" firstref="ref91"></nolink> <nolink nlid="nl43" bibid="bib38" firstref="ref92"></nolink> <nolink nlid="nl44" bibid="bib118" firstref="ref93"></nolink> <nolink nlid="nl45" bibid="bib56" firstref="ref94"></nolink> <nolink nlid="nl46" bibid="bib15" firstref="ref95"></nolink> <nolink nlid="nl47" bibid="bib53" firstref="ref96"></nolink> <nolink nlid="nl48" bibid="bib19" firstref="ref98"></nolink> <nolink nlid="nl49" bibid="bib48" firstref="ref99"></nolink> <nolink nlid="nl50" bibid="bib52" firstref="ref101"></nolink> <nolink nlid="nl51" bibid="bib12" firstref="ref102"></nolink> <nolink nlid="nl52" bibid="bib13" firstref="ref103"></nolink> <nolink nlid="nl53" bibid="bib74" firstref="ref104"></nolink> <nolink nlid="nl54" bibid="bib81" firstref="ref105"></nolink> <nolink nlid="nl55" bibid="bib79" firstref="ref107"></nolink> <nolink nlid="nl56" bibid="bib55" firstref="ref111"></nolink> <nolink nlid="nl57" bibid="bib30" firstref="ref112"></nolink> <nolink nlid="nl58" bibid="bib75" firstref="ref113"></nolink> <nolink nlid="nl59" bibid="bib89" firstref="ref114"></nolink> <nolink nlid="nl60" bibid="bib87" firstref="ref116"></nolink> <nolink nlid="nl61" bibid="bib88" firstref="ref117"></nolink> <nolink nlid="nl62" bibid="bib115" firstref="ref119"></nolink> <nolink nlid="nl63" bibid="bib58" firstref="ref120"></nolink> <nolink nlid="nl64" bibid="bib91" firstref="ref121"></nolink> <nolink nlid="nl65" bibid="bib17" firstref="ref125"></nolink> <nolink nlid="nl66" bibid="bib43" firstref="ref126"></nolink> <nolink nlid="nl67" bibid="bib44" firstref="ref128"></nolink> <nolink nlid="nl68" bibid="bib61" firstref="ref129"></nolink> <nolink nlid="nl69" bibid="bib31" firstref="ref131"></nolink> <nolink nlid="nl70" bibid="bib57" firstref="ref132"></nolink> <nolink nlid="nl71" bibid="bib49" firstref="ref133"></nolink> <nolink nlid="nl72" bibid="bib78" firstref="ref136"></nolink> <nolink nlid="nl73" bibid="bib11" firstref="ref138"></nolink> <nolink nlid="nl74" bibid="bib83" firstref="ref139"></nolink> <nolink nlid="nl75" bibid="bib40" firstref="ref140"></nolink> <nolink nlid="nl76" bibid="bib27" firstref="ref145"></nolink> <nolink nlid="nl77" bibid="bib69" firstref="ref146"></nolink> <nolink nlid="nl78" bibid="bib34" firstref="ref147"></nolink> <nolink nlid="nl79" bibid="bib35" firstref="ref149"></nolink> <nolink nlid="nl80" bibid="bib68" firstref="ref150"></nolink> <nolink nlid="nl81" bibid="bib92" firstref="ref152"></nolink> <nolink nlid="nl82" bibid="bib39" firstref="ref155"></nolink> <nolink nlid="nl83" bibid="bib99" firstref="ref164"></nolink> <nolink nlid="nl84" bibid="bib70" firstref="ref166"></nolink> <nolink nlid="nl85" bibid="bib71" firstref="ref170"></nolink> <nolink nlid="nl86" bibid="bib106" firstref="ref171"></nolink> <nolink nlid="nl87" bibid="bib29" firstref="ref172"></nolink> <nolink nlid="nl88" bibid="bib108" firstref="ref173"></nolink> <nolink nlid="nl89" bibid="bib16" firstref="ref175"></nolink> <nolink nlid="nl90" bibid="bib110" firstref="ref176"></nolink> <nolink nlid="nl91" bibid="bib76" firstref="ref178"></nolink> <nolink nlid="nl92" bibid="bib28" firstref="ref179"></nolink> <nolink nlid="nl93" bibid="bib22" firstref="ref181"></nolink> <nolink nlid="nl94" bibid="bib36" firstref="ref182"></nolink> <nolink nlid="nl95" bibid="bib60" firstref="ref184"></nolink> <nolink nlid="nl96" bibid="bib96" firstref="ref185"></nolink> <nolink nlid="nl97" bibid="bib114" firstref="ref186"></nolink> <nolink nlid="nl98" bibid="bib90" firstref="ref188"></nolink> <nolink nlid="nl99" bibid="bib46" firstref="ref191"></nolink> <nolink nlid="nl100" bibid="bib18" firstref="ref193"></nolink> <nolink nlid="nl101" bibid="bib21" firstref="ref194"></nolink> <nolink nlid="nl102" bibid="bib117" firstref="ref197"></nolink> <nolink nlid="nl103" bibid="bib82" firstref="ref201"></nolink> <nolink nlid="nl104" bibid="bib47" firstref="ref207"></nolink> <nolink nlid="nl105" bibid="bib51" firstref="ref208"></nolink> <nolink nlid="nl106" bibid="bib26" firstref="ref209"></nolink> <nolink nlid="nl107" bibid="bib97" firstref="ref210"></nolink> <nolink nlid="nl108" bibid="bib85" firstref="ref211"></nolink> <nolink nlid="nl109" bibid="bib41" firstref="ref212"></nolink> <nolink nlid="nl110" bibid="bib62" firstref="ref214"></nolink> <nolink nlid="nl111" bibid="bib66" firstref="ref218"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1466628 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Self-Regulated Learning in the Digitally Enhanced Science Classroom: Toward an Early Warning System – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Marcus+Kubsch%22">Marcus Kubsch</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-5497-8336">0000-0001-5497-8336</externalLink>)<br /><searchLink fieldCode="AR" term="%22Sebastian+Strauß%22">Sebastian Strauß</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-0647-1132">0000-0002-0647-1132</externalLink>)<br /><searchLink fieldCode="AR" term="%22Adrian+Grimm%22">Adrian Grimm</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0003-2701-3349">0000-0003-2701-3349</externalLink>)<br /><searchLink fieldCode="AR" term="%22Sebastian+Gombert%22">Sebastian Gombert</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-5598-9547">0000-0001-5598-9547</externalLink>)<br /><searchLink fieldCode="AR" term="%22Hendrik+Drachsler%22">Hendrik Drachsler</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0001-8407-5314">0000-0001-8407-5314</externalLink>)<br /><searchLink fieldCode="AR" term="%22Knut+Neumann%22">Knut Neumann</searchLink><br /><searchLink fieldCode="AR" term="%22Nikol+Rummel%22">Nikol Rummel</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-3187-5534">0000-0002-3187-5534</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Psychology+Review%22"><i>Educational Psychology Review</i></searchLink>. 2025 37(2). – Name: Avail Label: Availability Group: Avail Data: Springer. Available from: Springer Nature. One New York Plaza, Suite 4600, New York, NY 10004. Tel: 800-777-4643; Tel: 212-460-1500; Fax: 212-460-1700; e-mail: customerservice@springernature.com; Web site: https://link.springer.com/ – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 38 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Junior+High+Schools%22">Junior High Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Middle+Schools%22">Middle Schools</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Grade+7%22">Grade 7</searchLink><br /><searchLink fieldCode="EL" term="%22Grade+8%22">Grade 8</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Inquiry%22">Inquiry</searchLink><br /><searchLink fieldCode="DE" term="%22Science+Instruction%22">Science Instruction</searchLink><br /><searchLink fieldCode="DE" term="%22Electronic+Books%22">Electronic Books</searchLink><br /><searchLink fieldCode="DE" term="%22Workbooks%22">Workbooks</searchLink><br /><searchLink fieldCode="DE" term="%22Physics%22">Physics</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Collection%22">Data Collection</searchLink><br /><searchLink fieldCode="DE" term="%22Cognitive+Processes%22">Cognitive Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Metacognition%22">Metacognition</searchLink><br /><searchLink fieldCode="DE" term="%22Affective+Behavior%22">Affective Behavior</searchLink><br /><searchLink fieldCode="DE" term="%22Productivity%22">Productivity</searchLink><br /><searchLink fieldCode="DE" term="%22Self+Management%22">Self Management</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22At+Risk+Students%22">At Risk Students</searchLink><br /><searchLink fieldCode="DE" term="%22Middle+School+Students%22">Middle School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Grade+7%22">Grade 7</searchLink><br /><searchLink fieldCode="DE" term="%22Grade+8%22">Grade 8</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Germany%22">Germany</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1007/s10648-025-10011-9 – Name: ISSN Label: ISSN Group: ISSN Data: 1040-726X<br />1573-336X – Name: Abstract Label: Abstract Group: Ab Data: Recent research underscores the importance of inquiry learning for effective science education. Inquiry learning involves self-regulated learning (SRL), for example when students conduct investigations. Teachers face challenges in orchestrating and tracking student learning in such instruction; making it hard to adequately support students. Using AI methods such as machine learning (ML), the data that is generated when students interact in technology-enhanced classrooms can be used to track their learning and subsequently to inform teachers so that they can better support student learning. This study implemented digital workbooks in an inquiry-based physics unit, collecting cognitive, metacognitive, and affective data from 214 students. Using ML methods, an early warning system was developed to predict students' learning outcomes. Explainable ML methods were used to unpack these predictions and analyses were conducted for potential biases. Results indicate that an integration of cognitive, metacognitive, and affective data can predict students' productivity with an accuracy ranging from 60 to 100% as the unit progresses. Initially, affective and metacognitive variables dominate predictions, with cognitive variables becoming more significant later. Using only affective and metacognitive data, predictive accuracies ranged from 60 to 80% throughout. Bias was found to be highly dependent on the ML methods being used. The study highlights the potential of digital student workbooks to support SRL in inquiry-based science education, guiding future research and development to enhance instructional feedback and teacher insights into student engagement. Further, the study sheds new light on the data needed and the methodological challenges when using ML methods to investigate SRL processes in classrooms. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: Note Label: Notes Group: Note Data: https://osf.io/uv8tn – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1466628 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1466628 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1007/s10648-025-10011-9 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 38 Subjects: – SubjectFull: Inquiry Type: general – SubjectFull: Science Instruction Type: general – SubjectFull: Electronic Books Type: general – SubjectFull: Workbooks Type: general – SubjectFull: Physics Type: general – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Prediction Type: general – SubjectFull: Data Collection Type: general – SubjectFull: Cognitive Processes Type: general – SubjectFull: Metacognition Type: general – SubjectFull: Affective Behavior Type: general – SubjectFull: Productivity Type: general – SubjectFull: Self Management Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: At Risk Students Type: general – SubjectFull: Middle School Students Type: general – SubjectFull: Grade 7 Type: general – SubjectFull: Grade 8 Type: general – SubjectFull: Technology Uses in Education Type: general – SubjectFull: Germany Type: general Titles: – TitleFull: Self-Regulated Learning in the Digitally Enhanced Science Classroom: Toward an Early Warning System Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Marcus Kubsch – PersonEntity: Name: NameFull: Sebastian Strauß – PersonEntity: Name: NameFull: Adrian Grimm – PersonEntity: Name: NameFull: Sebastian Gombert – PersonEntity: Name: NameFull: Hendrik Drachsler – PersonEntity: Name: NameFull: Knut Neumann – PersonEntity: Name: NameFull: Nikol Rummel IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 1040-726X – Type: issn-electronic Value: 1573-336X Numbering: – Type: volume Value: 37 – Type: issue Value: 2 Titles: – TitleFull: Educational Psychology Review Type: main |
| ResultId | 1 |