Aligning Classroom Assessment with Engineering Practice: A Design-Based Research Study of a Two-Stage Exam with Authentic Assessment

Saved in:
Bibliographic Details
Title: Aligning Classroom Assessment with Engineering Practice: A Design-Based Research Study of a Two-Stage Exam with Authentic Assessment
Language: English
Authors: Koretsky, Milo D. (ORCID 0000-0002-6887-4527), McColley, Campbell J., Gugel, James L., Ekstedt, Thomas W.
Source: Journal of Engineering Education. Jan 2022 111(1):185-213.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 29
Publication Date: 2022
Sponsoring Agency: National Science Foundation (NSF)
Contract Number: 1519467
Document Type: Journal Articles
Reports - Research
Descriptors: Performance Based Assessment, Alignment (Education), Engineering Education, Computer Assisted Testing, Decision Making, Student Attitudes
DOI: 10.1002/jee.20436
ISSN: 1069-4730
Abstract: Background: Authentic assessment and two-stage exams have recently received attention; however, they are rarely used together. We reimagine assessment by integrating an authentic, computer-based assessment into the structure of a two-stage exam in a large engineering class. Purpose: We seek to identify ways that such assessment extends classroom testing to better align with engineering practice by examining the ways teams negotiate uncertainty to make engineering decisions. We also identify differing students' reactions to increased uncertainty during tests. Design/Method: Using the methodical framework of design-based research, we analyze performance and reflection data for 117 student teams through two design iterations to explore four design and theoretical conjectures. Results: Teams chose multiple solution paths to this authentic task, an aspect that aligns with the characteristics of engineering practice that we seek to assess. In addition, the technology tool allows the evaluation of procedural accuracy for many of the teams' chosen paths. The teams' decision-making performances correlate; however, decision-making and traditional assessments do not correlate, suggesting they measure different competencies. The computer-based second stage provides a holistic assessment that shifts the messages that students implicitly receive about valued practices in the classroom. However, not all students took up the authentic group assessment in desired ways. Conclusions: Technology-based two-stage exams with authentic assessment show promise to shift testing practices in large engineering classes to include decision-making. Such assessments better align with engineering practices that are valued in the profession, but more work is needed to develop systems for widespread implementation.
Abstractor: As Provided
Entry Date: 2022
Accession Number: EJ1322763
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwF3bcoWLkAjEKAYpahK_vxbAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDBuyGT1QHN_i61RiPgIBEICBm6F85i-65Ak3vKV8f4KO7KuwzBcQWWP5y5vWl5yVij91bSWB78vo1dUjl4TTZs-F628JXBHpNcGcdGfoF2vj82UnhH-gBtsHXoJn9HGBA2a1BBFGh8Np50eqphUS_QpJYj94lh5CqCsHgPD4xtML8fyz1jYJJ6pHHiweOvVGDh_tbTBxCTQGjcujhKbwRzbSq-gCMHeEVbPpIXa1
Text:
  Availability: 1
  Value: <anid>AN0154497207;6m401jan.22;2022Jan07.02:33;v2.2.500</anid> <title id="AN0154497207-1">Aligning classroom assessment with engineering practice: A design‐based research study of a two‐stage exam with authentic assessment </title> <p>Background: Authentic assessment and two‐stage exams have recently received attention; however, they are rarely used together. We reimagine assessment by integrating an authentic, computer‐based assessment into the structure of a two‐stage exam in a large engineering class. Purpose We seek to identify ways that such assessment extends classroom testing to better align with engineering practice by examining the ways teams negotiate uncertainty to make engineering decisions. We also identify differing students' reactions to increased uncertainty during tests. Design/Method: Using the methodical framework of design‐based research, we analyze performance and reflection data for 117 student teams through two design iterations to explore four design and theoretical conjectures. Results: Teams chose multiple solution paths to this authentic task, an aspect that aligns with the characteristics of engineering practice that we seek to assess. In addition, the technology tool allows the evaluation of procedural accuracy for many of the teams' chosen paths. The teams' decision‐making performances correlate; however, decision‐making and traditional assessments do not correlate, suggesting they measure different competencies. The computer‐based second stage provides a holistic assessment that shifts the messages that students implicitly receive about valued practices in the classroom. However, not all students took up the authentic group assessment in desired ways. Conclusions: Technology‐based two‐stage exams with authentic assessment show promise to shift testing practices in large engineering classes to include decision‐making. Such assessments better align with engineering practices that are valued in the profession, but more work is needed to develop systems for widespread implementation.</p> <p>Keywords: chemical engineering; design‐based research; higher education; learning technology; problem‐solving; student assessment</p> <hd id="AN0154497207-2">INTRODUCTION</hd> <p>Practicing engineers are assessed based on their ability to get the job done (Vincenti, 1990)—that is, develop and deliver a solution that satisfies a set of objectives and competing constraints. Engineers' work involves ambiguity where they often make decisions based on incomplete and uncertain knowledge, and this aspect leads to many possible solution paths (Murray et al., 2019; Trevelyan, 2014; Vincenti, 1990). In contrast, assessment practices in schools commonly place students in unrealistically limited contexts (Biggs, 2014; Boud, 1990; Jordan & Babrow, 2013). For example, they are asked to solve problems alone, allowed limited resources, and problems are designed to be deterministic and have one clearly appropriate solution path. In part, instructors design school‐based assessments this way as a matter of practicality, especially when teaching large enrollment classes. They want the assessments to be fair and without a large time burden to grade.</p> <p>In the design‐based research (DBR) study reported here, we investigate how the use of a computer simulation in a two‐stage exam (Efu, 2019) was able to align student assessment with engineering practice, an approach termed <emph>authentic assessment</emph> (Villarroel et al., 2020; Wiggins, 1990). Our approach builds upon identified differences between the nature of professional engineering practice and the work that students complete in many engineering classrooms. As Jonassen et al. (2006) state, "Workplace engineering problems are substantively different from the kinds of problems that engineering students most often solve in the classroom; therefore, learning to solve classroom problems does not necessarily prepare engineering students to solve workplace problems" (p. 139). In part, this difference arises as students are not asked to frame problems and make decisions (Douglas et al., 2012). In the assessment design reported here, the first exam stage contains a traditional individual assessment. In the second stage, student teams use a technology tool to complete an authentic, industrially situated engineering task. The second stage, therefore, elicits collaborative decision‐making. At the same time, this two‐stage exam fits within the instructional practices of many large undergraduate engineering classes.</p> <p>The DBR study sits within a broader organizational change initiative in a 4‐year engineering department. The initiative aims to shift learning activity to more closely resemble professional practice by using problem‐based learning activities within a studio structure (Koretsky, Keeler, et al., 2018). Revised studio problems position students in the role of engineers on teams where they need to identify core foundational principles as conceptual tools to progress on tasks that resemble realistic engineering work (Engle & Conant, 2002; Johri & Olds, 2011). In a previous study during this initiative, we used an activity systems framework (Boaler & Greeno, 2000; Engeström, 2001) to reveal elements that influence students' adoption of more or less productive approaches to learning. By far, exam expectations were the most common element cited by students to prompt undesired rote learning approaches (Michor & Koretsky, 2020). The two‐stage exam with authentic assessment reported here was developed to change students' exam experiences and thereby influence the ways students approach their learning.</p> <p>This study reports two design iterations where a two‐stage exam was used in a large (~200 students) undergraduate engineering science[<reflink idref="bib1" id="ref1">1</reflink>] course in an ABET‐accredited program in the United States. Using the technique of conjecture mapping (Sandoval, 2014), we explore the following design conjectures (DC) and theoretical conjectures (TC):</p> <p></p> <ulist> <item> DC1. Identifying teams' engineering decisions provides information about ways students are able to use core knowledge in practice.</item> <p></p> <item> DC2. Technology‐based assessments can automate the evaluation of procedural accuracy across different data sets and solution paths.</item> <p></p> <item> TC1. Authentic assessments measure teams' problem formulation and decision‐making competencies in ways not available in typical classroom assessments.</item> <p></p> <item> TC2. Students react to shifts in assessment practices based on their encultured ideas about schooling and testing.</item> </ulist> <p>In the following sections, we provide the background and theoretical framework for the study of the second stage's computer‐based authentic assessment and describe features of the technology tool. We then present findings and discuss each of the research conjectures.</p> <hd id="AN0154497207-3">BACKGROUND AND THEORETICAL FRAMEWORK</hd> <p>Assessment plays a key role in the engineering classroom. Assessments are used to certify mastery and determine grades and also to provide educators information about the effectiveness of instruction (Johnson & Johnson, 1999; Pellegrino et al., 2001). Importantly, they also indicate to students the valued knowledge, skills, and practices in the discipline (Boud, 1990; Hargreaves, 1997). In this section, we argue that computer‐based assessment tasks can place students in the role of engineers doing realistic work. This authentic assessment aligns with a practice‐centered perspective of learning (Manz et al., 2020) where managing uncertainty to make decisions is an important skill for engineers. Finally, we discuss how we adapted the two‐stage exam structure to accommodate such an authentic assessment within the constraints of a large enrollment class in engineering.</p> <hd id="AN0154497207-4">Technology‐enhanced assessment</hd> <p>Technology‐based assessments are increasingly being integrated into educational contexts (Bearman et al., 2020; Koretsky & Magana, 2019; Thomas, 2016). These assessments can take many forms including the use of data banks of multiple‐choice questions with automated marking and feedback (B. Chen, West, & Zilles, 2019; Ćukušić et al., 2014; Mitkov et al., 2006), the use of learning analytics to increase student success (Larrabee Sønderlund et al., 2019), and embedded assessments immersed in serious games (Caballero‐Hernández et al., 2017; Kim & Shute, 2015). Supporters see the potential of technology to enable learning environments that provide real‐time feedback and scalable and personalized support. For example, randomized computer‐generated "rolling problems" have been used to allow students to take exams asynchronously in large enrollment courses (Sud et al., 2019).</p> <p>However, others have critiqued common uses of these technologies stating they overemphasize practical instructional objectives such as efficiency, ease of grading, and record keeping (Wiggins, 1990). Such approaches often use highly specific and structured questions that form piecemeal assessments where instructors reduce a topic into several separate and discrete parts, assess their students' performance on each, and determine a score by adding those together (Haladyna et al., 2002). This type of assessment can reward rote learning: the memorization of facts and solution algorithms (Pellegrino & Quellmalz, 2010). Biggs (2014) advocates for instructors to replace piecemeal summative assessments with assessments that holistically address the "whole performance" (p. 8). Boud (1990) elaborates that those piecemeal assessment practices are antithetical to both the disciplinary practices instructors utilize in their own scholarly work and to the broad goals of university education.</p> <p>In this spirit, educators argue that technology‐based systems present an opportunity to reimagine assessment (Behrens et al., 2019; Bucciarelli, 2003; Timmis et al., 2016). They advocate for a paradigmatic shift from viewing the affordances of technology as primarily toward the efficient delivery at scale to rethinking the fundamental relationship between learning and assessment. Toward this end, the use of immersive simulations to create scenarios where students are placed in the role of practicing professionals doing complex and meaningful work becomes compelling (Barab et al., 2007; Clark et al., 2009; Quellmalz & Pellegrino, 2009; Siyahhan et al., 2017). Such technology‐based simulations have been used to assess important engineering competencies like open‐ended problem‐solving, critical thinking, and modeling and design skills (Koretsky et al., 2008; Rupp et al., 2010; Vieira et al., 2016; Xie et al., 2014; Xing et al., 2020). Thus, the core skills that are assessed shift from answering questions that only have value in a school context to doing meaningful science or engineering work. However, such an approach also has risks. In game‐based simulations, the assessment component may be tacit, leading some learners to be less careful about details as they complete their work (Winkley, 2010).</p> <p>The second stage of the two‐stage exam studied here builds on several areas that Timmis et al. (2016) identify to reimagine assessment, including using technology‐enhanced assessment to provide learners complex decision‐making opportunities where learners can work collaboratively and exercise more agency than with traditional assessment practices. This approach shifts the assessment activity from piecemeal practices to completing an engineering task with realistic goals. At the same time, the second stage is identified clearly as being part of an exam that students should seriously engage in.</p> <hd id="AN0154497207-5">Situated learning</hd> <p>Assessment practices are a powerful indicator to students of what is important to learn and, therefore, should be fundamentally grounded in principles and theories of learning (Hattie & Brown, 2007; Timmis et al., 2016). The computer‐based assessment design studied here draws on the work of learning scientists over the past 30 years, which shifts the view of learning from a cognitive approach that aligns with an acquisition metaphor to a practice‐centered approach that is more appropriately described through legitimate participation (Barab & Duffy, 2000; J. Brown et al., 1989; Engle & Conant, 2002; Lave & Wenger, 1991; Manz et al., 2020). According to the practice‐centered perspective, knowing and doing are intertwined, that is, what is learned is not separate from how it is learned. Learning depends on the content, context, and activity, and knowledge is situated in the experience (Barab & Duffy, 2000; Dewey, 1938; Turner & Nolen, 2015). This view fundamentally challenges the notion that concepts are self‐contained entities but rather positions concepts as tools, which can only be fully understood through use. Thus, learning involves more than "acquiring" conceptual understanding, but rather involves having students build an "increasingly rich implicit understanding of the world in which they use the (conceptual) tools and of the tools themselves" (J. Brown et al., 1989, p. 33). This understanding is framed by those situations in which the conceptual tools are learned and used. In the context of the computer‐based engineering assessment developed here, this perspective implies that an assessment tool should not only provide a set of tasks on which to evaluate a student but should also embed that assessment within a context where those tasks have a meaning that is connected to the engineering profession and its practices.</p> <hd id="AN0154497207-6">Uncertainty and decision‐making in engineering practice</hd> <p>Making progress in the face of uncertainty is fundamental to engineering and science practice (Pickering, 1995; Trevelyan, 2014; Vincenti, 1990). We follow Manz and Suárez's (2018) definition of uncertainty as to the aspects of an engineer's work that is "non‐obvious and contingent, which must be figured out by the scientist [or engineer] and negotiated in response to feedback from peers and the material world" (p. 288). In both engineering contexts (Douglas et al., 2012; Jonassen et al., 2006) and science contexts (Y. C. Chen, Benus, & Hernandez, 2019; Manz, 2015), researchers have identified fundamental differences between how uncertainty manifests in school and practice. In school, students are encultured to knowledge that is certain as measured by assessments with unambiguous solution paths. From this perspective, uncertainty is viewed as undesirable and effort should be taken to avoid it (Jordan & Babrow, 2013). In contrast, real engineering work contains knowledge that is initially unknown, is legitimately problematic, and only temporally emerges with experience and understanding (Manz et al., 2020; Pickering, 1995). By struggling to negotiate uncertainty and building on one another's perspectives to resolve it collaboratively, learners develop a deeper understanding of core scientific principles, even if their immediate resolution may not be normatively correct (Jordan & Babrow, 2013; Lotan, 2003; Manz, 2015).</p> <p>In the context of assessment, we argue that uncertainty is an important element of an authentic task, especially for group work. From this perspective, assessment in engineering should address the ability to work with others to navigate uncertainty. One way such uncertainty is manifest is by problems with many distinct potential solution paths. Borrowing from Jonassen's (2000) typology, the task studied here is constructed as a decision‐making problem where different engineering decisions can lead to different solution paths. Such problems involve identifying benefits and limitations, weighing options, selecting an alternative, and justifying those choices (Jonassen, 2000). While decision‐making problems do not embody the complete extent of uncertainty that professional engineers face, they do extend students further into authentic practice than they commonly experience in their engineering exams.</p> <hd id="AN0154497207-7">Authentic assessment</hd> <p>Equipping students with the skills to use their knowledge in professional practice is an important goal of many university programs. Correspondingly, the role of assessment is appropriate to measure how much learning has occurred and to provide feedback on the effectiveness of instruction (Johnson & Johnson, 1999; Pellegrino et al., 2001). Equally important but often less recognized, assessments also orient students to the nature of knowledge and the legitimate ways that knowledge is used in the profession (Biggs, 1996; Boud, 1990; Hargreaves, 1997; Wormald et al., 2009). The types of knowledge that students develop and the ways they learn to use it are mediated by the assessments by which they are evaluated. Therefore, assessments become critical in orienting students to take up desired engineering practices. As stated earlier, a key motivation for the shift in assessment studied here was based on student reports that traditional course assessments prompted undesired rote learning practices.</p> <p> <emph>Authentic assessment</emph> uses real‐world tasks that mirror the work required of professionals, in our case engineers (Villarroel et al., 2020; Wiggins, 1990). This type of assessment requires students to engage in the type of reasoning and problem‐solving needed in their field. Thus, students must demonstrate deep understanding, higher‐order thinking, and complex problem‐solving. Importantly, authentic interactions should also mirror the social interactions of engineering practice. Characteristics of authentic assessments include problem context, characteristics, and student–student interactions (Jonassen, 1997). The assessment should be set in a context that reflects how students will use their knowledge in the future and include self‐reflection (Villarroel et al., 2020). Problems should be realistic, partly ill‐defined (requiring the students to frame the problem), and complex with multiple solution paths possible. The work should be collaborative with students working in groups and other students' perspectives proving valuable resources (Bucciarelli, 2002; Lotan, 2003).</p> <hd id="AN0154497207-8">Two‐stage exams</hd> <p>In this study, we integrate the concept of authentic assessment into the structure of a two‐stage exam. Recently, two‐stage exams have gained attention as an assessment practice that can simultaneously improve learning (Brame & Biel, 2015; Efu, 2019; Rieger & Heiner, 2014; Stearns, 1996; Zipp, 2007). In the first stage, students complete a traditional exam individually; in the second stage, students work collaboratively in groups. Student scores are determined by a weighted average on both parts. Similar to the way students' collaboration in peer instruction builds understanding when concept questions are used in formative assessment (Mazur, 1997), during the second stage, collaborative sense‐making can lead to improved understanding during summative exams. Importantly, group exams have the potential to assess social competencies needed in the professional workplace (Jang, 2016) and send students the message that collaboration is a valued central practice of doing science and engineering (Boud, 1990; Gilley & Clarkston, 2014; James, 2014; Rieger & Heiner, 2014). However, while two‐stage exams appear to improve motivation and lower stress, there are mixed reports about whether they actually improve learning (Gilley & Clarkston, 2014; Kinnear, 2020; Leight et al., 2012).</p> <p>Most recent reports use two‐stage exam designs that ask student groups to solve the same constrained problems in the second stage as they did as individuals in the first stage. However, such an approach may not provide the fullest opportunity to leverage their team members' different ideas, perspectives, and competencies. Indeed, educators have argued that problems that are "context rich" (Heller & Hollabaugh, 1992) or "group‐worthy" (Lotan, 2003) provide tools that better align with the ways knowledge and skills are used by practicing professionals. Similarly, these problems can provide the basis for authentic assessment. In the assessment design reported in this study, we build on the idea of two‐stage exams by basing the group portion of the second stage on a more authentic group‐worthy task: a task enabled by a computer simulation as described in the next section.</p> <hd id="AN0154497207-9">INDUSTRIALLY SITUATED TECHNOLOGY DESIGN</hd> <p></p> <hd id="AN0154497207-10">Task design</hd> <p>The situated task for the second stage places students in teams as engineers where they must determine the rate constant for a sugar reaction and use the rate constant results to recommend a reactor run time for an industrial candy manufacturing process. The industrial context is provided by a 3D computer simulation (see Figure 1). Students navigate this environment in "first‐person" mode, engaging with equipment and instruments to run the reactors (left) and take measurements (right). After analyzing the results in collaboration with teammates, each individual student submits responses (see Figure 2). To support the context, the assignment (see Appendix A) is in the form of a memorandum from the Vice President of Engineering, and teams are provided an "equipment manual" that describes procedures to run the reactors and take measurements. They are also provided the rubric (see Appendix B) with guidelines for collaborative group work by which they will be assessed.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0001.jpg" title="1 Screenshots of the simulated 3D lab environment. The reactor bay (left) contains two hydrolysis reactors and associated control consoles. The analytical lab (right) contains an automated polarimeter and other analytical tools [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0002.jpg" title="2 Form used by students after the second design iteration to submit the responses for the project" /> </p> <p></p> <p>The simulation is built on a WebGL software platform and runs in standard web browsers. Reactor data are produced by a kinetic model with added random noise. The output provided by the kinetic model affords teams the opportunity for scientific reasoning, while added variation (noise) provides opportunity for statistical reasoning. Each team collects data from two reactors, and the simulation generates unique results for each team member. In effect, this feature provides the same function as randomized computer‐generated "rolling problems" as one team's reactors performance is fundamentally different from their classmates. In some cases, the two reactors have the same kinetic parameters, meaning that outside the process and measurement noise they will perform the same, while in other cases the two reactors are inherently different. The simulated reactor time–concentration data and the students' submitted responses are stored in a central database.</p> <hd id="AN0154497207-13">Iteration of technology design</hd> <p>We implemented two changes in the simulation for the second iteration of data collection based on analysis of the previous cohort's data. First, we enhanced the technology design to support team collaboration by enabling each student to view and access the results of their teammates along with their own results, as shown in Figure 3. Each student independently engages in the situated industrial environment; the intent is for them to each be able to collect data and then combine data with their teammates into a superset that they analyze collaboratively. Analysis of the previous cohort revealed that 59 students (27%) worked individually or reported analysis based only on their individual data set. After the redesign shown in Figure 3, only five students (3%) worked as individuals. To compare across cohorts, we use only data from the students whose submissions are consistent with collaborative work.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0003.jpg" title="3 Display of simulated reactant concentration measurements from Reactor 1 for the final design iteration. The second column represents the values of the student operating that reactor, while Columns 3 and 4 provide the measured values of their teammates, T1 and T2, on the same reactor" /> </p> <p></p> <p>Second, we altered the first question of the submission form to the one shown in Figure 2. For the previous cohort, question 1 had a response box the same size as question 2. This change was made as a cue to encourage teams to report interval estimates rather than point estimates, while not explicitly directing them to do so.</p> <hd id="AN0154497207-15">METHODS</hd> <p></p> <hd id="AN0154497207-16">Methodological framework</hd> <p>The study uses the methodological framework of DBR in education. In DBR, innovative educational systems are deployed in naturalistic settings while, simultaneously, experiments studying the innovative systems are systematically conducted (A. L. Brown, 1992; Collins, 1992; Sandoval, 2014). Through iteration, instructional designs are improved and theoretic conjectures are refined, that is, design of instruction and research on learning intermingle (Bakker, 2018; Minichiello & Caldwell, 2021). Like engineering design, itself, DBR uses a systems approach to create useful innovations within real constraints (Hjalmarson & Lesh, 2008; for examples of DBR studies in engineering, see Dasgupta, 2019; Newstetter, 2005). Importantly, just as Vincenti (1990) argues that the engineering design process produces new theoretical knowledge unique to engineering, DBR can provide new theoretical knowledge unique to learning and instruction. Indeed, Hjalmarson and Parsons (2021) argue that the epistemological differences between DBR and experimental or quasi‐experimental education studies are akin to the fundamental differences between engineering and science.</p> <p>Kelly (2004) describes a method's argumentative grammar as "the logic that guides the use of a method and that supports reasoning about its data" (p. 118). DBR operates within specific classroom learning ecologies that mandate a different grammar than experimental or quasi‐experimental designs (Cobb & Gravemeijer, 2008; Reimann, 2011). Sandoval (2014) developed the technique of conjecture mapping as a tool to articulate the argumentative grammar of a DBR study, that is, to identify the salient theoretical features of a learning environment and predict how they interact to produce desired outcomes.</p> <p>Figure 4 illustrates our adaptation of Sandoval's (2014) conjecture mapping to relate two high‐level conjectures—theoretically grounded ideas of how a technology tool can support desired assessment of learning—to more specific design and theoretical explorations pursued in this study. These high‐level conjectures are embodied within the tools and materials, task structures, and participant structures of the assessment process. This embodiment leads to system‐specific design and theoretical conjectures that connect data sources and data analyses to support empirically grounded claims.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0004.jpg" title="4 Conjecture map for this design‐based research study connecting high‐level conjectures to data sources and analyses [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <hd id="AN0154497207-18">Positionality</hd> <p>We (the four authors) are US born, able‐bodied, cisgender White males, three straight and one queer, who all have engineering degrees. We bring differing perspectives to the technology development, instruction, and research described here. Two of us have extensive industrial experience. The first author directs a research program that seeks to develop learning systems to allow students to be able to integrate and extend the knowledge developed in specific courses in the core curriculum to the more complex, authentic problems and projects they face in professional practice. He has previously conducted DBR studies of technology systems targeting professional practice (Koretsky et al., 2011) and connected conceptual understanding (Friedrichsen et al., 2017) and of organizational change initiatives focusing on higher‐level cognitive and social skills in engineering problem‐solving (Koretsky, Keeler, et al., 2018; Koretsky, Montfort, et al., 2018). The second author is a PhD student in environmental engineering. He brings interests in both technical and educational research to the project. His technical interests center on microplastic contaminant fate and transport in surface waters, while his educational interests target difference, power, and discrimination in environmental engineering. The third author participated in this project as an undergraduate student pursuing a BS in bioengineering. He completed the course studied here the year before the two‐phase exam was implemented and brought perspectives of an enrolled student. He has graduated and currently works as a brewer for a local microbrewery. The fourth author develops software solutions that support STEM education research in the areas of conceptual understanding, integrated learning, and active learning systems. Previously, he has worked in a variety of industrial settings developing IT systems that support manufacturing processes.</p> <p>As privileged White male engineers, we acknowledge the limitations of our collective experience. This work should be considered with that limitation in mind.</p> <hd id="AN0154497207-19">Participants and setting</hd> <p>This study was conducted at a public, research‐intensive university in the Pacific Northwest region of the United States. The two‐stage midterm exam was delivered in a large course required for biological, chemical, and environmental engineers. Data are reported from two cohorts during the second and third DBR iterations in consecutive spring terms. All enrolled students were invited to participate. In Cohort 1, 213 students consented to participate in the study, and in Cohort 2, 168 students consented. Students were mostly sophomores and juniors. Table 1 shows self‐reported demographic data. The research was approved by the institutional review board.</p> <p>1 TABLESelf‐reported demographic data by cohort based on a 69% response rate from study participants</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Classification</th><th align="left">Category</th><th align="left">Cohort 1 (%)</th><th align="left">Cohort 2 (%)</th></tr></thead><tbody valign="top"><tr><td>Major</td><td>Biological</td><td>22</td><td>26</td></tr><tr><td>Chemical</td><td>70</td><td>64</td></tr><tr><td>Environmental</td><td>8</td><td>10</td></tr><tr><td>Gender</td><td>Female</td><td>25</td><td>31</td></tr><tr><td>Male</td><td>73</td><td>66</td></tr><tr><td>Non‐binary/another gender</td><td>2</td><td>3</td></tr><tr><td>Race</td><td>Asian</td><td>14</td><td>18</td></tr><tr><td>Black</td><td>3</td><td>4</td></tr><tr><td>Hispanic/Latinx</td><td>6</td><td>9</td></tr><tr><td>Pacific Islander</td><td><1</td><td><1</td></tr><tr><td>White—non‐Hispanic</td><td>61</td><td>56</td></tr><tr><td>International</td><td>11</td><td>7</td></tr><tr><td>Two or more races</td><td>7</td><td>6</td></tr></tbody></table> </ephtml> </p> <p>Figure 5 shows a flowchart of class activity through the first 6 weeks, highlighting the assessment in Week 6. In Weeks 1 through 5, students participated in regular weekly class activity, which included two 2‐h lectures on Tuesday and Thursday, which the entire class attended. Lectures were interactive and utilized an audience response system called the Concept Warehouse (Koretsky et al., 2014). To prepare for lecture, students were assigned pre‐reading and asked to answer questions in the Concept Warehouse related to the reading assignment. Each student also attended a 2‐h studio section with a maximum enrollment of 24 where they completed activity‐based group work facilitated by a graduate student and an undergraduate student. The majority of the teams had three members (87%), with the remaining teams consisting of two or four members. The team composition for the second stage of the midterm exam was the same as that for the first 5 weeks in studio. Weekly homework assignments were submitted in Weeks 2 through 5 before the midterm exam.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0005.jpg" title="5 Flowchart of the first 6 weeks of activity in the engineering science course studied here" /> </p> <p></p> <p>The two‐stage midterm exam was delivered during the sixth week of the 10‐week term. Students completed the first stage, the individual portion, during lecture on Tuesday. They were allocated the entire class period and allowed to bring only a single 8 × 11″ note sheet and a non‐graphing calculator as resources. Five members of the instructional team circulated around the room to proctor and answer questions of clarification. Approximately half of the students finished within 90 min. Collected exams were scanned into Gradescope (a computer‐based grading tool) and then graded by the instructional team. Appendix C provides a sample item that assesses conceptual understanding and one that assesses procedural accuracy. The other items are similar and not presented for brevity.</p> <p>Student teams completed the second stage, the group portion, the following day during studio. A graduate and an undergraduate student instructor observed the teams and evaluated their engagement and collaboration using a rubric provided by the class instructor with explicit guidelines for collaboration (Appendix B). Students were also assessed on their submitted solutions (Figure 2). The second stage was "open everything but other humans." Team members used laptops to access the simulation, perform analysis in MATLAB or Excel, and could access course content and conduct more general internet searches. Most teams used the entire 2‐h allocation to complete this portion.</p> <p>Instead of the usual reading pre‐quiz for Thursday's lecture, students answered a brief reflection about the midterm. The class period was used to interactively explore the different possible solution paths and to reflect on group teaming strategies. At the end of Thursday's class, students were asked another reflection question. The question for Cohort 2 was different than for Cohort 1.</p> <hd id="AN0154497207-21">Data sources</hd> <p>Student performance and reflections formed the basis of the data analyzed in this study. These data sources relate to the explored design and research conjectures, as shown in Figure 4. The analysis of performance data focused on the second stage. The technology output from the sugar reactor task consisted of the following information for each student: their team numbers; the unique reactor parameter values provided to each team; the raw simulated data provided to each team member; and individuals' responses to the four questions in Figure 2. The responses allowed us to infer the decisions leading to their solution path and to assess their procedural accuracy by determining if the calculations were correct. Students also submitted supporting computational artifacts such as Excel and MATLAB files. Grades from the first stage of the midterm and the final exam were available from a technology‐based grading tool called Gradescope. Student scores for the Concept Warehouse audience response system also provided a measure of student performance.</p> <p>Students responded to a set of reflection prompts after the two‐stage midterm exam, including free‐response items and Likert scale items. After reading the three sets of written responses, the research team selected the following prompt for detailed coding of student perceptions:</p> <p> <emph>Exam Vent</emph>: This week was busy, with an in‐class individual Midterm and a group problem in studio. Sometimes it is useful to vent. In the space below, please provide any comments you feel like on these assessments.</p> <p>We selected this prompt as it addressed both stages and was framed in terms of "venting," which could probe ways that the exam disrupted expectations (theoretical conjecture 2). The other two prompts (write down one thing you learned from the Group Midterm; was there anything you could have done differently to help your group collaborate towards the Group Midterm solution?) were used for triangulation of the detailed coding results.</p> <hd id="AN0154497207-22">Data analysis</hd> <p></p> <hd id="AN0154497207-23">Student performance</hd> <p>The sugar reactor task is designed for teams to toggle between engineering and statistics principles as they complete their work. Teams need to work collaboratively to organize their approach and make engineering decisions. While each team made many choices during the task, we analyzed the following major decisions that were central to most teams' approaches:</p> <p></p> <ulist> <item> Mathematical form: How did they mathematize the kinetic rate expression?</item> <p></p> <item> Solution strategy: How did they organize the data collected from two reactors and several team members to perform their analysis?</item> <p></p> <item> Determination of rate constant: Did they provide a point estimate (single value) or an interval estimate (confidence interval) for the first‐order rate constant?</item> <p></p> <item> Time calculation: Did they account for process variation in selecting a process time for manufacturing? How?</item> <p></p> <item> Statistical analysis: Did they determine if the two reactors' kinetic response was statistically different?</item> </ulist> <p>The unit of analysis for student performance is the team. As described above, we only analyzed data from teams who submitted consistent values across at least two team members, indicating they worked collaboratively. Data from 53 valid teams from Cohort 1 and 64 teams from Cohort 2 were coded. Averages of team members' scores to individual assessments (Final exam, Stage 1, audience response system) were used for correlational analyses.</p> <p>The technology tool includes a set of prompts (Figure 2) in which students report their findings and engineering recommendations and justify them. We first coded the submissions to infer choices for the decisions identified above. The decision codes, mathematical form, solution strategy, determination of rate constant, time calculation, and reactor difference, are described in Table 2. These codes were developed during the first DBR iteration by examining the initial results and feedback from students during the interactive, in‐class reflection activities the following day. These codes remained fixed, and data from the second and third DBR iteration were reported. Using a set of key words for each code together with student values and explanations of the rate constant and production times, researchers identified each team's set of decisions. We counted a decision only if it was incorporated into their submitted recommendation; for example, if a team identified they would calculate a confidence interval but ran out of time, it was not coded as "confidence interval." Two raters independently analyzed 40 responses, selected at random. Fleiss's kappa was used to evaluate interrater reliability. Kappa values for mathematical form (0.906), solution strategy (1.000), confidence interval (0.950), and time calculation (0.922) indicate "almost perfect agreement" according to Landis and Koch (1977). All discrepancies were resolved by consensus. As a check, code results for 10 teams with different solution paths were compared with their submitted Excel or MATLAB supporting files. In all cases, the two matched. A single rater coded the remaining responses.</p> <p>2 TABLECategories for decision codes</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Decision category</th><th align="left">Code</th><th align="left">Description</th></tr></thead><tbody valign="top"><tr><td>Mathematical form</td><td>Linear regression model (linear)</td><td>Taking provided data and creating a linear model to test the dependence of the concentration on time</td></tr><tr><td>Exponential model (Exp)</td><td>Taking provided data and creating an exponential model to test the dependence of the concentration on time</td></tr><tr><td>Other</td><td>No information provided indicating mathematical form</td></tr><tr><td>Solution strategy</td><td>Pooled data</td><td>Team pools data from all team members and analyze the data as a single set. They may analyze reactors separately or together</td></tr><tr><td>k's per reactor</td><td>Team calculates the rate constant k for each reactor separately and averages the values to obtain the rate constant. They may analyze reactors separately or together</td></tr><tr><td>k's per time</td><td>Team calculates the rate constant k for each data point and averages the values of all the individual measurements</td></tr><tr><td>Other</td><td>No information provided indicating codes above</td></tr><tr><td>Rate constant determination</td><td>Confidence interval</td><td>The team provided an interval estimate based on a calculated confidence interval</td></tr><tr><td>Point estimate</td><td>The team provided a single value based on a fit or average of the data</td></tr><tr><td>Other</td><td>The rate constant was determined in another way or not explained</td></tr><tr><td>Time calculation</td><td>CI_Low</td><td>Time was calculated based on the value of the lower confidence interval level (CI_Low)</td></tr><tr><td>k</td><td>Time was calculated based on the reported rate constant</td></tr><tr><td>Other</td><td>Time was determined in another way or not explained</td></tr><tr><td>Reactor difference</td><td>Statistical test performed</td><td>The team statistical test performed a test for difference in reactor kinetics</td></tr><tr><td>No statistical test mentioned</td><td>The team did not indicate that they performed a test of difference</td></tr><tr><td>Rate constants reported separately</td><td>The team reported values for the two reactors separately when there was no difference</td></tr></tbody></table> </ephtml> </p> <p>We then assessed a team's procedural accuracy as follows. The database from the technology tool stored each team member's raw data. Based on a team's selected identified solution path, we calculated the values for rate constant and manufacturing time and compared those calculations with the team's submitted values. We used the team's calculated values for the rate constant determination to determine the correctness of their time calculation; thus, if they did not get the former correct, points did not get deducted twice. In this procedural assessment, there were several additional team decisions that we needed to identify, including if a team chose to analyze the two reactors separately and what significance level they chose to use (most chose <emph>α</emph> = .05, but some chose.01). If the submitted values agreed to our calculations within round off, it was labeled as correct. We were not able to determine correctness for the decisions coded "Other."</p> <hd id="AN0154497207-24">Student reflections</hd> <p>The unit of analysis for student reflections is the individual. Individual student‐written responses to the reflection question asked at the end of Thursday's class for Cohort 1 (Exam Vent) were analyzed using emergent coding and thematic analysis (Braun & Clarke, 2006; Riessman, 2008). After a stable set of code themes was established, axial coding was conducted by the second author (Strauss & Corbin, 1990).</p> <p>To develop the initial code sets, approximately 20 responses were coded collaboratively by the first, second, and fourth authors. The second author then coded the remaining responses, identifying issues with the codes and missing elements. The team then met to reconcile the code set. This process was repeated until the code set no longer changed. After a stable set of code themes was established, the second author coded the entire set.</p> <p>This coding process led to two sets of code categories. The first set (Table 3) focused on the content of the response and included the nature of the assessment problem, expectations of assessment, social context/group work, and value of experience. The second set (Table 4) was based on how the responses aligned with the intention of the assessment. Code categories included encultured in schooling, situated in engineering, and hybrid. When the initial set of coding was complete, the first and second authors read through the other reflective questions with an eye for confirming or disconfirming evidence and modified the summary results accordingly.</p> <p>3 TABLEContent categories for reflective responses</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Category</th><th align="left">Definition</th></tr></thead><tbody valign="top"><tr><td>Nature of assessment problem</td><td>Response addresses aspects of the sugar reactor task itself, including its open‐endedness, the need to make engineering decisions, and the uncertainty associated with those decisions</td></tr><tr><td>Expectations of assessment</td><td>Response comments on aspects of the two‐stage exam, including the appropriateness of the open‐ended task for assessment or being assessed as part of a team</td></tr><tr><td>Social context/group work</td><td>Response comments on the group format of the midterm, for example, if group members were a positive resource or deterred the individual's work</td></tr><tr><td>Value of experience</td><td>Response comments on benefits of the exam—or its lack of usefulness</td></tr></tbody></table> </ephtml> </p> <p>4 TABLEAlignment categories for reflective responses</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Category</th><th align="left">Definition</th></tr></thead><tbody valign="top"><tr><td>Encultured in schooling</td><td>Responses reflective of school focus and importance of grades</td></tr><tr><td>Situated in engineering</td><td>Responses reflective of doing work in an engineering context and may refer to engineering identity or understanding</td></tr><tr><td>Hybrid</td><td>Recognizing engineering context but still contains elements of a student's school encultured mindset</td></tr></tbody></table> </ephtml> </p> <hd id="AN0154497207-25">RESULTS</hd> <p>We present results to address the four research conjectures. We first characterize the solution paths of teams across the two cohorts showing teams make different decisions as they grapple with the uncertainty of the task. We then present the procedural accuracy of the teams' submitted results and show correlations between these performances and traditional assessments. Finally, we characterize student attitudes through their reflections, showing that some students did not engage in authentic group assessment in ways that we intended.</p> <hd id="AN0154497207-26">Team performance</hd> <p></p> <hd id="AN0154497207-27">Design conjecture 1: Engineering decisions</hd> <p>Figure 6 shows decision maps for the solution paths for Cohort 1 (Figure 6a) and Cohort 2 (Figure 6b) by characterizing four decisions: mathematical form, solution strategy, determination of rate constant, and time calculation. The width of the lines connecting decisions is proportional to the number of teams making that choice. For example, in Cohort 1, the highest proportion of teams that used a linear representation of the kinetic model used a "<emph>k</emph>'s per Reactor" solution strategy; however, some teams chose each of the other three solution strategies.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0006.jpg" title="6 Sankey diagrams showing the solution paths mapped for (a) Cohort 1 (n = 53 teams) and (b) Cohort 2 (n = 64 teams)" /> </p> <p></p> <p>In general, the teams followed many different solution paths in the sugar reactor task based on their collaborative decision‐making. A Kruskal–Wallis test shows the decision proportions for mathematical form (<emph>H</emph> = 0.098; <emph>p</emph> = .75) and time calculation (<emph>H</emph> = 0.024; <emph>p</emph> = .88) generally align among cohorts and do not show statistical difference. Cohort 2 selected less <emph>k</emph>'s per time (solution strategy; <emph>H</emph> = 17.8; <emph>p</emph> < .001) and more often decided to report an interval estimate than a point estimate (<emph>H</emph> = 13.0; <emph>p</emph> < .001). These choices align with desired engineering decisions. The latter change is also consistent with the second technology improvement described above (changing the size of the response box). Many teams labeled as "Other" for the time calculation recognized they needed to account in the observed variation in the experimental runs, but rather than using a statistical model to quantify that variation and make a resulting recommendation, they simply looked directly at the raw data from each run and used the slowest of that limited set to roughly estimate a recommended time.</p> <p>We present data on teams' decision to use a statistical test to compare reactors in Table 5. While some teams had reactors that were inherently different, others did not. This decision was not included in Figure 6 as the two cases are not symmetric. It is possible a team whose reactors were not different, performed the analysis and simply did not report it. On the other hand, teams whose reactors were different would more likely report it if that was investigated properly. Few teams whose reactors were different reported such an analysis (5 of 55 teams), even though they completed homework problems, a studio activity, and a Stage 1 exam question on statistical comparison for two samples. One team whose reactors had different rate kinetics explained they used Reactor 1 as it was slower but did not report doing statistical test of reactor difference. Thus, while many students demonstrated proficiency with the procedural calculation in traditional assessments, it did not occur to them to use this approach strategically when they needed to make an engineering decision in Stage 2. In addition, 10 of 62 teams whose reactors were not different chose to report rate constants for each reactor separately. Several of these teams reported confidence intervals that clearly overlapped.</p> <p>5 TABLEStatistical test of reactor difference</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Condition</th><th align="left">Action</th><th align="left">Cohort 1 (%) (<italic>n</italic> = 53 teams)</th><th align="left">Cohort 2 (%) (<italic>n</italic> = 64 teams)</th></tr></thead><tbody valign="top"><tr><td>Reactors different</td><td>Performed test for statistical difference</td><td>3.8</td><td>4.7</td></tr><tr><td>Did not perform test</td><td>41.5</td><td>43.7</td></tr><tr><td>Reactors not different</td><td>Single reactor rate constant reported</td><td>40.6</td><td>40.6</td></tr><tr><td>Reported rate constants separately for each reactor</td><td>5.6</td><td>10.9</td></tr></tbody></table> </ephtml> </p> <p>We can use these data to see the extent to which teams made decisions that align with appropriate engineering practice. Exactly 25% of teams in Cohort 1 and 59% in Cohort 2 chose to use an interval estimate rather than a point estimate to report the rate constant. Only 37% (Cohort 1) and 30% (Cohort 2) of teams used statistical modeling to account for variability in their time recommendation to manufacturing (CI_Low). In practice, not accounting for variability would lead to many batches in production not meeting specification. Finally, for the cases where the two reactors were different, less than 10% of teams reported performing a statistical analysis of difference.</p> <hd id="AN0154497207-29">Design conjecture 2: Procedural accuracy</hd> <p>Table 6 shows the calculation accuracy for each cohort for each of the two main deliverables, according to code category, with the other decision choices aggregated. The teams that chose an interval estimate for the rate constant determination or time calculation had a lower percentage correct than the more straightforward point estimate calculations. In several cases, we could identify a specific mistake a team made. For example, within the "<emph>k</emph>'s per Reactor" decision for solution strategy, seven teams used a <emph>z</emph>‐value rather than the appropriate <emph>t</emph>‐value to calculate the confidence interval. A few groups using this strategy included five runs in their average instead of all six and consequently did not report <emph>k</emph> correctly. Conversely, there were a few teams with legitimate creative approaches that were "outside the box" and needed to be evaluated accordingly. It would be challenging to automate the process for computer grading in such cases.</p> <p>6 TABLEStudent procedural performance for different solution paths</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Deliverable</th><th align="left">Rate constant determination</th><th align="left">Time calculation</th></tr><tr><th align="left">Code</th><th align="left">Confidence interval</th><th align="left">Point estimate</th><th align="left">CI_Low</th><th align="left"><italic>k</italic></th></tr><tr><th align="left">Cohort</th><th align="left">Cohort 1</th><th align="left">Cohort 2</th><th align="left">Cohort 1</th><th align="left">Cohort 2</th><th align="left">Cohort 1</th><th align="left">Cohort 2</th><th align="left">Cohort 1</th><th align="left">Cohort 2</th></tr></thead><tbody valign="top"><tr><td>Correct</td><td>6</td><td>19</td><td>28</td><td>18</td><td>11</td><td>15</td><td>22</td><td>29</td></tr><tr><td>Incorrect</td><td>7</td><td>18</td><td>12</td><td>9</td><td>8</td><td>4</td><td>3</td><td>2</td></tr><tr><td>Total</td><td>13</td><td>37</td><td>40</td><td>27</td><td>19</td><td>19</td><td>25</td><td>31</td></tr><tr><td>Percent correct (%)</td><td>46</td><td>51</td><td>70</td><td>67</td><td>58</td><td>79</td><td>88</td><td>94</td></tr></tbody></table> </ephtml> </p> <p>The percentage correct reported in Table 6 are comparable to the scores on the traditional Stage 1, which assessed procedural accuracy and conceptual understanding, but not engineering decision‐making (75.8 ± 11.0% [Cohort 1]; 68.4 ± 13.6% [Cohort 2]).</p> <hd id="AN0154497207-30">Theoretical conjecture 1: Student competencies</hd> <p>Table 7 shows the results of a Pearson's correlation analysis among components of Stage 2 assessment including decision‐making and procedural accuracy (calculation) for rate constant and time with traditional summative (Final exam, Stage 1) and formative (audience response system) assessment measures. In this analysis, the scores for the assessment measures of individual team members (Final exam, Stage 1, and audience response system) were averaged across the team to be consistent with the team data in Stage 2. The decision‐making components positively correlate to each other and the traditional assessments positively correlate to one another. However, the components of Stage 2 do not correlate to any of the traditional assessments. This result indicates the individual summative (Final exam, Stage 1) and formative (audience response system) assessments measure different competencies than the group decision‐making assessment (Stage 2). We infer from these correlations and the results from Figure 6 that the authentic assessment reported here measures teams' problem formulation and decision‐making competencies in ways not available in typical classroom assessments.</p> <p>7 TABLEPearson's correlation analysis of group assessments, individual assessments, and cohort</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left" /><th align="left">Decision time</th><th align="left">Calculation rate_<italic>k</italic></th><th align="left">Calculation time</th><th align="left">Final exam</th><th align="left">Midterm Stage 1</th><th align="left">Audience response system</th><th align="left">Cohort</th></tr></thead><tbody valign="top"><tr><td>Decision rate_k</td><td>Pearson correlation</td><td>.4693<xref ref-type="fn" rid="tfn1" /></td><td>−.1570</td><td>.1448</td><td>.0989</td><td>.1262</td><td>.1498</td><td>.3309<xref ref-type="fn" rid="tfn1" /></td></tr><tr><td>p value</td><td>.0000</td><td>.1030</td><td>.1330</td><td>.3060</td><td>.1908</td><td>.1199</td><td>.0004</td></tr><tr><td>Decision time</td><td>Pearson correlation</td><td /><td>−.0568</td><td>.1332</td><td>−.0861</td><td>.0042</td><td>.0786</td><td>.3853<xref ref-type="fn" rid="tfn1" /></td></tr><tr><td>p value</td><td /><td>.5576</td><td>.1675</td><td>.3732</td><td>.9651</td><td>.4163</td><td>.0000</td></tr><tr><td>Calculation rate_k</td><td>Pearson correlation</td><td /><td /><td>.1686</td><td>−.0006</td><td>.1309</td><td>.0430</td><td>−.0976</td></tr><tr><td>p value</td><td /><td /><td>.0797</td><td>.9951</td><td>.1748</td><td>.6572</td><td>.3128</td></tr><tr><td>Calculation time</td><td>Pearson correlation</td><td /><td /><td /><td>−.1094</td><td>.0216</td><td>−.0730</td><td>.0904</td></tr><tr><td>p value</td><td /><td /><td /><td>.2575</td><td>.8239</td><td>.4506</td><td>.3499</td></tr><tr><td>Final exam</td><td>Pearson correlation</td><td /><td /><td /><td /><td>.6453<xref ref-type="fn" rid="tfn1" /></td><td>.4127<xref ref-type="fn" rid="tfn1" /></td><td>−.0819</td></tr><tr><td>p value</td><td /><td /><td /><td /><td>.0000</td><td>.0000</td><td>.3970</td></tr><tr><td>Midterm Stage 1</td><td>Pearson correlation</td><td /><td /><td /><td /><td /><td>.4621<xref ref-type="fn" rid="tfn1" /></td><td>−.0679</td></tr><tr><td>p value</td><td /><td /><td /><td /><td /><td>.0000</td><td>.4832</td></tr><tr><td>Audience response system</td><td>Pearson correlation</td><td /><td /><td /><td /><td /><td /><td>−.0430</td></tr><tr><td>p value</td><td /><td /><td /><td /><td /><td /><td>.6573</td></tr></tbody></table> </ephtml> </p> <p>1 * Denotes significant correlation at <emph>p</emph> value <.01.</p> <p>The correlation of these assessment measures between cohorts is also shown. Only decision‐making measures correlate with cohort. This finding indicates the two technology changes described above may have supported the teams' decision‐making, providing support for the iterative DBR approach reported here as a means for improvement of educational technology.</p> <hd id="AN0154497207-31">Theoretical conjecture 2: Student assessment experience</hd> <p>We next summarize the main findings from the analysis of student reflection responses to the "Exam Vent" prompt. We observe two points of view in students' responses, encultured in schooling and situated in engineering. Table 8 presents sample quotations for each category with one column representative of encultured in schooling and the next situated in engineering. Responses coded as encultured focused on school work and the importance of grades, whereas those coded as situated were reflective of doing real engineering work and often connected to engineering identity or understanding. Not all students engaged in the group stage in the spirit of an authentic engineering task.</p> <p>8 TABLEExamples and number of responses (n) coded encultured in schooling and situated in engineering</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left" /><th align="left">Encultured in schooling</th><th align="left">Situated in engineering</th></tr><tr><th align="left" /><th align="left">Sample response</th><th align="left"><italic>n</italic></th><th align="left">Sample response</th><th align="left"><italic>n</italic></th></tr></thead><tbody valign="top"><tr><td>Nature of assessment problem</td><td>Apparently we all did the midterm exam wrong. Honestly a lot of things that you ask for us to compute are not straight forward and are really confusing. I think I would do much better in the course if you asked for what you are looking for more directly instead of being vague</td><td>39</td><td>The studio felt a bit rushed, but I really enjoyed the content. It made me feel knowledgeable, and it was fulfilling to actually apply my knowledge to a real world situation</td><td>12</td></tr><tr><td>Expectations of assessment</td><td>I understand that the studio section was able to add to the grade of the exam, but some useful direction as to what models to use may have helped me grasp the context of the assessment better</td><td>21</td><td>I thought that PART 2 was arguably the most helpful and fairest form of testing that I have had during my time at [this university] and I hope that future exams are similar. Being able to talk out loud what I'm thinking has always been a strong suit of mine, and an individual exam doesn't allow me to use this skill</td><td>9</td></tr><tr><td>Social context/group work</td><td>I did most of the work in my group. It is not fair that we all will get the same grade on it</td><td>7</td><td>It is helpful to discuss questions with team mates especially since sometimes I will approach a problem in a way another team mate won't so this gives me an opportunity to think about a problem from different perspective</td><td>10</td></tr><tr><td>Value of experience</td><td>I appreciate the opportunity to make my midterm grade better with the studio assignment</td><td>12</td><td>Honestly, I should have learned more from this studio. We were fast at making judgements, and though I had some of the conceptual knowledge, I didn't take enough time to effectively think about the problem. Therefore, I think I am learning to take things more slowly and make better judgements on these engineering problems. I'm bummed that I wasn't able to go through the problem as detailed as I feel that I should have, but happy that I recognize this now and want to try better next time</td><td>20</td></tr></tbody></table> </ephtml> </p> <p>Table 8 also shows the number of responses (<emph>n</emph>) coded for encultured and situated responses for each category. Of the total responses, 79 were identified as encultured and 51 as situated. In addition, there were six hybrid responses that contained elements of each. For the categories nature of the assessment problem and expectations of assessment, responses coded as encultured were approximately three times more frequent than those coded as situated. Encultured responses expressed expectations of a single, correct answer, did not see the connection to the course material, and conflated ill‐defined engineering work with unclear instructions (e.g., "some useful direction as to what models to use may have helped me grasp the context of the assessment better"). Conversely, situated responses generally identified that there were multiple approaches and not necessarily a "correct" answer. These responses identified the need to think creatively and connected the problem to professional settings.</p> <p>Students were more responsive to the shift in social context, with slightly more responses coded situated as encultured. The situated responses cited a value in learning from their teammates' perspectives and the way talking aloud helped their thinking while the encultured responses pointed more toward a single person driving the solution decisions. Similar views were expressed by students in Cohort 2, who were asked specifically about group work in their Thursday post‐class reflection.</p> <p>In remarking on the value of the assessment experience, approximately twice as many responses aligned with a situated perspective as an encultured perspective. The situated responses cited benefits of realistic engineering work expressing the idea that knowledge is iterative, progressive, and gained by experience. For example, the following responses juxtapose encultured and situated perspectives of reporting an interval estimate:</p> <p>Not a lot of direction was explicit in the instructions. If we had known that we were expected to make a confidence interval, we would have done so. [Encultured in schooling]</p> <p>I learned how to apply what we learned to real‐life situations. In this studio, my group used confidence intervals to ensure that our reaction rate constant was within a desirable range. [Situated in engineering]</p> <p>Students coded as situated often connected the value of the second stage to its assessment aim of having them make engineering decisions, as described by the following student:</p> <p>I learned that analyzing data is a lot more open‐ended than I expected. Throughout the studio, my group members and I understood what we needed to do and the general process on how to reach that conclusion. What I did not expect was the various ways to go about it and how multiple ways could be right (some being more right). I thought that determining which method was the best was the most difficult part of the analysis. [Situated in engineering]</p> <hd id="AN0154497207-32">DISCUSSION</hd> <p>Guided by a DBR methodology, we present two iterations of analysis of performance and reflection data for engineering students who completed a two‐stage midterm exam in a large engineering class required for graduation. The assessment design reported here differs from common manifestations of two‐stage exams reported in the literature (Gilley & Clarkston, 2014; Rieger & Heiner, 2014). Rather than having student teams revisit the questions they answered as individuals, we reimagined assessment using an authentic, computer‐based task designed to place teams in the role of engineers (Villarroel et al., 2020; Wiggins, 1990).</p> <p>As illustrated in Figure 4, this study was built upon two high‐level conjectures: (i) learning to solve classroom problems does not translate into solving engineering problems in practice (Jonassen et al., 2006) and (ii) authentic, technology‐based assessments can allow educators to reimagine assessment (Timmis et al., 2016; Wiggins, 1990). Through embodiment in the tools and materials, task structure, and participant structure of the second exam stage, we explored two design conjectures and two theoretical conjectures, which are discussed in the next sections.</p> <hd id="AN0154497207-33">Design conjecture 1: Assessment of engineering decisions</hd> <p>The different characteristic solution paths that teams used in responding to the task provide evidence that the task elicits their engineering decision‐making in the face of uncertainty (Pickering, 1995; Trevelyan, 2014; Vincenti, 1990). As illustrated by the Sankey diagrams in Figure 6, teams exhibited multiple solution paths based on their choices. Thus, the assessment aligned with important characteristics of engineering practice that we are trying to develop—engineering decision‐making. Students identified this aspect as different from their common testing experiences in their reflections.</p> <p>When faced with an authentic task, many teams struggled to make appropriate engineering decisions such as identifying to use interval estimates, accounting for process variation in making a time recommendation, testing for statistical difference among reactors, or analyzing reactors separately when there was clearly no difference. These results suggest students need more practice operationalizing core course content to make engineering decisions. Such struggles are consistent with situated perspectives of learning. Rather than viewing knowledge as an abstract entity to be "acquired" and then "used," a situated learning perspective suggests that knowing entails meaningful participation in activities situated within authentic practice (J. Brown et al., 1989; Engle & Conant, 2002; Manz et al., 2020). To the degree that open‐ended problem‐solving and decision‐making in the face of uncertainty are intrinsic to engineering practice itself, utilizing authentic engineering projects and computer technologies can better prepare students to operationalize knowledge in practical, real‐world situations. In addition, by including these practices as a core part of assessment, they likely become more central to engineering instruction throughout the curriculum (Biggs, 1996; James, 2014) and not merely a supporting objective achieved semi‐independently by students via internships and undergraduate research. We propose that with regular use, authentic assessments like the one presented here would encourage instructors to shift classroom activities in ways that better prepare students for the engineering profession (Jonassen et al., 2006).</p> <hd id="AN0154497207-34">Design conjecture 2: Assessment of procedural accuracy</hd> <p>In addition to assessing a team's engineering decisions, the computer system affords the ability to evaluate the procedural accuracy for many solution paths by determining if a team's submitted numerical answers are consistent with the data they were provided given the solution path that they selected. The numerical performance scores were similar to Stage 1, but it was difficult to compare procedural accuracy across teams as some paths inherently led to more challenging calculation procedures. For a subset of cases, we could not verify the team performed their calculations accurately, such as for the teams that estimated a reactor time directly from the raw data. In addition, we identified a few cases where the numerical answers did not match, but where the team pursued a creative solution outside those that we previously considered.</p> <p>There are challenges in integrating this aspect into an automated scoring system. The more open ended the task, the more possible alternatives that need to be considered. There is potential to make automated scoring more tractable with greater constraints (e.g., specifying the significance level). However, a better understanding is needed in the trade‐offs between posing such constraints and the broader pedagogical and epistemological goals of this type of open‐ended computer‐based authentic assessment.</p> <hd id="AN0154497207-35">Theoretical conjecture 1: Student competencies</hd> <p>A Pearson's correlation analysis suggests that the second stage measures different aspects of engineering knowledge than the first stage and the other individual assessments. Team performance in decision‐making for rate constant determination correlated with their decision‐making for production time (<emph>r</emph> = .47) but did not correlate with traditional summative (Final exam, Stage 1) or formative (audience response system) assessments. However, those traditional assessments all correlated with one another. The positive correlation in teams' decision‐making through a design iteration as the technology was improved from Cohort 1 to Cohort 2 adds to the trustworthiness. The method and context of the Stage 2 assessment align more closely with engineering practice than Stage 1.</p> <p>These correlational results are consistent with Jonassen et al.'s (2006) assertion that professional engineering problems are substantively different from typical in‐class problems. While many students demonstrated proficiency with the procedural calculations in traditional assessments, they were not able to operationalize that same content when they needed to make an engineering decision in Stage 2. This finding suggests that knowledge needed in engineering goes beyond procedural competency or even conceptual understanding, but students too need to develop understanding to appropriately apply that knowledge strategically to make decisions in the context of real engineering work (Shavelson et al., 2003). In that vein, it would be useful to investigate how Stage 1 and Stage 2 performances correlate to other professional performance experiences such as in capstone design, undergraduate research, or internships. For many teams, we were able to determine whether they completed their calculations accurately; however, their performance did not correlate with either decision‐making or traditional assessment. We believe the lack of correlation results from varying difficulty of calculations for different solution paths, but more research is needed.</p> <hd id="AN0154497207-36">Theoretical conjecture 2: Student assessment experience</hd> <p>Importantly, the computer‐based second stage provides a holistic assessment (Biggs, 2014) that shifts the messages that students implicitly receive about valued practices in the classroom (Boud, 1990). Consistent with Guzzomi et al. (2017), students expressed varying responses to the shift in assessment practice. On the one hand, some acknowledged how the assessment mirrored the legitimate ways knowledge is used in the profession; for others, however, it violated the norms of how learning should be measured in school. Like other forms of two‐stage exams reported in the literature, students identified using social competencies needed in the professional workplace (Jang, 2016). Specifically, some expressed the value of working in teams where members could bring diverse perspectives and expertise to bear on the complex task (Lotan, 2003; Smith, 1996). Some students embraced the creative and challenging aspects of the situated task and claimed this assessment aligned with their capabilities better than traditional sequestered exams. However, other students struggled with the shift in norms of the authentic assessment, especially around the open‐endedness and uncertainty, often in impassioned ways. They associated the uncertainty with a lack of clarity in how the task was presented to them rather than as an inherent aspect of engineering work. Prior school experiences enculture students to expect single correct answers and more obvious connections to the prior work in class (Doyle, 1988; Lampert, 1990; Schoenfeld, 1988). Clearly, this type of assessment disrupts the status quo. As the two‐stage exam structure modeled in this study could fit within many large enrollment courses, it provides opportunity to give students regular exposure to more open‐ended work. Such experiences can lead to shifts in students' notions of assessment and of engineering. More research is needed on how such programmatic shifts change students' perceptions of engineering practice and thereby influence their epistemological commitments.</p> <hd id="AN0154497207-37">Limitations</hd> <p>This study has several limitations. First, it was conducted at one institution and used a single problem. The studio structure of the course studied was particularly amenable to delivery of the second stage reported here, and problem characteristics are known to elicit different responses (Douglas et al., 2012; Jonassen, 2000). While two cohorts showed similar decisions, solution paths, and post‐exam reflections, implementation of this two‐stage exam structure in other settings with different problems is needed. Second, 59 students from Cohort 1 were excluded as they worked independently, while a technology modification led to only five students being excluded from Cohort 2. Even though both cohorts demonstrated similar decision‐making characteristics, this discrepancy leads to sampling concerns when comparing cohorts. Third, while this task was open‐ended compared with traditional assessments in engineering science classes, it was quite constrained relative to the messy work of engineering practice (Pickering, 1995; Vincenti, 1990). Fourth, the system reported here is in earlier stages of development, and teams' decisions were coded manually by researchers. The high reliability, as measured by Fleiss's kappa, contributes to the study's trustworthiness. However, our ultimate goal is to automate this type of authentic computer‐based assessment and enable widespread delivery in large enrollment engineering science courses. Such tools would assess students' ability to grapple with uncertainties in engineering work and correspondingly provide instructors a broad understanding of the engineering decisions that teams are making. While automated coding is necessary for widespread implementation, the high reliability obtained in this study combined with ongoing advances in lexical systems (Ha et al., 2011), machine learning (Jescovitch et al., 2020), and artificial intelligence (Johri, 2020; Roll & Wylie, 2016) make this aim appear tenable. Engagement of researchers and developers from these communities is needed. Fifth, while this format shows promise to assess social competencies associated with group work, more effort is needed to identify and refine the processes and practices associated with this aspect of the two‐stage exam.</p> <hd id="AN0154497207-38">Contributions</hd> <p>This study offers two major contributions. First, we implemented a technology‐based tool and explicated how the tool was used to reimagine assessment. Following the idea that assessment drives the learning (Felder & Brent, 2016; James, 2014; Kahn & O'Rourke, 2005), this study examined an innovative authentic assessment approach for a large engineering science class that placed student teams in the role of engineers doing realistic work. The focus of the assessment shifted from conceptual understanding and procedural accuracy to the measure of teams' ability to manage uncertainty to make decisions. A Pearson's correlation analysis suggests that the authentic assessment in the second stage measured different competencies than the traditional classroom assessments, supporting the high‐level conjecture that learning to solve classroom problems does not translate into solving engineering problems in practice. While some students embraced the creative and challenging aspects of the situated task, other students struggled with the open‐endedness and uncertainty, often in impassioned ways. In shifting assessment practices, instructors need to address explicitly classroom norms and student expectations.</p> <p>Second, we apply the methodological approach of DBR to a learning system in engineering education. DBR is particularly appropriate for instructional innovations pointed toward professional formation of engineers (Dasgupta, 2019; Diefes‐Dux et al., 2010; Gomez & Svihla, 2019; Minichiello & Caldwell, 2021; Newstetter, 2005; Weber et al., 2014) and where learning engineering is mediated by technology tools (Friedrichsen et al., 2017; Minichiello & Caldwell, 2021). Following Sandoval (2014), we used conjecture mapping to explore simultaneously design conjectures and theoretical conjectures through two design iterations. Our approach aligns with DBR criteria for computer‐based tools posed by Jeong et al. (2014) to be set in a specific context, grounded in theory, and be part of a larger DBR program. However, there are important differences between this study and DBR studies reported in the literature. Commonly, DBR studies addressing classroom learning focus on shifts in student reasoning and argumentation (e.g., Reimann, 2011; Sandoval, 2014). This orientation has appropriately led to the examination of the discursive practices between students with one another and with the instructor. We agree that these characteristics of classroom learning are important in the formation of engineers, and other studies from our larger initiative focus on discursive processes during classroom learning (Hirshfield & Koretsky, 2021; Koretsky et al., 2021). However, equally important, yet uncommon in DBR studies, are the assessment processes that drive the learning. At the same time, we see merit in examining discursive processes during the team‐based assessments described here as a fruitful area for future research.</p> <hd id="AN0154497207-39">ACKNOWLEDGMENTS</hd> <p>The authors gratefully acknowledge the support provided by the National Science Foundation through Grant EEC 1519467. Any opinions, findings, and conclusions or recommendations expressed in this material do not necessarily reflect the views of the National Science Foundation.</p> <p>A Appendix STAGE 2—PROBLEM STATEMENT FOR COHORT 2</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-gra-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-gra-0001.jpg" title="." /> </p> <p></p> <p>Our new OrangeCandy product line needs to go into volume production. For the process, we need a source for glucose (C<subs>6</subs>H<subs>12</subs>O<subs>6</subs>) and fructose (C<subs>6</subs>H<subs>12</subs>O<subs>6</subs>). We will produce these sugars through a hydrolysis reaction using sucrose (C<subs>12</subs>H<subs>22</subs>O<subs>11</subs>) as a reactant in aqueous solution (0.5 M HCl) with our proprietary RateEnhancer additive that is believed to catalyze the reaction.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-gra-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-gra-0002.jpg" title="." /> </p> <p></p> <p>The hydrolysis reaction is monitored by a polarimeter. In this technique, the angle of plane‐polarized laser light is measured as it is passed through the solution. The change in angle can be related to the concentration of sucrose in the solution.</p> <p>The biochemists from the consulting firm we hired report that the kinetics for this irreversible first‐order reaction can be described by the following ordinary differential equation:</p> <p> <ephtml> <math display="block" xmlns="http://www.w3.org/1998/Math/MathML"><msub><mi mathvariant="italic">dC</mi><mi>S</mi></msub><mi mathvariant="italic">dt</mi><mspace width=".5em" /><mo linebreak="goodbreak">=</mo><mspace width=".25em" /><mo linebreak="goodbreak">−</mo><msub><mi mathvariant="italic">kC</mi><mi>S</mi></msub><mo>,</mo></math> </ephtml> </p> <p>where <emph>C</emph><subs><emph>S</emph></subs> is the sucrose concentration in (mol/m<sups>3</sups>), <emph>k</emph> is the first‐order reaction rate constant in (h<sups>−1</sups>), and <emph>t</emph> is time in (h). However, they are not able to provide us a value for the rate constant, <emph>k</emph>, as we do not want to provide them access to our proprietary additive.</p> <p>The quality we provide to our customers is of utmost importance at Beaver Dam Sweet Treats. The process design team reports that it is critical that at least 70% of the initial sucrose has reacted to make the final product acceptable. Conversion of less than 70% requires reprocessing the entire batch. Due to production bottlenecks, we also need to run the process for as short a time as possible. Due to process flow requirements, our two batch reactors need to use the same process time.</p> <p>Please determine the rate constant and use it to make a process recommendation that you are confident will reach the 70% conversion requirement. I suggest you do the following analysis prior to experiments.</p> <p></p> <ulist> <item> First, the equation above must be solved to get a relationship between concentration and time. Please do this on your whiteboard.</item> <p></p> <item> Second, concentration versus time data are needed. It is helpful to draw a rough schematic of what the experimental equipment would look like. Please do this on your whiteboard and have it approved by one of your supervisors.</item> </ulist> <p>Since each run takes many hours, each team member will only have two experimental runs to collect data (one run on each batch reactor), but our <emph>ConceptWarehouse</emph> software will allow you to see the data from your teammates as well. Before you do your runs, you must receive approval from a supervisor.</p> <p>To do a run in the batch reactor, please log onto the Concept Warehouse. <emph>It is very important that you put in your correct section and team number to see your teammates' data and get credit for your run</emph>.</p> <p>Using the cumulative data set collected by your group (and only your group), please:</p> <p></p> <ulist> <item> How do you suggest to report the reaction rate constant, <emph>k</emph> , for the company databank? Please include numbers.</item> <p></p> <item> Recommend how long you think the operators on the production floor should run the batch reactors to get at least 70% conversion. Suggest a single process time.</item> <p></p> <item> Justify the values you provide and your recommendation, supporting your analysis with appropriate data.</item> <p></p> <item> As we change or modify our product line, the sugar specifications change. Develop a model that predicts the run time needed for <emph>any</emph> conversion the production supervisor may want in the future and report it.</item> </ulist> <p>As a team, please provide all software files you generated for analysis to Gradescope and as an individual report values for <emph>k</emph>, time, justification, and model to the Concept Warehouse.</p> <p>B Appendix STAGE 2—RUBRIC PROVIDED TO STUDENTS DURING THE SECOND STAGE</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-gra-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-gra-0003.jpg" title="." /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-gra-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-gra-0004.jpg" title="." /> </p> <p></p> <p>C Appendix SAMPLE ITEMS FROM THE CLASSROOM ASSESSMENTS</p> <p>Figures C1–C4 provide illustrative items from classroom assessments including the first stage of the midterm exam and the audience response system. For each case, an example focused on conceptual understanding and an example focused on procedural accuracy is shown.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0007.jpg" title="C1 Sample item from Stage 1 (individual) portion of the two‐stage exam focused on conceptual understanding [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0008.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0008.jpg" title="C2 Sample item from Stage 1 (individual) portion of the two‐stage exam focused on procedural accuracy" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0009.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0009.jpg" title="C3 Sample item from in‐class audience response system focused on conceptual understanding [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/6M4/01jan22/jee20436-fig-0010.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jee20436-fig-0010.jpg" title="C4 Sample item from in‐class audience response system focused on procedural accuracy [Color figure can be viewed at wileyonlinelibrary.com]" /> </p> <p></p> <ref id="AN0154497207-48"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> The term "engineering science" course is chosen to be consistent with the historical language in US ABET accreditation guidelines.</bibtext> </blist> <blist> <bibl id="bib2" type="bt">2</bibl> <bibtext> Funding information National Science Foundation of the United States, Grant/Award Number: 1519467</bibtext> </blist> </ref> <ref id="AN0154497207-49"> <title> REFERENCES </title> <blist> <bibtext> Bakker, A. (2018). Design research in education: A practical guide for early career researchers. Routledge. https://doi.org/10.4324/9780203701010</bibtext> </blist> <blist> <bibtext> Barab, S., Dodge, T., Tuzun, H., Job‐Sluder, K., Jackson, C., Arici, A., Job‐Sluder, L., Carteaux, R., Jr., Gilbertson, J., & Heiselt, C. (2007). The Quest Atlantis Project: A socially responsive play space for learning. In B. E. Shelton & D. Wiley (Eds.), The educational design and use of simulation computer games (pp. 159 – 186). Sense Publishers. https://doi.org/10.1163/9789087903121_011</bibtext> </blist> <blist> <bibl id="bib3" type="bt">3</bibl> <bibtext> Barab, S., & Duffy, T. (2000). From practice fields to communities of practice. Theoretical Foundations of Learning Environments, 1 (1), 25 – 55.</bibtext> </blist> <blist> <bibl id="bib4" type="bt">4</bibl> <bibtext> Bearman, M., Dawson, P., Ajjawi, R., Tai, J., & Boud, D. (2020). Re‐imagining university assessment in a digital world. Springer. https://doi.org/10.1007/978-3-030-41956-1</bibtext> </blist> <blist> <bibl id="bib5" type="bt">5</bibl> <bibtext> Behrens, J. T., DiCerbo, K. E., & Foltz, P. W. (2019). Assessment of complex performances in digital environments. The Annals of the American Academy of Political and Social Science, 683 (1), 217 – 232. https://doi.org/10.1177/0002716219846850</bibtext> </blist> <blist> <bibl id="bib6" type="bt">6</bibl> <bibtext> Biggs, J. (1996). Enhancing teaching through constructive alignment. Higher Education, 32 (3), 347 – 364. <ulink href="http://doi.org/10.1007/BF00138871">http://doi.org/10.1007/BF00138871</ulink></bibtext> </blist> <blist> <bibl id="bib7" type="bt">7</bibl> <bibtext> Biggs, J. (2014). Constructive alignment in university teaching. HERDSA Review of Higher Education, 1, 5 – 22.</bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Boaler, J., & Greeno, J. G. (2000). Identity, agency, and knowing in mathematics worlds. In J. Boaler (Ed.), Multiple perspectives on mathematics teaching and learning (pp. 171 – 200). Ablex.</bibtext> </blist> <blist> <bibl id="bib9" type="bt">9</bibl> <bibtext> Boud, D. (1990). Assessment and the promotion of academic values. Studies in Higher Education, 15 (1), 101 – 111. <ulink href="http://doi.org/10.1080/03075079012331377621">http://doi.org/10.1080/03075079012331377621</ulink></bibtext> </blist> <blist> <bibtext> Brame, C. J., & Biel, R. (2015). Test‐enhanced learning: The potential for testing to promote greater learning in undergraduate science courses. CBE—Life Sciences Education, 14 (2), es4. <ulink href="http://doi.org/10.1187/cbe.14-11-0208">http://doi.org/10.1187/cbe.14-11-0208</ulink></bibtext> </blist> <blist> <bibtext> Braun, V., & Clarke, V. (2006). Using thematic analysis in psychology. Qualitative Research in Psychology, 3 (2), 77 – 101. <ulink href="http://doi.org/10.1191/1478088706qp063oa">http://doi.org/10.1191/1478088706qp063oa</ulink></bibtext> </blist> <blist> <bibtext> Brown, A. L. (1992). Design experiments: Theoretical and methodological challenges in creating complex interventions in classroom settings. Journal of the Learning Sciences, 2 (2), 141 – 178. https://doi.org/10.1207/s15327809jls0202_2</bibtext> </blist> <blist> <bibtext> Brown, J., Collins, A., & Duguid, P. (1989). Situated cognition and the culture of learning. Educational Researcher, 18 (1), 32 – 42. https://doi.org/10.3102/0013189X018001032</bibtext> </blist> <blist> <bibtext> Bucciarelli, L. L. (2002). Between thought and object in engineering design. Design Studies, 23 (3), 219 – 231. https://doi.org/10.1016/S0142-694X(01)00035-7</bibtext> </blist> <blist> <bibtext> Bucciarelli, L. L. (2003). Engineering philosophy. Delft University Press.</bibtext> </blist> <blist> <bibtext> Caballero‐Hernández, J. A., Palomo‐Duarte, M., & Dodero, J. M. (2017). Skill assessment in learning experiences based on serious games: A systematic mapping study. Computers & Education, 113, 42 – 60. <ulink href="http://doi.org/10.1016/j.compedu.2017.05.008">http://doi.org/10.1016/j.compedu.2017.05.008</ulink></bibtext> </blist> <blist> <bibtext> Chen, B., West, M., & Zilles, C. (2019). Analyzing the decline of student scores over time in self‐scheduled asynchronous exams. Journal of Engineering Education, 108 (4), 574 – 594.</bibtext> </blist> <blist> <bibtext> Chen, Y. C., Benus, M. J., & Hernandez, J. (2019). Managing uncertainty in scientific argumentation. Science Education, 103 (5), 1235 – 1276. https://doi.org/10.1002/jee.20292</bibtext> </blist> <blist> <bibtext> Clark, D., Nelson, B., Sengupta, P., & D'Angelo, C. (2009). Rethinking science learning through digital games and simulations: Genres, examples, and evidence. Paper presented at the Learning Science: Computer Games, Simulations, and Education. A Workshop Sponsored by the National Academy of Sciences, Washington, DC. Retrieved from https://sites.nationalacademies.org/cs/groups/dbassesite/documents/webpage/dbasse_080068.pdf</bibtext> </blist> <blist> <bibtext> Cobb, P., & Gravemeijer, K. (2008). Experimenting to support and understand learning processes. In A. E. Kelly, R. A. Lesh, & J. Y. Baek (Eds.), Handbook of design research methods in education (pp. 68 – 95). Routledge. <ulink href="http://doi.org/10.4324/9781315759593.CH4">http://doi.org/10.4324/9781315759593.CH4</ulink></bibtext> </blist> <blist> <bibtext> Collins, A. (1992). Toward a design science of education. In E. Scanlon & T. O'Shea (Eds.), New directions in educational technology (pp. 15 – 22). Springer. <ulink href="http://doi.org/10.1007/978-3-642-77750-9%5f2">http://doi.org/10.1007/978-3-642-77750-9%5f2</ulink></bibtext> </blist> <blist> <bibtext> Ćukušić, M., Garača, Ž., & Jadrić, M. (2014). Online self‐assessment and students' success in higher education institutions. Computers & Education, 72, 100 – 109. https://doi.org/10.1016/j.compedu.2013.10.018</bibtext> </blist> <blist> <bibtext> Dasgupta, C. (2019). Improvable models as scaffolds for promoting productive disciplinary engagement in an engineering design activity. Journal of Engineering Education, 108 (3), 394 – 417. <ulink href="http://doi.org/10.1002/jee.20282">http://doi.org/10.1002/jee.20282</ulink></bibtext> </blist> <blist> <bibtext> Dewey, J. (1938). Experience and education. Macmillan.</bibtext> </blist> <blist> <bibtext> Diefes‐Dux, H. A., Zawojewski, J. S., & Hjalmarson, M. A. (2010). Using educational research in the design of evaluation tools for open‐ended problems. International Journal of Engineering Education, 26 (4), 807 – 819.</bibtext> </blist> <blist> <bibtext> Douglas, E. P., Koro‐Ljungberg, M., McNeill, N. J., Malcolm, Z. T., & Therriault, D. J. (2012). Moving beyond formulas and fixations: Solving open‐ended engineering problems. European Journal of Engineering Education, 37 (6), 627 – 651. https://doi.org/10.1080/03043797.2012.738358</bibtext> </blist> <blist> <bibtext> Doyle, W. (1988). Work in mathematics classes: The context of students' thinking during instruction. Educational Psychologist, 23 (2), 167 – 180. https://doi.org/10.1207/s15326985ep2302_6</bibtext> </blist> <blist> <bibtext> Efu, S. I. (2019). Exams as learning tools: A comparison of traditional and collaborative assessment in higher education. College Teaching, 67 (1), 73 – 83. https://doi.org/10.1080/87567555.2018.1531282</bibtext> </blist> <blist> <bibtext> Engeström, Y. (2001). Expansive learning at work: Toward an activity theoretical reconceptualization. Journal of Education and Work, 14 (1), 133 – 156. https://doi.org/10.1080/13639080123238</bibtext> </blist> <blist> <bibtext> Engle, R., & Conant, F. (2002). Guiding principles for fostering productive disciplinary engagement: Explaining an emergent argument in a community of learners classroom. Cognition and Instruction, 20 (4), 399 – 483. https://doi.org/10.1207/S1532690XCI2004_1</bibtext> </blist> <blist> <bibtext> Felder, R. M., & Brent, R. (2016). Teaching and learning STEM: A practical guide. Jossey‐Bass.</bibtext> </blist> <blist> <bibtext> Friedrichsen, D. M., Smith, C., & Koretsky, M. D. (2017). Propagation from the start: The spread of a concept‐based instructional tool. Educational Technology Research and Development, 65 (1), 177 – 202. https://doi.org/10.1007/S11423-016-9473-2</bibtext> </blist> <blist> <bibtext> Gilley, B. H., & Clarkston, B. (2014). Collaborative testing: Evidence of learning in a controlled in‐class study of undergraduate students. Journal of College Science Teaching, 43 (3), 83 – 91. https://doi.org/10.2505/4/jcst14_043_03_83</bibtext> </blist> <blist> <bibtext> Gomez, J., & Svihla, V. (2019). Building individual accountability through consensus. Chemical Engineering Education, 53 (2), 71 – 71. <ulink href="http://doi.org/10.18260/2-1-370.660-108007">http://doi.org/10.18260/2-1-370.660-108007</ulink></bibtext> </blist> <blist> <bibtext> Guzzomi, A. L., Male, S. A., & Miller, K. (2017). Students' responses to authentic assessment designed to develop commitment to performing at their best. European Journal of Engineering Education, 42 (3), 219 – 240. https://doi.org/10.1080/03043797.2015.1121465</bibtext> </blist> <blist> <bibtext> Ha, M., Nehm, R. H., Urban‐Lurain, M., & Merrill, J. E. (2011). Applying computerized‐scoring models of written biological explanations across courses and colleges: Prospects and limitations. CBE—Life Sciences Education, 10 (4), 379 – 393. <ulink href="http://doi.org/10.1187/cbe.11-08-0081">http://doi.org/10.1187/cbe.11-08-0081</ulink></bibtext> </blist> <blist> <bibtext> Haladyna, T. M., Downing, S. M., & Rodriguez, M. C. (2002). A review of multiple‐choice item‐writing guidelines for classroom assessment. Applied Measurement in Education, 15 (3), 309 – 333. https://doi.org/10.1207/S15324818AME1503_5</bibtext> </blist> <blist> <bibtext> Hargreaves, D. J. (1997). Student learning and assessment are inextricably linked. European Journal of Engineering Education, 22 (4), 401 – 409. https://doi.org/10.1080/03043799708923471</bibtext> </blist> <blist> <bibtext> Hattie, J. A., & Brown, G. T. (2007). Technology for school‐based assessment and assessment for learning: Development principles from New Zealand. Journal of Educational Technology Systems, 36 (2), 189 – 201. https://doi.org/10.2190/ET.36.2.g</bibtext> </blist> <blist> <bibtext> Heller, P., & Hollabaugh, M. (1992). Teaching problem solving through cooperative grouping. Part 2: Designing problems and structuring groups. American Journal of Physics, 60 (7), 637 – 644. https://doi.org/10.1119/1.17118</bibtext> </blist> <blist> <bibtext> Hirshfield, L. J., & Koretsky, M. D. (2021). Cultivating creative thinking in engineering student teams: Can a computer‐mediated virtual laboratory help? Journal of Computer Assisted Learning, 37 (2), 587 – 601. https://doi.org/10.1111/jcal.12509</bibtext> </blist> <blist> <bibtext> Hjalmarson, M. A., & Lesh, R. (2008). Engineering and design research: Intersections for education research and design. In A. E. Kelly, R. A. Lesh, & J. Y. Baek (Eds.), Handbook of design research methods in education (pp. 96 – 110). Routledge. https://doi.org/10.1017/CBO9781139013451.032</bibtext> </blist> <blist> <bibtext> Hjalmarson, M. A., & Parsons, A. W. (2021). Conjectures, cycles and contexts: A systematic review of design‐based research in engineering education. Studies in Engineering Education, 1 (2), 142 – 155. <ulink href="http://doi.org/10.21061/see.35">http://doi.org/10.21061/see.35</ulink></bibtext> </blist> <blist> <bibtext> James, D. (2014). Investigating the curriculum through assessment practice in higher education: The value of a 'learning cultures' approach. Higher Education, 67 (2), 155 – 169. https://doi.org/10.1007/s10734-013-9652-6</bibtext> </blist> <blist> <bibtext> Jang, H. (2016). Identifying 21st century STEM competencies using workplace data. Journal of Science Education and Technology, 25 (2), 284 – 301. https://doi.org/10.1007/s10956-015-9593-1</bibtext> </blist> <blist> <bibtext> Jeong, H., Hmelo‐Silver, C. E., & Yu, Y. (2014). An examination of CSCL methodological practices and the influence of theoretical frameworks 2005–2009. International Journal of Computer‐Supported Collaborative Learning, 9 (3), 305 – 334. https://doi.org/10.1007/s11412-014-9198-3</bibtext> </blist> <blist> <bibtext> Jescovitch, L. N., Scott, E. E., Cerchiara, J. A., Merrill, J., Urban‐Lurain, M., Doherty, J. H., & Haudek, K. C. (2020). Comparison of machine learning performance using analytic and holistic coding approaches across constructed response assessments aligned to a science learning progression. Journal of Science Education and Technology, 30, 150 – 167. https://doi.org/10.1007/s10956-020-09858-0</bibtext> </blist> <blist> <bibtext> Johnson, D. W., & Johnson, R. T. (1999). Cooperative learning and assessment [Paper No. FL 026 115]. In Cooperative learning: JALT applied materials (pp. 164 – 178). ERIC.</bibtext> </blist> <blist> <bibtext> Johri, A. (2020). Artificial intelligence and engineering education. Journal of Engineering Education, 109 (3), 358 – 361. https://doi.org/10.1002/jee.20326</bibtext> </blist> <blist> <bibtext> Johri, A., & Olds, B. M. (2011). Situated engineering learning: Bridging engineering education research and the learning sciences. Journal of Engineering Education, 100 (1), 151 – 185. https://doi.org/10.1002/j.2168-9830.2011.tb00007.x</bibtext> </blist> <blist> <bibtext> Jonassen, D. H. (1997). Instructional design models for well‐structured and III‐structured problem‐solving learning outcomes. Educational Technology Research and Development, 45 (1), 65 – 94. https://doi.org/10.1007/BF02299613</bibtext> </blist> <blist> <bibtext> Jonassen, D. H. (2000). Toward a design theory of problem solving. Educational Technology Research and Development, 48 (4), 63 – 85. https://doi.org/10.1007/BF02300500</bibtext> </blist> <blist> <bibtext> Jonassen, D. H., Strobel, J., & Lee, C. B. (2006). Everyday problem solving in engineering: Lessons for engineering educators. Journal of Engineering Education, 95 (2), 139 – 151. https://doi.org/10.1002/j.2168-9830.2006.tb00885.x</bibtext> </blist> <blist> <bibtext> Jordan, M. E., & Babrow, A. S. (2013). Communication in creative collaborations: The challenges of uncertainty and desire related to task, identity, and relational goals. Communication Education, 62 (2), 210 – 232. https://doi.org/10.1080/03634523.2013.769612</bibtext> </blist> <blist> <bibtext> Kahn, P., & O'Rourke, K. (2005). Understanding enquiry‐based learning. In T. Barret, I. Mac Labhrainn, & H. Fallon (Eds.), Handbook of enquiry & problem based learning (pp. 1 – 12). AIHSE.</bibtext> </blist> <blist> <bibtext> Kelly, A. (2004). Design research in education: Yes, but is it methodological? The Journal of the Learning Sciences, 13 (1), 115 – 128. https://doi.org/10.1207/s15327809jls1301_6</bibtext> </blist> <blist> <bibtext> Kim, Y. J., & Shute, V. J. (2015). The interplay of game elements with psychometric qualities, learning, and enjoyment in game‐based assessment. Computers & Education, 87, 340 – 356. https://doi.org/10.1016/j.compedu.2015.07.009</bibtext> </blist> <blist> <bibtext> Kinnear, G. (2020). Two‐stage collaborative exams have little impact on subsequent exam performance in undergraduate mathematics. International Journal of Research in Undergraduate Mathematics Education, 7, 1 – 28. https://doi.org/10.1007/s40753-020-00121-w</bibtext> </blist> <blist> <bibtext> Koretsky, M. D., Keeler, J., Ivanovitch, J., & Cao, Y. (2018). The role of pedagogical tools in active learning: A case for sense‐making. International Journal of STEM Education, 5 (1), 1 – 20. https://doi.org/10.1186/s40594-018-0116-5</bibtext> </blist> <blist> <bibtext> Koretsky, M. D., Kelly, C., & Gummer, E. (2011). Student perceptions of learning in the laboratory: Comparison of industrially situated virtual laboratories to capstone physical laboratories. Journal of Engineering Education, 100 (3), 540 – 573. https://doi.org/10.1002/j.2168-9830.2011.tb00026.x</bibtext> </blist> <blist> <bibtext> Koretsky, M. D., Montfort, D., Nolen, S. B., Bothwell, M., Davis, S., & Sweeney, J. (2018). Towards a stronger covalent bond: Pedagogical change for inclusivity and equity. Chemical Engineering Education, 52 (2), 117 – 127.</bibtext> </blist> <blist> <bibtext> Koretsky, M. D., Amatore, D., Barnes, C., & Kimura, S. (2008). Enhancement of student learning in experimental design using a virtual laboratory. IEEE Transactions on Education, 51 (1), 76 – 85. <ulink href="http://doi.org/10.1109/TE.2007.906894">http://doi.org/10.1109/TE.2007.906894</ulink></bibtext> </blist> <blist> <bibtext> Koretsky, M. D., Falconer, J. L., Brooks, B. J., Gilbuena, D. M., Silverstein, D. L., Smith, C., & Miletic, M. (2014). The AiChE Concept Warehouse: A web‐based tool to promote concept‐based instruction. Advances in Engineering Education, 4 (1), 1 – 27.</bibtext> </blist> <blist> <bibtext> Koretsky, M. D., & Magana, A. J. (2019). Using technology to enhance learning and engagement in engineering. Advances in Engineering Education, 7 (2), 1 – 53.</bibtext> </blist> <blist> <bibtext> Koretsky, M. D., Vauras, M., Jones, C., Iiskala, T., & Volet, S. (2021). Productive disciplinary engagement in high‐ and low‐outcome student groups: Observations from three collaborative science learning contexts. Research in Science Education, 51 (S1), 159 – 182. https://doi.org/10.1007/s11165-019-9838-8</bibtext> </blist> <blist> <bibtext> Lampert, M. (1990). When the problem is not the question and the solution is not the answer: Mathematical knowing and teaching. American Educational Research Journal, 27 (1), 29 – 63. https://doi.org/10.3102/00028312027001029</bibtext> </blist> <blist> <bibtext> Landis, J. R., & Koch, G. G. (1977). The measurement of observer agreement for categorical data. Biometrics, 33 (1), 159 – 174. https://doi.org/10.2307/2529310</bibtext> </blist> <blist> <bibtext> Larrabee Sønderlund, A., Hughes, E., & Smith, J. (2019). The efficacy of learning analytics interventions in higher education: A systematic review. British Journal of Educational Technology, 50 (5), 2594 – 2618. https://doi.org/10.1111/bjet.12720</bibtext> </blist> <blist> <bibtext> Lave, J., & Wenger, E. (1991). Situated learning: Legitimate peripheral participation. Cambridge University Press.</bibtext> </blist> <blist> <bibtext> Leight, H., Saunders, C., Calkins, R., & Withers, M. (2012). Collaborative testing improves performance but not content retention in a large‐enrollment introductory biology class. CBE—Life Sciences Education, 11 (4), 392 – 401. <ulink href="http://doi.org/10.1187/cbe.12-04-0048">http://doi.org/10.1187/cbe.12-04-0048</ulink></bibtext> </blist> <blist> <bibtext> Lotan, R. A. (2003). Group‐worthy tasks. Educational Leadership, 60 (6), 72 – 75.</bibtext> </blist> <blist> <bibtext> Manz, E. (2015). Representing student argumentation as functionally emergent from scientific activity. Review of Educational Research, 85 (4), 553 – 590. https://doi.org/10.3102/0034654314558490</bibtext> </blist> <blist> <bibtext> Manz, E., Lehrer, R., & Schauble, L. (2020). Rethinking the classroom science investigation. Journal of Research in Science Teaching, 57 (7), 1148 – 1174. https://doi.org/10.1002/tea.21625</bibtext> </blist> <blist> <bibtext> Manz, E., & Suárez, E. (2018). Supporting teachers to negotiate uncertainty for science, students, and teaching. Science Education, 102 (4), 771 – 795. https://doi.org/10.1002/sce.21343</bibtext> </blist> <blist> <bibtext> Mazur, E. (1997). Peer instruction. Prentice‐Hall.</bibtext> </blist> <blist> <bibtext> Michor, E., & Koretsky, M. (2020). Students' approaches to studying through a situative lens. Studies in Engineering Education, 1 (1), 38 – 57. <ulink href="http://doi.org/10.21061/see.3">http://doi.org/10.21061/see.3</ulink></bibtext> </blist> <blist> <bibtext> Minichiello, A., & Caldwell, L. (2021). A narrative review of design‐based research in engineering education: Opportunities and challenges. Studies in Engineering Education, 1 (2), 31 – 54. <ulink href="http://doi.org/10.21061/see.15">http://doi.org/10.21061/see.15</ulink></bibtext> </blist> <blist> <bibtext> Mitkov, R., Le An, H., & Karamanis, N. (2006). A computer‐aided environment for generating multiple‐choice test items. Natural Language Engineering, 12 (2), 177 – 194. https://doi.org/10.1017/S1351324906004177</bibtext> </blist> <blist> <bibtext> Murray, J. K., Studer, J. A., Daly, S. R., McKilligan, S., & Seifert, C. M. (2019). Design by taking perspectives: How engineers explore problems. Journal of Engineering Education, 108 (2), 248 – 275. https://doi.org/10.1002/jee.20263</bibtext> </blist> <blist> <bibtext> Newstetter, W. C. (2005). Designing cognitive apprenticeships for biomedical engineering. Journal of Engineering Education, 94 (2), 207 – 213. https://doi.org/10.1002/j.2168-9830.2005.tb00841.x</bibtext> </blist> <blist> <bibtext> Pellegrino, J. W., Chudowsky, N., & Glaser, R. (2001). Knowing what students know: The science and design of educational assessment. National Academy Press. https://doi.org/10.17226/10019</bibtext> </blist> <blist> <bibtext> Pellegrino, J. W., & Quellmalz, E. S. (2010). Perspectives on the integration of technology and assessment. Journal of Research on Technology in Education, 43 (2), 119 – 134. https://doi.org/10.1080/15391523.2010.10782565</bibtext> </blist> <blist> <bibtext> Pickering, A. (1995). The mangle of practice: Time, agency, and science. University of Chicago Press.</bibtext> </blist> <blist> <bibtext> Quellmalz, E. S., & Pellegrino, J. W. (2009). Technology and testing. Science, 323 (5910), 75 – 79. <ulink href="http://doi.org/10.1126/science.1168046">http://doi.org/10.1126/science.1168046</ulink></bibtext> </blist> <blist> <bibtext> Reimann, P. (2011). Design‐based research. In L. Markauskaite, P. Freebody, & J. Irwin (Eds.), Methodological choice and design for educational and social change: Linking scholarship, policy, and practice (pp. 37 – 50). Springer.</bibtext> </blist> <blist> <bibtext> Rieger, G. W., & Heiner, C. E. (2014). Examinations that support collaborative learning: The students' perspective. Journal of College Science Teaching, 43 (4), 41 – 47. <ulink href="http://doi.org/10.2505/4/jcst14%5f043%5f04%5f41">http://doi.org/10.2505/4/jcst14%5f043%5f04%5f41</ulink></bibtext> </blist> <blist> <bibtext> Riessman, C. K. (2008). Narrative methods for the human sciences. Sage Publications.</bibtext> </blist> <blist> <bibtext> Roll, I., & Wylie, R. (2016). Evolution and revolution in artificial intelligence in education. International Journal of Artificial Intelligence in Education, 26 (2), 582 – 599. https://doi.org/10.1007/s40593-016-0110-3</bibtext> </blist> <blist> <bibtext> Rupp, A. A., Gushta, M., Mislevy, R. J., & Shaffer, D. W. (2010). Evidence‐centered design of epistemic games: Measurement principles for complex learning environments. The Journal of Technology, Learning and Assessment, 8 (4). Retrieved from https://ejournals.bc.edu/index.php/jtla/article/view/1623</bibtext> </blist> <blist> <bibtext> Sandoval, W. (2014). Conjecture mapping: An approach to systematic educational design research. Journal of the Learning Sciences, 23 (1), 18 – 36. https://doi.org/10.1080/10508406.2013.778204</bibtext> </blist> <blist> <bibtext> Schoenfeld, A. H. (1988). When good teaching leads to bad results: The disasters of 'well‐taught' mathematics courses. Educational Psychologist, 23 (2), 145 – 166. https://doi.org/10.1207/s15326985ep2302_5</bibtext> </blist> <blist> <bibtext> Shavelson, R., Ruiz‐Primo, M. A., Li, M., & Ayala, C. C. (2003). Evaluating new approaches to assessing learning [Report No. 604]. National Center for Research on Evaluation.</bibtext> </blist> <blist> <bibtext> Siyahhan, S., Ingram‐Goble, A. A., Barab, S., & Solomou, M. (2017). Educational games to support caring and compassion among youth: A design narrative. International Journal of Gaming and Computer‐Mediated Simulations (IJGCMS), 9 (1), 61 – 76. <ulink href="http://doi.org/10.4018/IJGCMS.2017010104">http://doi.org/10.4018/IJGCMS.2017010104</ulink></bibtext> </blist> <blist> <bibtext> Smith, K. A. (1996). Cooperative learning: Making "group work" work. In T. E. Sutherland & C. C. Bonwell (Eds.), Using active learning in college classes: A range of options for faculty (pp. 71 – 82). Jossey‐Bass.</bibtext> </blist> <blist> <bibtext> Stearns, S. A. (1996). Collaborative exams as learning tools. College Teaching, 44 (3), 111 – 112. https://doi.org/10.1080/87567555.1996.9925564</bibtext> </blist> <blist> <bibtext> Strauss, A., & Corbin, J. M. (1990). Basics of qualitative research: Grounded theory procedures and techniques. Sage Publications.</bibtext> </blist> <blist> <bibtext> Sud, P., West, M., & Zilles, C. (2019). Reducing difficulty variance in randomized assessments. Paper presented at the ASEE Annual Conference and Exposition, Tampa, FL. Retrieved from https://strategy.asee.org/33228</bibtext> </blist> <blist> <bibtext> Thomas, S. (2016). Future ready learning: Reimagining the role of technology in education. US Department of Education.</bibtext> </blist> <blist> <bibtext> Timmis, S., Broadfoot, P., Sutherland, R., & Oldfield, A. (2016). Rethinking assessment in a digital age: Opportunities, challenges and risks. British Educational Research Journal, 42 (3), 454 – 476. https://doi.org/10.1002/berj.3215</bibtext> </blist> <blist> <bibtext> Trevelyan, J. (2014). The making of an expert engineer. CRC Press.</bibtext> </blist> <blist> <bibtext> Turner, J. C., & Nolen, S. B. (2015). Introduction: The relevance of the situative perspective in educational psychology. Educational Psychologist, 50 (3), 167 – 172. <ulink href="http://doi.org/10.1080/00461520.2015.1075404">http://doi.org/10.1080/00461520.2015.1075404</ulink></bibtext> </blist> <blist> <bibtext> Vieira, C., Goldstein, M. H., Purzer, Ş., & Magana, A. J. (2016). Using learning analytics to characterize student experimentation strategies in the context of engineering design. Journal of Learning Analytics, 3 (3), 291 – 317. https://doi.org/10.18608/jla.2016.33.14</bibtext> </blist> <blist> <bibtext> Villarroel, V., Boud, D., Bloxham, S., Bruna, D., & Bruna, C. (2020). Using principles of authentic assessment to redesign written examinations and tests. Innovations in Education and Teaching International, 57 (1), 38 – 49. https://doi.org/10.1080/14703297.2018.1564882</bibtext> </blist> <blist> <bibtext> Vincenti, W. G. (1990). What engineers know and how they know it: Analytical studies from aeronautical history. Johns Hopkins University Press.</bibtext> </blist> <blist> <bibtext> Weber, N. R., Strobel, J., Dyehouse, M. A., Harris, C., David, R., Fang, J., & Hua, I. (2014). First‐year students' environmental awareness and understanding of environmental sustainability through a life cycle assessment module. Journal of Engineering Education, 103 (1), 154 – 181. https://doi.org/10.1002/jee.20032</bibtext> </blist> <blist> <bibtext> Wiggins, G. (1990). The case for authentic assessment. Practical Assessment, Research, and Evaluation, 2 (2), 1 – 3. https://doi.org/10.7275/ffb1-mm19</bibtext> </blist> <blist> <bibtext> Winkley, J. (2010). E‐assessment and innovation. Emerging technologies. Becta.</bibtext> </blist> <blist> <bibtext> Wormald, B. W., Schoeman, S., Somasunderam, A., & Penn, M. (2009). Assessment drives learning: An unavoidable truth? Anatomical Sciences Education, 2 (5), 199 – 204. <ulink href="http://doi.org/10.1002/ase.102">http://doi.org/10.1002/ase.102</ulink></bibtext> </blist> <blist> <bibtext> Xie, C., Zhang, Z., Nourian, S., Pallant, A., & Hazzard, E. (2014). A time series analysis method for assessing engineering design processes using a CAD tool. International Journal of Engineering Education, 30 (1), 218 – 230.</bibtext> </blist> <blist> <bibtext> Xing, W., Li, C., Chen, G., Huang, X., Chao, J., Massicotte, J., & Xie, C. (2020). Automatic assessment of students' engineering design performance using a Bayesian network model. Journal of Educational Computing Research, 59 (2), 230 – 256. https://doi.org/10.1177/0735633120960422</bibtext> </blist> <blist> <bibtext> Zipp, J. F. (2007). Learning by exams: The impact of two‐stage cooperative tests. Teaching Sociology, 35 (1), 62 – 76. https://doi.org/10.1177/0092055X0703500105</bibtext> </blist> </ref> <aug> <p>By Milo D. Koretsky; Campbell J. McColley; James L. Gugel and Thomas W. Ekstedt</p> <p>Reported by Author; Author; Author; Author</p> <p></p> <p>Milo D. Koretsky is the McDonnell Family Bridge Professor in the Department of Chemical and Biological Engineering and the Department of Education at Tufts University, 574 Boston Avenue, Medford, MA 02153;.</p> <p>Campbell J. McColley is a Graduate Research Assistant in the Environmental Engineering Program in the School of Chemical, Biological, and Environmental Engineering at Oregon State University, 116 Johnson Hall, Corvallis, OR 97331‐2702;.</p> <p>James L. Gugel is an Undergraduate Research Assistant in the Bioengineering Program in the School of Chemical, Biological, and Environmental Engineering at Oregon State University, 116 Johnson Hall, Corvallis, OR 97331‐2702;.</p> <p>Thomas W. Ekstedt is a Software Developer in the School of Chemical, Biological, and Environmental Engineering at Oregon State University, 116 Johnson Hall, Corvallis, OR 97331‐2702;. [Correction added on 15 December 2021, after first online publication: Email address of Thomas W. Ekstedt was incorrect in the initial publication. It has been corrected.]</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1322763
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Aligning Classroom Assessment with Engineering Practice: A Design-Based Research Study of a Two-Stage Exam with Authentic Assessment
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Koretsky%2C+Milo+D%2E%22">Koretsky, Milo D.</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6887-4527">0000-0002-6887-4527</externalLink>)<br /><searchLink fieldCode="AR" term="%22McColley%2C+Campbell+J%2E%22">McColley, Campbell J.</searchLink><br /><searchLink fieldCode="AR" term="%22Gugel%2C+James+L%2E%22">Gugel, James L.</searchLink><br /><searchLink fieldCode="AR" term="%22Ekstedt%2C+Thomas+W%2E%22">Ekstedt, Thomas W.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Engineering+Education%22"><i>Journal of Engineering Education</i></searchLink>. Jan 2022 111(1):185-213.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 29
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: 1519467
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Performance+Based+Assessment%22">Performance Based Assessment</searchLink><br /><searchLink fieldCode="DE" term="%22Alignment+%28Education%29%22">Alignment (Education)</searchLink><br /><searchLink fieldCode="DE" term="%22Engineering+Education%22">Engineering Education</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Decision+Making%22">Decision Making</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Attitudes%22">Student Attitudes</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jee.20436
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1069-4730
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Background: Authentic assessment and two-stage exams have recently received attention; however, they are rarely used together. We reimagine assessment by integrating an authentic, computer-based assessment into the structure of a two-stage exam in a large engineering class. Purpose: We seek to identify ways that such assessment extends classroom testing to better align with engineering practice by examining the ways teams negotiate uncertainty to make engineering decisions. We also identify differing students' reactions to increased uncertainty during tests. Design/Method: Using the methodical framework of design-based research, we analyze performance and reflection data for 117 student teams through two design iterations to explore four design and theoretical conjectures. Results: Teams chose multiple solution paths to this authentic task, an aspect that aligns with the characteristics of engineering practice that we seek to assess. In addition, the technology tool allows the evaluation of procedural accuracy for many of the teams' chosen paths. The teams' decision-making performances correlate; however, decision-making and traditional assessments do not correlate, suggesting they measure different competencies. The computer-based second stage provides a holistic assessment that shifts the messages that students implicitly receive about valued practices in the classroom. However, not all students took up the authentic group assessment in desired ways. Conclusions: Technology-based two-stage exams with authentic assessment show promise to shift testing practices in large engineering classes to include decision-making. Such assessments better align with engineering practices that are valued in the profession, but more work is needed to develop systems for widespread implementation.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2022
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1322763
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1322763
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jee.20436
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 29
        StartPage: 185
    Subjects:
      – SubjectFull: Performance Based Assessment
        Type: general
      – SubjectFull: Alignment (Education)
        Type: general
      – SubjectFull: Engineering Education
        Type: general
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Decision Making
        Type: general
      – SubjectFull: Student Attitudes
        Type: general
    Titles:
      – TitleFull: Aligning Classroom Assessment with Engineering Practice: A Design-Based Research Study of a Two-Stage Exam with Authentic Assessment
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Koretsky, Milo D.
      – PersonEntity:
          Name:
            NameFull: McColley, Campbell J.
      – PersonEntity:
          Name:
            NameFull: Gugel, James L.
      – PersonEntity:
          Name:
            NameFull: Ekstedt, Thomas W.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2022
          Identifiers:
            – Type: issn-print
              Value: 1069-4730
          Numbering:
            – Type: volume
              Value: 111
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Journal of Engineering Education
              Type: main
ResultId 1