Frontier Model Chatbots Can Help Instructors Create, Improve, and Use Learning Objectives

Saved in:
Bibliographic Details
Title: Frontier Model Chatbots Can Help Instructors Create, Improve, and Use Learning Objectives
Language: English
Authors: Gregory J. Crowther (ORCID 0000-0003-0530-9130), Merrill D. Funk, Kelly M. Hennessey, Marcus M. Lawrence (ORCID 0000-0001-7106-574X)
Source: Advances in Physiology Education. 2025 49(1):219-229.
Availability: American Physiological Society. 9650 Rockville Pike, Bethesda, MD 20814-3991. Tel: 301-634-7164; Fax: 301-634-7241; e-mail: webmaster@the-aps.org; Web site: https://www.physiology.org/journal/advances
Peer Reviewed: Y
Page Count: 11
Publication Date: 2025
Sponsoring Agency: National Science Foundation (NSF), Research Coordination Networks in Undergraduate Biology Education (RCN-UBE)
Contract Number: 1624200
Document Type: Journal Articles
Reports - Research
Tests/Questionnaires
Education Level: Higher Education
Postsecondary Education
Descriptors: Artificial Intelligence, Learning Objectives, Technology Uses in Education, Curriculum Design, Curriculum Implementation, Best Practices, Evaluation Methods, Educational Quality, Undergraduate Study, Exercise Physiology, Anatomy, Perceptual Motor Learning, Taxonomy, Evaluation Criteria
DOI: 10.1152/advan.00159.2024
ISSN: 1043-4046
1522-1229
Abstract: Learning objectives (LOs) are a pillar of course design and execution and thus a focus of curricular reforms. This study explored the extent to which the creation and usage of LOs might be facilitated by three leading chatbots: ChatGPT-4o, Claude 3.5 Sonnet, and Google Gemini Advanced. We posed three main questions, as follows: "question A": when given course content, can chatbots create LOs that are consistent with five best practices in writing LOs?; "question B": when given LOs for a low level of the revised Bloom's taxonomy, can chatbots convert them to a higher level?; and "question C": when given LOs, can chatbots create assessment questions that meet six criteria of quality? We explored these questions in the context of four undergraduate courses: Applied Exercise Physiology, Human Anatomy, Human Physiology, and Motor Learning. According to instructor ratings, chatbots had a >70% success rate on most individual criteria for questions A--C. However, chatbots' "difficulties" with a few criteria (e.g., provision of appropriate context for an LO's action, assignment of an appropriate revised Bloom's taxonomy level) meant that, overall, only 38.3% of chatbot outputs fully met all criteria and thus were possibly ready for use with students. Our findings thus underscore the continuing need for instructor oversight of chatbot outputs but also illustrate chatbots' potential to expedite the design and improvement of LOs and LO-related curricular materials such as test question templates (TQTs), which directly align LOs with assessment questions.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1464069
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGlKTLoyXQM6kOf86QN96p-AAAA4TCB3gYJKoZIhvcNAQcGoIHQMIHNAgEAMIHHBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDNdfNj6UjUOxBxXlpAIBEICBmaaiDbX7NnS3u_6Kqijyff1kx6oyQQWy1I502B3-ySwFngeVFA0pXnhicDgg70fi2ahpt2F85cLh5d0xCuOrSzF9gRG95orSeAA5ARjZ0_Bm8ZRBKQ5tfDWkQ3ootjL0IxdGEDjvcbAkjmulq1mQci6QmHL14zYg3_TtyBZAXSLNE1-ZtQ1X_oz7C6NkTCue1mdF1wIYng_9Sw==
Text:
  Availability: 1
  Value: <anid>AN0183439570;apu01mar.25;2025Mar06.04:31;v2.2.500</anid> <title id="AN0183439570-1">Frontier model chatbots can help instructors create, improve, and use learning objectives </title> <sbt id="AN0183439570-2">INTRODUCTION</sbt> <p>Learning objectives (LOs) are a pillar of course design and execution and thus a focus of curricular reforms. This study explored the extent to which the creation and usage of LOs might be facilitated by three leading chatbots: ChatGPT-4o, Claude 3.5 Sonnet, and Google Gemini Advanced. We posed three main questions, as follows: question A: when given course content, can chatbots create LOs that are consistent with five best practices in writing LOs?; question B: when given LOs for a low level of the revised Bloom's taxonomy, can chatbots convert them to a higher level?; and question C: when given LOs, can chatbots create assessment questions that meet six criteria of quality? We explored these questions in the context of four undergraduate courses: Applied Exercise Physiology, Human Anatomy, Human Physiology, and Motor Learning. According to instructor ratings, chatbots had a >70% success rate on most individual criteria for questions A–C. However, chatbots' "difficulties" with a few criteria (e.g., provision of appropriate context for an LO's action, assignment of an appropriate revised Bloom's taxonomy level) meant that, overall, only 38.3% of chatbot outputs fully met all criteria and thus were possibly ready for use with students. Our findings thus underscore the continuing need for instructor oversight of chatbot outputs but also illustrate chatbots' potential to expedite the design and improvement of LOs and LO-related curricular materials such as test question templates (TQTs), which directly align LOs with assessment questions. NEW & NOTEWORTHY: Amid much recent interest in the impact of generative artificial intelligence on education, one relatively underexplored issue is the extent to which instructors can use chatbots to work more efficiently. This study determined whether the challenging tasks of writing learning objectives (LOs) and writing LO-linked assessment questions can be delegated to advanced chatbots. The chatbots' outputs were often impressive yet often imperfect and thus can be useful as solid drafts that still require instructor oversight.</p> <p>One of the most consequential responsibilities of teaching faculty is the design and/or redesign of academic courses. For biology and biology-adjacent fields, the Vision & Change (V&C) initiative ([<reflink idref="bib1" id="ref1">1</reflink>]) and the V&C-inspired Partnership for Undergraduate Life Sciences Education (PULSE) program ([<reflink idref="bib2" id="ref2">2</reflink>]) have encouraged individuals and departments to reform their curricula to better meet 21st-century needs and have provided valuable guidance for doing so.</p> <p>A recent essay ([<reflink idref="bib3" id="ref3">3</reflink>]) has argued that learning objectives (LOs) are a "natural anchor" (p. 1) for course transformations, since LOs define the highest priority functions that students should be able to perform by the end of a course ([<reflink idref="bib4" id="ref4">4</reflink>]). For a course's LOs to be maximally clear and useful, they should be consistent with both its learning activities and its assessments ([<reflink idref="bib3" id="ref5">3</reflink>]), a three-way match captured by the phrase "triadic course alignment" ([<reflink idref="bib5" id="ref6">5</reflink>]).</p> <p>Much of our recent work has focused on the issues of creating high-quality LOs ([<reflink idref="bib6" id="ref7">6</reflink>], [<reflink idref="bib7" id="ref8">7</reflink>]) and improving their alignment with assessments via test question templates (TQTs) ([<reflink idref="bib8" id="ref9">8</reflink>]). In brief, a TQT directly and explicitly links a LO with multiple specific, realistic examples of how that LO might be summatively assessed, thus clarifying assessment expectations for students and instructors alike ([<reflink idref="bib9" id="ref10">9</reflink>]). We believe, and have preliminary evidence, that TQTs offer important benefits to students ([<reflink idref="bib10" id="ref11">10</reflink>]). However, as is often the case with course-wide changes, implementing TQTs can be logistically challenging. For instance, TQTs are student-facing to teach students how they will be assessed, yet most students require repeated exposure and practice for them to truly understand TQTs and reap the benefits of them ([<reflink idref="bib10" id="ref12">10</reflink>]). Therefore, TQTs likely require a consistent course-wide implementation to yield meaningful benefits, so there is a high pedagogical "start-up cost" or "activation energy" for adding TQTs to a course.</p> <p>In contemplating transformative but time-consuming curricular reforms such as adding TQTs, we, like many others, have wondered whether we could harness the power of chatbots ([<reflink idref="bib11" id="ref13">11</reflink>]) to make our work more efficient ([<reflink idref="bib12" id="ref14">12</reflink>]). Large language model-based chatbots such as ChatGPT have recently become ubiquitous in most domains, including that of education ([<reflink idref="bib13" id="ref15">13</reflink>]). There are numerous reports documenting the often-impressive ability of these chatbots to pass challenging knowledge tests ([<reflink idref="bib14" id="ref16">14</reflink>], [<reflink idref="bib15" id="ref17">15</reflink>]) and to summarize technical papers ([<reflink idref="bib16" id="ref18">16</reflink>]) and qualitative data ([<reflink idref="bib17" id="ref19">17</reflink>]). Also under investigation are questions of whether chatbots can be useful as instructor-side assistants, e.g., in designing curricula ([<reflink idref="bib18" id="ref20">18</reflink>]), tutoring individuals ([<reflink idref="bib19" id="ref21">19</reflink>]), writing assessment questions ([<reflink idref="bib20" id="ref22">20</reflink>]), and grading students' free-form writing ([<reflink idref="bib21" id="ref23">21</reflink>]). However, we have yet to see much peer-reviewed research on chatbots' ability to help instructors write or use LOs, despite the potential centrality of LOs in transforming courses ([<reflink idref="bib3" id="ref24">3</reflink>]). Therefore, in this study, we leveraged our experience with LOs and TQTs to explore whether advanced chatbots could be effective in several fundamental LO-related tasks. In particular, we explored the following three main questions: <emph>question A</emph>: when given course content, can chatbots create LOs that are consistent with five best practices in writing LOs?; <emph>question B</emph>: when given LOs for a low level of the revised Bloom's taxonomy, can chatbots convert them to a higher level?; and <emph>question C</emph>: when given LOs, can chatbots create assessment questions that meet six criteria of quality?</p> <hd id="AN0183439570-3">METHODS</hd> <hd1 id="AN0183439570-4">Positionality</hd1> <p>The authors are all current (G.J.C., M.D.F., M.M.L.) or recent (K.M.H.) teaching faculty in Departments of Biology/Life Sciences (G.J.C., K.M.H.) or Kinesiology and Outdoor Recreation (M.D.F., M.M.L.). Since none of us has extensive experience in artificial intelligence (AI), chatbots, or related domains, we executed this project with an empirical trial-and-error approach broadly similar to the process of developing new experimental protocols, as we have done in previous research-focused positions. We offer the following report as potentially useful to other teaching faculty who, like us, seek to benefit from recent advances in AI despite a lack of AI-specific expertise.</p> <hd1 id="AN0183439570-5">Operational Definition of a LO</hd1> <p>Among the many instructor duties that could be performed in collaboration with chatbots, we focused on devising and using high-quality learning objectives (LOs).</p> <p>As in our previous work ([<reflink idref="bib7" id="ref25">7</reflink>]), we acknowledge that "learning objective" is only one of several closely related terms, with some instructors and researchers preferring the term "learning outcome" (also abbreviated LO), due to, e.g., its possible centering of learners rather than instructors ([<reflink idref="bib22" id="ref26">22</reflink>]). Here, following the precedent of Orr and colleagues ([<reflink idref="bib3" id="ref27">3</reflink>], [<reflink idref="bib4" id="ref28">4</reflink>]), we use "LO" to mean learning objective. LOs can be created for an institution as a whole, a program, a course, and a single day of instruction within a course ([<reflink idref="bib3" id="ref29">3</reflink>]). In this project, we focused on the latter, which Orr et al. refer to as instructional LOs. Here we use "LOs" as a shorthand for these relatively granular instructional LOs.</p> <hd1 id="AN0183439570-6">Use Cases</hd1> <p>Instructors may need to create their own LOs "from scratch" when developing a new course or when revising an existing course so extensively that the previous LOs are not that helpful even as a starting point. Alternatively, instructors may inherit LOs that are a good starting point but that can be improved to reflect personal, institutional, and/or community priorities. With these general situations in mind, we explored the following specific questions:</p> <p></p> <ulist> <item> <emph>Question A:</emph> when given course content, can chatbots create LOs that are consistent with five best practices in writing LOs? Drawing on the work of Orr and colleagues ([<reflink idref="bib4" id="ref30">4</reflink>]), we defined best practices in writing LOs as fulfillment of the following five criteria:</item> <p></p> <item> <emph>Criterion c1</emph>: the LO is clear, balancing conciseness with adequate detail.</item> <p></p> <item> <emph>Criterion c2</emph>: the LO uses an action verb corresponding to a visible performance.</item> <p></p> <item> <emph>Criterion c3</emph>: the LO specifies the context and conditions under which the action is performed.</item> <p></p> <item> <emph>Criterion c4</emph>: achievement of the LO is measurable via multiple-choice questions. [The generally applicable criterion is measurability; for simplicity, this study focuses on the more specific context of multiple-choice questions.]</item> <p></p> <item> <emph>Criterion c5</emph>: the LO is assigned to an appropriate level of the revised Bloom's taxonomy ([<reflink idref="bib23" id="ref31">23</reflink>]).</item> <p></p> <item> <emph>Question B:</emph> when given LOs for a low level of the revised Bloom's taxonomy, can chatbots convert them to a higher level? For example, can they change a "Remember" (level-1) LO to a related "Apply" (level-3) LO? We checked the same five best-practices criteria as above.</item> <p></p> <item> <emph>Question C:</emph> when given LOs, can chatbots create assessment questions that meet six criteria of quality? Drawing on the work of Brame ([<reflink idref="bib24" id="ref32">24</reflink>]) and Albano and colleagues ([<reflink idref="bib25" id="ref33">25</reflink>]), we defined best practices in writing multiple-choice questions as fulfillment of the following six criteria:</item> <p></p> <item> <emph>Criterion c1</emph>: the assessment question should match the content of the LO provided.</item> <p></p> <item> <emph>Criterion c2</emph>: the assessment question should match the revised Bloom's taxonomy level of the LO provided.</item> <p></p> <item> <emph>Criterion c3</emph>: the question stem should be meaningful by itself and should present a definite problem in the form of a question or a partial sentence.</item> <p></p> <item> <emph>Criterion c4</emph>: all answer choices should be reasonably concise and similar in length and format.</item> <p></p> <item> <emph>Criterion c5</emph>: all answer choices should be plausible.</item> <p></p> <item> <emph>Criterion c6</emph>: one answer choice should clearly be right, and the other answer choices should clearly be wrong.</item> </ulist> <p>Since these questions collectively cover both LOs (<emph>questions A</emph> and <emph>B</emph>) and LO-linked exam questions (<emph>question C</emph>), the two essential components of TQTs ([<reflink idref="bib8" id="ref34">8</reflink>]), the chatbots' overall performances represent an answer to the broad question (raised in the introduction) of whether chatbots can expedite the otherwise-laborious work of creating TQTs.</p> <p>In addition, when the chatbots scored relatively poorly on <emph>criterion c3</emph> for <emph>questions A</emph> and <emph>B</emph>, we piloted a possible solution: asking the chatbots to write each LO in a "given X, do Y" format ([<reflink idref="bib7" id="ref35">7</reflink>]). The "given X, do Y" format contrasts with the more common LO format of simply asking students to perform a task without any specification of context. For example, if one wanted to convert a LO of "Identify functions of muscles that affect the knee joint" into "given X, do Y" format, one might rewrite it as: "Given a marked muscle in a three-dimensional model of the lower limb, identify that muscle's effect on the knee joint" ([<reflink idref="bib7" id="ref36">7</reflink>]). By requiring that a "do Y" action be directly preceded by that action's context, this format might improve the communication of LO context as required by <emph>criterion c3</emph>. We thus asked the following offshoot of <emph>question A</emph>, which we dubbed <emph>question A3</emph>:</p> <p></p> <ulist> <item> <emph>Question A3</emph>: do requests for LOs in a "given X, do Y" format lead to successful specification of the context and conditions for the LO action (<emph>criterion c3</emph>)?</item> </ulist> <hd1 id="AN0183439570-7">Prompt Design</hd1> <p>Even for experienced chatbot users, the design of chatbot prompts is often an iterative process involving several rounds of trial and error ([<reflink idref="bib26" id="ref37">26</reflink>]), as reflected in the term "prompt engineering." Accordingly, the prompts used to answer <emph>questions A–C</emph> (above) were created via multiple good-faith attempts to help the chatbots provide the desired outputs. For each prompt, we started from the general principle of treating chatbots as humans ([<reflink idref="bib12" id="ref38">12</reflink>]), in particular, as naive but highly competent interns able to quickly assimilate and act upon instructions if their role is defined in detail ([<reflink idref="bib27" id="ref39">27</reflink>]). The final versions of each prompt are given in the appendix. Each prompt was roughly 250–300 words long, not counting linked or attached materials.</p> <hd1 id="AN0183439570-8">General Testing Matrix</hd1> <p>To improve the odds of reaching generalizable conclusions, we determined whether chatbots' answers to <emph>questions A–C</emph> (above) were robust to variation in the following variables (some of which are summarized in Table 1).</p> <p>Table 1. Courses where LOs were studied with ChatGPT-4o, Claude 3.5 Sonnet, and Google Gemini Advanced</p> <p> <ephtml> <table><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><thead><tr><th align="center" rowspan="1" colspan="1">Course (Level)</th><th align="center" rowspan="1" colspan="1">Course Topics (<italic>1–3</italic>)</th><th align="center" rowspan="1" colspan="1">Source of Material Given to Chatbots</th></tr></thead><tbody><tr><td align="left" rowspan="1" colspan="1">Applied Exercise Physiology (junior-senior)</td><td align="left" rowspan="1" colspan="1"><italic>1</italic>: Facility design (safe design); <italic>2</italic>: program design (needs analysis); <italic>3</italic>: periodization (central concepts/stress adaptation theories)</td><td align="left" rowspan="1" colspan="1">PowerPoint slides (7–10 per topic)</td></tr><tr><td align="left" rowspan="1" colspan="1">Human Anatomy (freshman-sophomore)</td><td align="left" rowspan="1" colspan="1"><italic>1</italic>: Epithelial tissue histology; <italic>2</italic>: lower-limb muscles; <italic>3</italic>: pulmonary circulation</td><td align="left" rowspan="1" colspan="1">PowerPoint slides (7–10 per topic)</td></tr><tr><td align="left" rowspan="1" colspan="1">Human Physiology (freshman-sophomore)</td><td align="left" rowspan="1" colspan="1"><italic>1</italic>: Heart anatomy; <italic>2</italic>: the respiratory system; <italic>3</italic>: the lungs</td><td align="left" rowspan="1" colspan="1">Textbook excerpts (about 2 long paragraphs per topic)</td></tr><tr><td align="left" rowspan="1" colspan="1">Motor Learning (junior-senior)</td><td align="left" rowspan="1" colspan="1"><italic>1</italic>: Factors influencing reaction time; <italic>2</italic>: the nature of motor skills; <italic>3</italic>: attentional capacity</td><td align="left" rowspan="1" colspan="1">Review guides (about 2 pages of text per topic)</td></tr></tbody></table> </ephtml> </p> <p>1 LOs, learning objectives.</p> <hd1 id="AN0183439570-9">Chatbot.</hd1> <p>We used three different chatbots: ChatGPT-4o, Claude 3.5 Sonnet, and Google Gemini Advanced. At the time of this study (summer 2024), these chatbots were considered frontier models ([<reflink idref="bib12" id="ref40">12</reflink>]), i.e., the most advanced and most capable models currently available, and each is available to members of the general public at a price ($20 per month) affordable by many individual instructors. For conciseness, we refer to these chatbots below as ChatGPT, Claude, and Gemini, respectively.</p> <hd1 id="AN0183439570-10">Academic course.</hd1> <p>We explored the application of chatbots to four undergraduate-level courses: Applied Exercise Physiology (for junior and senior kinesiology majors), Human Anatomy (for first- and second-year prenursing and preallied health students), Human Physiology (for first- and second-year biology prenursing and preallied health students), and Motor Learning (for junior and senior kinesiology majors).</p> <hd1 id="AN0183439570-11">Instructor.</hd1> <p>Each round of chatbot interaction was conducted by an author experienced in teaching the corresponding course (M.M.L. for Applied Exercise Physiology, G.J.C. for Human Anatomy, K.M.H. for Human Physiology, and M.D.F. for Motor Learning).</p> <hd1 id="AN0183439570-12">Chapter or module.</hd1> <p>Within each academic course, we performed tests for three different arbitrarily chosen course topics (as listed in Table 1 and in Figs. 1–3).</p> <p>PHOTO (COLOR): Figure 1. Instructor ratings of chatbot success for question A. Each 3-by-5 matrix represents testing of one chatbot for one academic course; each row is a course topics (topics 1–3), and each column is a learning objective best practices criterion (c1–c5). Each box represents an instructor's rating of whether a criterion was fulfilled: green/Y, Yes; yellow/M, Maybe; red/N, No. HOCS, higher-order cognitive skills. LOCS, lower-order cognitive skills.</p> <p>PHOTO (COLOR): Figure 2. Instructor ratings of chatbot success for question B. Each box represents an instructor's rating of whether a criterion was fulfilled: green/Y, Yes; yellow/M, Maybe; red/N, No (as in Fig. 1).</p> <p>PHOTO (COLOR): Figure 3. Instructor ratings of chatbot success for question C. Each 3-by-6 matrix represents testing of one chatbot for one academic course; each row is a course topic (topics 1–3), and each column is a multiple-choice question best practices criterion (c1–c6). Each box represents an instructor's rating of whether a criterion was fulfilled: green/Y, Yes; yellow/M, Maybe; red/N, No (as in Figs. 1 and 2). HOCS, higher-order cognitive skills. LOCS, lower-order cognitive skills.</p> <hd1 id="AN0183439570-13">Revised Bloom's taxonomy level.</hd1> <p>The six levels of the revised Bloom's taxonomy ([<reflink idref="bib23" id="ref41">23</reflink>]) are sometimes grouped into "lower-order cognitive skills" (LOCS) and "higher-order cognitive skills" (HOCS). Like several previous authors ([<reflink idref="bib6" id="ref42">6</reflink>], [<reflink idref="bib28" id="ref43">28</reflink>], [<reflink idref="bib29" id="ref44">29</reflink>]), we consider the Remember and Understand levels to constitute LOCS and the Apply, Analyze, Evaluate, and Synthesize levels to constitute HOCS. For each relevant task, we asked chatbots to create or use a LO at one of the LOCS levels and at one of the HOCS levels.</p> <p>One variable that we did not explore formally was that of chatbot response number. That is, once a prompt was finalized, we generally collected and analyzed only a chatbot's first response to that prompt, as opposed to <emph>1</emph>) repeatedly giving the chatbot the same prompt and collecting multiple rounds of answers ([<reflink idref="bib14" id="ref45">14</reflink>]) or <emph>2</emph>) "coaching" the chatbot on how to fulfill a request more successfully. (A partial exception was our use of <emph>question A3</emph> as described above, which gave the chatbots a second shot at achieving <emph>criterion c3</emph>: "The LO specifies the context and conditions under which the action is performed.") Moreover, we attempted to reduce the influence of previous prompts and responses by starting a new thread for each prompt and beginning each prompt with language suggested by Claude and Gemini: "Please disregard any previous context or conversations. Respond only to the following request..."</p> <hd1 id="AN0183439570-14">Rating Chatbots' Success</hd1> <p>For the questions above, after feeding the chatbots the prompts detailed in the appendix, we rated each response as Yes, Maybe, or No for each of the relevant criteria. While this rating system has low resolution and is somewhat subjective, it reflects the main goal of this study, i.e., to determine whether these chatbots produce ideas that instructors find useful and usable as they design or redesign academic courses. In this context, success cannot be defined with purely objective criteria; the ultimate readout is whether instructors are sufficiently satisfied with the chatbots' outputs to make use of them. Since each instructor is arguably the best judge of whether chatbot outputs would help them teach a specific class, instructors did not systematically cross-check each other's ratings. However, to prevent our results from being dominated by the quirks of a single instructor or a single course, we tested multiple courses taught by multiple instructors as noted in <emph>General Testing Matrix</emph>.</p> <hd1 id="AN0183439570-15">Statistical Analysis</hd1> <p>We combined ratings across instructors and chatbots to calculate overall success rates for the different above-listed criteria and for LOCS-level tasks as compared with HOCS-level tasks. We did so because our main goal was to explore general trends of chatbot success rather than to compare the success rates of different instructors or different chatbots.</p> <p>Since instructors' ratings of chatbot outputs were not formally validated in this exploratory study, and since we wished to avoid the impression of generalizing beyond the limits of the study, we declined to test them for statistically significant differences between groups.</p> <hd id="AN0183439570-16">RESULTS</hd> <hd1 id="AN0183439570-17">Optimized Testing Logistics: One Source of Course Content at a Time, One LO per Prompt</hd1> <p>In preliminary testing, chatbots produced more useful output when we gave them a single source of course content (e.g., PowerPoint slides, textbook excerpts, or lecture notes) rather than several such sources simultaneously. In addition, chatbot output was most useful when prompted with relatively small sections of content (e.g., a single textbook page, as opposed to several pages). Smaller chunks of content seemed preferable both because they led to less chatbot "wandering" into less relevant information and because the chatbots' more focused outputs were easier for us to evaluate. Therefore, with the exception of <emph>question A3</emph>, we limited each prompt to one LO covering one relatively small section of course content.</p> <hd1 id="AN0183439570-18">Instructors' Ratings of Chatbot Outputs</hd1> <p>Chatbots' responses to <emph>question A</emph> (on creating LOs from course material), <emph>question B</emph> (on converting LOs from LOCS to HOCS), and <emph>question C</emph> (on creating questions from LOs) were rated by instructors as shown in Figs. 1, 2, and 3, respectively. These figures provide an overall impression that the outputs of all three chatbots (ChatGPT, Claude, and Gemini) often successfully fulfilled many of the desired criteria. The many data points shown in these figures are summarized more compactly in Tables 2 and 3.</p> <p>Table 2. Summary of instructors' satisfaction with chatbot outputs, stratified by requested cognitive level</p> <p> <ephtml> <table><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><thead><tr><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1">Lower-Order Cognitive Skills (LOCS)</th><th align="center" rowspan="1" colspan="1">Higher-Order Cognitive Skills (HOCS)</th></tr></thead><tbody><tr><td align="left" rowspan="3" colspan="1"><italic>Question A</italic> (create LOs from course material)</td><td align="char" char="%" rowspan="1" colspan="1">86.1% Yes (155/180)</td><td align="char" char="%" rowspan="1" colspan="1">72.8% Yes (131/180)</td></tr><tr><td align="char" char="%" rowspan="1" colspan="1">12.2% Maybe (22/180)</td><td align="char" char="%" rowspan="1" colspan="1">16.1% Maybe (29/180)</td></tr><tr><td align="char" char="%" rowspan="1" colspan="1">1.7% No (3/180)</td><td align="char" char="%" rowspan="1" colspan="1">11.1% No (20/180)</td></tr><tr><td align="left" rowspan="3" colspan="1"><italic>Question C</italic> (create questions from LOs)</td><td align="char" char="%" rowspan="1" colspan="1">88.9% Yes (192/216)</td><td align="char" char="%" rowspan="1" colspan="1">78.2% Yes (169/216)</td></tr><tr><td align="char" char="%" rowspan="1" colspan="1">6.5% Maybe (14/216)</td><td align="left" rowspan="1" colspan="1">14.4% Maybe (31/216)</td></tr><tr><td align="char" char="%" rowspan="1" colspan="1">9.3% No (20/216)</td><td align="char" char="%" rowspan="1" colspan="1">7.4% No (16/216)</td></tr></tbody></table> </ephtml> </p> <p>2 See methods for more detail on LOCS and HOCS. Numbers given are summaries of data reported in Figs. 1 and 3. LOs, learning objectives.</p> <p>Table 3. Summary of instructors' satisfaction with chatbot outputs, stratified by criteria</p> <p> <ephtml> <table><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><col align="left" span="1" /><thead><tr><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1"><italic>Criterion 1</italic> (<italic>c1</italic>)</th><th align="center" rowspan="1" colspan="1"><italic>Criterion 2</italic> (<italic>c2</italic>)</th><th align="center" rowspan="1" colspan="1"><italic>Criterion 3</italic> (<italic>c3</italic>)</th><th align="center" rowspan="1" colspan="1"><italic>Criterion 4</italic> (<italic>c4</italic>)</th><th align="center" rowspan="1" colspan="1"><italic>Criterion 5</italic> (<italic>c5</italic>)</th><th align="center" rowspan="1" colspan="1"><italic>Criterion 6</italic> (<italic>c6</italic>)</th><th align="center" rowspan="1" colspan="1">Perfect</th></tr><tr><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1" /><th align="center" rowspan="1" colspan="1">Overall?</th></tr></thead><tbody><tr><td align="left" rowspan="3" colspan="1"><italic>Question A</italic> (create LOs from course material)</td><td align="char" char="." rowspan="1" colspan="1">76.4% Yes (55/72)</td><td align="left" rowspan="1" colspan="1">95.8% Yes (69/72)</td><td align="left" rowspan="1" colspan="1">68.1% Yes (49/72)</td><td align="left" rowspan="1" colspan="1">77.8% Yes (56/72)</td><td align="left" rowspan="1" colspan="1">79.2% Yes (57/72)</td><td align="left" rowspan="3" colspan="1">N/A</td><td align="left" rowspan="3" colspan="1">38.9% of outputs (28/72) were rated Yes for all 5 criteria</td></tr><tr><td align="char" char="." rowspan="1" colspan="1">19.4% Maybe (14/72)</td><td align="left" rowspan="1" colspan="1">4.2% Maybe (3/72)</td><td align="char" char="." rowspan="1" colspan="1">22.2% Maybe (16/72)</td><td align="left" rowspan="1" colspan="1">16.7% Maybe (12/72)</td><td align="left" rowspan="1" colspan="1">8.3% Maybe (6/72)</td></tr><tr><td align="char" char="." rowspan="1" colspan="1">4.2% No (3/72)</td><td align="left" rowspan="1" colspan="1">0.0% No (0/72)</td><td align="char" char="." rowspan="1" colspan="1">9.7% No (7/72)</td><td align="left" rowspan="1" colspan="1">5.6% No (4/72)</td><td align="left" rowspan="1" colspan="1">12.5% No (9/72)</td></tr><tr><td align="left" rowspan="3" colspan="1"><italic>Question B</italic> (convert LOs from LOCS to HOCS)</td><td align="char" char="." rowspan="1" colspan="1">63.9% Yes (23/36)</td><td align="left" rowspan="1" colspan="1">100% Yes (36/36)</td><td align="char" char="." rowspan="1" colspan="1">55.6% Yes (20/36)</td><td align="left" rowspan="1" colspan="1">72.2% Yes (26/36)</td><td align="left" rowspan="1" colspan="1">72.8% Yes (26/36)</td><td align="left" rowspan="3" colspan="1">N/A</td><td align="left" rowspan="3" colspan="1">25.0% of outputs (9/36) were rated Yes for all 5 criteria</td></tr><tr><td align="char" char="." rowspan="1" colspan="1">33.3% Maybe (12/36)</td><td align="left" rowspan="1" colspan="1">0% Maybe (0/36)</td><td align="char" char="." rowspan="1" colspan="1">38.9% Maybe (14/36)</td><td align="left" rowspan="1" colspan="1">27.8% Maybe (10/36)</td><td align="left" rowspan="1" colspan="1">13.9% Maybe (5/36)</td></tr><tr><td align="char" char="." rowspan="1" colspan="1">2.8% No (1/36)</td><td align="left" rowspan="1" colspan="1">0% No (0/36)</td><td align="char" char="." rowspan="1" colspan="1">5.6% No (2/36)</td><td align="left" rowspan="1" colspan="1">0% No (0/36)</td><td align="left" rowspan="1" colspan="1">13.9% No (5/36)</td></tr><tr><td align="left" rowspan="3" colspan="1"><italic>Question C</italic> (create questions from LOs)</td><td align="char" char="." rowspan="1" colspan="1">83.3% Yes (60/72)</td><td align="left" rowspan="1" colspan="1">63.4% Yes (46/72)</td><td align="char" char="." rowspan="1" colspan="1">90.3% Yes (65/72)</td><td align="left" rowspan="1" colspan="1">91.7% Yes (66/72)</td><td align="left" rowspan="1" colspan="1">84.7% Yes (61/72)</td><td align="left" rowspan="1" colspan="1">73.6% Yes (53/72)</td><td align="left" rowspan="3" colspan="1">44.4% of outputs (32/72) were rated Yes for all 6 criteria</td></tr><tr><td align="char" char="." rowspan="1" colspan="1">11.1% Maybe (8/72)</td><td align="left" rowspan="1" colspan="1">12.5% Maybe (9/72)</td><td align="char" char="." rowspan="1" colspan="1">9.7% Maybe (7/72)</td><td align="left" rowspan="1" colspan="1">4.2% Maybe (3/72)</td><td align="left" rowspan="1" colspan="1">12.5% Maybe (9/72)</td><td align="left" rowspan="1" colspan="1">12.5% Maybe (9/72)</td></tr><tr><td align="char" char="." rowspan="1" colspan="1">5.6% No (4/72)</td><td align="left" rowspan="1" colspan="1">23.6% No (17/72)</td><td align="left" rowspan="1" colspan="1">0% No (0/72)</td><td align="left" rowspan="1" colspan="1">4.2% No (3/72)</td><td align="left" rowspan="1" colspan="1">2.8% No (2/72)</td><td align="left" rowspan="1" colspan="1">13.9% No (10/72)</td></tr></tbody></table> </ephtml> </p> <p>3 See methods for definitions of <emph>criteria c1–c5</emph> (<emph>questions A</emph> and <emph>B</emph>) and <emph>criteria c1</emph>–<emph>c6</emph> (<emph>question C</emph>). Numbers given are summaries of data reported in Figs. 1–3. HOCS, higher-order cognitive skills. LOCS, lower-order cognitive skills; LOs, learning objectives.</p> <p>Table 2 considers whether the chatbots' success was different for LOCS-level and HOCS-level tasks. For both <emph>question A</emph> and <emph>question C</emph>, there was a tendency (not analyzed for statistical significance) toward somewhat less success on HOCS-level tasks. For creating LOs from course materials (<emph>question A</emph>), best practices criteria were fully met (i.e., with Yes ratings) 86.1% of the time for LOCS-level LOs but only 72.8% of the time for HOCS-level LOs. For creating assessment questions based on LOs (<emph>question C</emph>), overall success rates for fully meeting criteria (i.e., with Yes ratings) were 88.9% for LOCS questions and 78.2% for HOCS questions.</p> <p>Table 3 compares fulfillment of the different individual criteria for writing LOs (<emph>questions A</emph> and <emph>B</emph>) and writing assessment questions (<emph>question C</emph>). When creating LOs (<emph>questions A</emph> and <emph>B</emph>), the chatbots had little trouble providing appropriate action verbs (>90% Yes on <emph>criterion c2</emph> for both <emph>question A</emph> and <emph>question B</emph>), but they were less successful in defining the LO context and conditions (<70% Yes on <emph>criterion c3</emph> for both <emph>questions A</emph> and <emph>B</emph>).</p> <p>As one example of not meeting <emph>criterion c3</emph> for <emph>question A</emph>, for the Human Anatomy class, for the topic of lower limb muscles (<emph>topic 2</emph>) and the level of LOCS, both ChatGPT and Claude generated LOs saying that students should be able to identify major lower-limb muscles on a "labeled diagram." This LO is fairly reasonable overall but raises the question of what kinds of labels should be included in the diagram. (Presumably, the individual muscles themselves would not be labeled, which would reduce this LO to a matter of reading names off of the diagram. Perhaps what the chatbots "meant" is that the diagram would be labeled once the students completed the task, meaning that they would start with an unlabeled diagram?) In a Human Anatomy class, we consider it vital to convey the conditions in which structures must be identified; thus ChatGPT and Claude were both rated "Maybe" for <emph>criterion c3</emph> of <emph>question A</emph>.</p> <p>When <emph>criterion c3</emph> of <emph>question A</emph> and <emph>question B</emph> proved problematic for multiple chatbots in multiple courses (Table 3), we piloted the possible solution of requesting that each LO be written in the "given X, do Y" format ([<reflink idref="bib7" id="ref46">7</reflink>]). In <emph>question A</emph>3 we asked each chatbot to write ten LOs in "given X, do Y" format for each of the five human organ systems (nervous, endocrine, cardiovascular, digestive, and renal). However, this additional guidance on defining LO context did not solve the chatbots' problems with <emph>criterion c3</emph>. The chatbots' collective Yes rate for <emph>criterion 3</emph> was 67.3% (101/150) for <emph>question A</emph>3, essentially the same as their previous <emph>criterion 3</emph> Yes rate for <emph>question A</emph> (68.1%).</p> <p>For the <emph>question C</emph> task of writing assessment questions based on LOs, the chatbots were highly successful at writing good question stems with concise, similarly formatted answer choices (>90% Yes on <emph>criteria c3</emph> and <emph>c4</emph>). However, the chatbots were less successful at matching the requested Bloom's taxonomy level of the LO (<70% on <emph>criterion c2</emph>; Table 3).</p> <p>As one example of not meeting <emph>criterion c2</emph> for <emph>question C</emph>, for the Human Anatomy class, Claude produced essentially the same epithelial histology question in response to the LOCS (Remember) prompt and the HOCS (Apply) prompt; it asked which of four answer choices (A–D) best categorized an epithelial image (not actually provided). While different instructors might assign a different Bloom's taxonomy level to this type of question (especially given uncertainties about whether the image was a cartoon or a micrograph and whether the students had previously seen this particular image), the same question cannot reasonably be assigned both to the Remember level and to the Apply level.</p> <p>Overall, while the chatbots successfully met many criteria much of the time, in less than 50% of cases did the chatbot output receive a Yes for all criteria (Table 3, <emph>column 7</emph>). These rates of perfect compliance were 38.9% for <emph>question A</emph>, 25.0% for <emph>question B</emph>, and 44.4% for <emph>question C</emph>; combining all data, 38.3% of outputs (69 of 180) achieved perfect compliance with all criteria. Stated another way, less than half of the time was a chatbot output judged possibly good enough to not need any editing.</p> <hd id="AN0183439570-19">DISCUSSION</hd> <p>This study characterized the ability of three frontier-model chatbots to assist biology and biology-adjacent teaching faculty with tasks related to the design and usage of LOs.</p> <p>Overall, the results reported here are broadly consistent with our and others' previous findings (e.g., Refs. [<reflink idref="bib14" id="ref47">14</reflink>], [<reflink idref="bib30" id="ref48">30</reflink>]) that these chatbots are impressive but imperfect in their navigation of science pedagogy, thus underscoring the continuing need for faculty oversight and adjustment. The distinct contribution of this new study is the detailed characterization of chatbots' fulfillment of specific criteria in LO-related tasks. We hope that this characterization reminds readers of the centrality of LOs to curriculum (re)design ([<reflink idref="bib3" id="ref49">3</reflink>], [<reflink idref="bib4" id="ref50">4</reflink>]) and illustrates how chatbots can facilitate the development of LOs well-aligned to pedagogical priorities such as those emphasized in Vision & Change ([<reflink idref="bib1" id="ref51">1</reflink>]) and the Partnership for Undergraduate Life Sciences Education (PULSE) program ([<reflink idref="bib2" id="ref52">2</reflink>]).</p> <p>In both our preliminary trials and our formal investigations of <emph>questions A–C</emph>, we obtained the best chatbot outputs by balancing simplicity with specificity. Bombarding chatbots with huge amounts of information sometimes made the outputs worse; when one of us (M.M.L.) provided ChatGPT-4o with a comprehensive summary of course content by including slides, lecture scripts, and articles, the LOs produced were poor. In that case, our solution was to simplify the task by reducing the inputs to a single source of content. On the other hand, our final prompts were relatively long (∼250 words each; see appendix) because more guidance within the prompt led to better chatbot outputs. Simplifying the task of including appropriate context in LOs (<emph>criterion c3</emph> of <emph>questions A</emph> and <emph>B</emph>) with the "given X, do Y" format ([<reflink idref="bib7" id="ref53">7</reflink>]) seemed to help Claude achieve ≥80% success on <emph>criterion c3</emph> but did not lead to the same level of consistent success for ChatGPT or Gemini. Thus the ideal balance of simplicity and specificity will likely be highly context-dependent, so achieving it may require some trial and error. With that in mind, we encourage interested readers to use our prompts (see appendix) as useful starting points that they will probably want to adjust for their own purposes.</p> <p>In light of the chatbot outputs' numerous imperfections (sometimes reflecting poor clarity or poor LO alignment of the source materials rather than chatbot deficits per se), and the need to balance simplicity with specificity, we do not recommend that a course be (re)designed by prompting chatbots to bulk-process the entire course at once. While chatbots may be technically able to handle large volumes of input, we believe that the best course (re)designs will result from a more iterative and interactive process, in which smaller chunks (e.g., specific subtopics within a module) of chatbot output are sequentially reviewed and edited by instructors.</p> <p>This general recommendation can be applied to specific challenges such as that of creating TQTs ([<reflink idref="bib8" id="ref54">8</reflink>], [<reflink idref="bib10" id="ref55">10</reflink>]). Each TQT consists of a LO directly aligned to questions assessing the achievement of that LO. Since this study has demonstrated that chatbots can write LOs and write linked assessment questions reasonably well, we propose that the process of TQT development, which we have historically found to be worthwhile but time-intensive, can be expedited with frontier-level chatbots. For example, we have previously cautioned that an ideal TQT should not be "overly broad" nor "overly narrow" (Ref. [<reflink idref="bib9" id="ref56">9</reflink>], p. 210–211). To create a TQT that heeds this advice, choose a current LO, ask a chatbot to write several sample assessment questions for this LO, determine whether the range of questions is appropriate, and adjust the LO as needed.</p> <hd1 id="AN0183439570-20">Limitations of This Study</hd1> <p>This study has several obvious limitations, the most prominent of which is the authors' relative naivety with respect to chatbot technology. This limitation is also a strength in the sense that this paper may effectively speak to and resonate with readers who are similarly naive; however, AI experts would likely have even more success than we had in delegating tasks to chatbots. Therefore, despite our use of the best chatbots currently available for public use, our results are unlikely to represent the absolute peak of performance by these chatbots.</p> <p>Another key limitation of the study was that of the Yes/Maybe/No system for rating chatbots' compliance with our criteria for LO and assessment-question quality. We devised this system as a convenient ad hoc way to summarize our satisfaction as instructors with chatbots' outputs, and so our ratings are most useful as a readout of how close to "student-ready" we perceived these outputs to be. However, while we discussed a few ratings among ourselves, we did not conduct a formal calibration or validation of our rating system. Therefore, our ratings should not be treated as completely objective, unbiased metrics of chatbot performance.</p> <p>A third limitation of the study was that we did not extensively explore variations in chatbot output over different sessions and different time periods. Therefore, while the broad trends of our results (e.g., chatbots did not always agree with instructors regarding Bloom's taxonomy) should be meaningful, individual details (e.g., Gemini did not write good HOCS questions for Applied Exercise Physiology) should not be assumed stable and reproducible.</p> <p>Finally, since chatbot capabilities may continue to evolve rapidly ([<reflink idref="bib12" id="ref57">12</reflink>]), this study is limited in the sense that its results might soon become disconnected from the current state of chatbot capabilities. In light of this limitation, it is fair to ask whether there is value in reporting chatbots' capabilities at one particular moment in time, as we have done here. We believe that the answer is yes for at least two reasons. First, the already-impressive chatbot capabilities documented here and elsewhere may encourage otherwise reluctant faculty to consider boosting their efficiency in this way. Second, while the details of chatbot usage will continue to change, the general need for human oversight of chatbots is unlikely to evaporate anytime soon. In that context, our study may remain useful in its general illustration of how instructors can interact with chatbots to achieve desired goals.</p> <hd id="AN0183439570-21">DATA AVAILABILITY</hd> <p>Data will be made available upon reasonable request.</p> <hd id="AN0183439570-22">GRANTS</hd> <p>This study originated from collaborations between G.J.C. and M.M.L. that were facilitated by the Promoting Active Learning and Mentoring (PALM) Network [National Science Foundation Research Coordination Networks in Undergraduate Biology Education (RCN-UBE) Grant 1624200; principal investigator: Susan Wick of the University of Minnesota].</p> <hd id="AN0183439570-23">DISCLOSURES</hd> <p>No conflicts of interest, financial or otherwise, are declared by the authors.</p> <hd id="AN0183439570-24">AUTHOR CONTRIBUTIONS</hd> <p>G.J.C., M.D.F., and M.M.L. conceived and designed research; G.J.C., M.D.F., K.M.H., and M.M.L. performed experiments; G.J.C., M.D.F., K.M.H., and M.M.L. analyzed data; G.J.C., M.D.F., K.M.H., and M.M.L. interpreted results of experiments; G.J.C., M.D.F., and M.M.L. prepared figures; G.J.C. drafted manuscript; G.J.C., M.D.F., K.M.H., and M.M.L. edited and revised manuscript; G.J.C., M.D.F., K.M.H., and M.M.L. approved final version of manuscript.</p> <hd1 id="AN0183439570-25">APPENDIX</hd1> <p>Below are templates of chatbot prompts used in this project. Brackets indicate positions where instructors inserted context-specific details.</p> <hd1 id="AN0183439570-26">Question A: When Given Course Content, Can Chatbots Create LOs That Are Consistent with Five Best Practices in Writing LOs?</hd1> <p>Please disregard any previous context or conversations. Respond only to the following request. You are a professor teaching a <bold>[undergraduate student level,</bold> e.g.<bold>, first- and second-year or junior- and senior-level]</bold> course in <bold>[title of course,</bold> e.g.<bold>, Applied Exercise Physiology or Human Anatomy]</bold> at a <bold>[type of institution,</bold> e.g.<bold>, community college or regional</bold> 4-yr <bold>university]</bold>. The <bold>[content source materials,</bold> e.g.<bold>, lecture notes or PowerPoint slides]</bold> I provide are the primary source of information that students will learn and that you should use for the following task. Use straightforward language appropriate for <bold>[undergraduate student level]</bold>. Consider the following five criteria to be our best practices for writing a Learning Objective (LO), as elaborated by the attached Orr et al. 2022 paper and the associated instructor checklist:</p> <p></p> <p> <ephtml> <table border="0" width="95%"><tr><td valign="top">(1) </td><td colspan="5" valign="top"><p> The LO is clear, balancing conciseness with adequate detail.</p></td></tr><tr><td valign="top">(2) </td><td colspan="5" valign="top"><p> The LO uses an action verb corresponding to a visible performance.</p></td></tr><tr><td valign="top">(3) </td><td colspan="5" valign="top"><p> The LO specifies the context and conditions under which the action is performed.</p></td></tr><tr><td valign="top">(4) </td><td colspan="5" valign="top"><p> Achievement of the LO is measurable via multiple-choice questions.</p></td></tr><tr><td valign="top">(5) </td><td colspan="5" valign="top"><p> The LO is assigned to an appropriate revised Bloom's taxonomy level.</p></td></tr></table> </ephtml> </p> <p>Using the five criteria above, as well as the <bold>[content source materials]</bold>, create a Learning Objective for <bold>[course topic]</bold>. Your goal is to ensure that students can <bold>[revised Bloom's taxonomy level,</bold> e.g.<bold>, Remember or Apply]</bold> content using the <bold>[revised Bloom's taxonomy level,</bold> e.g.<bold>, Remember or Apply]</bold> construct from the revised Bloom's taxonomy as described on the following website: https://cft.vanderbilt.edu/guides-sub-pages/blooms-taxonomy/.</p> <hd1 id="AN0183439570-27">Question A3: Do Requests for Los in A "Given X, Do Y" Format Lead to Successful Specification of the Context and Conditions for the LO Action (criterion c3)?</hd1> <p>Please disregard any previous context or conversations. Respond only to the following request. You are a professor teaching a <bold>second-year</bold> course in <bold>Human Physiology</bold> at a <bold>community college</bold>. The <bold>chapter summary</bold> I provide is the primary source of information that students will learn and that you should use for the following task. Use straightforward language appropriate for <bold>sophomore-level students</bold>. Consider the following five criteria to be our best practices for writing a Learning Objective (LO), as elaborated by the attached Orr et al. 2022 paper and the associated instructor checklist:</p> <p></p> <p> <ephtml> <table border="0" width="95%"><tr><td valign="top">(1) </td><td colspan="5" valign="top"><p> The LO is clear, balancing conciseness with adequate detail.</p></td></tr><tr><td valign="top">(2) </td><td colspan="5" valign="top"><p> The LO uses an action verb corresponding to a visible performance.</p></td></tr><tr><td valign="top">(3) </td><td colspan="5" valign="top"><p> The LO specifies the context and conditions under which the action is performed. To provide such context and conditions, use the general format of "Given X, do Y." That is, on a written exam, when a student is given prompt X (which might be purely textual, or might include an image or diagram), the student should be able to respond in a way that demonstrates knowledge (Y). The desired response (Y) should NOT be a general dump of information, but instead should be specific to the context given (X).</p></td></tr><tr><td valign="top">(4) </td><td colspan="5" valign="top"><p> Achievement of the LO is measurable via multiple-choice questions.</p></td></tr><tr><td valign="top">(5) </td><td colspan="5" valign="top"><p> The LO is assigned to an appropriate revised Bloom's taxonomy level.</p></td></tr></table> </ephtml> </p> <p>Using the five criteria above, as well as the example LOs below and the attached <bold>chapter summary</bold>, create ten LOs for <bold>this chapter</bold> on the topic of <bold>[chapter topic]</bold> physiology. These LOs should cover a range of levels of the revised Bloom's taxonomy, as described on the following website: https://cft.vanderbilt.edu/guides-sub-pages/blooms-taxonomy/.</p> <p>Examples of high-quality LOs for a different organ system (respiratory):</p> <p></p> <p> <ephtml> <table border="0" width="95%"><tr><td valign="top">1. </td><td colspan="5" valign="top"><p> Given a flow of blood from point A to point B, predict whether the blood has gained, lost, or not changed its levels of O2 and/or CO2.</p></td></tr><tr><td valign="top">2. </td><td colspan="5" valign="top"><p> Given a hemoglobin percent saturation, convert it to an average number of O2 atoms per hemoglobin molecule, or vice versa.</p></td></tr><tr><td valign="top">3. </td><td colspan="5" valign="top"><p> Given an oxygen-hemoglobin dissociation curve and PO2's or percent saturations on the arterial and venous sides of a capillary bed, determine how much oxygen was picked up or dropped off in transit through that capillary bed.</p></td></tr><tr><td valign="top">4. </td><td colspan="5" valign="top"><p> Given data on a patient's arterial O2 saturation and/or hematocrit (HCT), determine whether their oxygen-carrying capacity is abnormally limited by anemia and/or a pulmonary problem.</p></td></tr><tr><td valign="top">5. </td><td colspan="5" valign="top"><p> Given information about air flow (into or out of lungs), lung pressure, thoracic cavity/lung volume, or respiratory muscle state (contracting or relaxing), make predictions about these other parameters.</p></td></tr><tr><td valign="top">6. </td><td colspan="5" valign="top"><p> Given a graph of lung volume versus. time, determine end-expiratory volume (EEV), end-inspiratory volume (EIV), residual volume (RV), or total lung capacity (TLC).</p></td></tr><tr><td valign="top">7. </td><td colspan="5" valign="top"><p> Given (graphical or numerical) spirometry data, estimate or calculate FEV1, FVC, RR, or TV.</p></td></tr><tr><td valign="top">8. </td><td colspan="5" valign="top"><p> Given (numerical or graphical) spirometry data, estimate or calculate FEV1/FVC ratio or minute ventilation (MV).</p></td></tr><tr><td valign="top">9. </td><td colspan="5" valign="top"><p> Given spirometry data, determine whether the data are consistent with obstructive pulmonary disease, restrictive pulmonary disease, both, or neither.</p></td></tr><tr><td valign="top">10. </td><td colspan="5" valign="top"><p> Given scenarios involving possible disturbances of arterial blood-gas homeostasis, make predictions about the body's responses.</p></td></tr></table> </ephtml> </p> <hd1 id="AN0183439570-28">Question B: Can Chatbots Convert a LO from One Level of Bloom's Taxonomy to Another?</hd1> <p>Please disregard any previous context or conversations. Respond only to the following request. You are a professor teaching a <bold>[undergraduate student level,</bold> e.g.<bold>, first- and second-year or junior- and senior-level]</bold> course in <bold>[title of course,</bold> e.g.<bold>, Applied Exercise Physiology or Human Anatomy]</bold> at a <bold>[type of institution,</bold> e.g.<bold>, community college or regional</bold> 4-yr <bold>university]</bold>. The <bold>[content source materials,</bold> e.g.<bold>, lecture notes or PowerPoint slides]</bold> I provide are the primary source of information that students will learn and that you should use for the following task. Use straightforward language appropriate for <bold>[undergraduate student level]</bold>. Consider the following five criteria to be our best practices for writing a Learning Objective (LO), as elaborated by the attached Orr et al. 2022 paper and the associated instructor checklist:</p> <p></p> <p> <ephtml> <table border="0" width="95%"><tr><td valign="top">(1) </td><td colspan="5" valign="top"><p> The LO is clear, balancing conciseness with adequate detail.</p></td></tr><tr><td valign="top">(2) </td><td colspan="5" valign="top"><p> The LO uses an action verb corresponding to a visible performance.</p></td></tr><tr><td valign="top">(3) </td><td colspan="5" valign="top"><p> The LO specifies the context and conditions under which the action is performed.</p></td></tr><tr><td valign="top">(4) </td><td colspan="5" valign="top"><p> Achievement of the LO is measurable via multiple-choice questions.</p></td></tr><tr><td valign="top">(5) </td><td colspan="5" valign="top"><p> The LO is assigned to an appropriate revised Bloom's taxonomy level.</p></td></tr></table> </ephtml> </p> <p>Using the five criteria above, as well as the <bold>[content source materials]</bold>, revise the following Learning Objective: "<bold>[fill in the LOCS LO generated by this chatbot for this topic in 1a]</bold>"</p> <p>Your job is to convert this Learning Objective from the <bold>[current level,</bold> e.g., <bold>Remember]</bold> level of the revised Bloom's taxonomy to the <bold>[target level,</bold> e.g., <bold>Analyze]</bold> level of the revised Bloom's taxonomy, as described on the following website: https://cft.vanderbilt.edu/guides-sub-pages/blooms-taxonomy/.</p> <hd1 id="AN0183439570-29">Question C: When Given LOs, Can Chatbots Create Assessment Questions That Meet Six Criteria of Quality?</hd1> <p>Please disregard any previous context or conversations. Respond only to the following request. You are a professor teaching a <bold>[undergraduate student level,</bold> e.g.<bold>, first- and second-year or junior- and senior-level]</bold> course in <bold>[title of course,</bold> e.g.<bold>, Applied Exercise Physiology or Human Anatomy]</bold> at a <bold>[type of institution,</bold> e.g.<bold>, community college or regional</bold> 4-yr <bold>university]</bold>. The <bold>[content source materials,</bold> e.g.<bold>, lecture notes or PowerPoint slides]</bold> I provide are the primary source of information that students will learn and that you should use for the following task.</p> <p>For a course that uses the attached <bold>[content source materials]</bold>, please provide a multiple-choice question that assesses fulfillment of the following Learning Objective: <bold>[insert LO Previously Provided by Chatbot]</bold></p> <p>This Learning Objective should be assessed at the <bold>[insert Bloom level of previous LO: Remember, Understand, Apply, or Analyze]</bold> level of the revised Bloom's taxonomy as described on the following website: https://cft.vanderbilt.edu/guides-sub-pages/blooms-taxonomy/.</p> <p>Your multiple-choice question should include three to five answer choices. You should identify the correct answer and provide a rationale for why the correct answer is correct and why the other answer choices are incorrect.</p> <p>Your multiple-choice question should also meet the following criteria for best practices in writing multiple-choice questions:</p> <p></p> <p> <ephtml> <table border="0" width="95%"><tr><td valign="top">(1) </td><td colspan="5" valign="top"><p> The assessment question should match the content of the Learning Objective provided.</p></td></tr><tr><td valign="top">(2) </td><td colspan="5" valign="top"><p> The assessment question should match the revised Bloom's taxonomy level of the Learning Objective provided.</p></td></tr><tr><td valign="top">(3) </td><td colspan="5" valign="top"><p> The question stem should be meaningful by itself and should present a definite problem in the form of a question or a partial sentence.</p></td></tr><tr><td valign="top">(4) </td><td colspan="5" valign="top"><p> All answer choices should be reasonably concise and similar in length and format.</p></td></tr><tr><td valign="top">(5) </td><td colspan="5" valign="top"><p> All answer choices should be plausible.</p></td></tr><tr><td valign="top">(6) </td><td colspan="5" valign="top"><p> One answer choice should clearly be right, and the other answer choices should clearly be wrong.</p></td></tr></table> </ephtml> </p> <ref id="AN0183439570-30"> <title> REFERENCES </title> <blist> <bibl id="bib1" idref="ref1" type="bt">1</bibl> <bibtext> Brewer CA, Smith D. Vision and Change in Undergraduate Biology Education: a Call to Action. Washington, DC: American Association for the Advancement of Science, 2011. https://<ulink href="http://www.aaas.org/sites/default/files/content%5ffiles/VC%5freport.pdf">www.aaas.org/sites/default/files/content%5ffiles/VC%5freport.pdf</ulink> [2024 Aug 10].Google Scholar</bibtext> </blist> <blist> <bibl id="bib2" idref="ref2" type="bt">2</bibl> <bibtext> Stavrianeas S, Bangera G, Bronson C, Byers S, Davis W, DeMarais A, Fitzhugh G, Linder N, Liston C, McFarland J, Otto J, Pape-Lindstrom P, Pollock C, Reiness CG, Offerdahl EG. Empowering faculty to initiate STEM education transformation: efficacy of a systems thinking approach. PLoS One 17: e0271123, 2022. doi:10.1371/journal.pone.0271123.Crossref | PubMed | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibl id="bib3" idref="ref3" type="bt">3</bibl> <bibtext> Orr RB, Gormally C, Brickman P. A road map for planning course transformation using learning objectives. CBE Life Sci Educ 23: es4, 2024. doi:10.1187/cbe.23-06-0114.Crossref | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibl id="bib4" idref="ref4" type="bt">4</bibl> <bibtext> Orr RB, Csikari MM, Freeman S, Rodriguez MC. Writing and using learning objectives. CBE Life Sci Educ 21: fe3, 2022. doi:10.1187/cbe.22-04-0073. Crossref | PubMed | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibl id="bib5" idref="ref6" type="bt">5</bibl> <bibtext> Pape-Zambito DA, Mostrom AM. Improving teaching through triadic course alignment. J Microbiol Biol Educ 19: 1642, 2018. doi:10.1128/jmbe.v19i3.1642.Crossref | PubMed | Google Scholar</bibtext> </blist> <blist> <bibl id="bib6" idref="ref7" type="bt">6</bibl> <bibtext> Hennessey KM, Freeman S. Nationally endorsed learning objectives to improve course design in introductory biology. PLoS One 19: e0308545, 2024. doi:10.1371/journal.pone.0308545.Crossref | PubMed | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibl id="bib7" idref="ref8" type="bt">7</bibl> <bibtext> Crowther GJ, VanHeel VL, Gradwell SD, Self CJ, Rompolski KL. General skills amidst the details: alternative learning objectives and a framework of competencies for human anatomy. Adv Physiol Educ 48: 799–807, 2024. doi:10.1152/advan.00076.2024.Link | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibl id="bib8" idref="ref9" type="bt">8</bibl> <bibtext> Crowther GJ, Wiggins BL, Jenkins LD. Testing in the age of active learning: test question templates help to align activities and assessments. HAPS Educ 24: 592–599, 2020. doi:10.21692/haps.2020.006.Crossref | Google Scholar</bibtext> </blist> <blist> <bibl id="bib9" idref="ref10" type="bt">9</bibl> <bibtext> Crowther GJ, Knight TA. Using test question templates to teach physiology core concepts. Adv Physiol Educ 47: 202–214, 2023. doi:10.1152/advan.00024.2022.Link | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Evans DP, Jenkins LD, Crowther GJ. Student perceptions of a framework for facilitating transfer from lessons to exams, and the relevance of this framework to published lessons. J Microbiol Biol Educ 24: e00200-22, 2023. doi:10.1128/jmbe.00200-22.Crossref | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Adamopoulou E, Moussiades L. An overview of chatbot technology. In: IFIP International Conference on Artificial Intelligence Applications and Innovations 2020. Cham, Switzerland: Springer, 2020, p. 373–383.Google Scholar</bibtext> </blist> <blist> <bibtext> Mollick E. Co-Intelligence: Living and Working with AI. New York: Penguin, 2024.Google Scholar</bibtext> </blist> <blist> <bibtext> Labadze L, Grigolia M, Machaidze L. Role of AI chatbots in education: systematic literature review. Int J Educ Technol High Educ 20: 56, 2023. doi:10.1186/s41239-023-00426-1.Crossref | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Crowther GJ, Sankar U, Knight LS, Myers DL, Patton KT, Jenkins LD, Knight TA. Chatbot responses suggest that hypothetical biology questions are harder than realistic ones. J Microbiol Biol Educ 24: e00153-23, 2023. doi:10.1128/jmbe.00153-23.Crossref | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Garabet R, Mackey BP, Cross J, Weingarten M. ChatGPT-4 performance on USMLE Step 1 style questions and its implications for medical education: a comparative study across systems and disciplines. Med Sci Educ 34: 145–152, 2024. doi:10.1007/s40670-023-01956-z. Crossref | PubMed | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Balabdaoui F, Dittmann-Domenichini N, Grosse H, Schlienger C, Kortemeyer G. A survey on students' use of AI at a technical university. Discov Educ 3: 51, 2024. doi:10.1007/s44217-024-00136-4.Crossref | Google Scholar</bibtext> </blist> <blist> <bibtext> Hitch D. Artificial Intelligence augmented qualitative analysis: the way of the future?Qual Health Res 34: 595–606, 2024. doi:10.1177/10497323231217392.Crossref | PubMed | Google Scholar</bibtext> </blist> <blist> <bibtext> Clark TM, Fhaner M, Stoltzfus M, Queen MS. Using ChatGPT to support lesson planning for the historical experiments of Thomson, Millikan, and Rutherford. J Chem Educ 101: 1992–1999, 2024. doi:10.1021/acs.jchemed.4c00200.Crossref | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Pardos ZA, Bhandari S. ChatGPT-generated help produces learning gains equivalent to human tutor-authored help on mathematics skills. PLoS One 19: e0304013, 2024. doi:10.1371/journal.pone.0304013.Crossref | PubMed | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Nasution NE. Using artificial intelligence to create biology multiple choice questions for higher education. Agricult Environ Educ 2: em002, 2023. doi:10.29333/agrenvedu/13071.Crossref | Google Scholar</bibtext> </blist> <blist> <bibtext> Allagui B. Chatbot feedback on students' writing: typology of comments and effectiveness. In: International Conference on Computational Science and Its Applications. Cham, Switzerland: Springer, 2023, p. 377–384. doi:10.1007/978-3-031-37129-5_31.Google Scholar</bibtext> </blist> <blist> <bibtext> Spiro K. The difference between "learning objectives" and "learning outcomes." EasyGenerator.com. https://www.easygenerator.com/en/blog/how-to/learning-objectives-vs-learning-outcomes/[2024 Dec 3].Google Scholar</bibtext> </blist> <blist> <bibtext> Krathwohl DR. A revision of Bloom's taxonomy: an overview. Theory Pract 41: 212–218, 2002. doi:10.1207/s15430421tip4104_2.Crossref | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Brame C. Writing good multiple choice questions. Nashville, TN: Vanderbilt University Center for Teaching, 2013. https://cft.vanderbilt.edu/guides-sub-pages/writing-good-multiple-choice-test-questions/[2024 Jul 11].Google Scholar</bibtext> </blist> <blist> <bibtext> Albano AD, Brickman P, Csikari M, Julian D, Orr RB, Rodriguez MC. Integrating Testing and Learning. Washington, DC: HHMI Biointeractive, 2020.Google Scholar</bibtext> </blist> <blist> <bibtext> Zamfirescu-Pereira JD, Wong RY, Hartmann B, Yang Q. Why Johnny can't prompt: how non-AI experts try (and fail) to design LLM prompts. In: Proceedings of the 2023 CHI Conference on Human Factors in Computing Systems. New York: Association for Computing Machinery, 2023. doi:10.1145/3544548.3581388.Google Scholar</bibtext> </blist> <blist> <bibtext> Reinhard P, Li MM, Peters C, Leimeister JM. Let employees train their own chatbots: design of generative AI-enabled delegation systems. In: European Conference on Information Systems (ECIS), Paphos, Cyprus. Charlottesville, VA: Association for Information Systems: 2024. doi:10.2139/ssrn.4807392.Google Scholar</bibtext> </blist> <blist> <bibtext> Momsen JL, Long TM, Wyse SA, Ebert-May D. Just the facts? Introductory undergraduate biology courses focus on low-level cognitive skills. CBE Life Sci Educ 9: 435–440, 2010. doi:10.1187/cbe.10-01-0001.Crossref | PubMed | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Semsar K, Casagrand J. Bloom's dichotomous key: a new tool for evaluating the cognitive difficulty of assessments. Adv Physiol Educ 41: 170–177, 2017. doi:10.1152/advan.00101.2016.Link | Web of Science | Google Scholar</bibtext> </blist> <blist> <bibtext> Al Husaeni DF, Haristiani N, Wahyudin W, Rasim R. Chatbot artificial intelligence as educational tools in science and engineering education: a literature review and bibliometric mapping analysis with its advantages and disadvantages. ASEAN J Sci Eng 4: 93–118, 2022. doi:10.17509/ajse.v4i1.67429.Crossref | Google Scholar</bibtext> </blist> </ref> <aug> <p>By Gregory J. Crowther; Merrill D. Funk; Kelly M. Hennessey and Marcus M. Lawrence</p> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib10" firstref="ref11"></nolink> <nolink nlid="nl2" bibid="bib11" firstref="ref13"></nolink> <nolink nlid="nl3" bibid="bib12" firstref="ref14"></nolink> <nolink nlid="nl4" bibid="bib13" firstref="ref15"></nolink> <nolink nlid="nl5" bibid="bib14" firstref="ref16"></nolink> <nolink nlid="nl6" bibid="bib15" firstref="ref17"></nolink> <nolink nlid="nl7" bibid="bib16" firstref="ref18"></nolink> <nolink nlid="nl8" bibid="bib17" firstref="ref19"></nolink> <nolink nlid="nl9" bibid="bib18" firstref="ref20"></nolink> <nolink nlid="nl10" bibid="bib19" firstref="ref21"></nolink> <nolink nlid="nl11" bibid="bib20" firstref="ref22"></nolink> <nolink nlid="nl12" bibid="bib21" firstref="ref23"></nolink> <nolink nlid="nl13" bibid="bib22" firstref="ref26"></nolink> <nolink nlid="nl14" bibid="bib23" firstref="ref31"></nolink> <nolink nlid="nl15" bibid="bib24" firstref="ref32"></nolink> <nolink nlid="nl16" bibid="bib25" firstref="ref33"></nolink> <nolink nlid="nl17" bibid="bib26" firstref="ref37"></nolink> <nolink nlid="nl18" bibid="bib27" firstref="ref39"></nolink> <nolink nlid="nl19" bibid="bib28" firstref="ref43"></nolink> <nolink nlid="nl20" bibid="bib29" firstref="ref44"></nolink> <nolink nlid="nl21" bibid="bib30" firstref="ref48"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1464069
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Frontier Model Chatbots Can Help Instructors Create, Improve, and Use Learning Objectives
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Gregory+J%2E+Crowther%22">Gregory J. Crowther</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0530-9130">0000-0003-0530-9130</externalLink>)<br /><searchLink fieldCode="AR" term="%22Merrill+D%2E+Funk%22">Merrill D. Funk</searchLink><br /><searchLink fieldCode="AR" term="%22Kelly+M%2E+Hennessey%22">Kelly M. Hennessey</searchLink><br /><searchLink fieldCode="AR" term="%22Marcus+M%2E+Lawrence%22">Marcus M. Lawrence</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-7106-574X">0000-0001-7106-574X</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Advances+in+Physiology+Education%22"><i>Advances in Physiology Education</i></searchLink>. 2025 49(1):219-229.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: American Physiological Society. 9650 Rockville Pike, Bethesda, MD 20814-3991. Tel: 301-634-7164; Fax: 301-634-7241; e-mail: webmaster@the-aps.org; Web site: https://www.physiology.org/journal/advances
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 11
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF), Research Coordination Networks in Undergraduate Biology Education (RCN-UBE)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: 1624200
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research<br />Tests/Questionnaires
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Objectives%22">Learning Objectives</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Curriculum+Design%22">Curriculum Design</searchLink><br /><searchLink fieldCode="DE" term="%22Curriculum+Implementation%22">Curriculum Implementation</searchLink><br /><searchLink fieldCode="DE" term="%22Best+Practices%22">Best Practices</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Quality%22">Educational Quality</searchLink><br /><searchLink fieldCode="DE" term="%22Undergraduate+Study%22">Undergraduate Study</searchLink><br /><searchLink fieldCode="DE" term="%22Exercise+Physiology%22">Exercise Physiology</searchLink><br /><searchLink fieldCode="DE" term="%22Anatomy%22">Anatomy</searchLink><br /><searchLink fieldCode="DE" term="%22Perceptual+Motor+Learning%22">Perceptual Motor Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Taxonomy%22">Taxonomy</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Criteria%22">Evaluation Criteria</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1152/advan.00159.2024
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1043-4046<br />1522-1229
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Learning objectives (LOs) are a pillar of course design and execution and thus a focus of curricular reforms. This study explored the extent to which the creation and usage of LOs might be facilitated by three leading chatbots: ChatGPT-4o, Claude 3.5 Sonnet, and Google Gemini Advanced. We posed three main questions, as follows: "question A": when given course content, can chatbots create LOs that are consistent with five best practices in writing LOs?; "question B": when given LOs for a low level of the revised Bloom's taxonomy, can chatbots convert them to a higher level?; and "question C": when given LOs, can chatbots create assessment questions that meet six criteria of quality? We explored these questions in the context of four undergraduate courses: Applied Exercise Physiology, Human Anatomy, Human Physiology, and Motor Learning. According to instructor ratings, chatbots had a >70% success rate on most individual criteria for questions A--C. However, chatbots' "difficulties" with a few criteria (e.g., provision of appropriate context for an LO's action, assignment of an appropriate revised Bloom's taxonomy level) meant that, overall, only 38.3% of chatbot outputs fully met all criteria and thus were possibly ready for use with students. Our findings thus underscore the continuing need for instructor oversight of chatbot outputs but also illustrate chatbots' potential to expedite the design and improvement of LOs and LO-related curricular materials such as test question templates (TQTs), which directly align LOs with assessment questions.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1464069
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1464069
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1152/advan.00159.2024
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 11
        StartPage: 219
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Learning Objectives
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Curriculum Design
        Type: general
      – SubjectFull: Curriculum Implementation
        Type: general
      – SubjectFull: Best Practices
        Type: general
      – SubjectFull: Evaluation Methods
        Type: general
      – SubjectFull: Educational Quality
        Type: general
      – SubjectFull: Undergraduate Study
        Type: general
      – SubjectFull: Exercise Physiology
        Type: general
      – SubjectFull: Anatomy
        Type: general
      – SubjectFull: Perceptual Motor Learning
        Type: general
      – SubjectFull: Taxonomy
        Type: general
      – SubjectFull: Evaluation Criteria
        Type: general
    Titles:
      – TitleFull: Frontier Model Chatbots Can Help Instructors Create, Improve, and Use Learning Objectives
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Gregory J. Crowther
      – PersonEntity:
          Name:
            NameFull: Merrill D. Funk
      – PersonEntity:
          Name:
            NameFull: Kelly M. Hennessey
      – PersonEntity:
          Name:
            NameFull: Marcus M. Lawrence
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1043-4046
            – Type: issn-electronic
              Value: 1522-1229
          Numbering:
            – Type: volume
              Value: 49
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Advances in Physiology Education
              Type: main
ResultId 1