A Mile High and an Inch Deep: Exploring ChatGPT as a Mathematics Curriculum Development Tool

Saved in:
Bibliographic Details
Title: A Mile High and an Inch Deep: Exploring ChatGPT as a Mathematics Curriculum Development Tool
Language: English
Authors: Amanda Gantt Sawyer (ORCID 0000-0001-5532-9047), Zareen Gul Aga
Source: School Science and Mathematics. 2026 126(1):60-74.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 15
Publication Date: 2026
Sponsoring Agency: National Science Foundation (NSF)
Contract Number: 2237151
Document Type: Journal Articles
Reports - Research
Education Level: Elementary Education
Descriptors: Artificial Intelligence, Technology Uses in Education, Curriculum Development, Mathematics Curriculum, Elementary School Mathematics, Curriculum Implementation, Common Core State Standards, Difficulty Level, Instructional Program Divisions
DOI: 10.1111/ssm.18348
ISSN: 0036-6803
1949-8594
Abstract: Artificial Intelligence (AI) technology has issues, including inherent bias and inaccurate information, yet it has not stifled individuals' use of this technology in the mathematics classroom. Thus, we conducted a document analysis to investigate ChatGPT, a generative AI chatbot, constructing elementary mathematical tasks to determine how mathematics teacher educators (MTEs) can support this implementation. We analyzed 191 text responses constructed by ChatGPT, each corresponding to a task created for one of the elementary Common Core mathematical standards from kindergarten to fifth Grade to answer the research questions: What are the associated levels of cognitive demand of the ChatGPT's text responses, is there a statistically significant difference between lower-level mathematical tasks and higher-level mathematical tasks, and what are the common characteristics of AI's tasks? We found significantly higher-level cognitive demands across the tasks, but many activities were repetitive and inappropriate for the specified grade level. Finally, we explored how MTEs could support teachers using this tool in their classrooms using Mathematical Critical Curation Questions.
Abstractor: As Provided
Entry Date: 2026
Accession Number: EJ1495594
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHIhdrK71Wl5AAoA7xwG998AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDAAv9TcB62SnytsWDgIBEICBmzyMDgt91AxR6t_rd6Mkw5Om2gfxmCxPr62vI0xizXARktKcT7vewj3ACrpansKSg0s6dgBrrD47DcMYz0FjIbtwzjm0Yv0t3R0dkaiIWz2Yr9FhCV3UfUFMhFirqQR8o7jRDzP8g_0RRdalYNAaeNEg3iEz5R9wmUlZSUfLfB-59MWpxuyHHt6ptYxA1qU5jutIRM2H_oyCsos9
Text:
  Availability: 1
  Value: <anid>AN0191298896;ssm01feb.26;2026Feb04.04:59;v2.2.500</anid> <title id="AN0191298896-1">A mile high and an inch deep: Exploring ChatGPT as a mathematics curriculum development tool </title> <p>Artificial Intelligence (AI) technology has issues, including inherent bias and inaccurate information, yet it has not stifled individuals' use of this technology in the mathematics classroom. Thus, we conducted a document analysis to investigate ChatGPT, a generative AI chatbot, constructing elementary mathematical tasks to determine how mathematics teacher educators (MTEs) can support this implementation. We analyzed 191 text responses constructed by ChatGPT, each corresponding to a task created for one of the elementary Common Core mathematical standards from kindergarten to fifth Grade to answer the research questions: What are the associated levels of cognitive demand of the ChatGPT's text responses, is there a statistically significant difference between lower‐level mathematical tasks and higher‐level mathematical tasks, and what are the common characteristics of AI's tasks? We found significantly higher‐level cognitive demands across the tasks, but many activities were repetitive and inappropriate for the specified grade level. Finally, we explored how MTEs could support teachers using this tool in their classrooms using Mathematical Critical Curation Questions.</p> <p>Keywords: artificial intelligence; ChatGPT; mathematics curriculum development</p> <p> <emph>Artificial Intelligence (AI) technology has issues, including inherent bias and inaccurate information, yet it has not stifled individuals' use of this technology in the mathematics classroom. Thus, we conducted a document analysis to investigate ChatGPT, a generative AI chatbot, constructing elementary mathematical tasks to determine how mathematics teacher educators (MTEs) can support this implementation. We analyzed 191 text responses constructed by ChatGPT, each corresponding to a task created for one of the elementary Common Core mathematical standards from Kindergarten to Fifth Grade to answer the research questions: What are the associated levels of cognitive demand of the ChatGPT's text responses, is there a statistically significant difference between lower‐level mathematical tasks and higher‐level mathematical tasks, and what are the common characteristics of AI's tasks? We found significantly higher‐level cognitive demands across the tasks, but many activities were repetitive and inappropriate for the specified grade level</emph>.</p> <p>Teachers and students utilize Artificial Intelligence (AI) chatbots to support writing emails, creating lesson plans, and answering homework questions. Artificial Intelligence (AI) is "a machine‐based system that can, for a given set of human‐defined objectives, make predictions, recommendations or decisions influencing real or virtual environments" (US Department of State, [<reflink idref="bib29" id="ref1">29</reflink>], p. 21). Researchers identified Artificial Intelligence (AI) technology as having inherent bias and inaccurate information (Wu, [<reflink idref="bib32" id="ref2">32</reflink>]), yet it has not stifled individuals' use of this technology in the mathematics classroom (Sawyer, [<reflink idref="bib18" id="ref3">18</reflink>]). In response to this issue, multiple organizations created position statements explaining their views on using AI tools in the classroom (NCTM, [<reflink idref="bib12" id="ref4">12</reflink>]; USDOE, [<reflink idref="bib29" id="ref5">29</reflink>]). For example, The United States Department of Education ([<reflink idref="bib29" id="ref6">29</reflink>]) identified seven areas of collaboration to learn more about how we can effectively and meaningfully use AI in the classroom: emphasizing humans‐in‐the‐loop, aligning AI models to a shared vision for education, designing AI using modern learning principles, prioritizing strengthening trust, informing and involving educators, focusing research and development on addressing context and enhancing trust and safety, and developing education‐specific guidelines and guardrails (USDOE, [<reflink idref="bib29" id="ref7">29</reflink>], paragraph 3).</p> <p>The National Council of Teachers of Mathematics (NCTM) provided a position statement identifying that AI should not be ignored. However, the statement suggests that mathematics teacher educators (MTEs) should best research how to support mathematics teaching in today's classrooms. NCTM stated, "Educators need to be involved in developing and testing AI tools in math education to stay up to date with current AI trends to best prepare students for an AI future" (NCTM, [<reflink idref="bib12" id="ref8">12</reflink>], paragraph 1). Because of these calls, we tested an AI chatbot for this paper to develop specific guidance for MTEs to support mathematics teachers' effective tool use. The need for this work is embedded in the increasing use of ChatGPT as a curriculum development tool, which is problematic because ChatGPT may share incorrect mathematical information (Sawyer & Aga, [<reflink idref="bib19" id="ref9">19</reflink>]). This warrants caution when using ChatGPT for curriculum development. While this paper focuses on mathematics, the implications could possibly impact other STEM education programs.</p> <p>In past research investigations, teachers searched for resources online from virtual resource pools like Teachers Pay Teachers (TpT) and Pinterest (Shapiro et al., [<reflink idref="bib24" id="ref10">24</reflink>]). Still, AI tools have created a whole new nonhuman curriculum developer. However, we need to find out if techniques designed to support the criticality of online resources would remain relevant for resources created by nonhuman curriculum developers. Therefore, we investigated the quality of resources available from an AI chatbot without training to develop specific recommendations for MTEs on supporting teachers using these tools.</p> <p>Most states use some version of Common Core or a modification of Common Core (Schmidt & Bush, [<reflink idref="bib22" id="ref11">22</reflink>]). Thus, we decided to investigate AI's construction of activities using Common Core standards from kindergarten to fifth grade to determine what MTEs need to know to help facilitate an effective mathematics curriculum.</p> <p>This investigation focused on ChatGPT 3.5, arguably the most popular free AI generative chatbot at that time (Tech Desk, [<reflink idref="bib28" id="ref12">28</reflink>]). We analyzed 191 ChatGPT‐generated tasks and coded them for the level of cognitive demand.</p> <hd id="AN0191298896-2">ChatGPT</hd> <p>In June 2020, OpenAI introduced its own generative AI chatbot, Chat Generative Pre‐trained Transformer (ChatGPT), which launched publicly on November 30, 2022 (OpenAI, [<reflink idref="bib14" id="ref13">14</reflink>]). A generative AI chatbot can simulate human responses and reply to users' questions in real time using natural language processing (NLP) (Marr, [<reflink idref="bib10" id="ref14">10</reflink>]).</p> <p>Other companies have constructed generative AI chatbots like Magic School AI, Google's Gemini, and Jasper AI. However, ChatGPT has become arguably one of the most popular, specifically on college campuses (Tech Desk, [<reflink idref="bib28" id="ref15">28</reflink>]). ChatGPT uses algorithms to build responses word by word to their users' questions by finding the best matches from its data, about 175 billion parameters of knowledge (Marr, [<reflink idref="bib10" id="ref16">10</reflink>]). This platform is ever‐growing, with the system being updated regularly. When this investigation occurred (October of 2023), ChatGPT identified that its data was collected before September 2021. As of September 2024, it identified its last training cut‐off as April 2023. The program also learns from its users and can be trained to create responses based on specific criteria provided by the user. Since ChatGPT learns from the questions a user asks, creating responses based on feedback from the user, we collected data for this investigation from a newly constructed account.</p> <hd id="AN0191298896-3">LITERATURE REVIEW</hd> <p></p> <hd id="AN0191298896-4">Artificial intelligence chatbots</hd> <p>Research into ChatGPT identified a need for awareness around specific issues. ChatGPT could not use discretion when selecting its answers; it can only create responses by building each sentence word by word as specified by its training (OpenAI, [<reflink idref="bib14" id="ref17">14</reflink>]). Therefore, ChatGPT will state items as facts without backing from other sources (Wu, [<reflink idref="bib32" id="ref18">32</reflink>]). Once that data has been constructed, it goes into the database as new facts, with ChatGPT digging itself into a "misinformation ouroboros" (Vincent, [<reflink idref="bib30" id="ref19">30</reflink>], para. 1). This hallucination of data has been identified as an issue, and researchers stated:</p> <p>Fixing this issue is challenging, as (<reflink idref="bib1" id="ref20">1</reflink>) during RL training, there is currently no source of truth; (<reflink idref="bib2" id="ref21">2</reflink>) training the model to be more cautious causes it to decline questions that it can answer correctly; and (<reflink idref="bib3" id="ref22">3</reflink>) supervised training misleads the model because the ideal answer depends on what the model knows, rather than what the human demonstrator knows. (Schulman et al., [<reflink idref="bib23" id="ref23">23</reflink>], para. 7)</p> <p>ChatGPT takes information from its database and selects its response based on specific details formed from its algorithm, which we need to learn more about. Therefore, it cannot determine which information is false or true (Newton, [<reflink idref="bib13" id="ref24">13</reflink>]).</p> <p>ChatGPT has specific implications in the mathematics classroom because ChatGPT will construct incorrect mathematical answers and portray them to its users as facts (Maiorca et al., [<reflink idref="bib9" id="ref25">9</reflink>]). For example, ChatGPT has difficulty multiplying large numbers (Sawyer & Aga, [<reflink idref="bib19" id="ref26">19</reflink>]), and to have the AI tool correct its mistake, prompt engineering techniques are needed. Prompt engineering is a strategy to support AI in understanding the user's question. Korzynski et al. ([<reflink idref="bib7" id="ref27">7</reflink>]) explained that prompt engineering includes providing AI with four specific information forms to understand your task better: context, instructions, input data, and expected output format. In our example, the new prompt allowed ChatGPT to find its error, but these techniques are only sometimes effective in reaching a mathematically valid solution.</p> <p>In investigating preservice teachers' use of ChatGPT, we found that preservice teachers were overconfident in the AI's abilities; thus, they were less critical when selecting resources from ChatGPT (Sawyer, [<reflink idref="bib18" id="ref28">18</reflink>]). Preservice teachers viewed ChatGPT as a "calculator for mathematics resources" (Sawyer, [<reflink idref="bib18" id="ref29">18</reflink>], p. 21). Therefore, they are more confident in the AI tool's response than their own understanding, which can be problematic when we know ChatGPT creates inaccurate information (Maiorca et al., [<reflink idref="bib9" id="ref30">9</reflink>]). We investigated ChatGPT to determine the quality of mathematics tasks created and what warnings teachers should be aware of while using this tool.</p> <hd id="AN0191298896-5">Virtual resource pools</hd> <p>Teachers believe in using new technologies in the classroom. Thus, many researchers investigated teachers' resources based on technology trends (Bill and Melinda Gates Foundation, [<reflink idref="bib1" id="ref31">1</reflink>]; Shapiro et al., [<reflink idref="bib24" id="ref32">24</reflink>]). Researchers found that teachers needed more time to create their activities. Thus, 91% of the teachers identified finding resources online as a way to save planning time (Bill and Melinda Gates Foundation, [<reflink idref="bib1" id="ref33">1</reflink>]). However, even before the pandemic, 99% of elementary mathematics teachers surveyed reported using online teacher‐sharing websites in their classrooms regardless of years of teaching experience, with, on average, mathematics teachers searching online weekly for elementary mathematics activities (Shapiro et al., [<reflink idref="bib25" id="ref34">25</reflink>]). Since the pandemic, Carpenter et al. ([<reflink idref="bib4" id="ref35">4</reflink>]) found that teachers regularly use online resource‐sharing websites, only changing how they select classroom activities.</p> <p>Researchers investigated where these teachers are going to find activities, and in 2014, teachers identified YouTube (76%), Discovery.com (59%), Scholastic.com (56%), PBS.org (55%), and Pinterest (46%) as the most common websites (Bill and Melinda Gates Foundation, [<reflink idref="bib1" id="ref36">1</reflink>]). Shapiro et al. ([<reflink idref="bib25" id="ref37">25</reflink>]) recently found that 89% of elementary teachers use TpT, followed by 74% using Pinterest. However, researchers identified that these websites will change again because virtual resource pools are an ever‐changing landscape contingent upon the popularity of websites at specific times (Sawyer et al., [<reflink idref="bib17" id="ref38">17</reflink>]).</p> <p>Researchers have also investigated the quality of the materials found in virtual resource pools (Shapiro et al., [<reflink idref="bib24" id="ref39">24</reflink>]; Shelton & Archambault, [<reflink idref="bib26" id="ref40">26</reflink>]). Dick et al. ([<reflink idref="bib5" id="ref41">5</reflink>]) decided to examine the quality of the materials using Stein and Smith's ([<reflink idref="bib27" id="ref42">27</reflink>]) Task Analysis Guide Framework (TAG).</p> <p>Stein and Smith's ([<reflink idref="bib27" id="ref43">27</reflink>]) TAG categorizes resources into four levels of cognitive demand based on how much thinking is involved in the mathematics task. The lowest level of cognitive demand is Memorization, where the tasks require students to state facts without giving any meaning behind the understanding. The next level is Procedures without Connections, which has students following an algorithmic process to solve a mathematical procedure but does not require any connection to the mathematical concept. Next is Procedures with Connections, where the task requires students to follow an algorithm, but mathematical connections are made through discussion, representations, or differing strategies. Doing mathematics is the highest level of cognitive demand, which requires supporting students' self‐regulation through a problem‐solving process that connects to mathematical ideas (Stein & Smith, [<reflink idref="bib27" id="ref44">27</reflink>]).</p> <p>The level of cognitive demand has been studied by multiple researchers (Stein & Smith, [<reflink idref="bib27" id="ref45">27</reflink>]; Wilhelm, [<reflink idref="bib31" id="ref46">31</reflink>]). Stein and Smith ([<reflink idref="bib27" id="ref47">27</reflink>]) identified that the level can change based on how a teacher sets up the task and how a student interprets the implementation. Stein and Smith ([<reflink idref="bib27" id="ref48">27</reflink>]) stated that "the nature of tasks often changes as they pass from one phase to another" (p. 270). Therefore, the level of cognitive demand for a task will change based on individuals' perceptions. Stein and Smith ([<reflink idref="bib27" id="ref49">27</reflink>]) also identified that teachers should use various strategies. Since different lessons have different goals, multiple levels of cognitive demand are needed in the classroom.</p> <p>Wilhelm ([<reflink idref="bib31" id="ref50">31</reflink>]) also investigated the TAG framework and viewed mathematical tasks as having three stages. The first stage was the written curricular or instructional material, the second stage was a setup and implementation by the teacher and the student in the classroom, and the final product was the student learning. Shapiro et al. ([<reflink idref="bib25" id="ref51">25</reflink>]) and Sawyer et al. ([<reflink idref="bib17" id="ref52">17</reflink>]) adopted Wilhelm's ([<reflink idref="bib31" id="ref53">31</reflink>]) interpretation to investigate virtual resource pools. They focused on the first stage of the mathematical task, which is the written instructional materials. Wilhelm states, "The cognitive demand of the selected task sets the stage for the cognitive demand over the course of the lesson" (2014, p. 640). We know that the initial phase will influence the final impact on student learning. Still, most tasks from virtual resource pools like Pinterest are lower‐level cognitive demand, with Doing Mathematics accounting for less than one percent (Sawyer et al., [<reflink idref="bib17" id="ref54">17</reflink>]).</p> <p>Dick et al. ([<reflink idref="bib5" id="ref55">5</reflink>]) also investigated the top 1000 resources (500 free vs. 500 less than five dollars) from TpT to determine the resources' level of cognitive demand. Free resources had a significantly lower cognitive demand than resources for a price. As identified in Figure 1, 50% of the activities were considered to have a higher level of cognitive demand for free resources. However, 69% of the tasks below five dollars were considered higher‐level cognitive demand. They also found that the level of cognitive demand was significantly lower in specific Common Core domains, such as Geometry, but higher in areas like Numbers and Operations with Fractions.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0001.jpg" title="1 Level of cognitive demand of TpT free versus <$5." /> </p> <p></p> <p>Dick et al. ([<reflink idref="bib5" id="ref56">5</reflink>]) also researched Teachers Pay Teachers resources created for kindergarten through second‐grade students and then compared what was found to resources created for third through fifth‐grade students (see Figure 2). For free resources available for K‐2, 40% could be identified as higher‐level cognitive demand, but 50% of K‐2 resources for <$5 were identified as higher‐level cognitive demand. Forty‐nine percent of free resources constructed for 3rd through 5th grade were categorized as high‐level tasks, and 64% of higher‐level resources were priced higher. In both situations, resources identified as Doing Mathematics were significantly lower than all other categories. When comparing the cognitive demand across the Common Core domains, they found that specific resources included lower cognitive demand, like Geometry for all grade levels.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0002.jpg" title="2 Comparison of the TpT free versus <$5 activities and their associated mode levels of cognitive demand." /> </p> <p></p> <p>Since virtual resource pools showed that the items available online tended to have a lower level of cognitive demand, the process of critical curation was created to support teachers in thinking critically about what is implemented in the classroom. Critical curation is the thoughtful selection of resources based on one's pedagogical knowledge, content knowledge, personal experiences, and the lesson's purpose (Sawyer, Dredger, et al., [<reflink idref="bib21" id="ref57">21</reflink>]). In our mathematics pedagogical courses, preservice teachers are instructed on how to analyze mathematics resources based on Mathematical Critical Curation Questions:</p> <p></p> <ulist> <item> How does the resource support the goals of the lesson?</item> <p></p> <item> Does the resource provide mathematically valid concepts (i.e., does it perpetuate misconceptions)?</item> <p></p> <item> Is the resource's content appropriate for my students?</item> <p></p> <item> Who is doing the mathematics?</item> <p></p> <item> What is the level of cognitive demand for your resource?</item> <p></p> <item> Should you provide any visuals to help others learn?</item> <p></p> <item> Does your resource support any biases or discriminatory practices?</item> <p></p> <item> Can your resource support social justice issues? (Sawyer, [<reflink idref="bib18" id="ref58">18</reflink>]).</item> </ulist> <p>However, we need to determine if these questions encompass all the inherent problems with a nonhuman‐created task or if more questions might be required to support AI‐generated resources.</p> <p>We know that the ChatGPT tool uses resources collected from the Internet (Marr, [<reflink idref="bib10" id="ref59">10</reflink>]), but we need to understand how it influences the level of cognitive demand for activities created. We speculated that the level of cognitive demand could be similar to what was found in (Dick et al., [<reflink idref="bib5" id="ref60">5</reflink>]) research, and this investigation looked into whether the tool produced identical results. In particular, our research questions were:</p> <p></p> <ulist> <item> What are the associated levels of cognitive demand of ChatGPT's text responses for each category below:</item> <p></p> <item> All mathematical tasks</item> <p></p> <item> Each grade level between kindergarten and fifth grade</item> <p></p> <item> Each common core domain.</item> <p></p> <item> Is there a statistically significant difference between lower‐level mathematical tasks and higher‐level mathematical tasks for each category?</item> <p></p> <item> What characteristics were shared across the ChatGPTs tasks?</item> </ulist> <p>Based on the findings, we created recommendations for MTEs to support the critical curation of AI resources in the elementary classroom.</p> <hd id="AN0191298896-8">METHODS</hd> <p>We conducted a document analysis on 191 text responses constructed by ChatGPT. Each text corresponded to an AI task created for one of the elementary Common Core mathematical standards. We analyzed the data using a thematic analysis (Braun & Clarke, [<reflink idref="bib2" id="ref61">2</reflink>]) for categorizing the level of cognitive demand. We used chi‐squared (Pearson, [<reflink idref="bib15" id="ref62">15</reflink>]) statistical analysis because it helped determine how likely any observed differences in the level of cognitive demand occurred by chance or if they were due to a relationship between the variables.</p> <hd id="AN0191298896-9">Data collection</hd> <p>Between October 30, 2023 and November 6, 2023, we asked ChatGPT to "Create a mathematical task for [insert Common Core Standard]." For example, we asked, "Create a mathematical task for CCSS.MATH.CONTENT.1.MD.C.4." The standard numbering came from https://<ulink href="http://www.thecorestandards.org/Math/">www.thecorestandards.org/Math/</ulink>. We completed this prompt from the same new user account 191 times, each prompt corresponding with one Common Core Standard. Kindergarten had 25 standards, first grade had 24 standards, second grade had 28 standards, third grade had 37 standards, fourth grade had 37 standards, and fifth grade had 40 standards. Each grade level's standards were subdivided into Common Core Domains: Counting and Cardinality, Operations and Algebraic Thinking, Number and Operations in Base Ten, Measurement and Data, Geometry, and Number and Operation ‐ Fractions.</p> <hd id="AN0191298896-10">Data analysis</hd> <p>The prompt created 191 distinctly different tasks, meaning no two tasks had the same content, which we analyzed using a Thematic Analysis (Braun & Clarke, [<reflink idref="bib2" id="ref63">2</reflink>]) approach by categorizing each task using Stein and Smith's ([<reflink idref="bib27" id="ref64">27</reflink>]) Task Analysis Guide Framework. Each task was identified as one of the four levels of cognitive demand: Memorization, Procedures without Connections, Procedures with Connections, or Doing Mathematics. As seen in Figure 3, the tasks were coded as the lowest level of cognitive demand, Memorization, if they focused on reproducing previously learned formulas, facts, or definitions or were non‐ambiguous without any meaningful connections to underlying mathematical concepts. The tasks were coded as Procedures without Connections if they were algorithmic, required little thought to complete, and had no specific connection to the mathematics or elements associated with understanding the mathematics. The following two levels, Procedures with Connections and Doing Mathematics, were considered high levels of cognitive demand. The tasks coded as Procedures with Connections still had algorithmic elements; however, they usually had students represent multiple strategies, or they had students explain their mathematical thinking through the process. Finally, at the highest level of cognitive demand, Doing Mathematics tasks had complex structures that required students to self‐monitor, self‐regulate, and analyze the understanding of the mathematics behind the concepts.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0003.jpg" title="3 Examples of levels of cognitive demand task from ChatGPT's responses." /> </p> <p></p> <p>Both researchers individually coded the data for one of the levels of cognitive demand. If they came up with a discrepancy in coding, they discussed the text until both came to a 100% agreement on its category. After this coding process was complete, chi‐squared (Pearson, [<reflink idref="bib15" id="ref65">15</reflink>]) statistical analysis was run, looking at maximum levels of cognitive demand for all mathematical tasks, grade level, and each Common Core Domain. Because of the low frequency of Memorization and Doing Mathematics tasks, statistical analysis was conducted across two categories of levels of demand: low‐level cognitive demand was identified as the combination of both Memorization tasks and Procedures without Connection, and high‐level cognitive demand was categorized as the combination of Procedures with Connections tasks and Doing Mathematics tasks. We selected our alpha level to be 0.05. Thus, we determined if our data was statistically significant if our p‐value was less than our alpha.</p> <p>We used a thematic analysis approach (Braun & Clarke, [<reflink idref="bib2" id="ref66">2</reflink>]) to find similar characteristics across the ChatGPT data. Both researchers observed the data as a whole and came together to discuss common themes across the data. Multiple concepts were identified, and we used keyword searches of the data to support our findings. For example, we found repetition of phrases in the data, such as "real‐world applications," for which we used Google Docs Find and Replace function to help calculate the number of times it occurred across the data.</p> <hd id="AN0191298896-12">FINDINGS</hd> <p>In this section, we share the main themes from our data analysis. First, we share our findings on the level of cognitive demand. Then, we discuss the statistical significance of those results across all tasks, grade levels, and Common Core Domains. Finally, we describe our observed common characteristics across ChatGPT's created tasks.</p> <hd id="AN0191298896-13">Q1 level of cognitive demand for all ChatGPT tasks</hd> <p>As seen in Figure 4, 70.1%, or 134 of the 191 constructed tasks, could be considered higher levels of cognitive demand, and of those, 67% were Procedures with Connections tasks. These activities had the students create diagrams, connect to real‐world situations, and have students explain their mathematical thinking. For example, one ChatGPT response stated, "After solving the word problems, gather the students to discuss the solutions as a class. Ask students to explain their thought process, including how they used drawings and equations to solve the problems." These statements increased the level of cognitive demand of the activity because of the connections the task explicitly asked for in its instructions. However, 3.1% of the tasks were Doing Mathematics, a lower value but consistent with the research on virtual research pools, which found that less than 1% of tasks typically were Doing Mathematics (Shapiro et al., [<reflink idref="bib24" id="ref67">24</reflink>]). Memorization tasks were only 2.1%, even lower than the highest level of cognitive demand. Finally, Procedures without Connections only accounted for about 28% of the tasks constructed by ChatGPT.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0004.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0004.jpg" title="4 Total level of cognitive demand on ChatGPT's tasks." /> </p> <p></p> <p>When we break the data down across grade levels, as seen in Figure 5, we notice that the lowest level of cognitive demand is consistent across kindergarten through fifth grade, as is Doing Mathematics. However, the Procedures with and without Connections have a different ratio. Notice how the first and second‐grade tasks typically have more of a 50/50 consistency between the two areas, while third through fifth grade have a higher number of tasks categorized as Procedures with Connections versus Procedures without Connections. This is evident in Figure 5, where there is a shift from the number of higher‐level cognitive demand tasks for third‐ to fifth‐grade tasks versus the number of higher‐level cognitive demand tasks for first through second grade.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0005.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0005.jpg" title="5 Grade's level of cognitive demand." /> </p> <p></p> <p>Figure 6 shows the difference in cognitive demand between the Common Core domains. Most Common Core domains—except Number and Operations‐Base Ten and Counting and Cardinality—significantly differ between higher‐level and lower‐level cognitive demands. Specifically, notice that the Number and Operations ‐ Fractions unit has 33 of the 37 (89%) resources identified as having a higher level of cognitive demand. Also, the data shows that all Doing Mathematics tasks came from Common Core domains: Operations and Algebraic Thinking.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0006.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0006.jpg" title="6 Common core domain's level of cognitive demand." /> </p> <p></p> <hd id="AN0191298896-17">Q2 statistical significance between high level of cognitive demand and low level of cognitive...</hd> <p></p> <hd id="AN0191298896-18">The level of cognitive demand and all of ChatGPT's responses</hd> <p>Our overall data analysis pointed towards a more remarkable instance of higher‐level cognitive demand tasks. The data indicated a statistically significant difference between the higher‐level and lower‐level cognitive demand tasks produced by ChatGPT's responses (<emph>p</emph>‐value <0.001). This significance level indicates that there was a greater likelihood of an individual receiving a higher‐level cognitive demand response for an elementary mathematics activity than a lower‐level cognitive demand response.</p> <hd id="AN0191298896-19">The level of cognitive demand and grade level</hd> <p>We identified that the level of cognitive demand of tasks in kindergarten (<emph>p</emph> = 0.07), first grade (<emph>p</emph> = 0.68), and second grade (<emph>p</emph> = 0.71) was not statistically significant. However, third grade (<emph>p</emph> < 0.001), fourth grade (<emph>p</emph> = 0.002), and fifth grade (<emph>p</emph> = 0.002) did have a statistically significant difference between the level of cognitive demand. This indicates that there is more of a possibility of receiving a higher‐level cognitive demand task in a higher‐level classroom than in a lower‐level classroom. This could be because of the nature of how the standards were written. Higher‐level class standards typically have the students explain their understanding behind properties, while kindergarten through second‐grade students are just beginning to discover the initial elements of mathematics, thus exploring definitions, naming, and identifying figures.</p> <hd id="AN0191298896-20">The level of cognitive demand and common core domains</hd> <p>The data also indicated a statistical difference in the level of cognitive demand of tasks for all the Common Core domains except for two areas: Counting and Cardinality and Number and Operations in Base Ten. We found statistical significance for Operations and Algebraic Thinking (<emph>p</emph> = 0.002), Measurement and Data (<emph>p</emph> = 0.04), Geometry (<emph>p</emph> = 0.004), and Number and Operations ‐ Fractions (<emph>p</emph> < 0.001). Number and Operations in Base Ten (<emph>p</emph> = 0.87) and Counting and Cardinality (<emph>p</emph> = 0.53) did not have a statistically significant difference, indicating almost a 50% split between higher‐level and lower‐level demand tasks. ChatGPT's response in both of the categories tended to be more algorithmic without having the students explain their thinking. For example, ChatGPT's response exploring the understanding of place value had students identify each number's place value without explaining why that place value should be used.</p> <hd id="AN0191298896-21">Q3 shared characteristics across ChatGPT's tasks</hd> <p>When exploring these tasks, specific themes emerged around the characteristics of the activities created by ChatGPT. We noticed the repetition of activities and phrases, inappropriate content, and a lack of mathematical information constructed by the nonhuman creator.</p> <hd id="AN0191298896-22">Repetitive nature of the tasks</hd> <p>One of the major themes that surfaced from the data was that tasks were almost identical across multiple grade levels, with only the content changing. For example, as seen in Figure 7, mathematics was taught by following the steps: "Have the students solve these problems, showing their work and explaining their thought processes." The content is different, but the activities are very similar, showing that despite higher‐level cognitive demand, there is much repetition and less ingenuity in teaching those concepts.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0007.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0007.jpg" title="7 Examples of the repetitive nature of tasks." /> </p> <p></p> <hd id="AN0191298896-24">Real‐world applications</hd> <p>Another repetitive factor was that specific statements were added at the end of the task. We found "discuss real‐life scenarios," "discuss the real‐world application," and other variations of that exact phrase repeated in multiple tasks with "real‐life" and "real‐world" appearing 139 times in the data. For example, ChatGPT's response to CCSS.MATH.CONTENT.3.OA.A.4. stated, "Discuss real‐life scenarios where solving for the unknown is useful, such as calculating the number of groups needed for a specific quantity or determining the quantity of items in a set." The AI chatbot was not trained to add real‐world elements to the data, and the tasks were collected on a new user account. However, the data demonstrates that the AI chatbot could recognize a mathematical task as including real‐world elements based on its training.</p> <hd id="AN0191298896-25">Create their own</hd> <p>The ChatGPT responses also included elements for teachers to have their students create examples of problems. "Create their own" was repeated 80 times in the data. For example, ChatGPT's response for CCSS.MATH.CONTENT.K.CC.C.6. stated, "As a follow‐up, you can have students create their own shape cards with drawings or cut‐outs of shapes, labeling them correctly." This phrase tended to be an expansion activity at the lesson's end. However, it increased the rigor of some activities because it required the students to devise ways of solving the problem while explaining their thinking.</p> <hd id="AN0191298896-26">Inappropriate responses</hd> <p>Despite the regularity of activities involving real‐world scenarios and students creating their problems, there were situations where the activities should not be promoted in the elementary classroom. These activities focused on "correct" solutions and had students penalized or removed from activities when their answers were not "right." For example, ChatGPT's response for CCSS.MATH.CONTENT.K.CC.A.2. taught counting in unison and stated, "If a student makes a mistake or hesitates, they are out, and the game continues with the remaining students." This example demonstrates that despite a majority of ChatGPT's responses having a higher level of cognitive demand, there still exist activities created by the AI chatbot that should not be promoted in an elementary classroom because they could harm students' development of confidence and meaningful mathematical connections.</p> <hd id="AN0191298896-27">Lack of mathematical explanation</hd> <p>One commonality across all 191 tasks was the lack of the mathematics being explained or described in the ChatGPT task. The language in the task included students or instructors "explaining their thought processes," but the task never included what mathematics should be the focus of the explanation. Sawyer and Aga ([<reflink idref="bib19" id="ref68">19</reflink>]) found that ChatGPT has difficulty with mathematical reasoning. Thus, it would not necessarily be able to accurately describe that process (Korzynski et a., 2023). We noticed that the ChatGPT knew to ask for students' thought processes and knew that the teacher should explain the topic. However, if the teacher did not know the mathematics, they could have difficulty implementing these activities.</p> <hd id="AN0191298896-28">DISCUSSION</hd> <p>MTEs need to be aware of the following results of this investigation. First, we noticed that 70% of all responses constructed were higher levels of cognitive demand, strikingly different from the research on virtual resource pools. We also noticed that the standards influenced the kinds of cognitive demand produced by the AI chatbot, and the repetitive nature of these tasks meant that teachers needed to use other forms of resources when implementing classroom activities.</p> <hd id="AN0191298896-29">Biased from virtual resource pools</hd> <p>Since AI chatbots collect their data points from the Internet, we found some trends in the data similar to what was observed in the virtual resource pool research. We conjecture that online content could bias AI's data. The research on TpT's top 1000 resources identified a lower‐level cognitive demand for tasks (Dick et al., [<reflink idref="bib5" id="ref69">5</reflink>]). Our overall findings from our investigation indicate that if a teacher randomly selects a task from ChatGPT, they are likely to select a higher‐level cognitive demand task. However, for TpT's resources, when comparing the level of cognitive demand of K–2nd grade mathematical tasks to 3rd–5th grade mathematics tasks have similar results to what we found with ChatGPT. The tasks for K–2 across both TpT and AI chatbot were lower‐level cognitive demands compared to the 3rd–5th grade tasks. The research on TpT also identified that the level of cognitive demand for resources was statistically significantly different across Common Core Domains, which we also found with AI‐generated tasks. We recognize that ChatGPT uses more than just the Internet to produce responses to user queries but hypothesize that existing mathematical tasks could influence the characteristics of tasks produced by ChatGPT.</p> <hd id="AN0191298896-30">Standard influencing the level of cognitive demand</hd> <p>We speculate that there was a possible connection between how the Common Core standard was written and the tasks constructed by AI chatbots. As mathematics educators, we know that standards constructed for kindergarten through second grade typically aim to introduce mathematical concepts and have students begin to find the meaning behind operations. Third through fifth grade standards tend to dig deeper into practicing the skills learned in kindergarten through second grade, specifically in place value and Number and Operations. This difference in the goal of standards may have influenced the cognitive demand of the generated tasks. Data revealed that kindergarten through second‐grade tasks had lower cognitive demand than grades 3–5. We observed a relationship between what the standard was asking the teacher to implement in the classroom and the cognitive demand of the task created by AI tools.</p> <p>Data also revealed that AI‐created tasks have the highest cognitive demand for operations and algebraic thinking. These standards focus on understanding and applying properties of operations and the relationship between mathematical concepts. The way the standards were written was more associated with students regulating their thought processes without an algorithmic solution provided. While tasks generated for all standards that ask for students to understand a concept were not identified as doing mathematics, we hypothesize that there could be a relationship between the intended goal of the standard and the cognitive demand of the AI‐generated task.</p> <hd id="AN0191298896-31">Lack of mathematical description and repetitive tasks</hd> <p>One thing that surprised us in the data was that we found the activities to be repetitive without explaining the mathematics descriptions. For example, when asked to create tasks for fraction word problems in fifth grade and whole‐number word problems in fourth grade, ChatGPT created activities having students solve real‐world word problems that were almost identical, with only the mathematical values changed, without explaining how they should solve the problems. This showed us that even though the tasks created were of higher cognitive demand, implementing these in the classroom could become repetitive and boring. Also, the mathematics teacher would need more pedagogical knowledge to implement these activities and determine what thinking would be needed to solve the problems. Therefore, humans are still needed in this curriculum development process to understand what is expected to be taught.</p> <p>This data did not give teachers different lessons, but the same lesson had minor changes. We saw this specifically whenever analyzing the level of cognitive demand. Because we had multiple lessons with the same activity, it was easy to code the exact activity for multiple grade levels. Since this occurred in mathematics lessons, this investigation could imply that this could occur in other subjects. Therefore, ChatGPT should not be used as the sole form of any STEM curriculum because of its possible repetitive nature.</p> <p>Research on teachers' use of curricula has shown that experienced teachers use their discretion when adopting curricula, editing it to make it relevant for their students (Bümen & Yazicilar, [<reflink idref="bib3" id="ref70">3</reflink>]; Li & Harfitt, [<reflink idref="bib8" id="ref71">8</reflink>]; Remillard, [<reflink idref="bib16" id="ref72">16</reflink>]). Karataş et al. ([<reflink idref="bib6" id="ref73">6</reflink>]) found the same curriculum adoption pattern when teachers use AI‐generated curriculum. Our findings support the use of teacher discretion when using online educational material.</p> <hd id="AN0191298896-32">LIMITATIONS</hd> <p>This investigation has a few clear limitations that need to be addressed. First, the way the data were collected focused only on creating tasks using the numerical values of the Common Core standards rather than copying and pasting the mathematical standard into ChatGPT. Therefore, the data could be biased because ChatGPT had to find resources from its database that used Common Core numerical representations. This could mean that individuals should use numerical representations when looking for a task to find resources with higher‐level cognitive demand. However, we need to find out what aspects of the algorithm prompted the AI chatbot to construct these higher‐level tasks.</p> <p>The second limitation is that the data was constructed from October through November 2023. We know that the data is ever‐changing based on the system's updates. Our data might be limited to a specific time frame. Thus, more research needs to be done to see if the high level of cognitive demand would change based on updates in the AI system.</p> <p>The third limitation is that ChatGPT is only one AI chatbot. At the same time, multiple other types of chatbots are available specifically related to educational curriculum, like Magic School AI, which creates lesson plans based on the standards you implement. This data is related only to ChatGPT, and the results from different AI chatbots would find a different level of cognitive demand.</p> <hd id="AN0191298896-33">IMPLICATIONS</hd> <p>Based on this investigation, we have the following recommendations. MTEs need to discuss with their mathematics preservice teachers the use of AI as a curriculum development tool in their mathematics content or methods courses. Since the data was created using a single prompt, MTEs should understand prompt engineering techniques to get complex responses. Finally, MTEs need to provide critical insight to their teachers to help them understand the biased nature of this tool.</p> <hd id="AN0191298896-34">MTEs teaching prompt engineering</hd> <p>Since the data indicated that there was less than a 3% chance of getting a Doing Mathematics task, it implies that if a teacher wants a task to be the highest level of cognitive demand, they must be able to implement prompt engineering or ask AI in a particular way to gain these resources using the process (Korzynski et al., [<reflink idref="bib7" id="ref74">7</reflink>]). Thus, MTEs need to know prompt engineering techniques to teach their preservice teachers how to use AI chatbots in the classrooms effectively.</p> <p>Prompt engineering involves adding more context, instructions, input, or expected output data to help the AI better understand the task (Korzynski et al., [<reflink idref="bib7" id="ref75">7</reflink>]). Thus, MTEs can instruct preservice teachers to create prompts that provide more context, telling AI what data could be used or in what context it should view the preservice teachers' questions. The prompt could include more specific instructions, including sub‐questions that it must implement first to construct the final product better. Preservice teachers can provide AI with example input data to analyze and construct similar results. Finally, MTEs can teach their preservice teachers to identify the expected output format, such as, "Please display your answer in a table including each response as one row." In this investigation, we only provided simple instructions without any more information. Thus, MTEs need to include either context, input data, or more information about expected output data to change the level of cognitive demand.</p> <p>ChatGPT does not recognize specific terminology like Level of Cognitive Demand or Smith and Stein's Task Analysis Guide Framework (Sawyer & Aga, [<reflink idref="bib19" id="ref76">19</reflink>]). For example, if you do ask for a "Doing Mathematics task," it creates an inaccurate response of:</p> <p>In educational research and practice, tasks are often categorized into different levels of cognitive demand based on the type and complexity of thinking they require. These levels are commonly associated with Bloom's Taxonomy or other similar frameworks.</p> <p>Then, it went on to describe Bloom's Taxonomy. We decided to see if it knew about Stein and Smith's ([<reflink idref="bib27" id="ref77">27</reflink>]) Task Analysis Guide framework, and it inaccurately stated, "The framework includes three main components: (<reflink idref="bib1" id="ref78">1</reflink>) Access to Relevant Knowledge, (<reflink idref="bib2" id="ref79">2</reflink>) Mathematical Habits of Mind, and (<reflink idref="bib3" id="ref80">3</reflink>) Appropriateness for Grades K–12."</p> <p>Therefore, MTEs need to provide new prompts for context for a Doing Mathematics task using its definition, as seen in Figure 8. Notice that when you use the definition as a requirement, ChatGPT's response brings a real‐world context to create a non‐algorithmic activity exploring number deconstruction.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/SSM/01feb26/ssm18348-fig-0008.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="ssm18348-fig-0008.jpg" title="8 Example of ChatGPT response to create a doing math task." /> </p> <p></p> <hd id="AN0191298896-36">Be critical because everything is biased on the internet, including ChatGPT</hd> <p>MTEs should also share with their preservice teachers why they should be critical. Nothing on the Internet is free from bias. When you search on TpT for resources, the creators of the materials have to pay for the website to be featured and placed at the top of search menus (Sawyer, Dick, & Sutherland, [<reflink idref="bib20" id="ref81">20</reflink>]). Humans create the material online, and humans create the algorithms for these technologies. Thus, we cannot assume that what is produced is bias‐free (Wu, [<reflink idref="bib32" id="ref82">32</reflink>]). Therefore, teachers must always be critical of anything we find online (Sawyer et al., [<reflink idref="bib17" id="ref83">17</reflink>]), and ChatGPT is no exception to this rule.</p> <p>We should teach criticality to anyone who uses the Internet because this technology is currently being used and will continue to be used by individuals all around us. We need to ensure they are aware of these issues and what could happen if AI is misused. We want teachers to refrain from taking this research and assuming that what is created by ChatGPT is better than something that they could get from a publishing company. Even though there is a higher level of cognitive demand for tasks created by ChatGPT, other considerations must be made to ensure that the task is appropriate for individual classrooms. ChatGPT is biased; the Internet is biased. Thus, teachers need to be critical consumers of their output.</p> <hd id="AN0191298896-37">CONCLUSION</hd> <p>While this investigation might make it seem that AI chatbots are the way to go when thinking about activities, it must be noted that this is not the magical solution to this complex problem in education. We know that teachers are overconfident with the abilities of ChatGPT and need help making the significant adaptations necessary to make the materials usable in an elementary classroom (Sawyer, [<reflink idref="bib18" id="ref84">18</reflink>]).</p> <p>Teaching is so much more than just activities! Effective teachers incorporate the context of the individuals, learners' knowledge, mathematical content, pedagogical knowledge, and experiences to build relationships with their students (NCTM, [<reflink idref="bib11" id="ref85">11</reflink>]). Since AI chatbots do not know the individual learners in the classrooms and they are not able to build relationships with people, we believe AI chatbots will not replace the human element in an education classroom. The computer cannot replicate a teacher's discretion, knowledge, and experience of their students for what is needed to teach the whole person. Mathematics teachers can use AI as a tool. However, they need to be the critical curators of their classrooms because their knowledge alone can support the students' emotional and academic needs.</p> <hd id="AN0191298896-38">DATA AVAILABILITY STATEMENT</hd> <p>The data that support the findings of this study are available on request from the corresponding author. The data are not publicly available due to privacy or ethical restrictions.</p> <ref id="AN0191298896-39"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref20" type="bt">1</bibl> <bibtext> This article does not contain a study of human participants conducted by any authors.</bibtext> </blist> </ref> <ref id="AN0191298896-40"> <title> REFERENCES </title> <blist> <bibtext> Bill and Melinda Gates Foundation. (2014). America's teachers on teaching in an era of change. In Primary sources (3rd ed.). Scholastic Inc. and Bill & Melinda Gates Foundation. <ulink href="http://www.scholastic.com/primarysources/index.htm">http://www.scholastic.com/primarysources/index.htm</ulink></bibtext> </blist> <blist> <bibl id="bib2" idref="ref21" type="bt">2</bibl> <bibtext> Braun, V., & Clarke, V. (2012). Thematic analysis. In H. Cooper (Ed.), The APA handbook of research methods in psychology (pp. 57 – 71). APA Books.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref22" type="bt">3</bibl> <bibtext> Bümen, N., & Yazicilar, Ü. (2020). A case study on the teachers' curriculum adaptations: Differences in state and private high school. Gazi Üniversitesi Gazi Eğitim Fakültesi Dergisi, 40 (1), 183 – 224. https://doi.org/10.17152/gefad.595058</bibtext> </blist> <blist> <bibl id="bib4" idref="ref35" type="bt">4</bibl> <bibtext> Carpenter, J., Shelton, C. C., & Mitchell, L. (2021, July). "Like buying time": Educators' use of online education marketplaces during the COVID‐19 pandemic. In EdMedia+ innovate learning (pp. 644 – 653). Association for the Advancement of Computing in Education (AACE).</bibtext> </blist> <blist> <bibl id="bib5" idref="ref41" type="bt">5</bibl> <bibtext> Dick, L. K., Sawyer, A. G., MacNeille, M., Shapiro, E., & Wismer, T. A. (2023). Exploring grades 3–5 mathematics activities found online, mathematics teacher. Learning and Teaching PK‐12, 116 (3), 174 – 183. Retrieved on 22 March, 2023, from. https://doi.org/10.5951/MTLT.2021.0330</bibtext> </blist> <blist> <bibl id="bib6" idref="ref73" type="bt">6</bibl> <bibtext> Karataş, F., Eriçok, B., & Tanrikulu, L. (2024). Reshaping curriculum adaptation in the age of artificial intelligence: Mapping teachers' AI‐driven curriculum adaptation patterns. British Educational Research Journal, 51 (1), 154 – 180.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref27" type="bt">7</bibl> <bibtext> Korzynski, P., Mazurek, G., Krzypkowska, P., & Kurasinski, A. (2023). Artificial intelligence prompt engineering as a new digital competence: Analysis of generative AI technologies such as ChatGPT. Entrepreneurial Business and Economics Review, 11 (3), 25 – 37.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref71" type="bt">8</bibl> <bibtext> Li, Z., & Harfitt, G. J. (2016). An examination of language teachers' enactment of curriculum materials in the context of a centralised curriculum. Pedagogy, Culture and Society, 25 (3), 403 – 416. https://doi.org/10.1080/14681366.2016.12709</bibtext> </blist> <blist> <bibl id="bib9" idref="ref25" type="bt">9</bibl> <bibtext> Maiorca, C., Burton, M., Ivy, J., & Roberts, T. (2024, March). Developing mathematics lessons and assessments with chatbots for learning in teacher education: Innovation and challenges. In Society for information technology & teacher education international conference (pp. 1783 – 1791). Association for the Advancement of Computing in Education.</bibtext> </blist> <blist> <bibtext> Marr, B. (2023, May 19). A short history of ChatGPT: How we got to where we are today. Forbes. https://<ulink href="http://www.forbes.com/sites/bernardmarr/2023/05/19/a‐short‐history‐of‐chatgpt‐how‐we‐got‐to‐where‐we‐are‐today/?sh=3aa87014674f">www.forbes.com/sites/bernardmarr/2023/05/19/a‐short‐history‐of‐chatgpt‐how‐we‐got‐to‐where‐we‐are‐today/?sh=3aa87014674f</ulink></bibtext> </blist> <blist> <bibtext> National Council of Teachers of Mathematics. (2014). Principles to actions: Ensuring mathematics success for all. National Council of Teachers of Mathematics.</bibtext> </blist> <blist> <bibtext> National Council of Teachers of Mathematics. (2024). Artificial intelligence and mathematics teaching. NCTM. https://<ulink href="http://www.nctm.org/standards‐and‐positions/Position‐Statements/Artificial‐Intelligence‐and‐Mathematics‐Teaching/">www.nctm.org/standards‐and‐positions/Position‐Statements/Artificial‐Intelligence‐and‐Mathematics‐Teaching/</ulink></bibtext> </blist> <blist> <bibtext> Newton, C. (2023). The AI is eating itself ‐ by Casey Newton. Platformer. Retrieved on 20 September, 2023, from https://www.platformer.news/p/the‐ai‐is‐eating‐itself?utm_source=post‐email‐title&publication_id=7976&post_id=131409572&isFreemail=false&utm_medium=email</bibtext> </blist> <blist> <bibtext> OpenAI. (2023). Privacy policy. OpenAI. https://openai.com/policies/privacy-policy/</bibtext> </blist> <blist> <bibtext> Pearson, K. (1992). On the criterion that a given system of deviations from the probable in the case of a correlated system of variables is such that it can be reasonably supposed to have arisen from random sampling. In Breakthroughs in statistics (pp. 11 – 28). Springer New York. https://doi.org/10.1007/978-1-4612-4380-9_2</bibtext> </blist> <blist> <bibtext> Remillard, J. T. (2005). Examining key concepts in research on teachers' use of mathematics curricula. Review of Educational Research, 75 (2), 211 – 246. https://doi.org/10.3102/00346543075002211</bibtext> </blist> <blist> <bibtext> Sawyer, A., Dick, L., Shapiro, E., & Wismer, T. (2019). The top 500 mathematics pins: An analysis of elementary mathematics activities on Pinterest. Journal of Technology and Teacher Education, 27 (2), 235 – 263.</bibtext> </blist> <blist> <bibtext> Sawyer, A. G. (2024). Artificial intelligence chatbot as a mathematics curriculum developer: Discovering preservice teachers' overconfidence in ChatGPT. International Journal on Responsibility, 7 (1). https://doi.org/10.62365/2576-0955.1106</bibtext> </blist> <blist> <bibtext> Sawyer, A. G., & Aga, Z. (2024). Counterexamples to demonstrate artificial intelligence chatbot's lack of knowledge in the mathematics education classrooms. Association of Mathematics Teacher Education's Connections. https://amte.net/sites/amte.net/files/Connections%20%28Sawyer%29.pdf</bibtext> </blist> <blist> <bibtext> Sawyer, A. G., Dick, L. K., & Sutherland, P. (2020). Online mathematics teacherpreneurs developers on teachers pay teachers: Who are they and why are they popular? Education in Science, 10 (9), 248. https://doi.org/10.3390/educsci10090248</bibtext> </blist> <blist> <bibtext> Sawyer, A. G., Dredger, K., Myers, J., Barnes, S., Wilson, D. R., Sullivan, J., & Sawyer, D. (2020). Teachers as curators: Investigating elementary preservice teachers' inspiration for lesson planning. Journal of Teacher Education, 9 (1), 1 – 19. https://doi.org/10.1177/0022487119879894</bibtext> </blist> <blist> <bibtext> Schmidt & Bush. (2024). Are current K‐5 state mathematics standards really the common Core in disguise? [Conference session]. Association of Mathematic Teacher Educators 2024 Conference.</bibtext> </blist> <blist> <bibtext> Schulman, J., Zoph, B., & Kim, C. (2022). Introducing ChatGPT. OpenAI. Retrieved on 20 September, 2023, from https://openai.com/blog/chatgpt</bibtext> </blist> <blist> <bibtext> Shapiro, E., Dick, L., Sawyer, A. G., & Wismer, T. (2021). Critical consumption of online resources. Texas Mathematics Educator, 67 (1), 6 – 11.</bibtext> </blist> <blist> <bibtext> Shapiro, E., Sawyer, A. G., Dick, L., & Wismer, T. (2019). Just what online resources are elementary mathematics teachers using? Contemporary Issues in Teacher Education, 19 (4), 670 – 686. https://<ulink href="http://www.citejournal.org/proofing/just‐whatonline‐resources‐are‐elementary‐mathematics‐teachers‐using/">www.citejournal.org/proofing/just‐whatonline‐resources‐are‐elementary‐mathematics‐teachers‐using/</ulink></bibtext> </blist> <blist> <bibtext> Shelton, C. C., & Archambault, L. M. (2019). Who are online teacherpreneurs and what do they do? A survey of content creators on TeachersPayTeachers.com. Journal of Research on Technology in Education, 51 (4), 398 – 414.</bibtext> </blist> <blist> <bibtext> Stein, M. K., & Smith, M. S. (1998). Mathematical tasks as a framework for reflection: From research to practice. Mathematics Teaching in the Middle School, 3 (4), 268 – 275.</bibtext> </blist> <blist> <bibtext> Tech Desk. (2023). ChatGPT most popular AI tool. https://indianexpress.com/article/technology/artificial‐intelligence/chatgpt‐most‐popular‐ai‐tool‐study‐bard‐midjourney‐trail‐9026116/#:~:text=ChatGPT%20has%20emerged%20as%20the,September%202022%20and%20August%202023</bibtext> </blist> <blist> <bibtext> United States Department of Education. (2023). U.S. Department of Education shares insights and recommendations for artificial intelligence. Press Office. May 24, 2023. https://<ulink href="http://www.ed.gov/news/press‐releases/us‐department‐education‐shares‐insights‐and‐recommendations‐artificial‐intelligence">www.ed.gov/news/press‐releases/us‐department‐education‐shares‐insights‐and‐recommendations‐artificial‐intelligence</ulink></bibtext> </blist> <blist> <bibtext> Vincent, J. (2023). AI is killing the old web, and the new web struggles to be born. The Verge. Retrieved on 20 September, 2023, from https://<ulink href="http://www.theverge.com/2023/6/26/23773914/ai-large-language-models-data-scraping-generation-remaking-web">www.theverge.com/2023/6/26/23773914/ai-large-language-models-data-scraping-generation-remaking-web</ulink></bibtext> </blist> <blist> <bibtext> Wilhelm, A. G. (2014). Mathematics teachers' enactment of cognitively demanding tasks: Investigating links to teachers' knowledge and conceptions. Journal for Research in Mathematics Education, 45 (5), 636 – 674.</bibtext> </blist> <blist> <bibtext> Wu, G. (2023). 8 big problems with OpenAI's ChatGPT. MakeUseOf. https://<ulink href="http://www.makeuseof.com/openai-chatgpt-biggest-probelms/">www.makeuseof.com/openai-chatgpt-biggest-probelms/</ulink></bibtext> </blist> </ref> <aug> <p>By Amanda Gantt Sawyer and Zareen Gul Aga</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib29" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib32" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib18" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib12" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib19" firstref="ref9"></nolink> <nolink nlid="nl6" bibid="bib24" firstref="ref10"></nolink> <nolink nlid="nl7" bibid="bib22" firstref="ref11"></nolink> <nolink nlid="nl8" bibid="bib28" firstref="ref12"></nolink> <nolink nlid="nl9" bibid="bib14" firstref="ref13"></nolink> <nolink nlid="nl10" bibid="bib10" firstref="ref14"></nolink> <nolink nlid="nl11" bibid="bib30" firstref="ref19"></nolink> <nolink nlid="nl12" bibid="bib23" firstref="ref23"></nolink> <nolink nlid="nl13" bibid="bib13" firstref="ref24"></nolink> <nolink nlid="nl14" bibid="bib25" firstref="ref34"></nolink> <nolink nlid="nl15" bibid="bib17" firstref="ref38"></nolink> <nolink nlid="nl16" bibid="bib26" firstref="ref40"></nolink> <nolink nlid="nl17" bibid="bib27" firstref="ref42"></nolink> <nolink nlid="nl18" bibid="bib31" firstref="ref46"></nolink> <nolink nlid="nl19" bibid="bib21" firstref="ref57"></nolink> <nolink nlid="nl20" bibid="bib15" firstref="ref62"></nolink> <nolink nlid="nl21" bibid="bib16" firstref="ref72"></nolink> <nolink nlid="nl22" bibid="bib20" firstref="ref81"></nolink> <nolink nlid="nl23" bibid="bib11" firstref="ref85"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1495594
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: A Mile High and an Inch Deep: Exploring ChatGPT as a Mathematics Curriculum Development Tool
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Amanda+Gantt+Sawyer%22">Amanda Gantt Sawyer</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-5532-9047">0000-0001-5532-9047</externalLink>)<br /><searchLink fieldCode="AR" term="%22Zareen+Gul+Aga%22">Zareen Gul Aga</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22School+Science+and+Mathematics%22"><i>School Science and Mathematics</i></searchLink>. 2026 126(1):60-74.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 15
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2026
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: 2237151
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Technology+Uses+in+Education%22">Technology Uses in Education</searchLink><br /><searchLink fieldCode="DE" term="%22Curriculum+Development%22">Curriculum Development</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Curriculum%22">Mathematics Curriculum</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Mathematics%22">Elementary School Mathematics</searchLink><br /><searchLink fieldCode="DE" term="%22Curriculum+Implementation%22">Curriculum Implementation</searchLink><br /><searchLink fieldCode="DE" term="%22Common+Core+State+Standards%22">Common Core State Standards</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink><br /><searchLink fieldCode="DE" term="%22Instructional+Program+Divisions%22">Instructional Program Divisions</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1111/ssm.18348
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0036-6803<br />1949-8594
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Artificial Intelligence (AI) technology has issues, including inherent bias and inaccurate information, yet it has not stifled individuals' use of this technology in the mathematics classroom. Thus, we conducted a document analysis to investigate ChatGPT, a generative AI chatbot, constructing elementary mathematical tasks to determine how mathematics teacher educators (MTEs) can support this implementation. We analyzed 191 text responses constructed by ChatGPT, each corresponding to a task created for one of the elementary Common Core mathematical standards from kindergarten to fifth Grade to answer the research questions: What are the associated levels of cognitive demand of the ChatGPT's text responses, is there a statistically significant difference between lower-level mathematical tasks and higher-level mathematical tasks, and what are the common characteristics of AI's tasks? We found significantly higher-level cognitive demands across the tasks, but many activities were repetitive and inappropriate for the specified grade level. Finally, we explored how MTEs could support teachers using this tool in their classrooms using Mathematical Critical Curation Questions.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2026
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1495594
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1495594
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1111/ssm.18348
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 15
        StartPage: 60
    Subjects:
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Technology Uses in Education
        Type: general
      – SubjectFull: Curriculum Development
        Type: general
      – SubjectFull: Mathematics Curriculum
        Type: general
      – SubjectFull: Elementary School Mathematics
        Type: general
      – SubjectFull: Curriculum Implementation
        Type: general
      – SubjectFull: Common Core State Standards
        Type: general
      – SubjectFull: Difficulty Level
        Type: general
      – SubjectFull: Instructional Program Divisions
        Type: general
    Titles:
      – TitleFull: A Mile High and an Inch Deep: Exploring ChatGPT as a Mathematics Curriculum Development Tool
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Amanda Gantt Sawyer
      – PersonEntity:
          Name:
            NameFull: Zareen Gul Aga
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 02
              Type: published
              Y: 2026
          Identifiers:
            – Type: issn-print
              Value: 0036-6803
            – Type: issn-electronic
              Value: 1949-8594
          Numbering:
            – Type: volume
              Value: 126
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: School Science and Mathematics
              Type: main
ResultId 1