Examining the Idea Density and Semantic Distance of Responses Given by AI to Tests of Divergent Thinking
Saved in:
| Title: | Examining the Idea Density and Semantic Distance of Responses Given by AI to Tests of Divergent Thinking |
|---|---|
| Language: | English |
| Authors: | Mark A. Runco (ORCID |
| Source: | Journal of Creative Behavior. 2025 59(3). |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 11 |
| Publication Date: | 2025 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Artificial Intelligence, Creative Thinking, Models, Differences, Responses, Creativity |
| DOI: | 10.1002/jocb.1528 |
| ISSN: | 0022-0175 2162-6057 |
| Abstract: | Research suggests that generative AI (GAI) responds to divergent thinking (DT) prompts with multiple ideas, some of which seem to be original. The present investigation administered 55 DT tasks to three GAI services (Bard, GPT 3.5, and GPT 4.0). Instead of examining individual responses, an Idea Density algorithm was used to assess the output. This algorithm quantifies the ideas within responses, controlling for the number of words. A subset of the DT tests administered to the GAI were also scored for Semantic Distance, which estimates originality. Results indicated that the three GAI models differed in the Idea Density of the output. There were also significant differences between Realistic and Nonrealistic DT tasks. As has been the case in human samples, directions given when the GAI received the prompts also had a significant impact, with more Idea Density following directions that explicitly prompted original responses. Adjusted scores removed all verbiage in the output, which did not actually address the questions conveyed by the prompts. These corrected scores shared approximately 50% of the variance with the uncorrected "raw" responses, implying that the typical output of GAI is not always relevant. This was interpreted in the context of the standard definition of creativity, which emphasizes effectiveness, as well as originality. |
| Abstractor: | As Provided |
| Entry Date: | 2025 |
| Accession Number: | EJ1482985 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFJO59w_43wWfu5hxFzsERjAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDFHB7CkpP8az41c8rAIBEICBmxiyyT-3_npJZQFAC9nTcM-hMjN66lnxC-yEL4bbRvj4Hw-wW6kTv2ONomMTqz5p2n_bgkZZbldrFUlb1LqTwCUuXcU6PHjDQf4qixj8dR5d9k0NPKgQzkD3QOno8ZJMCuW3v1zQ1CywHQRnv_mnYNHleqjlCO12zIEXVVa0sTAkYem2seOhWNtA6ycIdtHhKp5yecKt5_gRQxSd Text: Availability: 1 Value: <anid>AN0187843392;3u701sep.25;2025Sep11.06:34;v2.2.500</anid> <title id="AN0187843392-1">Examining the Idea Density and Semantic Distance of Responses Given by AI to Tests of Divergent Thinking </title> <p>Research suggests that generative AI (GAI) responds to divergent thinking (DT) prompts with multiple ideas, some of which seem to be original. The present investigation administered 55 DT tasks to three GAI services (Bard, GPT 3.5, and GPT 4.0). Instead of examining individual responses, an Idea Density algorithm was used to assess the output. This algorithm quantifies the ideas within responses, controlling for the number of words. A subset of the DT tests administered to the GAI were also scored for Semantic Distance, which estimates originality. Results indicated that the three GAI models differed in the Idea Density of the output. There were also significant differences between Realistic and Nonrealistic DT tasks. As has been the case in human samples, directions given when the GAI received the prompts also had a significant impact, with more Idea Density following directions that explicitly prompted original responses. Adjusted scores removed all verbiage in the output, which did not actually address the questions conveyed by the prompts. These corrected scores shared approximately 50% of the variance with the uncorrected "raw" responses, implying that the typical output of GAI is not always relevant. This was interpreted in the context of the standard definition of creativity, which emphasizes effectiveness, as well as originality.</p> <p>Keywords: AI; artificial creativity; creative potential; divergent thinking; idea density; semantic distance</p> <p>Claims abound about the creativity of AI, especially in the popular press. Molik ([<reflink idref="bib23" id="ref1">23</reflink>]), for example, described discussions with AI that "are intensely creative." Another recent article was titled, "Artists Beware: Game Changer Test Results Show AI is More Creative Than 99% of Humans" (Plain [<reflink idref="bib29" id="ref2">29</reflink>]). This claim was a reaction to the empirical results of Guzik, Byrge, and Gilde ([<reflink idref="bib15" id="ref3">15</reflink>]). They administered the <emph>Torrance Tests of Creative Thinking</emph> (TTCT) to AI and did indeed find that it earned very high scores. These results need to be interpreted very carefully. The TTCT is a good test of creative potential, but it is a test, and all tests are samples of some larger universe of actual behavior (usually that which occurs in the natural environment). Furthermore, tests are administered to examinees. They are not initiated by the examinee. This is problematic for creativity, given its association with intrinsic motivation and spontaneity (Brainard [<reflink idref="bib10" id="ref4">10</reflink>]; Runco [<reflink idref="bib32" id="ref5">32</reflink>]). All of this supports the view that the TTCT (and other similar tests) are not really indicators of creativity per se. All tests of divergent thinking (DT), including the TTCT, must be interpreted as estimates of the potential for creative problem solving, not as guarantee of authentic creativity (Runco and Acar [<reflink idref="bib34" id="ref6">34</reflink>]). The results from Guzik, Byrge, and Gilde ([<reflink idref="bib15" id="ref7">15</reflink>]) are quite interesting, but even if qualified by the above (namely, that the TTCT merely provides estimates of potential), generalizations are limited. There are differences among DT tests (Runco, Alabbasi, and Paek [<reflink idref="bib35" id="ref8">35</reflink>]; Silvia, Martin, and Nusbaum [<reflink idref="bib45" id="ref9">45</reflink>]; Torrance [<reflink idref="bib47" id="ref10">47</reflink>]), so other tests might provide different results.</p> <p>The present investigation was designed to examine the creative potential of generative AI (GAI) with a large set of diverse DT tests. It was also designed to compare three AI tools, namely Bard, GPT 3.5, and GPT 4.0. One prediction was that GPT 4.0 would outperform GPT 3.5. This followed from descriptions of the engineering advances used to develop 4.0 and from some tentative evidence (Koubaa [<reflink idref="bib21" id="ref11">21</reflink>]; Taloni et al. [<reflink idref="bib46" id="ref12">46</reflink>]). We also predicted that prompts directing AI to give original responses would have the same significant impact on ideation as they do with humans in studies of explicit instructions (Hong, O'Neil Jr., and Peng [<reflink idref="bib19" id="ref13">19</reflink>]; Runco, Illies, and Eisenman [<reflink idref="bib38" id="ref14">38</reflink>]; Runco, Illies, and Reiter‐Palmon [<reflink idref="bib39" id="ref15">39</reflink>]). In research with humans, originality scores increase when examinees are told that they should be original, just to name one example (Abdulla Alabbasi et al. [<reflink idref="bib2" id="ref16">2</reflink>]; Acar and Runco [<reflink idref="bib6" id="ref17">6</reflink>]). Scores increase quite significantly if the instructions provide a definition of originality and procedural information about methods that increase the probability of finding an unusual idea.</p> <p>The TTCT verbal is scored for fluency, originality, and flexibility. Each is a measure of ideation. The TTCT relies on norms, which are updated on a regular basis (Abdulla Alabbasi, Paek, and Cramond [<reflink idref="bib3" id="ref18">3</reflink>]). There is some subjectivity in the scoring, as there is often the case when DT tests are used (Hass, Rivera, and Silvia [<reflink idref="bib17" id="ref19">17</reflink>]; Runco and Mraz [<reflink idref="bib41" id="ref20">41</reflink>]; Silvia [<reflink idref="bib44" id="ref21">44</reflink>]). The present research used maximally objective ideational scores. No humans were involved in the scoring of the responses from the GAI; scoring was fully automated. More specifically, all 55 of the tasks presented to the GAI scored for <emph>Idea Density</emph> (Brown et al. [<reflink idref="bib11" id="ref22">11</reflink>]; Covington [<reflink idref="bib12" id="ref23">12</reflink>]; Runco et al. [<reflink idref="bib42" id="ref24">42</reflink>]). Idea Density indicates how many ideas are produced, relative to the amount of verbiage. Previous research has reported good reliability for the Idea Density algorithm (0.80 &lt; rs &lt; 0.97). Details about the linguistic basis for the elements of the algorithm can be found in Brown et al. ([<reflink idref="bib11" id="ref25">11</reflink>]). A subset of the DT tests was also scored for <emph>semantic distance</emph>. This is a measure of relatedness, where concepts which are strongly related are close and concepts that are only remotely related are distant from one another. Concepts are close together when the relatedness is easy to see and apparent to most people. Concepts are less strongly related when it is uncommon to associate one with the other. This is entirely consistent with the theory of remote associates (Mednick [<reflink idref="bib22" id="ref26">22</reflink>]; Russ and Hoffmann [<reflink idref="bib43" id="ref27">43</reflink>]) and relevant because greater semantic distance is indicative of more original thinking (Acar et al. [<reflink idref="bib5" id="ref28">5</reflink>]; Heinen and Johnson [<reflink idref="bib18" id="ref29">18</reflink>]).</p> <p>Idea density was used in this investigation, first, because many creative activities rely on ideation, as do all DT tests. Ideation is not synonymous with creative cognition, but it is an important part of it, as is evidenced by the large body of research that has used DT tests as measures of creative potential. Second, the Idea Density algorithm minimizes subjectivity. It is fully automated. Third, Idea Density has been correlated with DT and several measures of creative performance in the natural environment (Runco et al. [<reflink idref="bib42" id="ref30">42</reflink>]). Runco et al. ([<reflink idref="bib42" id="ref31">42</reflink>]) found correlations of Idea Density with DT test scores, the number of online "hits" of TedTalks, and the citation impact of published research. Semantic distance similarly has been used in previous studies of creative potential (Acar and Runco [<reflink idref="bib6" id="ref32">6</reflink>]; Heinen and Johnson [<reflink idref="bib18" id="ref33">18</reflink>]). It is a useful automated estimate for the originality component of creative potential.</p> <p>There are pros and cons to using Idea Density. To begin with, Idea Density is automated and provides one score for whole paragraphs. This is advantageous because it means that scores are based on a large amount of information, which is not true of scoring which examines ideas one at a time (Runco and Mraz [<reflink idref="bib41" id="ref34">41</reflink>]; Silvia, Martin, and Nusbaum [<reflink idref="bib45" id="ref35">45</reflink>]). Additionally, as stated above, Idea Density has been correlated with various meaningful indices, including DT. Idea Density is entirely based on the quantity (or density) of ideas, however, unlike the TTCT and other DT batteries, which can provide originality scores. Then again, these originality scores have their own limitations (e.g., they may be based on norms, which are indicative of past performances and no guarantee of anything in the future). Recall here that Idea Density is especially useful with DT tests in that both focus on ideas and ideation. The present study had Idea Density scores for all of the DT tasks presented to the AI systems, but it also had originality scores (based on semantic distance) for the 12 Uses and Instances tasks presented to the AI.</p> <p>A fairly large number of DT tasks (<reflink idref="bib55" id="ref36">55</reflink>) was administered in the present investigation to check the generalizability of the results (across tasks). Realistic DT tests are especially pertinent to the question of generalizability. Too often DT tests contain questions about somewhat unrealistic things, such as how a potato and a carrot are alike (the Similarities Test), or exemplars of the category of square things (the Instances Test). Realistic DT tasks instead ask about situations that are indicative of what the individual might encounter in the natural environment (e.g., "it starts to rain but you do not have a hat or umbrella..."). A second prediction was, therefore, that GAI will do well on realistic tasks. This kind of task would seem to be fairly well aligned with the information that AI will find when searching for options in existing data bases (i.e., the large language models). Problem generation DT tasks were also administered to the GAI systems. These were included because one criticism of AI is that it solves problems well, but unlike humans, does not by itself find problems to solve (Csikszentmihalyi [<reflink idref="bib13" id="ref37">13</reflink>]; Runco [<reflink idref="bib32" id="ref38">32</reflink>]). Problems must be presented to the AI. This is an important point because problem finding plays a consequential role in the creative process (Getzels and Csikszentmihalyi [<reflink idref="bib14" id="ref39">14</reflink>]; Mumford, Reiter‐Palmon, and Redmond [<reflink idref="bib24" id="ref40">24</reflink>]). Problem finding is actually an umbrella term and has been operationalized in various ways, including problem discovery, problem formulation, problem generation, problem construction, and problem identification (Abdulla and Cramond [<reflink idref="bib1" id="ref41">1</reflink>]; Runco [<reflink idref="bib31" id="ref42">31</reflink>]). The problem generation tasks used here are not all‐encompassing measures of problem finding, but previous research has demonstrated that they provide reliable estimates (Abdulla Alabbasi et al. [<reflink idref="bib4" id="ref43">4</reflink>]; Ayoub et al. [<reflink idref="bib8" id="ref44">8</reflink>]; Okuda, Runco, and Berger [<reflink idref="bib25" id="ref45">25</reflink>]). The prediction here was that GAI would do poorly on problem generation tasks.</p> <p>To summarize, the objectives of this project were (a) to compare the three GAI systems in terms of the Idea Density and semantic distance of their responses; (b) to compare two instructional (prompt) conditions (one just asking for ideas and the other explicitly prompting original ideas); and (c) to compare the various DT tasks. The last of these comparisons involved two contrasts to test predictions mentioned above. One contrast compared Realistic and Standard (Nonrealistic) DT, and one comparing Problem Generation with the other DT tasks. (The DT tasks that do not ask for Problem Generation can be seen as problem solving rather than problem finding tasks, or as presented rather than discovered problems.) One particular comparison also helped to address the first objective (a). It contrasted raw scores from the GAI with adjusted scores. The adjustments removed verbiage that did not answer the question that was presented via the prompt. More details about the adjusted scores are given below, in the Method section. This adjustment was examined because it allowed a check of the effectiveness of the output. In that light, the adjustment has implications for the standard definition of creativity (Runco and Jaeger [<reflink idref="bib40" id="ref46">40</reflink>]; Runco [<reflink idref="bib33" id="ref47">33</reflink>]). That definition requires originality and effectiveness for creativity. Each of these has some latitude: originality may be novelty, for example, and effectiveness may be utility or appropriateness. Analyses of adjusted scores, reported below, were expected to offer more information about possible differences among the AI models but also provide information about how well the standard definition of creativity holds up with GAI. The final objective (d) was to compare responses from GAI with humans in terms of the idea density of their ideation. This final objective was secondary but justified by the research cited above, where GAI outperformed humans on various tests of creative potential (Guzik, Byrge, and Gilde [<reflink idref="bib15" id="ref48">15</reflink>]; Haase and Hanel [<reflink idref="bib16" id="ref49">16</reflink>]). The data we collected to address the first three objectives also provided an opportunity to compare of AI with humans, and given how important that comparison is (and what was reported in the previous research), it did seem to be worthwhile to determine what the present data said about AI versus humans.</p> <hd id="AN0187843392-2">Method</hd> <p>Tasks were presented to Bard, GPT 3.5, and GPT 4.0 in the Summer of 2023. GPT 3.5 was date stamped "ChatGPT 24 May 2023" and Bard "July 2023." These three tools were chosen for the same reasons Haase and Hanel ([<reflink idref="bib16" id="ref50">16</reflink>]) selected their systems: "based on their free usability and similar functions to ensure comparability" (p. 2). Additionally, Bard, GPT 4.0, and GPT 3.5 are all widely available, meaning that results should apply to the activities of broad audiences. The DT tasks (described below) were administered via different computers. Browsers were alternated for various tasks. The order of administration was varied such that it was unlikely that the AI systems would learn from the interaction history. Responses were captured verbatim and pasted into a word processing document, which was later scored for Idea Density, adjusted Idea Density, and Semantic Distance, each of which is explained below.</p> <hd id="AN0187843392-3">Divergent Thinking Tasks</hd> <p>Fifty‐five DT tasks were administered. All of them have been used previously in published empirical research and are available in those publications (Abdulla Alabbasi et al. [<reflink idref="bib4" id="ref51">4</reflink>]; Ames and Runco [<reflink idref="bib7" id="ref52">7</reflink>]; Ayoub et al. [<reflink idref="bib8" id="ref53">8</reflink>]; Okuda, Runco, and Berger [<reflink idref="bib25" id="ref54">25</reflink>]; Runco [<reflink idref="bib30" id="ref55">30</reflink>]). An overview of all tasks is presented in Table 1. Each task is open‐ended and allows multiple responses. Each also allowed a simple revision for the condition where the GAI systems were explicitly prompted to be creative (details below). This allowed us to address the second objective (b) of this investigation, described earlier. All DT tasks were given with prompts that did not mention creativity, but later, with controls to preclude AI learning (i.e., varied devices, varied browsers, and order of administration), were also given with a prompt that required that the responses were creative (e.g., "what are some creative names...," "be creative,"). This kind of prompting is consistent with the research on explicit instructions given to humans when they receive DT tests (Hong, O'Neil Jr., and Peng [<reflink idref="bib19" id="ref56">19</reflink>]; Runco, Illies, and Eisenman [<reflink idref="bib38" id="ref57">38</reflink>]; Runco, Illies, and Reiter‐Palmon [<reflink idref="bib39" id="ref58">39</reflink>]).</p> <p>1 TABLE Overview of the tasks administered to GPT 3.5, GPT 4.0, and Bard.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Test&lt;/th&gt;&lt;th align="center"&gt;Number of items&lt;/th&gt;&lt;th align="center"&gt;Example questions&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Realistic DT&lt;/td&gt;&lt;td align="center"&gt;Seven tasks&lt;/td&gt;&lt;td align="center"&gt;"I am walking several miles from home, and it started raining but I do not have a hat or umbrella. How can I stay dry?""I am riding my bike, far from home, and I get a flat tire. What can I do?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Problem Generation&lt;/td&gt;&lt;td align="center"&gt;Thirteen tasks&lt;/td&gt;&lt;td align="center"&gt;"What problems might someone have with his or her spouse?""What problems might occur while traveling?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Instances&lt;/td&gt;&lt;td align="center"&gt;Six tasks&lt;/td&gt;&lt;td align="center"&gt;"What things are really strong?" "What things are really loud?""What things are fast?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Uses&lt;/td&gt;&lt;td align="center"&gt;Six tasks&lt;/td&gt;&lt;td align="center"&gt;"What are possible uses for a spoon?""What are possible uses for Chopsticks?""What are possible uses for aluminum foil?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Similarities&lt;/td&gt;&lt;td align="center"&gt;Seven tasks&lt;/td&gt;&lt;td align="center"&gt;"How are a ranch and a farm alike?""How are the Sky and the Ocean alike?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SWOT&lt;/td&gt;&lt;td align="center"&gt;Four tasks&lt;/td&gt;&lt;td align="center"&gt;"What are the advantages of working for oneself?""What are the potential drawbacks or weaknesses of working for oneself?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Social Games&lt;/td&gt;&lt;td align="center"&gt;Seven tasks&lt;/td&gt;&lt;td align="center"&gt;"What are some ways of conveying to someone that they are not dressed appropriately?""What are some ways to convey the idea that you like someone but only as a friend and not in a romantic way?"&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Name Game&lt;/td&gt;&lt;td align="center"&gt;Five tasks&lt;/td&gt;&lt;td align="center"&gt;"What are some good names for a boat?""What are good names for your sweatheart?"&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 Abbreviations: DT, divergent thinking; SWOT, strengths, weaknesses, opportunities, and threats.</p> <p>Semantic distance for the six Instances items and the six Uses items was found by entering the AI responses into the Open Creativity Scoring (OCS) platform (Organisciak and Dumas [<reflink idref="bib27" id="ref59">27</reflink>]), https://openscoring.du.edu/scoring. A second set of semantic distance scores were found by asking GPT 3.5 for the semantic distance of responses. This required that GPT 3.5 was given the prompt as well as the specific response. GPT 4.0 does not provide semantic distances.</p> <hd id="AN0187843392-4">Adjusting the AI Responses for Relevance</hd> <p>Some of the AI responses contained verbiage which was not relevant to the question posed by the DT task (and the prompt given to the GAI). Raw responses, which were verbatim output from GAI, were analyzed, but a second set of responses was prepared as a way to approximate how effectively AI solved the problems posed by the DT tasks. This was only an approximation of effectiveness, but given that AI is sometimes accused of being ineffective (e.g., it may hallucinate), we felt it was worthwhile to explore the relevance of the responses. Additionally, effectiveness is part of the standard definition of creativity (Runco and Jaeger [<reflink idref="bib40" id="ref60">40</reflink>]), which also made it worthwhile to examine the adjusted scores. The analyses summarized below used both raw Idea Density scores and adjusted Idea Density scores.</p> <p>As noted above, the first objective of this research (a) was in part addressed by comparing raw and adjusted scores, which were based on the ideational output. The adjusted scores removed superfluous information—information that was not actually part of a response to the problem administered. The irrelevance of the verbiage removed by the adjustments is obvious from the detailed list of what was involved in the adjustments, which is given in the Data S1. Summarizing the things that were removed for the adjustments: (a) repetition of the question in the response; (b) solutions given when asked about problems (in the Problem Generation tasks) (because the prompt asked for problems, not solutions); (c) differences given in the Similarities task (because the prompt asked for similarities, not differences); and (d) repetition at the end of a response, as a sort of summary. In this manner, there were raw Idea Density scores, which contained everything given in the GAI responses, and adjusted Idea Density scores, which had removed wording that was superfluous (as in exact repetitions) or did not answer the question.</p> <hd id="AN0187843392-5">Results</hd> <p></p> <hd id="AN0187843392-6">Reliability and Descriptive Statistics</hd> <p>Table 2 presents the reliability and descriptive statistics for the Idea Density scores.</p> <p>2 TABLE Means, standard deviations, and inter‐item reliabilities (alphas) for the three GAI models (Bard, GPT3.5, and GPT4) and the two prompt conditions (explicit vs. standard).</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="center"&gt;Bard&lt;/th&gt;&lt;th align="center"&gt;GPT3.5&lt;/th&gt;&lt;th align="center"&gt;GPT 4.0&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="center"&gt;Explicit&lt;/th&gt;&lt;th align="center"&gt;Standard&lt;/th&gt;&lt;th align="center"&gt;Explicit&lt;/th&gt;&lt;th align="center"&gt;Standard&lt;/th&gt;&lt;th align="center"&gt;Explicit&lt;/th&gt;&lt;th align="center"&gt;Standard&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;M&lt;/td&gt;&lt;td align="center"&gt;0.516&lt;/td&gt;&lt;td align="center"&gt;0.492&lt;/td&gt;&lt;td align="center"&gt;0.503&lt;/td&gt;&lt;td align="center"&gt;0.492&lt;/td&gt;&lt;td align="center"&gt;0.484&lt;/td&gt;&lt;td align="center"&gt;0.483&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;SD&lt;/td&gt;&lt;td align="center"&gt;0.046&lt;/td&gt;&lt;td align="center"&gt;0.050&lt;/td&gt;&lt;td align="center"&gt;0.038&lt;/td&gt;&lt;td align="center"&gt;0.053&lt;/td&gt;&lt;td align="center"&gt;0.039&lt;/td&gt;&lt;td align="center"&gt;0.051&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Cronbach's &amp;#945;&lt;/td&gt;&lt;td align="center"&gt;0.65&lt;/td&gt;&lt;td align="center"&gt;0.73&lt;/td&gt;&lt;td align="center"&gt;0.72&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>2 <emph>Note:</emph> These means and SDs were calculated across all 55 divergent thinking tasks.</p> <hd id="AN0187843392-7">Idea Density</hd> <p>A hierarchical multiple regression analysis was conducted using the raw (unadjusted) Idea Density score as the criterion. Predictors were entered in four steps of the analysis. The first step of the analysis entered two dichotomous variables to test the magnitude of the differences among the DT tests. Given claims that AI is unable to find problems (Csikszentmihalyi [<reflink idref="bib13" id="ref61">13</reflink>]; Runco [<reflink idref="bib32" id="ref62">32</reflink>]), one of the dichotomous variables allowed a contrast of the Problem Generation tasks with all of the other DT tasks. Recall here that one prediction was that GAI would not do well with problem generation. The second dichotomous DT variable allowed a contrast of all of the Realistic DT tests (Realistic DT, Problem Generation, SWOT, Social Games) with the Nonrealistic DT tests (Instances, Uses, and Similarities). The prediction here, mentioned earlier, was that GAI would do well with realistic tasks. The second step of the analysis entered a dichotomous variable allowing a comparison of the directions that explicitly prompted unusual and original responses with directions which did not prompt originality. This provided a test of the prediction that explicit instructions given to GAI would boost originality. The third step of the analysis entered two dichotomous variables which allowed comparisons of (a) Bard versus the two GPT systems and (b) GPT 3.5 versus GPT 4.0. The prediction here was that 4.0 would provide more ideas than 3.5. The last step of the analysis entered the adjusted Idea Density score (a continuous scale variable). The adjusted score was entered to determine whether the raw scores contained superfluous information or whether the adjusted scores were more fitting and appropriate. Recall here that the standard definition requires appropriateness (sometimes called effectiveness), in addition to originality (Runco and Jaeger [<reflink idref="bib40" id="ref63">40</reflink>]).</p> <p>Results indicated that there were significant differences among the DT tests (<emph>R</emph> = 0.45, <emph>F</emph>(<reflink idref="bib2" id="ref64">2</reflink>, 327) = 40.40, <emph>p</emph> &lt; 0.001). The second step of the analysis confirmed that there were significant differences in Idea Density resulting from the prompts (explicit vs. nonexplicit) (<emph>R</emph> = 0.46, <emph>ΔR</emph><sups>2</sups> = 0.017, <emph>F</emph>(<reflink idref="bib1" id="ref65">1</reflink>, 326) = 7.01, <emph>p</emph> = 0.008). The third step confirmed differences among the AI models (<emph>R</emph> = 0.50, <emph>ΔR</emph><sups>2</sups> = 0.032, <emph>F</emph>(<reflink idref="bib2" id="ref66">2</reflink>, 325) = 6.85, <emph>p</emph> = 0.001). The last step of the analysis confirmed a significant relationship between the corrected and the raw Idea Density scores (<emph>R</emph> = 0.82, <emph>ΔR</emph><sups>2</sups> = 0.43, <emph>F</emph>(<reflink idref="bib1" id="ref67">1</reflink>, 323) = 423.0, <emph>p</emph> &lt; 0.001). The squared coefficient 0.43 indicates that the adjusted Idea Density scores were associated with the raw Idea Density scores, with 43% shared variance. This will be explored in the Discussion, but that correlation does imply that the two Idea Density scores are far from redundant. Put differently, a prediction of the raw Idea Density scores from the adjusted scores would be only moderately accurate. A product moment correlation (testing only the raw (<emph>M</emph> = 0.4954, SD = 0.048) and the adjusted (<emph>M</emph> = 0.4831, SD = 0.078). Idea Density scores and none of the other IVs or contrasts) confirmed this; it indicated that the shared variance was 55% (<emph>r</emph> = 0.744).</p> <p>Two of the objectives of this investigation mentioned earlier (which were labeled (a) and (c)) were addressed by examining the regression model with all predictors included. Only one of the two DT contrasts was significant. This was the contrast of the Realistic (<emph>M</emph> = 0.4586) versus Nonrealistic (<emph>M</emph> = 0.4961) DT tests (beta = 0.293, <emph>t</emph>(<reflink idref="bib323" id="ref68">323</reflink>) = 8.28, <emph>p</emph> &lt; 0.001). The Problem Generation tasks (<emph>M</emph> = 0.5059, SD = 0.045) were not significantly different from the all of the other tasks (<emph>M</emph> = 0.4761, SD = 0.085). Both AI contrasts were significant (betas = 0.266 and 0.350, <emph>ts</emph>(<reflink idref="bib323" id="ref69">323</reflink>) = 4.17 and 5.51, <emph>ps</emph> &lt; 0.001). Bard (<emph>M</emph> = 0.4752) differed significantly from the two GPTs (<emph>M</emph> = 0.4990), and as predicted, GPT 3.5 (<emph>M</emph> = 0.4670) differed from GPT 4.0 (<emph>M</emph> = 0.4879). Only one variable included in the regression analysis had a VIF value above 4.0, and it was below 4.10. Regression results are summarized in Table 3.</p> <p>3 TABLE Hierarchical regression examining the type of DT task, type of prompt, AI tools, and the adjusted Idea Density with raw Idea Density Scores (N = 330).</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left"&gt;Variable&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;&amp;#946;&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;t&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;p&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;&lt;italic&gt;2&lt;/italic&gt;&lt;/sup&gt;&lt;/th&gt;&lt;th align="center"&gt;&amp;#916;&lt;italic&gt;R&lt;/italic&gt;&lt;sup&gt;&lt;italic&gt;2&lt;/italic&gt;&lt;/sup&gt;&lt;/th&gt;&lt;th align="center"&gt;&amp;#916;&lt;italic&gt;F&lt;/italic&gt;&lt;/th&gt;&lt;th align="center"&gt;95% CI&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Step 1&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Constant&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;64.10&lt;/td&gt;&lt;td align="center"&gt;0.007&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;0.20&lt;/td&gt;&lt;td align="center"&gt;0.20&lt;/td&gt;&lt;td align="center"&gt;40.40&lt;/td&gt;&lt;td align="center"&gt;[0.457, 0.486]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (Realistic vs. Unrealistic)&lt;/td&gt;&lt;td align="center"&gt;0.43&lt;/td&gt;&lt;td align="center"&gt;7.86&lt;/td&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[0.032, 0.054]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (PG vs. Other DT)&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.82&lt;/td&gt;&lt;td align="center"&gt;0.006&lt;/td&gt;&lt;td align="center"&gt;0.412&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.017, 0.007]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Step 2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Constant&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;60.72&lt;/td&gt;&lt;td align="center"&gt;0.008&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;0.22&lt;/td&gt;&lt;td align="center"&gt;0.02&lt;/td&gt;&lt;td align="center"&gt;7.01&lt;/td&gt;&lt;td align="center"&gt;[0.450, 0.480]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (Realistic vs. Unrealistic)&lt;/td&gt;&lt;td align="center"&gt;0.43&lt;/td&gt;&lt;td align="center"&gt;7.93&lt;/td&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[0.032, 0.054]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (PG vs. Other DT)&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.83&lt;/td&gt;&lt;td align="center"&gt;0.006&lt;/td&gt;&lt;td align="center"&gt;0.408&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.017, 0.007]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Instructions (Explicit vs. Standard)&lt;/td&gt;&lt;td align="center"&gt;0.13&lt;/td&gt;&lt;td align="center"&gt;2.65&lt;/td&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;0.008&lt;/td&gt;&lt;td align="center"&gt;[0.003, 0.022]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Step 3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Constant&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;46.23&lt;/td&gt;&lt;td align="center"&gt;0.010&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;0.25&lt;/td&gt;&lt;td align="center"&gt;0.03&lt;/td&gt;&lt;td align="center"&gt;6.85&lt;/td&gt;&lt;td align="center"&gt;[0.441, 0.480]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (Realistic vs. Unrealistic)&lt;/td&gt;&lt;td align="center"&gt;0.43&lt;/td&gt;&lt;td align="center"&gt;8.07&lt;/td&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[0.033, 0.054]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (PG vs. Other DT)&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.04&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.84&lt;/td&gt;&lt;td align="center"&gt;0.006&lt;/td&gt;&lt;td align="center"&gt;0.400&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.017, 0.007]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Instructions (Explicit vs. Standard)&lt;/td&gt;&lt;td align="center"&gt;0.13&lt;/td&gt;&lt;td align="center"&gt;2.70&lt;/td&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;0.007&lt;/td&gt;&lt;td align="center"&gt;[0.003, 0.022]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;AI Platform (Bard vs. GPT3.5 &amp; GPT4)&lt;/td&gt;&lt;td align="center"&gt;0.07&lt;/td&gt;&lt;td align="center"&gt;0.705&lt;/td&gt;&lt;td align="center"&gt;0.010&lt;/td&gt;&lt;td align="center"&gt;0.481&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.012, 0.026]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;AI Platform (GPT3.5 vs. GPT4)&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.23&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;2.44&lt;/td&gt;&lt;td align="center"&gt;0.006&lt;/td&gt;&lt;td align="center"&gt;0.015&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.025, &amp;#8722;0.003]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Step 4&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Constant 4&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;20.47&lt;/td&gt;&lt;td align="center"&gt;0.012&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;0.67&lt;/td&gt;&lt;td align="center"&gt;0.43&lt;/td&gt;&lt;td align="center"&gt;423.00&lt;/td&gt;&lt;td align="center"&gt;[0.225, 0.273]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (Realistic vs. Unrealistic)&lt;/td&gt;&lt;td align="center"&gt;0.29&lt;/td&gt;&lt;td align="center"&gt;8.28&lt;/td&gt;&lt;td align="center"&gt;0.004&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[0.023, 0.037]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT (PG vs. Other DT)&lt;/td&gt;&lt;td align="center"&gt;0.01&lt;/td&gt;&lt;td align="center"&gt;0.364&lt;/td&gt;&lt;td align="center"&gt;0.004&lt;/td&gt;&lt;td align="center"&gt;0.716&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.006, 0.009]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Instructions (Explicit vs. Standard)&lt;/td&gt;&lt;td align="center"&gt;0.10&lt;/td&gt;&lt;td align="center"&gt;3.01&lt;/td&gt;&lt;td align="center"&gt;0.003&lt;/td&gt;&lt;td align="center"&gt;0.003&lt;/td&gt;&lt;td align="center"&gt;[0.003, 0.015]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;AI Platform (Bard vs. GPT3.5 &amp; GPT4)&lt;/td&gt;&lt;td align="center"&gt;0.27&lt;/td&gt;&lt;td align="center"&gt;4.17&lt;/td&gt;&lt;td align="center"&gt;0.007&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[0.014, 0.040]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;AI Platform (GPT3.5 vs. GPT4)&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;0.35&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;5.51&lt;/td&gt;&lt;td align="center"&gt;0.004&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[&amp;#8722;0.028, &amp;#8722;0.013]&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Idea Density&lt;/td&gt;&lt;td align="center"&gt;0.68&lt;/td&gt;&lt;td align="center"&gt;20.57&lt;/td&gt;&lt;td align="center"&gt;0.020&lt;/td&gt;&lt;td align="center"&gt;&amp;#60;&amp;#8201;0.001&lt;/td&gt;&lt;td align="center"&gt;[0.380, 0.461]&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>3 Abbreviations: AI, artificial intelligence; DT, divergent thinking; PG, problem generation.</p> <hd id="AN0187843392-8">Semantic Distance</hd> <p>Two of the DT tests, namely Instances and Uses, could be scored for semantic distance using (a) the OCS online software, (b) Bard, and (c) GPT 3.5. (GPT 4.0 does not provide semantic distance). GPT 3.5 provided scores with little variability (only 0.2 &lt; 0.7). These were included in analyses but the low variability must be taken into account when interpreting the results. Low variability would, for instance, attenuate any correlations. Each of the DT tests had six prompts. Figures 1–3 show the distributions of the scores. These suggest that the OCS semantic distances scores were well distributed. GPT 3.5 scores were poorly distributed, as expected, given the limited range. Bard was mostly well‐distributed, but there was a spike. These distributions could influence other analyses using the semantic distance scores.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/3U7/01sep25/jocb1528-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jocb1528-fig-0001.jpg" title="1 Frequency distribution of semantic distance scores as calculated By Bard." /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/3U7/01sep25/jocb1528-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jocb1528-fig-0002.jpg" title="2 Frequency distribution of semantic distance scores as calculated by Open Creativity Scoring." /> </p> <p></p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/3U7/01sep25/jocb1528-fig-0003.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jocb1528-fig-0003.jpg" title="3 Frequency distribution of semantic distance scores as calculated by ChatGPT 3.5." /> </p> <p></p> <p>Correlations among the three semantic distance scores were examined. There was a positive correlation between OCS and Bard (<emph>r</emph> = 0.177) and negative correlations of GPT 3.5 with both OCS and Bard (<emph>r</emph>s = −0.50 and −0.024). Based on the issues of restricted range and negative correlations (lack of convergent validity), GPT 3.5 was not included in the subsequent multilevel analyses.</p> <p>Multilevel modeling was next applied to the semantic distance scores from OCS and Bard. The responses were nested under the DT prompts. The first analysis used OCS semantic distance with random effects. The intraclass correlation was 0.278. This indicates a substantial amount of clustering, justifying the need for a multilevel model. Starting with the unconditional model, three predictors (type of instructions, type of DT test, and GAI platform) were tested. Each predictor was added as a fixed effect, and if significant, its random slope was tested. Models were assessed on the basis of AIC and BIC values. As presented in Table 4, the type of instructions (explicit vs. nonexplicit) and type of DT test were added to the model in respective steps. They were not significant, and thus, their random slopes were not tested. Then, the type of GAI was added to the model, where Bard was used as a reference category for GPT 3.5 and GPT 4.0 Type of GAI predicted semantic distance such that it was significantly <emph>lower</emph> when the responses were generated by GPT 3.5 than Bard. When the random slope was added, it was significant, yet fixed effects were not. Following the suggestions by Bosker and Snijders ([<reflink idref="bib9" id="ref70">9</reflink>]), the simpler model (Model 1) was selected as the final model. This points to Bard as a tenable platform of semantic distance measurement and a solid GAI generation platform to produce ideas that are likely to be recognized by OCS.</p> <p>4 TABLE Multilevel models predicting semantic distance (SD) scores measured by Open Creativity Scoring (OCS) and Bard.</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;th align="left" /&gt;&lt;th align="center"&gt;OCS&amp;#8208;SD&lt;/th&gt;&lt;th align="center"&gt;Bard&amp;#8208;SD&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="center"&gt;Model 0&lt;/th&gt;&lt;th align="center"&gt;Model 1&lt;/th&gt;&lt;th align="center"&gt;Model 0&lt;/th&gt;&lt;th align="center"&gt;Model 1&lt;/th&gt;&lt;th align="center"&gt;Model 2&lt;/th&gt;&lt;th align="center"&gt;Model 3&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th align="center"&gt;Estimates&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;th align="center"&gt;Estimates&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;th align="center"&gt;Estimates&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;th align="center"&gt;Estimates&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;th align="center"&gt;Estimates&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;th align="center"&gt;Estimates&lt;/th&gt;&lt;th align="center"&gt;SE&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody valign="top"&gt;&lt;tr&gt;&lt;td align="left"&gt;Fixed effects&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Intercept&lt;/td&gt;&lt;td align="center"&gt;0.749&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.02&lt;/td&gt;&lt;td align="center"&gt;0.758&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.021&lt;/td&gt;&lt;td align="center"&gt;0.587&lt;/td&gt;&lt;td align="center"&gt;0.023&lt;/td&gt;&lt;td align="center"&gt;0.558&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.022&lt;/td&gt;&lt;td align="center"&gt;0.556&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.028&lt;/td&gt;&lt;td align="center"&gt;0.546&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.029&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Instructions&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;0.095&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.012&lt;/td&gt;&lt;td align="center"&gt;0.097&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.035&lt;/td&gt;&lt;td align="center"&gt;0.087&lt;xref ref-type="fn" rid="tfn6" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.035&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;DT test type&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;GPT3.5&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;&amp;#8722;0.022&lt;xref ref-type="fn" rid="tfn6" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.010&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;&amp;#8722;0.012&lt;/td&gt;&lt;td align="center"&gt;0.016&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;GPT4&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;0.011&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;0.058&lt;xref ref-type="fn" rid="tfn7" /&gt;&lt;/td&gt;&lt;td align="center"&gt;0.017&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Random effects&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Instructions&lt;/td&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center" /&gt;&lt;td align="center"&gt;0.014 (0.118)&lt;/td&gt;&lt;td align="center"&gt;0.014 (0.119)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Level 2&lt;/td&gt;&lt;td align="center"&gt;0.005 (0.070)&lt;/td&gt;&lt;td align="center"&gt;0.005&lt;/td&gt;&lt;td align="center"&gt;0.069&lt;/td&gt;&lt;td align="center"&gt;0.006 (0.078)&lt;/td&gt;&lt;td align="center"&gt;0.005 (0.074)&lt;/td&gt;&lt;td align="center"&gt;0.010 (0.098)&lt;/td&gt;&lt;td align="center"&gt;0.008 (0.091)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;Residual&lt;/td&gt;&lt;td align="center"&gt;0.013 (0.113)&lt;/td&gt;&lt;td align="center"&gt;0.013&lt;/td&gt;&lt;td align="center"&gt;0.113&lt;/td&gt;&lt;td align="center"&gt;0.037 (0.193)&lt;/td&gt;&lt;td align="center"&gt;0.035 (0.188)&lt;/td&gt;&lt;td align="center"&gt;0.033 (0.181)&lt;/td&gt;&lt;td align="center"&gt;0.032 (0.179)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;AIC&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;1525&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;1518&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;425&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;470&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;523&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;533&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td align="left"&gt;BIC&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;1510&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;1494&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;410&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;450&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;493&lt;/td&gt;&lt;td align="center"&gt;&amp;#8722;494&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>4 <emph>Note:</emph> Values in parentheses are standard deviations. OCS (Organisciak and Dumas [<reflink idref="bib27" id="ref71">27</reflink>]) https://openscoring.du.edu/scoring.</item> <item>5 Abbreviations: AIC, Akaike information criterion; BIC, Bayesian information criterion; DT, divergent thinking; OCS, Open Creativity Scoring.</item> <item>6 * <emph>p</emph> &lt; 0.05.</item> <item>7 ** <emph>p</emph> &lt; 0.01.</item> </ulist> <p>The Bard‐scored semantic distance scores were then examined as the dependent variable. The intraclass correlation was 0.125. The key models are presented in Table 4. DT test type was not a significant predictor whereas type of instructions was. Explicit instructions led to significantly higher semantic distance scores as measured by Bard, compared with the inexplicit instructions. Furthermore, the random slope of the explicit instructions was significant, showing that the impact of explicit instructions varied across the different DT prompts. That means that the explicit instructions helped the semantic distance scores of certain prompts more than the others. Finally, the type of GAI was added to the model, which was significant for GPT 4.0. That means that responses generated by GPT 4.0 had a greater semantic distance than those generated by Bard.</p> <p>Putting these findings together, the multilevel analyses indicated that Bard generated more semantically distant responses than GPT 3.5 when semantic distance was measured by OCS, whereas GPT 4.0 produced more semantically distant responses than Bard when the semantic distance was measured by Bard. Furthermore, semantic distance scores, as measured by Bard, were more sensitive to the explicit instructions than OCS‐measured semantic distance scores.</p> <hd id="AN0187843392-12">Idea Density From GAI vs. Humans</hd> <p>Although the primary objectives of this investigation only required comparisons of the three GAI systems, GAI responses elicited by standard versus explicit prompts, and raw versus adjusted GAI Idea Density Scores (reported below), one of the DT tests used here was the same as that used in an earlier investigation of the Idea Density of humans (Runco et al. [<reflink idref="bib42" id="ref72">42</reflink>]). Thus, the GAI data collected here could be compared with that of humans, to address the last (albeit secondary) objective (d) of this investigation. These data were specifically from the Uses DT test (not all 55 DT tasks). The highest Uses mean Idea Density in the present investigation was from GPT 3.5 (<emph>M</emph> = 0.482, SD = 0.030), but a <emph>t</emph>‐test indicated that even it was significantly lower than the average from the 41 humans (university students) who had received the Uses test in the previous investigation (<emph>M</emph> = 0.599, SD = 052), <emph>t</emph>(<reflink idref="bib145" id="ref73">145</reflink>) = 16.61, <emph>p</emph> &lt; 0.001, <emph>d</emph> = 3.54; <emph>g</emph> = 3.53.</p> <hd id="AN0187843392-13">Are AI Answers Appropriate?</hd> <p>The regression analysis above indicated that the raw and the adjusted AI responses shared 43% of their variance. The product moment analysis gave a slightly higher figure: 55% shared variance. One other method, used in previous creativity research, provided an estimate of the appropriateness of AI responses. This method was developed in an investigation of children's DT. Runco and Charles ([<reflink idref="bib37" id="ref74">37</reflink>]) used it to assess the appropriateness or aptness of answers give to two DT questions. One question was "Name all of the square things that you can think of." Runco and Charles reasoned that an entirely appropriate answer to that question would be something that was a literal square (four equal sides) and not a rectangle (e.g., a door) nor metaphorically square (a square meal, old music, a square root). This method could also be used with one other question, namely "list things that move on wheels." The criteria for a literally appropriate answer to this question included (a) whatever was listed had to actually move from point A to point B (and was not stationary, like a clock) and (b) the wheels had to be actual wheels and not just parts that allowed movement (e.g., the treads of a military tank or tractor). These are quite stringent requirements which only give credit to literally correct ideas, but Runco and Charles ([<reflink idref="bib37" id="ref75">37</reflink>]) did find this method to produce reliable scoring, and it does offer information about the effectiveness criterion (from the standard definition of creativity), given that effectiveness is sometimes defined in terms of utility, aptness, usefulness, or appropriateness. With this in mind the same method was used in the present investigation. The same two Instances questions (Square things, things that move on Wheels) were among the 55 questions administered to the AI in the present research.</p> <p>Two judges independently examined all ideas from the three GAI systems and identified those which did not fit the criteria summarized above. One judge was one of the co‐authors of this paper, but he was unaware of which platform or instructional condition elicited which ideas. Additionally, the ideas were randomly presented to the two judges. The judges agreed 95% of the time. One thing that came up in the present data that did not show up in the earlier investigation of appropriateness: In the responses to the Instances question about Square things, AI sometimes added the adjective, "square," as in "square coins" and "square apples." In these cases, the adjective was ignored. Judgments were based on objects being "always or typically square." A comparison of the raw and adjusted AI responses indicated that 67.2% of the ideas give to the Square and the Move on Wheels Instances tasks were appropriate and relevant.</p> <hd id="AN0187843392-14">Discussion</hd> <p>This investigation of AI and creative potential had four objectives: (a) to compare three GAI tools in terms of the Idea Density and semantic distance of their responses; (b) to compare two instructional (prompt) conditions (one just asking for ideas, and the other explicitly prompting original ideas); (c) to compare various DT tasks (including Realistic vs. Standard, Nonrealistic DT tasks, and Problem Generation with standard problem solving DT tasks); and (d) to determine the degree to which GAI produced output which was in fact relevant to the DT task at hand. Several results from the analyses were quite clear. There was, for example, a difference (as predicted) between the Idea Density scores of Realistic DT and the Nonrealistic DT tests. (The latter are not entirely unrealistic, so nonrealistic is probably the best label.) There are two ways to look at this. First, GAI draws on existing information, and very likely more of that is Realistic than Nonrealistic. Second, assuming that AI can in fact make connections among bits of information that it finds (and in that sense goes beyond what it finds in the LLM), a proposal from Runco and Albert ([<reflink idref="bib36" id="ref76">36</reflink>]) may apply. They had compared figural and verbal DT tests and found that the former elicited more original ideas than the latter. They hypothesized that this was because rote or existing associations can usually be used when responding to verbal tasks. Figural tasks, on the contrary, are dissimilar to what has been previously experienced so ideas must be truly imaginative and are more difficult to find. The same reasoning may explain the difference between Realistic and Nonrealistic tasks. It may be easier to find ideas in existing data bases when dealing with Realistic tasks.</p> <p>The relationship of the raw and the corrected Idea Density scores was moderate, with 43% of the variance shared. This was determined by the hierarchical regression, after the other predictors (i.e., type of DT test, instructions, and GAI platform) had been added to the equation and their variance controlled. Recall here that adjusted Idea Density had been calculated by removing superfluous information from the AI output. Some of the time that information served to make the responses conversational, which may make things easier for some users, but the superfluous information did make the responses inefficient in the sense that they contained unnecessary information. Worse, quite a few of the raw responses contained information that was entirely inapplicable to the tasks conveyed by the prompts. The prompts presented problems to the AI, and the responses were in a sense solution, and clearly irrelevant information is ineffective in a solution. It does not help solve the problem. This was the clearest in the Similarities questions (e.g., How are the sky and the ocean alike?), when the AI gave similarities but then also listed differences, and in the Problem Generation prompts, where the AI gave problems but then also courses of action to deal with those problems. Certainly, the adjustment used here is not the only way to estimate the effectiveness (or ineffectiveness) of AI output, but we felt that it was worth exploring this one estimation, given how important effectiveness is in the standard definition of creativity (Runco and Jaeger [<reflink idref="bib40" id="ref77">40</reflink>]). The measurement of appropriateness reported above (resulting in the 67% estimate of appropriate GAI output) is also a rough approximation of effectiveness and complements the regression results. In a sense, the hallucinations and confabulations of AI uncovered in other research (cf. Østergaard and Nielbo [<reflink idref="bib28" id="ref78">28</reflink>]) might be interpreted as another kind of ineffectiveness.</p> <p>The multilevel analyses indicated that explicit instructions elicited larger semantic distances from the GAI than the inexplicit instructions, at least when the semantic scores were obtained from GPT 3.5 or Bard. This was not true of the semantic distances obtained from OCS. Semantic distance was larger when responses were generated by GPT 4.0 than Bard. Interestingly, the OCS website has added an AI (LLM) score as an option to semantic distance and the AI score seems to work even better than semantic distance (Organisciak et al. [<reflink idref="bib26" id="ref79">26</reflink>]). As noted above, the present results suggested that Bard was a useful platform for semantic distance measurement and a useful AI generation platform to produce ideas when they were measured by OCS.</p> <p>A <emph>t</emph>‐test reported in the Results indicated that none of the GAI Idea Density scores was as high as that of humans, but this was based on a small sample of university students and only the Uses test. Any conclusions based on this comparison are tentative, but it may be that the Idea Density for AI was lower because it controls for the number of words used. Humans may have been frugal with the number of words because of the effort required to write longer responses. There is no such expense for AI, so it can be verbose, relative to humans, leading to a lower ratio (fewer ideas after controlling for the number of words). The difference between GAI and humans in terms of Idea Density might be explored in future research. Perhaps that research could tweak the Idea Density algorithm such that it focuses on original ideas instead of ideas regardless of their novelty. Research also might continue to investigate humans working along with GAI, or what Ivcevic and Grandinetti ([<reflink idref="bib20" id="ref80">20</reflink>]) called <emph>co‐creation</emph>. The present investigation was focused on the idea density and semantic distance of GAI when it worked alone, but it would be interesting to examine the output when GAI worked with humans. This is likely how GAI will most often be used, together with humans.</p> <p>There are limitations to the research reported here. There are other AI models, for example. The present research only compared three of them. Then, there is the fact that Idea Density is not a measure of creativity. It only measures ideation. Ideas are useful for creativity, but there is more to creativity than ideation, as is implied by the standard definition and other research on the things that contribute to authentic creativity (Runco [<reflink idref="bib32" id="ref81">32</reflink>]). We did measure semantic distance in this investigation, but only for 12 of the 55 tasks. As mentioned above, future research might adjust the Idea Density algorithm such that it recognizes only original ideas instead of all ideas. Perhaps future research could examine GAI performance resulting from prompts in other languages. Also, GAI is changing quite rapidly. Hence, the results reported here should be replicated with new GAI when it is released. The changes in GAI suggest a limitation vis‐à‐vis reliability. Because GAI is changing, there might be a question about test–retest reliability. This form of reliability requires stability across time, and changes in GAI would indicate a kind of instability. Then again, it is easy to explain this kind of instability, and there are other forms of reliability that support the use of GAI. The present investigation relied on inter‐item reliability. Future research might explore different indices of reliability. It should also look beyond mere ideation, perhaps using more tasks that allow the calculation of semantic distance. There certainly is a benefit to relying on automated scoring when investigating GAI, as we did here.</p> <hd id="AN0187843392-15">Acknowledgments</hd> <p>The authors have nothing to report.</p> <hd id="AN0187843392-16">Conflicts of Interest</hd> <p>The authors declare no conflicts of interest.</p> <hd id="AN0187843392-17">Data Availability Statement</hd> <p>Data are available from the correspondence author.</p> <p>GRAPH: Data S1</p> <ref id="AN0187843392-18"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref41" type="bt">1</bibl> <bibtext> Funding: The authors received no specific funding for this work.</bibtext> </blist> </ref> <ref id="AN0187843392-19"> <title> References </title> <blist> <bibtext> Abdulla, A. M., and B. Cramond. 2018. " The Creative Problem Finding Hierarchy: A Suggested Model for Understanding Problem Finding." Creativity. Theories‐ Research‐ Applications 5, no. 2 : 197 – 229. https://doi.org/10.1515/ctra‐2018‐0019.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref16" type="bt">2</bibl> <bibtext> Abdulla Alabbasi, A. M., S. Acar, R. Al‐Shehri, and F. Aljasim. 2024. " Testing the Effects of Time‐On‐Task and Instructions to "Be Creative" on Gifted Students." Gifted Education International 40, no. 1 : 67 – 91. https://doi.org/10.1177/02614294231173783.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref18" type="bt">3</bibl> <bibtext> Abdulla Alabbasi, A. M., S. Paek, and B. Cramond. 2022. " What Do Educators Need to Know About the Torrance Tests of Creative Thinking? A Comprehensive Review." Frontiers in Psychology 13 : 1000385. https://doi.org/10.3389/fpsyg.2022.1000385.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref43" type="bt">4</bibl> <bibtext> Abdulla Alabbasi, A. M., R. Reiter‐Palmon, Z. Sultan, and A. Ayoub. 2021. " Which Divergent Thinking Index Is More Associated With Problem Finding Ability? The Role of Flexibility and Task Nature." Frontiers in Psychology 12 : 671146. https://doi.org/10.3389/fpsyg.2021.671146.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref28" type="bt">5</bibl> <bibtext> Acar, S., K. Berthiaume, K. Grajzel, D. Dumas, C. T. Flemister, and P. Organisciak. 2023. " Applying Automated Originality Scoring to the Verbal Form of Torrance Tests of Creative Thinking." Gifted Child Quarterly 67, no. 1 : 3 – 17. https://doi.org/10.1177/00169862211061874.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref17" type="bt">6</bibl> <bibtext> Acar, S., and M. A. Runco. 2014. " Assessing Associative Distance Among Ideas Elicited by Tests of Divergent Thinking." Creativity Research Journal 26, no. 2 : 229 – 238. https://doi.org/10.1080/10400419.2014.901095.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref52" type="bt">7</bibl> <bibtext> Ames, M., and M. A. Runco. 2005. " Predicting Entrepreneurship From Ideation and Divergent Thinking." Creativity and Innovation Management 14, no. 3 : 311 – 315. https://doi.org/10.1111/j.1467‐8691.2004.00349.x.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref44" type="bt">8</bibl> <bibtext> Ayoub, A. E. A., A. M. Abdulla Alabbasi, A. M. Alsubaie, M. A. Runco, and S. Acar. 2022. " Enhanced Open‐Mindedness and Problem Finding Among Gifted Female Students Involved in Future Robotics Design." Roeper Review 44, no. 2 : 85 – 93. https://doi.org/10.1080/02783193.2022.2043500.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref70" type="bt">9</bibl> <bibtext> Bosker, R., and T. A. Snijders. 2012. Multilevel Analysis: An Introduction to Basic and Advanced Multilevel Modeling. London, United Kingdom : Sage.</bibtext> </blist> <blist> <bibtext> Brainard, L. 2023. " The Curious Case of Uncurious Creation." Inquiry : 1 – 31. https://doi.org/10.1080/0020174X.2023.2261503.</bibtext> </blist> <blist> <bibtext> Brown, C., T. Snodgrass, S. J. Kemper, R. Herman, and M. A. Covington. 2008. " Automatic Measurement of Propositional Idea Density From Part‐Of‐Speech Tagging." Behavior Research Methods 40 : 540 – 545. https://doi.org/10.3758/BRM.40.2.540.</bibtext> </blist> <blist> <bibtext> Covington, M. A. 2009. "Idea Density — A Potentially Informative Characteristic of Retrieved Documents." IEEE 978‐1‐4244‐3978‐2/09. https://doi.org/10.1109/SECON.2009.5174076.</bibtext> </blist> <blist> <bibtext> Csikszentmihalyi, M. 1988. " Solving a Problem Is Not Finding a New One: A Reply to Simon." Journal New Ideas in Psychology 6 : 183 – 186. https://doi.org/10.1016/0732‐118X(88)90003‐7.</bibtext> </blist> <blist> <bibtext> Getzels, J. W., and M. Csikszentmihalyi. 1976. The Creative Vision: A Longitudinal Study of Problem Finding in Art. New York, NY : Wiley.</bibtext> </blist> <blist> <bibtext> Guzik, E., C. Byrge, and C. Gilde. 2023. " The Originality of Machines: AI Takes the Torrance Test." Journal of Creativity 33, no. 3 : 1 – 8. https://doi.org/10.1016/j.yjoc.2023.100065.</bibtext> </blist> <blist> <bibtext> Haase, J., and P. H. P. Hanel. 2023. " Artificial Muses: Generative Artificial Intelligence Chatbots Have Risen to Human‐Level Creativity." Journal of Creativity 33, no. 3 : 100066. https://doi.org/10.1016/j.yjoc.2023.100066.</bibtext> </blist> <blist> <bibtext> Hass, R. W., M. Rivera, and P. J. Silvia. 2018. " On the Dependability and Feasibility of Layperson Ratings of Divergent Thinking." Frontiers in Psychology 9 : 1 – 13. https://doi.org/10.3389/fpsyg.2018.01343.</bibtext> </blist> <blist> <bibtext> Heinen, D. J. P., and D. R. Johnson. 2018. " Semantic Distance: An Automated Measure of Creativity That Is Novel and Appropriate." Psychology of Aesthetics, Creativity, and the Arts 12, no. 2 : 144 – 156. https://doi.org/10.1037/aca0000125.</bibtext> </blist> <blist> <bibtext> Hong, E., H. F. O'Neil Jr., and Y. Peng. 2016. " Effects of Explicit Instructions, Metacognition, and Motivation on Creative Performance." Creativity Research Journal 28, no. 1 : 33 – 45. https://doi.org/10.1080/10400419.2016.1125252.</bibtext> </blist> <blist> <bibtext> Ivcevic, Z., and M. Grandinetti. 2024. " Artificial Intelligence as a Tool for Creativity." Journal of Creativity 34, no. 2 : 1 – 5. https://doi.org/10.1016/j.yjoc.2024.100079.</bibtext> </blist> <blist> <bibtext> Koubaa, A. 2023. "GPT‐4 vs. GPT‐3.5: A Concise Showdown." https://doi.org/10.36227/techrxiv.22312330.v2.</bibtext> </blist> <blist> <bibtext> Mednick, S. 1962. " The Associative Basis of the Creative Process." Psychological Review 69, no. 3 : 220 – 232. https://doi.org/10.1037/h0048850.</bibtext> </blist> <blist> <bibtext> Molik, E. 2023, August. "Automating Creativity: There is Now Strong Evidence That AI Can Help Us to be More Innovative." https://substack.com/@oneusefulthing.</bibtext> </blist> <blist> <bibtext> Mumford, M. D., R. Reiter‐Palmon, and M. R. Redmond. 1994. " Problem Construction and Cognition: Applying Problem Representations in Ill‐Defined Domains." In Problem Finding, Problem Solving, and Creativity, edited by M. A. Runco, 3 – 39. Norwood, NJ : Ablex.</bibtext> </blist> <blist> <bibtext> Okuda, S. M., M. A. Runco, and D. E. Berger. 1991. " Creativity and the Finding and Solving of Real‐World Problems." Journal of Psychoeducational Assessment 9, no. 1 : 45 – 53. https://doi.org/10.1177/073428299100900104.</bibtext> </blist> <blist> <bibtext> Organisciak, P., S. Acar, D. Dumas, and K. Berthiaume. 2023. " Beyond Semantic Distance: Automated Scoring of Divergent Thinking Greatly Improves With Large Language Models." Thinking Skills and Creativity 49 : 101356. https://doi.org/10.1016/j.tsc.2023.101356.</bibtext> </blist> <blist> <bibtext> Organisciak, P., and D. Dumas. 2020. "Open Creativity Scoring [Computer Software]." Denver, CO: University of Denver. https://openscoring.du.edu/scoring.</bibtext> </blist> <blist> <bibtext> Østergaard, S. D., and K. L. Nielbo. 2023. " False Responses From Artificial Intelligence Models Are Not Hallucinations." Schizophrenia Bulletin 49 : 1105 – 1107. https://doi.org/10.1093/schbul/sbad068.</bibtext> </blist> <blist> <bibtext> Plain, C. 2023. "Artists Beware: "Game Changer" Test Results Show AI Is More Creative Than 95% of Humans." The Debrief. https://thedebrief.org/artists‐beware‐game‐changer‐test‐results‐show‐ai‐is‐more‐creative‐than‐99‐of‐humans/.</bibtext> </blist> <blist> <bibtext> Runco, M. A., ed. 1991. Divergent Thinking. Norwood, NJ : Ablex.</bibtext> </blist> <blist> <bibtext> Runco, M. A., ed. 1994. Problem Finding, Problem Solving, and Creativity. Norwood, NJ : Ablex.</bibtext> </blist> <blist> <bibtext> Runco, M. A. 2023a. " AI Can Only Produce Artificial Creativity." Journal of Creativity 33, no. 3 : 1 – 7. https://<ulink href="http://www.sciencedirect.com/science/article/pii/S2713374523000225">www.sciencedirect.com/science/article/pii/S2713374523000225</ulink>.</bibtext> </blist> <blist> <bibtext> Runco, M. A. 2023b. " Updating the Standard Definition of Creativity to Account for the Artificial Creativity of AI." Creativity Research Journal : 1 – 5. https://doi.org/10.1080/10400419.2023.2257977.</bibtext> </blist> <blist> <bibtext> Runco, M. A., and S. Acar. 2012. " Divergent Thinking as an Indicator of Creative Potential." Creativity Research Journal 24, no. 1 : 66 – 75. https://doi.org/10.1080/10400419.2012.652929.</bibtext> </blist> <blist> <bibtext> Runco, M. A., A. M. A. Alabbasi, and S.‐H. Paek. 2016. " Which Test of Divergent Thinking Is Best? " Creativity: Theories‐Research‐Applications 3 : 4 – 18.</bibtext> </blist> <blist> <bibtext> Runco, M. A., and R. S. Albert. 1985. " The Reliability and Validity of Ideational Originality in the Divergent Thinking of Academically Gifted and Nongifted Children." Educational and Psychological Measurement 45 : 483 – 501.</bibtext> </blist> <blist> <bibtext> Runco, M. A., and R. E. Charles. 1993. " Judgments of Originality and Appropriateness as Predictors of Creativity." Personality and Individual Differences 15, no. 5 : 537 – 546. https://doi.org/10.1016/0191‐8869(93)90337‐3.</bibtext> </blist> <blist> <bibtext> Runco, M. A., J. J. Illies, and R. Eisenman. 2005. " Creativity, Originality, and Appropriateness: What Do Explicit Instructions Tell Us About Their Relationships? " Journal of Creative Behavior 39, no. 2 : 137 – 148. https://doi.org/10.1002/j.2162‐6057.2005.tb01255.x.</bibtext> </blist> <blist> <bibtext> Runco, M. A., J. J. Illies, and R. Reiter‐Palmon. 2005. " Explicit Instructions to Be Creative and Original: A Comparison of Strategies and Criteria as Targets With Three Types of Divergent Thinking Tests." Korean Journal of Thinking and Problem Solving 15 : 5 – 15.</bibtext> </blist> <blist> <bibtext> Runco, M. A., and G. J. Jaeger. 2012. " The Standard Definition of Creativity." Creativity Research Journal 24, no. 1 : 92 – 96. https://doi.org/10.1080/10400419.2012.650092.</bibtext> </blist> <blist> <bibtext> Runco, M. A., and W. Mraz. 1992. " Scoring Divergent Thinking Tests Using Total Ideational Output and a Creativity Index." Educational and Psychological Measurement 52, no. 1 : 213 – 221. https://doi.org/10.1177/001316449205200126.</bibtext> </blist> <blist> <bibtext> Runco, M. A., B. Turkman, S. Acar, and M. V. Nural. 2017. " Idea Density and the Creativity of Written Works." Journal of Genius and Eminence 2 : 26 – 31.</bibtext> </blist> <blist> <bibtext> Russ, S. W., and J. D. Hoffmann. 2020. " Associative Theory." In Encyclopedia of Creativity, edited by M. A. Runco and S. R. Pritzker, 3rd ed., 76 – 82. Oxford, UK : Elsevier.</bibtext> </blist> <blist> <bibtext> Silvia, P. J. 2011. " Subjective Scoring of Divergent Thinking: Examining the Reliability of Unusual Uses, Instances, and Consequences Tasks." Thinking Skills and Creativity 6, no. 1 : 24 – 30. https://doi.org/10.1016/j.tsc.2010.06.001.</bibtext> </blist> <blist> <bibtext> Silvia, P. J., C. Martin, and E. C. Nusbaum. 2009. " A Snapshot of Creativity: Evaluating a Quick and Simple Method for Assessing Divergent Thinking." Thinking Skills and Creativity 4, no. 2 : 79 – 85.</bibtext> </blist> <blist> <bibtext> Taloni, A., M. Borselli, V. Scarsi, et al. 2023. " Comparative Performance of Humans Versus GPT‐4.0 and GPT‐3.5 in the Self‐Assessment Program of American Academy of Ophthalmology." Scientific Reports 13 : 18562. https://doi.org/10.1038/s41598‐023‐45837‐2.</bibtext> </blist> <blist> <bibtext> Torrance, E. P. 1995. Why Fly? Cresskill, NJ : Hampton Press.</bibtext> </blist> </ref> <aug> <p>By Mark A. Runco; Burak Turkman; Selcuk Acar and Ahmed M. Abdulla Alabbasi</p> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib23" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib29" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib15" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib10" firstref="ref4"></nolink> <nolink nlid="nl5" bibid="bib32" firstref="ref5"></nolink> <nolink nlid="nl6" bibid="bib34" firstref="ref6"></nolink> <nolink nlid="nl7" bibid="bib35" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib45" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib47" firstref="ref10"></nolink> <nolink nlid="nl10" bibid="bib21" firstref="ref11"></nolink> <nolink nlid="nl11" bibid="bib46" firstref="ref12"></nolink> <nolink nlid="nl12" bibid="bib19" firstref="ref13"></nolink> <nolink nlid="nl13" bibid="bib38" firstref="ref14"></nolink> <nolink nlid="nl14" bibid="bib39" firstref="ref15"></nolink> <nolink nlid="nl15" bibid="bib17" firstref="ref19"></nolink> <nolink nlid="nl16" bibid="bib41" firstref="ref20"></nolink> <nolink nlid="nl17" bibid="bib44" firstref="ref21"></nolink> <nolink nlid="nl18" bibid="bib11" firstref="ref22"></nolink> <nolink nlid="nl19" bibid="bib12" firstref="ref23"></nolink> <nolink nlid="nl20" bibid="bib42" firstref="ref24"></nolink> <nolink nlid="nl21" bibid="bib22" firstref="ref26"></nolink> <nolink nlid="nl22" bibid="bib43" firstref="ref27"></nolink> <nolink nlid="nl23" bibid="bib18" firstref="ref29"></nolink> <nolink nlid="nl24" bibid="bib55" firstref="ref36"></nolink> <nolink nlid="nl25" bibid="bib13" firstref="ref37"></nolink> <nolink nlid="nl26" bibid="bib14" firstref="ref39"></nolink> <nolink nlid="nl27" bibid="bib24" firstref="ref40"></nolink> <nolink nlid="nl28" bibid="bib31" firstref="ref42"></nolink> <nolink nlid="nl29" bibid="bib25" firstref="ref45"></nolink> <nolink nlid="nl30" bibid="bib40" firstref="ref46"></nolink> <nolink nlid="nl31" bibid="bib33" firstref="ref47"></nolink> <nolink nlid="nl32" bibid="bib16" firstref="ref49"></nolink> <nolink nlid="nl33" bibid="bib30" firstref="ref55"></nolink> <nolink nlid="nl34" bibid="bib27" firstref="ref59"></nolink> <nolink nlid="nl35" bibid="bib323" firstref="ref68"></nolink> <nolink nlid="nl36" bibid="bib145" firstref="ref73"></nolink> <nolink nlid="nl37" bibid="bib37" firstref="ref74"></nolink> <nolink nlid="nl38" bibid="bib36" firstref="ref76"></nolink> <nolink nlid="nl39" bibid="bib28" firstref="ref78"></nolink> <nolink nlid="nl40" bibid="bib26" firstref="ref79"></nolink> <nolink nlid="nl41" bibid="bib20" firstref="ref80"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1482985 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Examining the Idea Density and Semantic Distance of Responses Given by AI to Tests of Divergent Thinking – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Mark+A%2E+Runco%22">Mark A. Runco</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-5043-7900">0000-0002-5043-7900</externalLink>)<br /><searchLink fieldCode="AR" term="%22Burak+Turkman%22">Burak Turkman</searchLink><br /><searchLink fieldCode="AR" term="%22Selcuk+Acar%22">Selcuk Acar</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4044-985X">0000-0003-4044-985X</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ahmed+M%2E+Abdulla+Alabbasi%22">Ahmed M. Abdulla Alabbasi</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-4773-4955">0000-0002-4773-4955</externalLink>) – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Journal+of+Creative+Behavior%22"><i>Journal of Creative Behavior</i></searchLink>. 2025 59(3). – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 11 – Name: DatePubCY Label: Publication Date Group: Date Data: 2025 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Creative+Thinking%22">Creative Thinking</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Differences%22">Differences</searchLink><br /><searchLink fieldCode="DE" term="%22Responses%22">Responses</searchLink><br /><searchLink fieldCode="DE" term="%22Creativity%22">Creativity</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1002/jocb.1528 – Name: ISSN Label: ISSN Group: ISSN Data: 0022-0175<br />2162-6057 – Name: Abstract Label: Abstract Group: Ab Data: Research suggests that generative AI (GAI) responds to divergent thinking (DT) prompts with multiple ideas, some of which seem to be original. The present investigation administered 55 DT tasks to three GAI services (Bard, GPT 3.5, and GPT 4.0). Instead of examining individual responses, an Idea Density algorithm was used to assess the output. This algorithm quantifies the ideas within responses, controlling for the number of words. A subset of the DT tests administered to the GAI were also scored for Semantic Distance, which estimates originality. Results indicated that the three GAI models differed in the Idea Density of the output. There were also significant differences between Realistic and Nonrealistic DT tasks. As has been the case in human samples, directions given when the GAI received the prompts also had a significant impact, with more Idea Density following directions that explicitly prompted original responses. Adjusted scores removed all verbiage in the output, which did not actually address the questions conveyed by the prompts. These corrected scores shared approximately 50% of the variance with the uncorrected "raw" responses, implying that the typical output of GAI is not always relevant. This was interpreted in the context of the standard definition of creativity, which emphasizes effectiveness, as well as originality. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2025 – Name: AN Label: Accession Number Group: ID Data: EJ1482985 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1482985 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1002/jocb.1528 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 11 Subjects: – SubjectFull: Artificial Intelligence Type: general – SubjectFull: Creative Thinking Type: general – SubjectFull: Models Type: general – SubjectFull: Differences Type: general – SubjectFull: Responses Type: general – SubjectFull: Creativity Type: general Titles: – TitleFull: Examining the Idea Density and Semantic Distance of Responses Given by AI to Tests of Divergent Thinking Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Mark A. Runco – PersonEntity: Name: NameFull: Burak Turkman – PersonEntity: Name: NameFull: Selcuk Acar – PersonEntity: Name: NameFull: Ahmed M. Abdulla Alabbasi IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 09 Type: published Y: 2025 Identifiers: – Type: issn-print Value: 0022-0175 – Type: issn-electronic Value: 2162-6057 Numbering: – Type: volume Value: 59 – Type: issue Value: 3 Titles: – TitleFull: Journal of Creative Behavior Type: main |
| ResultId | 1 |