Automated Scoring of Scientific Creativity in German

Saved in:
Bibliographic Details
Title: Automated Scoring of Scientific Creativity in German
Language: English
Authors: Benjamin Goecke (ORCID 0000-0002-3050-1848), Paul V. DiStefano (ORCID 0009-0002-9638-3220), Wolfgang Aschauer (ORCID 0000-0002-1221-6509), Kurt Haim (ORCID 0000-0003-4093-0512), Roger Beaty (ORCID 0000-0001-6114-5973), Boris Forthmann (ORCID 0000-0001-9755-7304)
Source: Journal of Creative Behavior. 2024 58(3):321-327.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 7
Publication Date: 2024
Sponsoring Agency: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
National Science Foundation (NSF), Division of Undergraduate Education (DUE)
Document Type: Journal Articles
Reports - Research
Descriptors: Creativity, Creative Thinking, Scoring, Automation, German, Prediction, Models, Correlation, Sciences, Thinking Skills
DOI: 10.1002/jocb.658
ISSN: 0022-0175
2162-6057
Abstract: Automated scoring is a current hot topic in creativity research. However, most research has focused on the English language and popular verbal creative thinking tasks, such as the alternate uses task. Therefore, in this study, we present a large language model approach for automated scoring of a scientific creative thinking task that assesses divergent ideation in experimental tasks in the German language. Participants are required to generate alternative explanations for an empirical observation. This work analyzed a total of 13,423 unique responses. To predict human ratings of originality, we used XLM-RoBERTa (Cross-lingual Language Model-RoBERTa), a large, multilingual model. The prediction model was trained on 9,400 responses. Results showed a strong correlation between model predictions and human ratings in a held-out test set (n = 2,682; r = 0.80; CI-95% [0.79, 0.81]). These promising findings underscore the potential of large language models for automated scoring of scientific creative thinking in the German language. We encourage researchers to further investigate automated scoring of other domain-specific creative thinking tasks.
Abstractor: As Provided
Notes: https://osf.io/aw95p
Entry Date: 2024
Accession Number: EJ1444134
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwGoyucWrWLO3IMEfCLTh-tqAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDKBPOZNX1OviR8SX5AIBEICBmyUtSq3AwGMgNfxIXcwfUme38nNoH9N2D5z_G8ClBlq7MMUgNh6lAERJWX40eqkUwxznuJcE-V8WW-akkcEqwHVuGhBGuutYMtypAOBJgoANRzdbMZBZrQl90rinlpsV9HvRLIOtNgwwt4p9QdaV9MHM00cvPYcTHCgSD3GaCxwphwfXMhCpf2w40Rv_og7FWFAO9ZG-CcbIAS2c
Text:
  Availability: 1
  Value: <anid>AN0180293695;3u701sep.24;2024Oct18.02:58;v2.2.500</anid> <title id="AN0180293695-1">Automated Scoring of Scientific Creativity in German </title> <p>Automated scoring is a current hot topic in creativity research. However, most research has focused on the English language and popular verbal creative thinking tasks, such as the alternate uses task. Therefore, in this study, we present a large language model approach for automated scoring of a scientific creative thinking task that assesses divergent ideation in experimental tasks in the German language. Participants are required to generate alternative explanations for an empirical observation. This work analyzed a total of 13,423 unique responses. To predict human ratings of originality, we used XLM‐RoBERTa (Cross‐lingual Language Model‐RoBERTa), a large, multilingual model. The prediction model was trained on 9,400 responses. Results showed a strong correlation between model predictions and human ratings in a held‐out test set (n = 2,682; r = 0.80; CI‐95% [0.79, 0.81]). These promising findings underscore the potential of large language models for automated scoring of scientific creative thinking in the German language. We encourage researchers to further investigate automated scoring of other domain‐specific creative thinking tasks.</p> <p>Keywords: creativity; automated scoring; scientific creativity; large language models</p> <hd id="AN0180293695-2">AUTOMATED SCORING OF CREATIVE THINKING TASKS</hd> <p>Attempts of automated scoring of creative thinking tasks are surprisingly old (Paulus, [<reflink idref="bib31" id="ref1">31</reflink>]). At the end of the 1960s, a regression‐based prediction model of scores for the Torrance Test of Creative Thinking was proposed, which is one of the most widely used tests of divergent thinking. Simple text mining statistics such as the number of words in a response were used to predict scoring by human raters and their approach has provided reasonable predictions of human ratings in recent efforts (Forthmann & Doebler, [<reflink idref="bib18" id="ref2">18</reflink>]). Benefits of automated scoring are self‐evident: for example, lower associated costs due to less labor, less prone to human biases (although this assumption can be considered contested), availability of full computerized assessment. Application of creativity measures for individual purposes and possibly personnel selection (compare responses against huge reference groups of prior data).</p> <p>Generally, 15 years ago the idea of automated scoring revived with most works relying on latent semantic analysis or other word vector models of meaning (Bossomaier, Harré, Knittel, & Snyder, [<reflink idref="bib6" id="ref3">6</reflink>]; Dumas & Dunbar, [<reflink idref="bib14" id="ref4">14</reflink>]; Forster & Dunbar, [<reflink idref="bib17" id="ref5">17</reflink>]; Green, Kraemer, Fugelsang, Gray, & Dunbar, [<reflink idref="bib21" id="ref6">21</reflink>]). Initially, validity findings were rather inconsistent which is most likely attributable to technical issues such as the known elaboration‐bias inherent in latent semantic analysis. Focusing on other word vector models such as GloVe (Dumas, Organisciak, & Doherty, [<reflink idref="bib15" id="ref7">15</reflink>]), using multiplicative compositional approaches (Beaty & Johnson, [<reflink idref="bib5" id="ref8">5</reflink>]), or maximum semantic distance (Yu et al., [<reflink idref="bib36" id="ref9">36</reflink>]) seemed to clearly improve automated scoring quality. Most recent work switched to Large Language Models (LLMs) which can closely mimic the scoring of human raters (Organisciak, Acar, Dumas, & Berthiaume, [<reflink idref="bib28" id="ref10">28</reflink>]).</p> <p>LLMs can score divergent thinking tasks automatically (Organisciak et al., [<reflink idref="bib28" id="ref11">28</reflink>]) and already achieve impressive prediction of human ratings without additional training (i.e., at zero‐shot). The models' capacity to score divergent thinking tasks improves substantially with small amounts of training data, resulting in correlations of up to <emph>r</emph> ~ .80 at the response level, which was argued to be the possible ceiling. After all, inter‐rater reliabilities between any two or more human scorers rarely do exceed such values as well.</p> <hd id="AN0180293695-3">COMPLEMENTING EFFORTS OF SCORING CREATIVITY AUTOMATICALLY</hd> <p>Most studies that use automated scoring techniques rely on data from the English language and a limited set of creativity indicators, primarily the alternate uses task (e.g., Guilford, [<reflink idref="bib22" id="ref12">22</reflink>]). However, recent advancements have expanded our ability to automatically score creativity across different languages (Forthmann & Doebler, [<reflink idref="bib18" id="ref13">18</reflink>]; Patterson, Merseal, et al., [<reflink idref="bib30" id="ref14">30</reflink>]; Zielińska, Organisciak, Dumas, & Karwowski, [<reflink idref="bib38" id="ref15">38</reflink>]) and tasks (Acar, Organisciak, & Dumas, [<reflink idref="bib1" id="ref16">1</reflink>]; Cropley & Marrone, [<reflink idref="bib11" id="ref17">11</reflink>]; Patterson, Barbot, Lloyd‐Cox, & Beaty, [<reflink idref="bib29" id="ref18">29</reflink>]).</p> <p>Although some progress has been made in transferring knowledge gained in the English language domain to other languages, such as German (e.g., Forthmann & Doebler, [<reflink idref="bib18" id="ref19">18</reflink>]; Patterson, Merseal, et al., [<reflink idref="bib30" id="ref20">30</reflink>]), further evidence is needed to evaluate the predictive performance of automated scoring approaches for this language. With roughly an estimated 100–160 million individuals speaking German as their first or second language in the world, German belongs to the 20 most spoken languages in the world (c.f., https://de.statista.com; full link provided in the Supplementary Materials in the OSF). Due to the large potential of automated creativity scoring and the considerable number of individuals on earth not using English as their lingua franca in their daily lives, it makes sense to extend previous efforts of automated creativity scoring to further languages, such as German as well.</p> <p>Second, today a large arsenal of measurement instruments of creativity is available (Weiss, Wilhelm, & Kyllonen, [<reflink idref="bib35" id="ref21">35</reflink>]) and new measurement approaches are continuously developed and evaluated. This includes both new measurement instruments, but also refined constructs that can be measured. For example, measuring domain‐specific creative abilities has become increasingly popular over the last few years, and new measurement instruments for domains like emotions (Weiss, Olderbak, & Wilhelm, [<reflink idref="bib34" id="ref22">34</reflink>]), music (Merseal et al., [<reflink idref="bib26" id="ref23">26</reflink>]), or scientific creative thinking (Aschauer, Haim, & Weber, [<reflink idref="bib3" id="ref24">3</reflink>]) were developed and validated. These developments can also be understood as a call for broadening our capabilities of automatically evaluating the creativity of test‐takers, as future multivariate studies could benefit from less labor‐intensive scoring practices. Hence, testing existing or new algorithms for automatically scoring creativity tasks would benefit from extending the scope of tasks to which such algorithms can be applied.</p> <hd id="AN0180293695-4">AIM OF THE CURRENT STUDY</hd> <p>Most previous work regarding automated scoring of creativity has focused mostly on the English language, on the classical Alternate Uses Task, or other verbal divergent thinking tasks such as the Consequences Task. While there exists now psychometric work on automated scoring of creative writing tasks (Johnson et al., [<reflink idref="bib23" id="ref25">23</reflink>]) or figural divergent thinking tasks (Cropley & Marrone, [<reflink idref="bib11" id="ref26">11</reflink>]; Patterson, Barbot, et al., [<reflink idref="bib29" id="ref27">29</reflink>]), work on domain‐specific creative thinking tasks in other languages than English is still lacking. Hence, the main aim of this work is to address this gap in the literature by reporting a study that shows that automated scoring using a LLM of a scientific creative thinking task (i.e., a domain‐specific creative thinking task) in the German language is feasible. Scientific creative thinking can be conceptualized as a <emph>divergent problem solving ability in science</emph> (DPAS; Aschauer et al., [<reflink idref="bib3" id="ref28">3</reflink>]). In this paper, the operationalization focuses on originality (e.g., uncommon, remote, and clever ideas). Given the increasing importance of creativity in science education and research on fostering creative thinking in education, it is necessary to improve both our measurement instruments for assessing this ability and our scoring methods for evaluating test takers' responses. The research objectives of the present study were not pre‐registered.</p> <hd id="AN0180293695-5">METHOD</hd> <p></p> <hd id="AN0180293695-6">SAMPLE AND PROCEDURE</hd> <p>The sample for the current study consisted of <emph>N</emph> = 1,272 students from 5th to 12th grade from 52 classes (middle schools; academic‐track schools, and vocational‐track schools) altogether in Austria. The age range was approximately 11 to 18 (please note that this information was not collected, but that the grade is a good proxy for the age of a student in [please will be disclosed after peer‐review]). Approximately, 51% of the students were females. The measurement was part of a larger multivariate study consisting of two measurement time points. As these design‐related characteristics of the study are not inherently associated with our research aims, we will refrain from going into more detail here (please see Aschauer et al., [<reflink idref="bib3" id="ref29">3</reflink>] for more comprehensive information).</p> <hd id="AN0180293695-7">MEASURES</hd> <p>For the current study, we use one specific item of a test battery designed to measure scientific creativity by DPAS (please see Aschauer et al., [<reflink idref="bib3" id="ref30">3</reflink>] for a comprehensive description of the test battery). The employed item is part of the DPAS subscale for divergent ideation in experimental tasks. Students were asked to come up with creative reasons for the following hypothetical phenomenon: "In 2000, the room temperature of a classroom was measured for 1 year, and the average temperature at that time was 18°C. Today, 19 years later, the mean temperature of this classroom was 22°C. Give as many different reasons as possible why the room temperature in this classroom is 4°C higher today." Students were informed that they were participating in a creativity test that will test the richness of their ideas (e.g., "find many different ideas"). They were instructed to think of as many responses as possible. The test was administered computerized. In total, <emph>N</emph> = 18,653 responses were generated and considered for analysis in this work.</p> <hd id="AN0180293695-8">HUMAN SCORING RATING DESIGN</hd> <p>The rating design for the human scoring was based on a planned missingness design (Forthmann, Goecke, & Beaty, [<reflink idref="bib19" id="ref31">19</reflink>]) that we obtained through a simulation‐based approach in R (R Core Team, [<reflink idref="bib32" id="ref32">32</reflink>]). Specifically, we based our rating design on a simulation based on three available empirical data sets: the data of an Alternate Uses Task (<emph>N</emph> = 209 participants with <emph>n</emph> = 3,236 responses; c.f., Patterson, Barbot, et al., [<reflink idref="bib29" id="ref33">29</reflink>]; Patterson, Merseal, et al., [<reflink idref="bib30" id="ref34">30</reflink>]), and two data sets of a scientific creativity test (<emph>N</emph><subs>1</subs> = 1,147 with <emph>n</emph><subs>1</subs> = 4,369 responses; <emph>N</emph><subs>2</subs> = 146 with <emph>n</emph><subs>2</subs> = 7,923 responses; c.f., Beaty et al., under review). We applied a generalized partial credit model (GPCM; Muraki, [<reflink idref="bib27" id="ref35">27</reflink>]) for the simulation and tested five competing scenarios in our simulations with different parameters in order to find the best planned missingness design regarding cost‐reliability trade‐off for our purposes. Please see the Supplementary Materials for more information regarding the data generation process of the simulation.</p> <p>We found that a design with <emph>N</emph> = 40 human raters, and 4 ratings per response for 50% of all responses would yield an average reliability of.822 which was deemed appropriate considering the trade‐off between costs, human labor and reliability as a good cutoff is considered to be.80 (c.f., the factor determinacy index, Ferrando & Lorenzo‐Seva, [<reflink idref="bib16" id="ref36">16</reflink>]). Explicitly, this rater design meant that on average, each rater had to rate 1,569 responses (range: 1,551–1,701 responses).</p> <p>We used the obtained rater design to prepare 40 single rating sheets containing, in total, all the available responses according to our planned missingness rater design. We then recruited 40 human raters with sufficient German skills via Prolific, who rated the German responses. Each rating sheet was assigned to one rater (we provide one example rating sheet including instructions for raters via the OSF). The raters received detailed instructions regarding scoring with the rating sheets and were motivated to avoid missingness. The instructions were available to them at all times. In addition to that, three of the co‐authors picked 3 example responses for each possible response category and explained why each response should be rated this way. Raters were reimbursed with 30$ (12$/h), as we estimated that it would take them about 2.5 hours to score the responses assigned to them.</p> <p>Once the ratings were completed, we tested the data not only based on the initially imposed GPCM but also decided to test competing measurement models, such as the Graded Response Model (GRM; Samejima, [<reflink idref="bib33" id="ref37">33</reflink>]). The reliabilities and correlations for both models are depicted in Table 1. The GRM fitted the data slightly better, so we decided to use the factor scores obtained by this model for further evaluation (i.e., human scored "creativity ratings"). Please note, however, that the correlation between both models' parameters amounted to <emph>r</emph> = .99. These parameters were the basis for the automated scoring approach which will be outlined next. Please note that for validity purposes, we provide correlations between the human ratings (i.e., factor scores) and fluency and flexibility scores of the scientific creativity task in the Supplementary Materials (Figure S2). We provide all materials needed to replicate our approach in an open repository: https://osf.io/aw95p/.</p> <p>1 Table Fit Indices, Empirical Reliabilities, and Correlations of Applied Rater Models</p> <p> <ephtml> <table><thead valign="bottom"><tr><th align="left">Model</th><th align="center">AIC</th><th align="center">BIC</th><th align="center"><italic>r</italic><sub>θθ</sub></th><th align="center">√<italic>r</italic><sub>θθ</sub></th></tr></thead><tbody valign="top"><tr><td align="left">GCPM</td><td align="char" char=".">124348</td><td align="char" char=".">125895.6</td><td align="char" char=".">0.731</td><td align="char" char=".">0.855</td></tr><tr><td align="left">GRM</td><td align="char" char=".">124150</td><td align="char" char=".">125697.5</td><td align="char" char=".">0.742</td><td align="char" char=".">0.861</td></tr></tbody></table> </ephtml> </p> <p>1 <emph>Note. r</emph><subs>θθ</subs> = Empirical Reliability; √<emph>r</emph><subs>θθ</subs> = Correlation between estimated latent response scores and the true response scores.</p> <hd id="AN0180293695-9">AUTOMATED SCORING APPROACH</hd> <p>We fine‐tuned a large language model (LLM) to predict the human ratings of the problem‐solving task. The LLM we used was XLM‐RoBERTa‐base (Conneau et al., [<reflink idref="bib10" id="ref38">10</reflink>]), which is an open‐source pre‐trained multi‐lingual language model with 125 M parameters based on Meta's RoBERTa model (Liu et al., [<reflink idref="bib24" id="ref39">24</reflink>]). This model was trained using a Transformer architecture (Devlin, Chang, Lee, & Toutanova, [<reflink idref="bib12" id="ref40">12</reflink>]), as is the case with models of the Bidirectional Encoder Representations from Transformers (BERT) architecture. XLM‐RoBERTa specifically was trained in over 100 languages, thus it can predict ratings for text in German.</p> <p>To identify the model settings (hyperparameters) that elicit the best performance, we engaged in a hyperparameter search using the Optuna Python package (Akiba, Sano, Yanase, Ohta, & Koyama, [<reflink idref="bib2" id="ref41">2</reflink>]). Prior to this search, the data were deduplicated, reducing the data set size from 18,314 to 13,423 unique responses. Subsequently, the dataset was randomly split into training, validation, and held‐out test data sets, utilizing a 70/10/20 ratio, consistent with the best practices in machine learning (Zhou, [<reflink idref="bib37" id="ref42">37</reflink>]). The model input was the task prompt "Ein kreativer Grund für Temperaturveränderungen in einem Klassenzimmer ist" ("A Creative Reason for Temperature Changes in a Classroom is") immediately followed by the participant response. An example of a highly creative response was "Ein übernatürlicher Mensch aus einer anderen Dimension machte die Sonne heißer, so dass es bei uns auch heißer ist" ("A supernatural person from another dimension made the sun hotter so that ours is hotter too.").</p> <p>Across 230 trials, we searched over three hyperparameters (learning rate, batch size, and the number of training epochs). Learning rate corresponds to the extent to which training episodes influence the model weights, where higher rates result in greater change during each training batch. The range of learning rates we searched over was 5e‐07 to 5e‐02. Batch size is the number of responses that the model receives feedback on at a time and we tested values of 4, 8, 16, and 32. Lastly, we surveyed an epoch range of 10 to 150 to determine the optimal number of passes through the training dataset. Optuna searches over these hyperparameter settings for the configuration that minimizes the mean square error (MSE) between the provided human ratings and the model predicted‐ratings of the validation set. The trial with the best performance on the validation set (<emph>n</emph> = 1,341) employed a learning rate of 4.37e‐06, ran for 11 epochs, and utilized a training batch size of 4. Using these hyperparameters, rating predictions generated by XLM‐RoBERTa correlated positively with human ratings (<emph>r</emph> = 0.81; CI 95% [0.79, 0.83]). The settings of this trial were used to evaluate the held‐out test set. We removed outliers in the model predictions that were 3 standard deviations from the mean (please note, however, that we report the results of this analysis without outlier removal in the Figure S1).</p> <hd id="AN0180293695-10">RESULTS</hd> <p>When evaluating the held‐out test set (<emph>n</emph> = 2,682), rating predictions generated by XLM‐RoBERTa correlated positively with the human ratings (<emph>r</emph> = 0.80; CI 95% [0.79, 0.81]); see Figure 1.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/3U7/01sep24/jocb658-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="jocb658-fig-0001.jpg" title="1 Correlation between Human‐Rated Solutions and Model Predictions.Note. Model predictions and human ratings are z‐transformed. n = 2682; 3 outliers were removed." /> </p> <p></p> <hd id="AN0180293695-12">DISCUSSION</hd> <p>Taken together, we have shown that automatically scoring responses on a scientific creativity task in German using a fine‐tuned LLM is feasible and results in a very high correlation with human ratings. Leveraging the XLM‐RoBERTa model, our study accurately predicts human ratings for a scientific creativity task in German. The consistency between our validation and test set performance indicates a robust training procedure and no evidence of overfitting. The achieved correlations between model predictions and human ratings are in line with other research using LLMs (e.g., Organisciak et al., [<reflink idref="bib28" id="ref43">28</reflink>]) and thus validate the effectiveness of these automated assessment methods. With that, our study closes a gap in research regarding automated scoring methods aiming to evaluate scientific creative thinking tasks in the German language.</p> <p>Most obviously, by targeting the German language, we limited our work to this language, and we thus encourage researchers to reevaluate our approach for scientific creative thinking tasks available for other languages such as Turkish (Ayas & Sak, [<reflink idref="bib4" id="ref44">4</reflink>]), among others. In addition, our work is limited by the task used to assess divergent ideation in experimental tasks. Future work should extend this effort to other items aiming at this specific scientific thinking ability as well as other scientific creative thinking abilities such as divergent ideation in science tasks (Aschauer et al., [<reflink idref="bib3" id="ref45">3</reflink>]). Moreover, future work could investigate not only rater agreement but also rater disagreement (e.g., Forthmann et al., [<reflink idref="bib20" id="ref46">20</reflink>]), because previous work showed that raters' residual disagreement is predictive of controversial responses, which, in turn, might be considered important for broader applications, such as educational settings (Dumas et al., [<reflink idref="bib13" id="ref47">13</reflink>]).</p> <p>Most recent efforts of automated scoring of creativity tasks emphasize the use of LLMs. However, we want to point out that competing prediction models such as xgboost (Chen & Guestrin, [<reflink idref="bib9" id="ref48">9</reflink>]) or random forests (Breiman, [<reflink idref="bib7" id="ref49">7</reflink>]), for example, have also been shown to be capable of predicting human ratings for prototypical divergent thinking tasks like the alternate uses task (Buczak, Huang, Forthmann, & Doebler, [<reflink idref="bib8" id="ref50">8</reflink>]). Although all of these models perform well, it is important to note their inherent differences. Unlike LLMs, prediction models like xboost or random forests do not work with plain text (i.e., natural language), but rather with meta‐information (i.e., features) extracted from texts. For example, Buczak et al. ([<reflink idref="bib8" id="ref51">8</reflink>]) constructed a set of simple text mining features such as the number of words in a response and the average word length in a response. These features are then used to inform the prediction models. These differences between models beg the question whether we could enhance our understanding of automated scoring of domain‐specific creative thinking tasks in German and possibly other languages by comparing the prediction models. Future work should thus consider comparing a variety of automated scoring procedures.</p> <p>Finally, we would like to point out potential issues of fairness in machine learning, for example with regards to demographic biases (Mehrabi, Morstatter, Saxena, Lerman, & Galstyan, [<reflink idref="bib25" id="ref52">25</reflink>]) or biases that might have been carried over from the training data. Although no demographic data were used to train the current LLM and the existence of reliable human biases in our data is unlikely due to the large number of raters, it cannot be fully ruled out that rating distortions exist in our data.</p> <p>In conclusion, our study marks a meaningful cumulative step towards automating the assessment of domain‐specific creative thinking tasks in languages beyond English, presenting validated approaches that closely mimic human raters' judgments.</p> <hd id="AN0180293695-13">Acknowledgment</hd> <p>Open Access funding enabled and organized by Projekt DEAL.</p> <hd id="AN0180293695-14">DATA AVAILABILITY STATEMENT</hd> <p>All files and data for analyses are available at Open Science Framework: https://osf.io/aw95p.</p> <p>GRAPH: Figure S1. Correlation between Human‐Rated Solutions and Model Predictions without outlier removal.Figure S2. Correlation between Human‐Rated Solutions, Fluency, and Flexibility.Table S1. Simulation Designs for the Planned Missingness Rater Design.</p> <ref id="AN0180293695-15"> <title> Footnotes </title> <blist> <bibl id="bib1" idref="ref16" type="bt">1</bibl> <bibtext> The study was conducted in accordance with the declaration of Helsinki. The authors have no conflicts of interest to disclose.</bibtext> </blist> </ref> <ref id="AN0180293695-16"> <title> REFERENCES </title> <blist> <bibtext> Acar, S., Organisciak, P., & Dumas, D. (2023). A comparison of supervised and unsupervised learning methods in automated scoring of figural tests of creativity. ResearchGate.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref41" type="bt">2</bibl> <bibtext> Akiba, T., Sano, S., Yanase, T., Ohta, T., & Koyama, M. (2019). Optuna: A next‐generation hyperparameter optimization framework. Proceedings of the 25th ACM SIGKDD International Conference on Knowledge Discovery & Data Mining, Pp. 2623–2631. Available from: 10.1145/3292500.3330701.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref24" type="bt">3</bibl> <bibtext> Aschauer, W., Haim, K., & Weber, C. (2022). A contribution to scientific creativity: a validation study measuring divergent problem solving ability. Creativity Research Journal, 34 (2), 195 – 212 ; doi: 10.1080/10400419.2021.1968656.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref44" type="bt">4</bibl> <bibtext> Ayas, M.B., & Sak, U. (2014). Objective measure of scientific creativity: Psychometric validity of the Creative Scientific Ability Test. Thinking Skills and Creativity, 13, 195 – 205 ; doi: 10.1016/j.tsc.2014.06.001.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref8" type="bt">5</bibl> <bibtext> Beaty, R.E., & Johnson, D.R. (2021). Automating creativity assessment with SemDis: An open platform for computing semantic distance. Behavior Research Methods, 53 (2), Article 2 ; doi: 10.3758/s13428‐020‐01453‐w.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref3" type="bt">6</bibl> <bibtext> Bossomaier, T., Harré, M., Knittel, A., & Snyder, A. (2009). A semantic network approach to the creativity quotient (CQ). Creativity Research Journal, 21 (1), 64 – 71 ; doi: 10.1080/10400410802633517.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref49" type="bt">7</bibl> <bibtext> Breiman, L. (2001). Random forests. Machine Learning, 45, 5 – 32.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref50" type="bt">8</bibl> <bibtext> Buczak, P., Huang, H., Forthmann, B., & Doebler, P. (2023). The machines take over: A comparison of various supervised learning approaches for automated scoring of divergent thinking tasks. The Journal of Creative Behavior, 57 (1), 17 – 36 ; doi: 10.1002/jocb.559.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref48" type="bt">9</bibl> <bibtext> Chen, T., & Guestrin, C. (2016). XGBoost: A scalable tree boosting system. Proceedings of the 22nd ACM SIGKDD International Conference on Knowledge Discovery and Data Mining, Pp. 785–794. Available from: https://doi.org/10.1145/2939672.2939785.</bibtext> </blist> <blist> <bibtext> Conneau, A., Khandelwal, K., Goyal, N., Chaudhary, V., Wenzek, G., Guzmán, F., ... Stoyanov, V. (2020). Unsupervised cross‐lingual representation learning at scale (arXiv:1911.02116). arXiv. Available from: <ulink href="http://arxiv.org/abs/1911.02116">http://arxiv.org/abs/1911.02116</ulink>.</bibtext> </blist> <blist> <bibtext> Cropley, D.H., & Marrone, R.L. (2022). Automated scoring of figural creativity using a convolutional neural network. Psychology of Aesthetics, Creativity, and the Arts. Advance online publication; doi: 10.1037/aca0000510.</bibtext> </blist> <blist> <bibtext> Devlin, J., Chang, M.‐W., Lee, K., & Toutanova, K. (2019). BERT: Pre‐training of deep bidirectional transformers for language understanding (arXiv:1810.04805). arXiv. Available from: <ulink href="http://arxiv.org/abs/1810.04805">http://arxiv.org/abs/1810.04805</ulink>.</bibtext> </blist> <blist> <bibtext> Dumas, D., Acar, S., Berthiaume, K., Organisciak, P., Eby, D., Grajzel, K., ... Carrera, M. (2023). What makes children's responses to creativity assessments difficult to judge reliably? The Journal of Creative Behavior, 57 (3), 419 – 438 ; doi: 10.1002/jocb.588.</bibtext> </blist> <blist> <bibtext> Dumas, D., & Dunbar, K.N. (2014). Understanding fluency and originality: A latent variable perspective. Thinking Skills and Creativity, 14, 56 – 67 ; doi: 10.1016/j.tsc.2014.09.003.</bibtext> </blist> <blist> <bibtext> Dumas, D., Organisciak, P., & Doherty, M. (2021). Measuring divergent thinking originality with human raters and text‐mining models: A psychometric comparison of methods. Psychology of Aesthetics, Creativity, and the Arts, 15 (4), 645 – 663 ; doi: 10.1037/aca0000319.</bibtext> </blist> <blist> <bibtext> Ferrando, P.J., & Lorenzo‐Seva, U. (2018). Assessing the quality and appropriateness of factor solutions and factor score estimates in exploratory item factor analysis. Educational and Psychological Measurement, 78 (5), 762 – 780 ; doi: 10.1177/0013164417719308.</bibtext> </blist> <blist> <bibtext> Forster, E.A., & Dunbar, K.N. (2009). Creativity evaluation through latent semantic analysis. Proceedings of the Annual Meeting of the Cognitive Science Society, 31. Available from: https://escholarship.org/uc/item/4wp633ph.</bibtext> </blist> <blist> <bibtext> Forthmann, B., & Doebler, P. (2022). Fifty years later and still working: rediscovering Paulus et al.'s (1970) automated scoring of divergent thinking tests. PsyArXiv. Available from: https://doi.org/10.31234/osf.io/byj8c.</bibtext> </blist> <blist> <bibtext> Forthmann, B., Goecke, B., & Beaty, R.E. (2023). Planning missing data designs for human ratings in creativity research: A practical guide. Creativity Research Journal, 1–12 ; doi: 10.1080/10400419.2023.2250976.</bibtext> </blist> <blist> <bibtext> Forthmann, B., Holling, H., Zandi, N., Gerwig, A., Çelik, P., Storme, M., & Lubart, T. (2017). Missing creativity: The effect of cognitive workload on rater (dis‐)agreement in subjective divergent‐thinking scores. Thinking Skills and Creativity, 23, 129 – 139 ; doi: 10.1016/j.tsc.2016.12.005.</bibtext> </blist> <blist> <bibtext> Green, A.E., Kraemer, D.J.M., Fugelsang, J.A., Gray, J.R., & Dunbar, K.N. (2012). Neural correlates of creativity in analogical reasoning. Journal of Experimental Psychology: Learning, Memory, and Cognition, 38 (2), 264 – 272 ; doi: 10.1037/a0025764.</bibtext> </blist> <blist> <bibtext> Guilford, J.P. (1967). The nature of human intelligence. New York : McGraw‐Hill, Inc.</bibtext> </blist> <blist> <bibtext> Johnson, D.R., Kaufman, J.C., Baker, B.S., Patterson, J.D., Barbot, B., Green, A.E., ... Beaty, R.E. (2022). Divergent semantic integration (DSI): Extracting creativity from narratives with distributional semantic modeling. Behavior Research Methods, 55 (7), 3726 – 3759 ; doi: 10.3758/s13428‐022‐01986‐2.</bibtext> </blist> <blist> <bibtext> Liu, Y., Ott, M., Goyal, N., Du, J., Joshi, M., Chen, D., ... Stoyanov, V. (2019). RoBERTa: A robustly optimized BERT pretraining approach (arXiv:1907.11692). arXiv. Available from: <ulink href="http://arxiv.org/abs/1907.11692">http://arxiv.org/abs/1907.11692</ulink>.</bibtext> </blist> <blist> <bibtext> Mehrabi, N., Morstatter, F., Saxena, N., Lerman, K., & Galstyan, A. (2022). A survey on bias and fairness in machine learning. ACM Computing Surveys, 54 (6), 1 – 35 ; doi: 10.1145/3457607.</bibtext> </blist> <blist> <bibtext> Merseal, H.M., Beaty, R.E., Kenett, Y.N., Lloyd‐Cox, J., De Manzano, Ö., & Norgaard, M. (2023). Representing melodic relationships using network science. Cognition, 233, 105362 ; doi: 10.1016/j.cognition.2022.105362.</bibtext> </blist> <blist> <bibtext> Muraki, E. (1992). A generalized partial credit model: Application of an EM algorithm. ETS Research Report Series, 1992, 1 – 33 ; doi: 10.1002/j.2333‐8504.1992.tb01436.x.</bibtext> </blist> <blist> <bibtext> Organisciak, P., Acar, S., Dumas, D., & Berthiaume, K. (2023). Beyond semantic distance: Automated scoring of divergent thinking greatly improves with large language models. Thinking Skills and Creativity, 49, 101356 ; doi: 10.1016/j.tsc.2023.101356.</bibtext> </blist> <blist> <bibtext> Patterson, J.D., Barbot, B., Lloyd‐Cox, J., & Beaty, R.E. (2023). AuDrA: An automated drawing assessment platform for evaluating creativity. Behavior Research Methods ; doi: 10.3758/s13428‐023‐02258‐3.</bibtext> </blist> <blist> <bibtext> Patterson, J.D., Merseal, H.M., Johnson, D.R., Agnoli, S., Baas, M., Baker, B.S., ... Beaty, R.E. (2023). Multilingual semantic distance: Automatic verbal creativity assessment in many languages. Psychology of Aesthetics, Creativity, and the Arts, 17 (4), 495 – 507 ; doi: 10.1037/aca0000618.</bibtext> </blist> <blist> <bibtext> Paulus, D.H. (1970). Computer simulation of human ratings of creativity. Final Report.</bibtext> </blist> <blist> <bibtext> R Core Team. (2022). R: A language and environment for statistical computing [Software]. Vienna : R Foundation for Statistical Computing. Available from: https://www.R‐project.org/.</bibtext> </blist> <blist> <bibtext> Samejima, F. (1968). Estimation of latent ability using a response pattern of graded scores. ETS Research Bulletin Series, 1968 (1), 1 – 169 ; doi: 10.1002/j.2333‐8504.1968.tb00153.x.</bibtext> </blist> <blist> <bibtext> Weiss, S., Olderbak, S., & Wilhelm, O. (2023). Conceptualizing and measuring ability emotional creativity. Psychology of Aesthetics, Creativity, and the Arts. Advance online publication; doi: 10.1037/aca0000585.</bibtext> </blist> <blist> <bibtext> Weiss, S., Wilhelm, O., & Kyllonen, P. (2021). An improved taxonomy of creativity measures based on salient task attributes. Psychology of Aesthetics, Creativity, and the Arts. Advance online publication; doi: 10.1037/aca0000434.</bibtext> </blist> <blist> <bibtext> Yu, Y., Beaty, R.E., Forthmann, B., Beeman, M., Cruz, J.H., & Johnson, D. (2023). A MAD method to assess idea novelty: Improving validity of automatic scoring using maximum associative distance (MAD). Psychology of Aesthetics, Creativity, and the Arts. Advance online publication; doi: 10.1037/aca0000573.</bibtext> </blist> <blist> <bibtext> Zhou, Z.‐H. (2021). Machine Learning. Singapore : Springer ; doi: 10.1007/978‐981‐15‐1967‐3.</bibtext> </blist> <blist> <bibtext> Zielińska, A., Organisciak, P., Dumas, D., & Karwowski, M. (2023). Lost in translation? Not for Large Language Models: Automated divergent thinking scoring performance translates to non‐English contexts. Thinking Skills and Creativity, 50, 101414 ; doi: 10.1016/j.tsc.2023.101414.</bibtext> </blist> </ref> <aug> <p>By Benjamin Goecke; Paul V. DiStefano; Wolfgang Aschauer; Kurt Haim; Roger Beaty and Boris Forthmann</p> <p>Reported by Author; Author; Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib31" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib18" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib14" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib17" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib21" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib15" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib36" firstref="ref9"></nolink> <nolink nlid="nl8" bibid="bib28" firstref="ref10"></nolink> <nolink nlid="nl9" bibid="bib22" firstref="ref12"></nolink> <nolink nlid="nl10" bibid="bib30" firstref="ref14"></nolink> <nolink nlid="nl11" bibid="bib38" firstref="ref15"></nolink> <nolink nlid="nl12" bibid="bib11" firstref="ref17"></nolink> <nolink nlid="nl13" bibid="bib29" firstref="ref18"></nolink> <nolink nlid="nl14" bibid="bib35" firstref="ref21"></nolink> <nolink nlid="nl15" bibid="bib34" firstref="ref22"></nolink> <nolink nlid="nl16" bibid="bib26" firstref="ref23"></nolink> <nolink nlid="nl17" bibid="bib23" firstref="ref25"></nolink> <nolink nlid="nl18" bibid="bib19" firstref="ref31"></nolink> <nolink nlid="nl19" bibid="bib32" firstref="ref32"></nolink> <nolink nlid="nl20" bibid="bib27" firstref="ref35"></nolink> <nolink nlid="nl21" bibid="bib16" firstref="ref36"></nolink> <nolink nlid="nl22" bibid="bib33" firstref="ref37"></nolink> <nolink nlid="nl23" bibid="bib10" firstref="ref38"></nolink> <nolink nlid="nl24" bibid="bib24" firstref="ref39"></nolink> <nolink nlid="nl25" bibid="bib12" firstref="ref40"></nolink> <nolink nlid="nl26" bibid="bib37" firstref="ref42"></nolink> <nolink nlid="nl27" bibid="bib20" firstref="ref46"></nolink> <nolink nlid="nl28" bibid="bib13" firstref="ref47"></nolink> <nolink nlid="nl29" bibid="bib25" firstref="ref52"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1444134
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Automated Scoring of Scientific Creativity in German
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Benjamin+Goecke%22">Benjamin Goecke</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3050-1848">0000-0002-3050-1848</externalLink>)<br /><searchLink fieldCode="AR" term="%22Paul+V%2E+DiStefano%22">Paul V. DiStefano</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0002-9638-3220">0009-0002-9638-3220</externalLink>)<br /><searchLink fieldCode="AR" term="%22Wolfgang+Aschauer%22">Wolfgang Aschauer</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-1221-6509">0000-0002-1221-6509</externalLink>)<br /><searchLink fieldCode="AR" term="%22Kurt+Haim%22">Kurt Haim</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4093-0512">0000-0003-4093-0512</externalLink>)<br /><searchLink fieldCode="AR" term="%22Roger+Beaty%22">Roger Beaty</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6114-5973">0000-0001-6114-5973</externalLink>)<br /><searchLink fieldCode="AR" term="%22Boris+Forthmann%22">Boris Forthmann</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9755-7304">0000-0001-9755-7304</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Creative+Behavior%22"><i>Journal of Creative Behavior</i></searchLink>. 2024 58(3):321-327.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 7
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)<br />National Science Foundation (NSF), Division of Undergraduate Education (DUE)
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Creativity%22">Creativity</searchLink><br /><searchLink fieldCode="DE" term="%22Creative+Thinking%22">Creative Thinking</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22German%22">German</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Correlation%22">Correlation</searchLink><br /><searchLink fieldCode="DE" term="%22Sciences%22">Sciences</searchLink><br /><searchLink fieldCode="DE" term="%22Thinking+Skills%22">Thinking Skills</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/jocb.658
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0022-0175<br />2162-6057
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Automated scoring is a current hot topic in creativity research. However, most research has focused on the English language and popular verbal creative thinking tasks, such as the alternate uses task. Therefore, in this study, we present a large language model approach for automated scoring of a scientific creative thinking task that assesses divergent ideation in experimental tasks in the German language. Participants are required to generate alternative explanations for an empirical observation. This work analyzed a total of 13,423 unique responses. To predict human ratings of originality, we used XLM-RoBERTa (Cross-lingual Language Model-RoBERTa), a large, multilingual model. The prediction model was trained on 9,400 responses. Results showed a strong correlation between model predictions and human ratings in a held-out test set (n = 2,682; r = 0.80; CI-95% [0.79, 0.81]). These promising findings underscore the potential of large language models for automated scoring of scientific creative thinking in the German language. We encourage researchers to further investigate automated scoring of other domain-specific creative thinking tasks.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: Note
  Label: Notes
  Group: Note
  Data: https://osf.io/aw95p
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1444134
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1444134
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/jocb.658
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 7
        StartPage: 321
    Subjects:
      – SubjectFull: Creativity
        Type: general
      – SubjectFull: Creative Thinking
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: German
        Type: general
      – SubjectFull: Prediction
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Correlation
        Type: general
      – SubjectFull: Sciences
        Type: general
      – SubjectFull: Thinking Skills
        Type: general
    Titles:
      – TitleFull: Automated Scoring of Scientific Creativity in German
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Benjamin Goecke
      – PersonEntity:
          Name:
            NameFull: Paul V. DiStefano
      – PersonEntity:
          Name:
            NameFull: Wolfgang Aschauer
      – PersonEntity:
          Name:
            NameFull: Kurt Haim
      – PersonEntity:
          Name:
            NameFull: Roger Beaty
      – PersonEntity:
          Name:
            NameFull: Boris Forthmann
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 09
              Type: published
              Y: 2024
          Identifiers:
            – Type: issn-print
              Value: 0022-0175
            – Type: issn-electronic
              Value: 2162-6057
          Numbering:
            – Type: volume
              Value: 58
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Journal of Creative Behavior
              Type: main
ResultId 1