Automated Scoring of Scientific Creativity in German

Saved in:
Bibliographic Details
Title: Automated Scoring of Scientific Creativity in German
Language: English
Authors: Benjamin Goecke (ORCID 0000-0002-3050-1848), Paul V. DiStefano (ORCID 0009-0002-9638-3220), Wolfgang Aschauer (ORCID 0000-0002-1221-6509), Kurt Haim (ORCID 0000-0003-4093-0512), Roger Beaty (ORCID 0000-0001-6114-5973), Boris Forthmann (ORCID 0000-0001-9755-7304)
Source: Journal of Creative Behavior. 2024 58(3):321-327.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 7
Publication Date: 2024
Sponsoring Agency: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
National Science Foundation (NSF), Division of Undergraduate Education (DUE)
Document Type: Journal Articles
Reports - Research
Descriptors: Creativity, Creative Thinking, Scoring, Automation, German, Prediction, Models, Correlation, Sciences, Thinking Skills
DOI: 10.1002/jocb.658
ISSN: 0022-0175
2162-6057
Abstract: Automated scoring is a current hot topic in creativity research. However, most research has focused on the English language and popular verbal creative thinking tasks, such as the alternate uses task. Therefore, in this study, we present a large language model approach for automated scoring of a scientific creative thinking task that assesses divergent ideation in experimental tasks in the German language. Participants are required to generate alternative explanations for an empirical observation. This work analyzed a total of 13,423 unique responses. To predict human ratings of originality, we used XLM-RoBERTa (Cross-lingual Language Model-RoBERTa), a large, multilingual model. The prediction model was trained on 9,400 responses. Results showed a strong correlation between model predictions and human ratings in a held-out test set (n = 2,682; r = 0.80; CI-95% [0.79, 0.81]). These promising findings underscore the potential of large language models for automated scoring of scientific creative thinking in the German language. We encourage researchers to further investigate automated scoring of other domain-specific creative thinking tasks.
Abstractor: As Provided
Notes: https://osf.io/aw95p
Entry Date: 2024
Accession Number: EJ1444134
Database: ERIC
Full text is not displayed to guests.
Description
Abstract:Automated scoring is a current hot topic in creativity research. However, most research has focused on the English language and popular verbal creative thinking tasks, such as the alternate uses task. Therefore, in this study, we present a large language model approach for automated scoring of a scientific creative thinking task that assesses divergent ideation in experimental tasks in the German language. Participants are required to generate alternative explanations for an empirical observation. This work analyzed a total of 13,423 unique responses. To predict human ratings of originality, we used XLM-RoBERTa (Cross-lingual Language Model-RoBERTa), a large, multilingual model. The prediction model was trained on 9,400 responses. Results showed a strong correlation between model predictions and human ratings in a held-out test set (n = 2,682; r = 0.80; CI-95% [0.79, 0.81]). These promising findings underscore the potential of large language models for automated scoring of scientific creative thinking in the German language. We encourage researchers to further investigate automated scoring of other domain-specific creative thinking tasks.
ISSN:0022-0175
2162-6057
DOI:10.1002/jocb.658