Planning Missing Data Designs for Human Ratings in Creativity Research: A Practical Guide

Saved in:
Bibliographic Details
Title: Planning Missing Data Designs for Human Ratings in Creativity Research: A Practical Guide
Language: English
Authors: Boris Forthmann (ORCID 0000-0001-9755-7304), Benjamin Goecke (ORCID 0000-0002-3050-1848), Roger E. Beaty (ORCID 0000-0001-6114-5973)
Source: Creativity Research Journal. 2025 37(1):167-178.
Availability: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
Peer Reviewed: Y
Page Count: 12
Publication Date: 2025
Sponsoring Agency: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
National Science Foundation (NSF), Division of Undergraduate Education (DUE)
Contract Number: 1920653
2155070
Document Type: Journal Articles
Reports - Research
Descriptors: Creativity, Research, Researchers, Research Methodology, Psychometrics, Simulation, Measurement Techniques, Item Response Theory, Evaluators, Models, Matrices, Data, Computation, Cost Effectiveness
DOI: 10.1080/10400419.2023.2250976
ISSN: 1040-0419
1532-6934
Abstract: Human ratings are ubiquitous in creativity research. Yet, the process of rating responses to creativity tasks -- typically several hundred or thousands of responses, per rater -- is often time-consuming and expensive. Planned missing data designs, where raters only rate a subset of the total number of responses, have been recently proposed as one possible solution to decrease overall rating time and monetary costs. However, researchers also need ratings that adhere to psychometric standards, such as a certain degree of reliability, and psychometric work with planned missing designs is currently lacking in the literature. In this work, we introduce how judge response theory and simulations can be used to fine-tune planning of missing data designs. We provide open code for the community and illustrate our proposed approach by a cost-effectiveness calculation based on a realistic example. We clearly show that fine-tuning helps to save time (to perform the ratings) and monetary costs, while simultaneously targeting expected levels of reliability.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1458414
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwFvgCgIeO6O_crM0Wiycif_AAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDLgokbY3FcyKNGDWGQIBEICBmwIhXju18UmYnlBnrAJsm-Gk2wg9J472t5VKKgbNXKK5yVPyFKBCBMOXXR3qGVblMgZdg5ibH_7Gi-ixg2BI0ZCm27YJ34kPRcMe006o-mlH9X_lwabMVPD_c62lL4coAKKmJ7CEcOlXNGX9nNhKQcmeEK9kBw9jJjGkQ3wY2k0zzXeFRDaT4wjGnNyakrRApSz5n7qL1ji9LCSZ
Text:
  Availability: 1
  Value: <anid>AN0182340424;7lo01jan.25;2025Jan23.01:32;v2.2.500</anid> <title id="AN0182340424-1">Planning Missing Data Designs for Human Ratings in Creativity Research: A Practical Guide </title> <p>Human ratings are ubiquitous in creativity research. Yet, the process of rating responses to creativity tasks – typically several hundred or thousands of responses, per rater – is often time-consuming and expensive. Planned missing data designs, where raters only rate a subset of the total number of responses, have been recently proposed as one possible solution to decrease overall rating time and monetary costs. However, researchers also need ratings that adhere to psychometric standards, such as a certain degree of reliability, and psychometric work with planned missing designs is currently lacking in the literature. In this work, we introduce how judge response theory and simulations can be used to fine-tune planning of missing data designs. We provide open code for the community and illustrate our proposed approach by a cost-effectiveness calculation based on a realistic example. We clearly show that fine-tuning helps to save time (to perform the ratings) and monetary costs, while simultaneously targeting expected levels of reliability.</p> <p>Human ratings are ubiquitous in creativity research. Beginning with early work by Guilford, which required human judges for the scoring of creative thinking performance (Christensen, Guilford, & Wilson, [<reflink idref="bib9" id="ref1">9</reflink>]; Wilson, Guilford, & Christensen, [<reflink idref="bib46" id="ref2">46</reflink>]), to Amabile's Consensual Assessment Technique (CAT) for the assessment of creative products (Amabile, [<reflink idref="bib2" id="ref3">2</reflink>]), and more recently, to creative thinking assessment of divergent thinking (Benedek, Mühlmann, Jauk, & Neubauer, [<reflink idref="bib6" id="ref4">6</reflink>]; Forthmann et al., [<reflink idref="bib14" id="ref5">14</reflink>]; Silvia et al., [<reflink idref="bib41" id="ref6">41</reflink>]), creative metaphor (Beaty & Silvia, [<reflink idref="bib5" id="ref7">5</reflink>]; Primi, [<reflink idref="bib30" id="ref8">30</reflink>]) or humor production (Christensen, Silvia, Nusbaum, & Beaty, [<reflink idref="bib10" id="ref9">10</reflink>]; Nusbaum, Silvia, & Beaty, [<reflink idref="bib28" id="ref10">28</reflink>]), human ratings seemingly follow a long tradition of being indispensable for the assessment of creativity. In a sense, they are considered a gold standard in creativity research. This status of human ratings is further emphasized by the fact that the most recent attempts of automated scoring of creative thinking aim at the prediction of human ratings (Beaty & Johnson, [<reflink idref="bib4" id="ref11">4</reflink>]; Buczak, Huang, Forthmann, & Doebler, [<reflink idref="bib7" id="ref12">7</reflink>]; Dumas, Organisciak, & Doherty, [<reflink idref="bib12" id="ref13">12</reflink>]; Stevenson et al., [<reflink idref="bib42" id="ref14">42</reflink>]). Automated scoring is motivated by the fact that human ratings are associated with two drawbacks: scoring might be affected by various idiosyncrasies of raters (Mouchiroud & Lubart, [<reflink idref="bib24" id="ref15">24</reflink>]; Robitzsch & Steinfeld, [<reflink idref="bib34" id="ref16">34</reflink>]) and scoring of creativity tasks by human raters does not always result in agreements between raters (Forthmann et al., [<reflink idref="bib14" id="ref17">14</reflink>]), but is very laborious and may even take weeks (Benedek, Mühlmann, Jauk, & Neubauer, [<reflink idref="bib6" id="ref18">6</reflink>]; Shaw, [<reflink idref="bib37" id="ref19">37</reflink>]; Silvia, Martin, & Nusbaum, [<reflink idref="bib40" id="ref20">40</reflink>]).</p> <p>The current work addresses the latter issue by demonstrating how Judge Response Theory (JRT; Myszkowski, [<reflink idref="bib26" id="ref21">26</reflink>]; Myszkowski & Storme, [<reflink idref="bib27" id="ref22">27</reflink>])—i.e., the application of item response modeling to human ratings in creativity research—and simulation techniques can be leveraged for effective distribution of rater work (i.e., by means of a planned missing data design), while at the same time a sufficient degree of measurement precision of the final scores is to be expected. Such a careful planning approach to rating designs can reduce the amount of coding work put on the shoulder of raters (preventing adverse effects of rater fatigue; e.g. Forthmann et al., [<reflink idref="bib14" id="ref23">14</reflink>]), allows the overall coding task being finished in comparably less time (i.e., compared to having all products rated by all available raters), and provides an empirical rationale for saving valuable project money. We argue that such work will be highly useful for the field as long as automated scorings are not yet fully available for all creativity measures that typically require human ratings.</p> <hd id="AN0182340424-2">Human ratings in creativity research</hd> <p>Roughly speaking, creativity refers to the novelty and usefulness of a perceptible product (Plucker, Beghetto, & Dow, [<reflink idref="bib29" id="ref24">29</reflink>]). This notion of creativity as a property of perceptible products is important for the current work, because it is open to encompass rather small expressions of thought. From this perspective, products include responses generated in a divergent thinking task (Runco, Plucker, & Lim, [<reflink idref="bib35" id="ref25">35</reflink>]), a list of ideas obtained from a small-group brainstorming session (Reinig & Briggs, [<reflink idref="bib33" id="ref26">33</reflink>]), a written story (Kornilov, Kornilova, & Grigorenko, [<reflink idref="bib20" id="ref27">20</reflink>]; Taylor & Barbot, [<reflink idref="bib44" id="ref28">44</reflink>]), responses to a scientific creative thinking task (Long, [<reflink idref="bib21" id="ref29">21</reflink>]; Long & Pang, [<reflink idref="bib22" id="ref30">22</reflink>]), as well as drawings or a design for furniture. Indeed, all these kinds of products are rated by human judges in creativity research which highlights the broad range of usage contexts of human rater scores.</p> <p>Beyond the product, it is important to look at the samples of raters used. For example, the famous consensual assessment technique (Amabile, [<reflink idref="bib2" id="ref31">2</reflink>]) is considered a valid measure of creativity because experts of the respective product domain assess a product's creativity (e.g., furniture designers rate the creativity of furniture designs). However, quasi-experts (e.g., design students which are not yet fully developed experts; Kaufman, Baer, Cropley, Reiter-Palmon, & Sinnett, [<reflink idref="bib18" id="ref32">18</reflink>]) have also been sampled for providing creativity ratings as well as laypersons (also named novices or naïve raters; Hass, Rivera, & Silvia, [<reflink idref="bib17" id="ref33">17</reflink>]; Kaufman, Baer, Cropley, Reiter-Palmon, & Sinnett, [<reflink idref="bib18" id="ref34">18</reflink>]). While researchers have cautioned against using other than expert samples, empirical work suggests that laypersons may provide valid and reliable ratings when adequately prepared for the rating task (Hass, Rivera, & Silvia, [<reflink idref="bib17" id="ref35">17</reflink>]; Storme, Myszkowski, Çelik, & Lubart, [<reflink idref="bib43" id="ref36">43</reflink>]). Either way, rater-characteristics should be taken into account when considering the provided responses in sophisticated statistical models (Myszkowski & Storme, [<reflink idref="bib27" id="ref37">27</reflink>]; Primi, Silvia, Jauk, & Benedek, [<reflink idref="bib31" id="ref38">31</reflink>]; Robitzsch & Steinfeld, [<reflink idref="bib34" id="ref39">34</reflink>]).</p> <p>When considering the full product range, it becomes further clear that for some products (e.g., responses generated in divergent thinking tasks) no group of experts is readily available. Who can be reasonably considered being an expert on how to creatively use a spoon? In such situations, raters are typically equipped with a more extensive coding guide instructing them that more creative responses tend to be uncommon, remote, and clever (Silvia et al., [<reflink idref="bib41" id="ref40">41</reflink>]). These three classical indicators of originality were already used by Guilford and colleagues for the scoring of divergent thinking tasks (Wilson, Guilford, & Christensen, [<reflink idref="bib46" id="ref41">46</reflink>]). For example, responses for the Plot Titles task were scored for cleverness, whereas responses to the Consequences task were judged by human raters for their remoteness (Christensen, Guilford, & Wilson, [<reflink idref="bib9" id="ref42">9</reflink>]). Similar scores as obtained by human ratings were used for tasks requiring the production of creative metaphors or humor (Beaty & Silvia, [<reflink idref="bib5" id="ref43">5</reflink>]; Christensen, Silvia, Nusbaum, & Beaty, [<reflink idref="bib10" id="ref44">10</reflink>]; Nusbaum, Silvia, & Beaty, [<reflink idref="bib28" id="ref45">28</reflink>]; Primi, [<reflink idref="bib30" id="ref46">30</reflink>]).</p> <p>Beyond these various characteristics of human ratings such as the type of products to rate, the sample of raters, or the scoring dimensions which differ from study to study, there is one common aspect of all such rating tasks that should be emphasized: the high workload put on the raters. For example, in a study focusing on divergent thinking, thousands of ratings might be needed (Kleinkorres, Forthmann, & Holling, [<reflink idref="bib19" id="ref47">19</reflink>]) and researchers have already considered approaches to reduce the amount of work this brings along. For example, rating all responses generated by a participant at once (i.e., the full response set and not each response separately) has been proposed to reduce overall rating time (Shaw, [<reflink idref="bib37" id="ref48">37</reflink>]; Silvia, Martin, & Nusbaum, [<reflink idref="bib40" id="ref49">40</reflink>]), but scoring all responses at once comes along with the need to rate a comparably more complex product (i.e., all responses vs. only one response) and it has been shown that increasing complexity of sets of responses is associated with larger rater disagreement (Forthmann et al., [<reflink idref="bib14" id="ref50">14</reflink>]). In addition, scoring all responses is not an option if the focus of a study is the level of responses (Silvia, Martin, & Nusbaum, [<reflink idref="bib40" id="ref51">40</reflink>]).</p> <p>A more sophisticated approach that reduces the burden on raters is to rely on planned missing data designs (Graham, [<reflink idref="bib16" id="ref52">16</reflink>]) which allow for unbiased estimation of creativity scores based on the response level. In fact, these designs take advantage of a reduced amount of information per rater without relying on a single data point by treating each response for what it is: an individual behavioral outcome to a given task. In the case of creativity research, these can be referred to as products. Notably, there are applications of planned missing data designs in creativity research (Barbot, [<reflink idref="bib3" id="ref53">3</reflink>]; Fürst, [<reflink idref="bib15" id="ref54">15</reflink>]; Primi, Silvia, Jauk, & Benedek, [<reflink idref="bib31" id="ref55">31</reflink>]). For example, Barbot ([<reflink idref="bib3" id="ref56">3</reflink>]) merged different datasets in which some measures were in common and added another dataset, as a type of linking sample to increase covariance coverage. Hence, his approach of integrative data analysis was highly similar with a planned missing data design. In addition, Fürst ([<reflink idref="bib15" id="ref57">15</reflink>]) employed a planned missing data design for a more efficient, yet comprehensive, assessment of creative potential. While Barbot ([<reflink idref="bib3" id="ref58">3</reflink>]) and Fürst ([<reflink idref="bib15" id="ref59">15</reflink>]) used structural equation modeling and full information maximum likelihood estimation to handle missing data, Primi, Silvia, Jauk, and Benedek ([<reflink idref="bib31" id="ref60">31</reflink>]) examined simulated missing data patterns in the context of many facet Rasch modeling.</p> <hd id="AN0182340424-3">Psychometric modeling of human ratings</hd> <p>Judge response theory (JRT) refers to the adaptation of polytomous item response theory models [e.g., the graded response model (Samejima, [<reflink idref="bib36" id="ref61">36</reflink>]) or the generalized partial credit model (Muraki, [<reflink idref="bib25" id="ref62">25</reflink>])] for human ratings in the context of creativity research (Myszkowski, [<reflink idref="bib26" id="ref63">26</reflink>]; Myszkowski & Storme, [<reflink idref="bib27" id="ref64">27</reflink>]). JRT explicitly models differences in the rating behavior of human judges as reflected by severity (or leniency) effects and differences between raters with respect to their discrimination parameter. The considered unit of measurement in JRT is situated at the product level. Products must be conceptualized very broadly for the context of the current work. For example, one might consider more complex products such as newly designed electronic devices or very simple expressions of thought (e.g., a response generated within the context of creative thinking testing). Following Plucker, Beghetto, and Dow ([<reflink idref="bib29" id="ref65">29</reflink>]) and Runco, Plucker, and Lim ([<reflink idref="bib35" id="ref66">35</reflink>]), we use the term "product" in its broadest sense referring to something that is perceptible. With the availability of ratings for each of the products in a given dataset, JRT as implemented in the R package jrt (Myszkowski, [<reflink idref="bib26" id="ref67">26</reflink>]) is a powerful tool for estimation of rater parameters and reliability. The jrt package uses the mirt package (Chalmers, [<reflink idref="bib8" id="ref68">8</reflink>]) for multidimensional item response theory modeling as estimation engine. The mirt package provides fast estimation algorithms (also when missing data are present), and it also includes highly efficient functions for simulation studies.</p> <p>Specifically, the default setting of the jrt() function (i.e., the main function of the jrt package) fits various unidimensional polytomous item response theory models to the rating data and chooses the best fitting model (Myszkowski, [<reflink idref="bib26" id="ref69">26</reflink>]). The best fitting model is chosen based on the Akaike information criterion (AIC; Akaike, [<reflink idref="bib1" id="ref70">1</reflink>]) which combines a model's likelihood and a penalty for model complexity to emphasize that a useful model should be parsimonious and fit to the data. Smaller AIC values imply better model fit, while model parsimony is also taken into account. Consequently, in the jrt() function the model with the lowest AIC is chosen. In addition, it provides classical inter-rater agreement statistics such as intra-class correlations and also model-based empirical reliability estimates along with estimates of the latent creative quality for each of the products in the datasets. Here, reliability refers to the estimated squared correlation between the estimated latent creative quality [i.e., the factor scores provided by the jrt() function] and true latent creative quality. The square-root of this estimate is also known as factor determinacy index (FDI) in the literature on factor analysis. Hence, the FDI is an estimate of the correlation between the factor scores obtained for the rated products and the true factor scores. For both types of indices (i.e., reliability and FDI), cutoffs for research and practical assessment contexts exist and such cutoffs should be considered when planning a study involving creativity ratings. For a research project, the rating design could be planned toward a target FDI of.80 (Ferrando & Lorenzo-Seva, [<reflink idref="bib13" id="ref71">13</reflink>]), for example. Although we appreciate that such cutoffs are frequently subject of debate, and rightfully so, they offer a pragmatic benchmark that can be easily tested against.</p> <p>In general, the recruitment of human raters is hard, and especially so in the case of rating creativity data, because rating creativity data is less of a trivial task as compared to other rating ventures. Further, the recruitment of human judges might be limited due to monetary reasons, for example, because the available project budget might be already exhausted. Regardless, researchers will be interested in getting the job done in the most pragmatic way: with the least amount of time and money spent, while still obtaining high-quality data (i.e., reliable ratings that approximate the true ability of a participant).</p> <p>Such a dataset that needs rating might, for example, include 5000 products (e.g., divergent thinking test responses). If we assume 500 responses can be rated (without rushing) in 1 hour (i.e., on average 7.2 seconds per response) and each rater receives $10 per hour of work, this would equal $100 per rater for rating the full response set. Obviously, costs increase with the number of raters that are employed for the task, and as the raters will usually not work in full synchronicity, the temporal aspect need not be forgotten either.</p> <p>The number of raters depends, as mentioned above, on various considerations and, in a best-case scenario, should be chosen on the basis of empirical reasoning. For example, an arguably well-defined criterion would be to adhere to the afore-mentioned.80 FDI cutoff (i.e., FDI >.80), in order to ensure sufficient reliability for further analysis. If the most likely model and typical model parameters for the target rater population and the target product population are known (at least reasonable "guesstimates" are needed, which could also rely on experiences with similar data), it is possible to leverage mirt's simulation functionality to answer such questions of rater design (i.e., how many raters are needed?).</p> <p>Based on a simulation study, an expected FDI can be obtained. In the same vein, we obtain information regarding the number of raters and how many responses are rated by how many raters. Hence, the expected measurement precision and how much it would cost can be determined a priori. For example, based on the above calculations, it might be that a cutoff of.80 for the FDI will be surpassed with 3 raters which implies costs of $300. However, it might take 10 raters to surpass a cutoff of.90 and a much higher budget of $1000 would then be needed. Although the costs for the respective number of raters could have been easily calculated before, the simulation study extends the provided information by an estimate regarding the FDI, and thus the reliability of the obtained ratings. This information provides real value to researchers, as this enables them to consider trade-offs between monetary and temporal costs and measurement precision.</p> <p>To further refine planning of rater design, it is also highly useful that mirt models can be estimated for planned missing data designs (e.g., Fürst, [<reflink idref="bib15" id="ref72">15</reflink>]). For example, it is possible to simulate data that account for specific levels of missingness (e.g., one rating less for 20% of all responses). This functionality can thus be used in a pragmatic way to further reduce costs without sacrificing too much measurement precision. By means of planned missing data designs, rating designs can be efficiently and effectively planned toward both a target level of measurement precision and an available budget. Similarly, missing data designs could be used for planning toward a given due date at which the ratings must be available for further data analysis.</p> <hd id="AN0182340424-4">The present research</hd> <p>The goal of the current work is to introduce simulation-based planning of rating designs that (a) incorporate planned missingness designs and (b) allow for an effective outweighing of a target level of measurement precision and monetary costs (or costs in terms of time, when approaching a due date). We argue that this work will be helpful for researchers studying creativity, as it provides a practical example on how to use planned missing rating designs for their own purposes. To this end, we illustrate the usefulness of this approach based on a realistic planning scenario when layperson ratings are to be used for scoring of an Alternate Uses Task (e.g., Hass, Rivera, & Silvia, [<reflink idref="bib17" id="ref73">17</reflink>]).</p> <hd id="AN0182340424-5">Method</hd> <p>The empirical part of this work comprises (a) an initial analysis of rating data and (b) a simulation study to inform planning of a missing data design. The first part was needed to derive a realistic simulation model and ranges for model parameters, whereas the second part was needed to see which planned missing data designs work well with respect to psychometric as well as cost-effectiveness criteria. Importantly, the empirical part should be understood as providing proof-of-concept on how to implement the approach for planning a missing data design for human ratings. As such, all reported findings are limited to the rater population that was sampled and the Alternate Uses Task as a measure of creative thinking. Thus, we strongly recommend caution when interpreting the current findings. Especially, for other rater populations, other populations from which participants are sampled from, and other creativity measures, we recommend to contextualize and redo all steps outlined in this work.</p> <hd id="AN0182340424-6">Dataset</hd> <p>We first tested different rater models on an available dataset. The dataset included 3236 responses generated by <emph>N</emph> = 209 participants on two different Alternate Uses Tasks (using the words <emph>box</emph> and <emph>rope</emph>, respectively). Participants had 2 minutes to complete each of the tasks and were instructed to be creative. Each response was rated by three raters (undergraduate students majoring in psychology) using the subjective scoring method guidelines for divergent thinking (https://osf.io/vie7s). Following these guidelines, the raters used a 5-point Likert-scale. According to Cicchetti's criteria (Cicchetti, [<reflink idref="bib11" id="ref74">11</reflink>]), inter-rater reliabilities were fair in terms of absolute agreement (ICC =.42, 95%-CI: [.02,.63]) and consistency (ICC =.56, 95%-CI: [.53,.58]). The study was approved by the Institutional Review Board of The Pennsylvania State University. All participants gave informed consent to participate in the study.</p> <hd id="AN0182340424-7">Obtaining a realistic rater model</hd> <p>Before a simulation can be set up for planning of a rating design, a reasonable simulation model and a realistic range of rater parameters (i.e., parameters related to raters' severity and discrimination between products) ought to be found. We used the jrt() function from the jrt package (Myszkowski, [<reflink idref="bib26" id="ref75">26</reflink>]) which is implemented in the statistical software R (R Core Team, [<reflink idref="bib32" id="ref76">32</reflink>]). This way, the best fitting polytomous IRT model was determined based on the Akaike Information Criterion and AIC-based model weights (Wagenmakers & Farrell, [<reflink idref="bib45" id="ref77">45</reflink>]). The parameter estimates from the best fitting model were then used to construct an empirically justified simulation setup. Model comparison results of all computed models can be found in Table S1 in the online supplemental material (https://osf.io/7b9z5/).</p> <p>The best fitting model was the generalized partial credit model (Muraki, [<reflink idref="bib25" id="ref78">25</reflink>]). As this model will be used for the simulation below, it is worthwhile to consider its model equation</p> <p>(<reflink idref="bib1" id="ref79">1</reflink>)</p> <p>Graph</p> <p> <ephtml> <math xmlns="http://www.w3.org/1998/Math/MathML"><mi mathvariant="italic">P</mi><mo stretchy="false">(</mo><mi mathvariant="italic">X</mi><mo>=</mo><mi mathvariant="italic">k</mi><mo fence="false" stretchy="false">|</mo><msub><mi>θ</mi><mi>i</mi></msub><mo>,</mo><msub><mi>a</mi><mi>j</mi></msub><mo>,</mo><msub><mi>d</mi><mi>j</mi></msub><mo stretchy="false">)</mo><mo>=</mo><mrow><mfrac><mrow><mrow><mi mathvariant="normal">e</mi></mrow><mrow><mi mathvariant="normal">x</mi></mrow><mrow><mi mathvariant="normal">p</mi></mrow><mrow><mo fence="false" stretchy="false">[</mo><mi mathvariant="italic">a</mi><msub><mi mathvariant="italic">k</mi><mrow><mi mathvariant="italic">k</mi><mo>−</mo><mn>1</mn></mrow></msub><mrow><mo stretchy="false">(</mo><mi mathvariant="italic">a</mi><mo>∗</mo><mi mathvariant="italic">θ</mi><mo stretchy="false">)</mo></mrow><mo>+</mo><msub><mi>d</mi><mrow><mi mathvariant="italic">k</mi><mo>−</mo><mn>1</mn></mrow></msub><mo fence="false" stretchy="false">]</mo></mrow></mrow><mfrac><mrow /><mrow /></mfrac></mfrac><mrow><mrow><munderover><mo>∑</mo><mrow><mi mathvariant="italic">v</mi><mo>=</mo><mn>1</mn></mrow><mi>K</mi></munderover></mrow><mrow><mi mathvariant="normal">e</mi></mrow><mrow><mi mathvariant="normal">x</mi></mrow><mrow><mi mathvariant="normal">p</mi></mrow><mrow><mo fence="false" stretchy="false">[</mo><mi mathvariant="italic">a</mi><msub><mi>k</mi><mrow><mi mathvariant="italic">v</mi><mo>−</mo><mn>1</mn></mrow></msub><mrow><mo stretchy="false">(</mo><mi mathvariant="italic">a</mi><mo>∗</mo><mi mathvariant="italic">θ</mi><mo stretchy="false">)</mo></mrow><mo>+</mo><msub><mi>d</mi><mrow><mi mathvariant="italic">v</mi><mo>−</mo><mn>1</mn></mrow></msub><mo fence="false" stretchy="false">]</mo></mrow></mrow></mrow><mo>,</mo></math> </ephtml> </p> <p>with θ<subs><emph>i</emph></subs> being the latent score for response <emph>i</emph>, <emph>a</emph><subs><emph>j</emph></subs> being the discrimination parameter of rater <emph>j</emph>, <emph>d</emph><subs><emph>j</emph></subs> being the intercept vector of rater <emph>j</emph>, and <emph>ak</emph><subs><emph>k</emph></subs> being constraint to 0, 1, ... , <emph>K</emph>-1 (with <emph>K</emph> being the number of response categories). In addition, <emph>I</emph> refers to the number of responses in the context of this work, and <emph>J</emph> to the number of raters. This parameterization of the GPCM is implemented in the mirt package (Chalmers, [<reflink idref="bib8" id="ref80">8</reflink>]) and commonly referred to as the slope-intercept parameterization (e.g., Matlock, Turner, & Gitchel, [<reflink idref="bib23" id="ref81">23</reflink>]). The model is further identified by assuming that the latent response scores are <emph>N</emph>(0,1) distributed. Given that the data at hand had five response categories (i.e., <emph>K</emph> = 5), there were <emph>K</emph>-1 = 4 intercept parameters for each of the three raters (i.e., <emph>d</emph><subs>1<emph>j</emph></subs>, <emph>d</emph><subs>2<emph>j</emph></subs>, <emph>d</emph><subs>3<emph>j</emph></subs>, and <emph>d</emph><subs>4<emph>j</emph></subs> with <emph>j</emph> = 1, ... , <emph>J</emph>), and three discrimination parameters (i.e., <emph>a</emph><subs>1</subs>, <emph>a</emph><subs>2</subs>, and <emph>a</emph><subs>3</subs>). The estimated model parameters for each rater are shown in Table 1. For example, Rater 2 was found to have a much higher discrimination parameter as compared to Rater 1 and Rater 3 which means that this rater was much better at distinguishing highly creative responses from less creative responses. In addition, for Rater 2 by far the lowest intercept parameters were obtained which means that this rater was the most severe during the rater process. Rater 1 and Rater 3 were much more lenient in their ratings with Rater 1 being the most lenient one in this rater sample (c.f., Table 1).</p> <p>Table 1. Rater parameter estimates based on the generalized partial credit model.</p> <p> <ephtml> <table><thead><tr><td>Parameter</td><td>Rater 1</td><td>Rater 2</td><td>Rater 3</td></tr></thead><tbody><tr><td><italic>a</italic><sub><italic>j</italic></sub></td><td><italic>a</italic><sub>1</sub> = 0.72</td><td><italic>a</italic><sub>2</sub> = 1.83</td><td><italic>a</italic><sub>3</sub> = 0.77</td></tr><tr><td><italic>d</italic><sub>1<italic>j</italic></sub></td><td><italic>d</italic><sub>11</sub> = 3.84</td><td><italic>d</italic><sub>12</sub> = −1.66</td><td><italic>d</italic><sub>13</sub> = 1.67</td></tr><tr><td><italic>d</italic><sub>2<italic>j</italic></sub></td><td><italic>d</italic><sub>21</sub> = 3.53</td><td><italic>d</italic><sub>22</sub> = −3.76</td><td><italic>d</italic><sub>23</sub> = 1.03</td></tr><tr><td><italic>d</italic><sub>3<italic>j</italic></sub></td><td><italic>d</italic><sub>31</sub> = 1.92</td><td><italic>d</italic><sub>32</sub> = −7.52</td><td><italic>d</italic><sub>33</sub> = −0.89</td></tr><tr><td><italic>d</italic><sub>4<italic>j</italic></sub></td><td><italic>d</italic><sub>41</sub> = −1.16</td><td><italic>d</italic><sub>42</sub> = −13.42</td><td><italic>d</italic><sub>43</sub> = −3.79</td></tr></tbody></table> </ephtml> </p> <p>1 Note. <emph>a</emph> = discrimination parameter, <emph>d</emph> = intercept. The response category curves for all three raters can be found in Figure S1 in the online supplemental material (https://osf.io/7b9z5/).</p> <hd id="AN0182340424-8">Simulations</hd> <p></p> <hd id="AN0182340424-9">Construction of planned missing data matrices</hd> <p>In this work, we focus on matrix planned missing data designs (Silvia, Kwapil, Walsh, & Myin-Germeys, [<reflink idref="bib39" id="ref82">39</reflink>]) because they allow equal distribution of work across available raters. We constructed the designs the following way:</p> <p></p> <ulist> <item> We obtained all possible combinations of raters based on the overall number of raters and the target number of ratings per response. The combinations were calculated by means of the CombSet() function from the DescTools R package (Signorell, [<reflink idref="bib38" id="ref83">38</reflink>]). For example, for three available raters and two ratings per response there would be three possible combinations of raters: {{Rater 1, Rater 2}, {Rater 1, Rater 3}, {Rater 2, Rater 3}}.</item> <p></p> <item> We determined how many rows in the planned missing data matrix should be rated by each of the combinations obtained from Step 1. This was obtained by the floor function of the ratio of overall number of responses and the number of combinations obtained from Step 1. If the number of responses exceeded this number, the last combinations were randomly sampled with replacement from all possible combinations.</item> <p></p> <item> The matrix planned missing data designs obtained from Step 2 were further reduced or increased as a final optional step. A reduced design was obtained by randomly setting a planned rating to a planned missing value for a fixed percentage of responses. Analogously, an increased design was obtained by randomly setting a planned missing value to a planned rating for a fixed percentage of responses. Responses were also chosen randomly for both types of designs.</item> </ulist> <hd id="AN0182340424-10">Design</hd> <p>In our simulation design, we varied the number of raters (2 vs. 3 vs. 4 vs. 5) and the number of responses rated by each rater (2 vs. 3 vs. 4 vs. 5) resulting in 10 possible design cells (i.e., 5 ratings were only possible with 5 raters; see also Figure 1). The number of possible raters adheres to numbers of raters usually used for research purposes. In addition, we crossed this design with different percentages (20% vs. 40% vs. 60% vs. 80%) to reduce the number of ratings needed which resulted in 40 additional design cells. Decreasing a design with 3 ratings per response by 20%, for example, means that 20% of the responses will receive only 2 ratings per response. The responses and the rater who would not rate the response anymore were chosen randomly. Analogously, we crossed the design with different percentages (20% vs. 40% vs. 60% vs. 80%) to increase the number of ratings needed which resulted in another 24 additional design cells. Thus, here we did not combine cells with increasing percentages in which the number of raters equals the number of ratings (i.e., all complete designs). It should be noted, however, that increasing a complete design with two raters by 20% would also result from reducing a complete three rater design by 80% (which is already included). Increasing a design with 3 ratings per response by 20%, for example, means that 20% of the responses will receive 4 ratings per response. The responses and the rater who would rate this additional response were also chosen randomly. Thus, overall 74 different design cells were simulated.</p> <p>Graph: Figure 1. Results of Full Design Simulations.</p> <hd id="AN0182340424-11">Data generation</hd> <p>We used the simdata() function from the mirt package (Chalmers, [<reflink idref="bib8" id="ref84">8</reflink>]) for data generation. First, we sampled latent response scores from a <emph>N</emph>(0, 1) distribution. Discrimination parameters were sampled from a <emph>U</emph>(0.72, 1.83) distribution (i.e., the range was taken from the estimates reported in Table 1). The intercept parameters were sampled as follows: first, we calculated the average across each rater's intercept parameters to reflect rater easiness. Then, we sampled from a <emph>U</emph>(−6.59, 2.03) to reflect rater easiness. Next, we subtracted each rater's easiness from their four intercept parameters for centered intercept parameters. Each of the four centered intercept parameters was averaged across raters and used to construct a sampling rationale for the four intercept parameters. The <emph>d</emph><subs>1</subs> parameter was sampled from <emph>U</emph>(2.97, 4.93) with 2.97 being the average centered <emph>d</emph><subs>1</subs> parameter across raters and 4.93 being the maximum of the centered <emph>d</emph><subs>1</subs> parameters. The <emph>d</emph><subs>2</subs> parameter was sampled from <emph>U</emph>(1.95, 2.97) and the <emph>d</emph><subs>3</subs> parameter from <emph>U</emph>(−0.48, 1.94) with the lower bounds here being the average centered intercept parameters, respectively. Finally, the <emph>d</emph><subs>4</subs> parameter was sampled from <emph>U</emph>(−6.83, −0.49) with −6.83 being the minimum of the centered <emph>d</emph><subs>4</subs> parameters. The sampled easiness and the sampled centered intercept parameters were added up to yield the intercept parameters for data generation. For each cell, we simulated 500 replications of 1000 AUT responses (e.g., approximating an assessment context in which for <emph>n</emph> = 100 participants 10 responses are to be expected on average). The average correlation across replications between the estimated latent response scores (based on the expected a-posteriori method; EAP) and the true latent response scores was our main dependent variable in this simulation. We further obtained the standard deviations and the standard errors of the correlations as an indicator of sampling variability. The R code to reproduce all reported results in this work is openly available via the Open Science Framework (https://osf.io/7b9z5/).</p> <hd id="AN0182340424-12">Results and discussion</hd> <p></p> <hd id="AN0182340424-13">Simulation-based planning of a rater design</hd> <p>The reported findings from our simulation study serve the purpose of making readers familiar with interpreting findings obtained by the proposed approach for planning of missing data designs. In addition, we report the findings quite comprehensively so that interested researchers get an impression of how one might adjust simulation-based planning (e.g., by decreasing or increasing the number of rated responses for a proportion of raters) in ways that improve initially unsuccessful designs. For example, a design could be considered as unsuccessful when reliability is far above a target cutoff (the efficiency of the design can still be improved) or still below such a target cutoff (the expected psychometric quality must be improved).</p> <p>First, we present the results of the full-data simulations, where either all raters rated all responses or all responses were rated by <emph>n</emph>-1 raters in Figure 1. We observed a clear main effect on the number of ratings per response. Although the confidence intervals for all simulations supposing three ratings per response include the defined target level (<emph>r</emph> =.8) of the correlation between latent score estimates and their true values, on average this specific target level is not surpassed under this condition (i.e., three ratings per response). This holds independent of the number of raters that were specified. In order to exceed the defined target level of <emph>r</emph> =.8, at least four ratings per response would be needed, which corresponds to employing at least four independent raters.</p> <p>Next, we present the results of the planned-missingness data simulations, where in each simulation the ratings per response of the full dataset were reduced by either 20%, 40%, 60%, or 80% (Figure 2). Again, decreasing a design with 3 ratings per response by 20%, for example, means that 20% of the responses will receive only 2 ratings per response. Again, we observed a clear main effect on the number of ratings per response. However, although again at least four ratings per response yield the best results in terms of surpassing the a priori defined correlation of.80, further reducing the relative amount of responses that need to be rated at least four times, does not impair the estimated correlations very much. On the contrary, reducing the responses needed to be rated by all four raters by 60% still leaves enough information in the data to surpass the target level of <emph>r</emph> =.80. This finding can be readily translated to a monetary advantage, as not all raters have to rate all of the responses, but sufficient reliability is still achieved.</p> <p>Graph: Figure 2. Results of reduced design simulations.</p> <p>Lastly for this section, we show the results of the planned-missingness data simulations, where in each simulation the ratings per response of the full dataset were increased by either 20%, 40%, 60%, or 80% (Figure 3). In these simulations, the previously observed main effect of number of ratings per response remained. We were not able to identify any substantial effects that go beyond this main effect; the confidence intervals of all remaining simulation cells were overlapping. Increasing the responses needed to be rated by a given set of raters slightly increases the observed correlation, but according to the here provided data the differences might be negligible.</p> <p>Graph: Figure 3. Results of increased design simulations.</p> <hd id="AN0182340424-14">Cost-effectiveness calculations</hd> <p>In this section, we provide some insights into possible cost-effectiveness calculations, that is, considerations regarding a trade-off between measurement precision and monetary costs. To do so, we first provide a set of assumptions for our calculations: We assume that a layperson rater, who is properly trained via a short, written instruction regarding what is expected of them, can rate about 500 responses per hour. This equals 7.2 seconds per response, but this estimate seems reasonable given that raters will usually accelerate the rating process over time and with every response. For the sake of the argument, we also assume that raters are paid $10 per every hour of work. In the current case, we further assume that instructing a rater does not count as time spent working; after all, we would like to provide relatively pure estimations only regarding the rating process itself. In addition to that, we base but do not constrain our calculations to the assumption that any given number of raters can work perfectly parallel to each other. Although this assumption will be rarely met in reality, it will help to illustrate the inherent advantages of using certain planned-missingness rater designs. For our first example, we will further assume that four human raters can be appointed to rating data of an Alternate Uses Task with 1000 responses in total. We aim at illustrating the process that can be applied to decide for one or the other planned-missingness rater design.</p> <p>In Table 2, we provide an overview of all relevant parameters important for deciding for a rater design. Each row of the table refers to a unique (planned-missingness) rater design. We explicitly report the number of total raters; the given ratings per response; whether the full, an increased, or a reduced dataset was used; how many responses were assigned to each rater; what the mean and the standard deviation of the obtained correlation was; how much money a specific design translates; and the estimated rating time in total and per rater (which would also equal the total rating time for all four raters, if all of them would be working perfectly parallel).</p> <p>Table 2. Example of cost-effectiveness calculations.</p> <p> <ephtml> <table><thead><tr><td><italic>N</italic><sub>Raters</sub></td><td>Ratings per Response</td><td>Condition</td><td>Range<sub>Responses</sub><sub>per Rater</sub></td><td><italic>M</italic><sub><italic>r</italic></sub></td><td><italic>SD</italic><sub><italic>r</italic></sub></td><td>Estimated Costs</td><td>Estimated Time<sub>total</sub> in h</td><td>Estimated Time<sub>per Rater</sub> in h</td></tr></thead><tbody><tr><td>4</td><td>4</td><td>Full</td><td>1000</td><td>.832</td><td>.116</td><td>$80.00</td><td>8.00</td><td>2.00</td></tr><tr><td>4</td><td>4</td><td>reduction (20%)</td><td>940–957</td><td>.828</td><td>.113</td><td>$76.00</td><td>7.60</td><td>1.90</td></tr><tr><td>4</td><td>4</td><td>reduction (40%)</td><td>891–913</td><td>.819</td><td>.109</td><td>$72.00</td><td>7.20</td><td>1.80</td></tr><tr><td>4</td><td>4</td><td>reduction (60%)</td><td>831–864</td><td>.808</td><td>.109</td><td>$68.00</td><td>6.80</td><td>1.70</td></tr><tr><td>4</td><td>4</td><td>reduction (80%)</td><td>782–809</td><td>.804</td><td>.105</td><td>$64.00</td><td>6.40</td><td>1.60</td></tr><tr><td>4</td><td>3</td><td>increase (20%)</td><td>758–802</td><td>.801</td><td>.105</td><td>$62.12</td><td>6.21</td><td>1.55</td></tr><tr><td>4</td><td>3</td><td>increase (40%)</td><td>776–860</td><td>.804</td><td>.106</td><td>$64.20</td><td>6.42</td><td>1.61</td></tr><tr><td>4</td><td>3</td><td>increase (60%)</td><td>779–914</td><td>.808</td><td>.107</td><td>$66.50</td><td>6.65</td><td>1.66</td></tr><tr><td>4</td><td>3</td><td>increase (80%)</td><td>804–956</td><td>.810</td><td>.109</td><td>$68.70</td><td>6.87</td><td>1.72</td></tr></tbody></table> </ephtml> </p> <p>2 Note. Reduction = % of total responses that are rated by n − 1 raters. Increase = % of total responses that are rated by n + 1 raters. An extended version of this table including much more simulated conditions can be found in Table S2 in the online supplemental material file in the OSF repository (https://osf.io/7b9z5/).</p> <p>For example, while having the full data set rated by all four raters (i.e., 1000 responses per rater) would cost $80 and, on average, yield a correlation of.83 between latent score estimates and their true scores; using a design that supposes only 3 ratings per response, with an increase of one more rating per response for only 20% of the data, would reduce the total estimated costs by > ⅕ (i.e., ~22.5%), and still yields a correlation of <emph>r</emph> =.80. This reduction in monetary costs is obviously also reflected in the time that is needed to obtain all necessary ratings; that is, instead of 8 h of scoring for the full data, implementing the planned missingness design of the provided example results in a total time of 6.2 h.</p> <p>It can be argued that this reduction of monetary and temporal costs by 22.5% could be understood as both a relative and an absolute increase in cost-effectiveness. Whereas in our example with 1000 responses, the absolute cost reduction of the planned missingness rating design can seem negligible in the light of huge research grants, or when researchers only plan on rating one creativity task like the Alternate Uses Task, the inherent benefit of these planned designs becomes clearer, when a larger scale is considered.</p> <p>For example, consider a large online-panel study assessing creativity by means of a two-item Alternate Uses Task with 1000 participants. If we assume that each participant, on average, provides 10 responses per item, a huge dataset with 20,000 responses would be obtained. Appointing four raters to rate all of the responses would result in costs of $1,600 (40 h of work per rater) and take a considerable amount of time, as rating creativity responses is usually not a full-time job and moreover exhausting for the raters (fatigue). If the design mentioned above would be applied to this situation (3 ratings per response, with an increase of one more rating per response for only 20% of the data), the relative decrease in costs would of course remain the same, but in terms of absolute numbers the cost decrease would add up to $360, which sometimes is the price of attending a conference to present the results of a study. In addition, of course the time for each rater working on their rating would decrease considerably (9 h – which is longer than the time spent working in an ordinary 9–5 job).</p> <hd id="AN0182340424-15">Summary and recommendations</hd> <p>In this work, we proposed a simulation-based approach for effective planning of rater designs with missing data. We demonstrated in an empirical proof of concept illustration how a reasonable simulation model can be obtained from existing data, how simulations can be used to fine-tune the planned design, and how based on these simulations cost-benefit analyses can be done when project budget and/or time are limited resources. Specifically, we used available rating data for responses on the Alternate Uses task and found by means of the jrt package (Myszkowski, [<reflink idref="bib26" id="ref85">26</reflink>]) that the GPCM fitted these data best. Hence, we used the GPCM and the obtained parameter estimates for informing simulation-based planning. Then, our simulation implies strategies that are useful for fine-tuning the planned missing values design: run simulations with full data designs and varying numbers of raters, identify the full data designs that are closest to a target level of the correlation between estimated and true factor scores (e.g.,.80), and finally increase or decrease the number of ratings per response for a certain proportion of randomly chosen responses.</p> <p>Importantly, we have shown that even in situations in which ratings might not be too expensive in terms of monetary costs, enough money could be saved that allows a doctoral student, for example, to go to a conference. Clearly, in case that experts are needed as raters for a study the planning approach outlined in this paper is expected to result in even greater savings, because expert raters are much more expensive; for example, architects that would be hired to rate construction designs provided by participants of a study on architectural creativity.</p> <p>We strongly recommend that researchers use this approach – adapted to the context of their studies – for the case that human ratings of creative products are involved to ensure the quality of final scores based on planned missing data designs. We provide the needed R code for simulation-based planning in an openly accessible repository (https://osf.io/7b9z5/) to facilitate this step for researchers who are not yet familiar with the software used in this work. However, having planned a missing data rating design implies that further steps are needed.</p> <p>As a final step, one would reevaluate the best fitting model of the obtained ratings by means of the jrt package. Of course, the more is known about the target rater population (e.g., laypersons for rating divergent thinking responses), the unlikelier it will be that the JRT model fitting the data best will deviate from the anticipated model in the planning phase. However, as in our illustration here one might have only three raters available for setting up a reasonable simulation model (or even no data at all). In such situations, the final data could better fit to a different model which should then be used for deriving final latent scores. Furthermore, also the finally achieved reliability of the scores should be reevaluated to check if the rating process resulted in the anticipated level of measurement precision and/or if the level of measurement precision is high enough for the purpose of measurement (Ferrando & Lorenzo-Seva, [<reflink idref="bib13" id="ref86">13</reflink>]). We recommend to check the square-root of empirical reliability of the final scores which provides an estimate of the correlation between the estimated latent response scores and the true responses.</p> <hd id="AN0182340424-16">Limitations and future directions</hd> <p>The dataset we used in our study for illustration and for a hypothetical planning scenario might not have been comprehensive. Additional complexities are expected to arise when, for example, model parameters for each rater differ as a function of the task for which the ratings are needed. For example, in the used dataset, participants generated responses for two different AUT objects (i.e., <emph>box</emph> and <emph>rope</emph>) and we ignored that discrimination and intercept parameters in the GPCM could differ between both objects. Such differences could be considered during the simulation by means of using a multiple-group model with as many groups as there are tasks in the planned study. However, such a more complex simulation would only make sense when enough empirical evidence for a mostly non-overlapping parameter range between the tasks is available. Of course, this knowledge can only be gained if such differences in parameters are evaluated and this can be nicely done at the stage of reevaluating model fit and reliability of the final ratings.</p> <p>The outlined empirical example is further limited to a range of matrix planned missing data designs (Silvia, Kwapil, Walsh, & Myin-Germeys, [<reflink idref="bib39" id="ref87">39</reflink>]). This design type is attractive as it will likely result in a well-linked sample which guarantees unbiased parameter and latent score estimation. However, there are other designs that might be as attractive for a rating study. For example, Fürst ([<reflink idref="bib15" id="ref88">15</reflink>]) used a design in which two raters (out of five for one of the tasks) rated all responses, whereas three other raters provided two ratings per response in a full matrix design. Of course, such designs come with their own disadvantages, namely that at least some raters are experiencing the full burden of the rating task. With that being said, we recommend researchers choosing their rating design also based on the expected workload of the single raters and take into consideration the experience of their raters. In addition, it is arguably a good idea to check a simulation even in situations that allows all raters to rate all responses, just to make sure study planning is sound. The open material we provide along with this paper can be easily extended to such other designs.</p> <p>Furthermore, the designs considered in this work can be easily extended by anticipating other missing value issues (e.g., missing values because of study drop-out of participants). For example, when assuming that a certain proportion of responses will be missing completely at random, it is possible to incorporate this in the simulation to identify a "safe" rater design. Furthermore, it is important that not all studies will focus on comparably large numbers of products to be rated. Some studies may require only very few (or at least much fewer) ratings instead. For such situations one might further consider technical issues such as problems with model convergence, for example. In such situations, the target model might not be estimable and it would be very useful to anticipate approaches to deal with such issues. For example, one could increase the number of iterations, focus on less complex models (i.e., models that require less parameters to be estimated), or use Bayesian estimation with somewhat informative priors. As a final remark, it should be noted that there could be a trade-off between the number of Likert-points used by the raters and model complexity. While more scale points provide more information and potentially increase reliability, the estimated models would incorporate more parameters (e.g., intercept parameters in the GPCM) to be estimated.</p> <hd id="AN0182340424-17">Conclusion</hd> <p>Human ratings are ubiquitous in creativity research which makes running studies a laborious endeavor. In this work, we have demonstrated how information obtained from JRT and simulations can be used for a fine-tuned planning of missing data designs that reduce the amount of work needed for reliable scoring. We have further shown how such a careful planning further translates into cost-effectiveness considerations. Hence, we anticipate that the outlined approach will be of great practical value for the field and invite interested researchers to explore and use the material we made available. This way, research money – and a lot of time – will be saved for all of us.</p> <hd id="AN0182340424-18">Disclosure statement</hd> <p>No potential conflict of interest was reported by the author(s).</p> <ref id="AN0182340424-19"> <title> References </title> <blist> <bibl id="bib1" idref="ref70" type="bt">1</bibl> <bibtext> Akaike, H. (1973). Information theory and an extension of the maximum likelihood principle. In B. N. Petrov & F. Csáki (Eds.), 2nd international symposium on information theory (pp. 267 – 281). Budapest, Hungary : Akadémiai Kiadó.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref3" type="bt">2</bibl> <bibtext> Amabile, T. M. (1982). Social psychology of creativity: A consensual assessment technique. Journal of Personality and Social Psychology, 43 (5), 997 – 1013. doi: 10.1037/0022-3514.43.5.997</bibtext> </blist> <blist> <bibl id="bib3" idref="ref53" type="bt">3</bibl> <bibtext> Barbot, B. (2020). Creativity and self‐esteem in adolescence: A study of their domain‐specific, multivariate relationships. The Journal of Creative Behavior, 54 (2), 279 – 292. doi: 10.1002/jocb.365</bibtext> </blist> <blist> <bibl id="bib4" idref="ref11" type="bt">4</bibl> <bibtext> Beaty, R. E., & Johnson, D. R. (2021). Automating creativity assessment with SemDis: An open platform for computing semantic distance. Behavior Research Methods, 53 (2), 757 – 780. doi: 10.3758/s13428-020-01453-w</bibtext> </blist> <blist> <bibl id="bib5" idref="ref7" type="bt">5</bibl> <bibtext> Beaty, R. E., & Silvia, P. (2013). Metaphorically speaking: Cognitive abilities and the production of figurative language. Memory & Cognition, 41 (2), 255 – 267. doi: 10.3758/s13421-012-0258-5</bibtext> </blist> <blist> <bibl id="bib6" idref="ref4" type="bt">6</bibl> <bibtext> Benedek, M., Mühlmann, C., Jauk, E., & Neubauer, A. C. (2013). Assessment of divergent thinking by means of the subjective top-scoring method: Effects of the number of top-ideas and time-on-task on reliability and validity. Psychology of Aesthetics, Creativity, and the Arts, 7 (4), 341 – 349. doi: 10.1037/a0033644</bibtext> </blist> <blist> <bibl id="bib7" idref="ref12" type="bt">7</bibl> <bibtext> Buczak, P., Huang, H., Forthmann, B., & Doebler, P. (2022). The machines take over: A comparison of various supervised learning approaches for automated scoring of divergent thinking tasks. The Journal of Creative Behavior, 57 (1), 17 – 36. doi: 10.1002/jocb.559</bibtext> </blist> <blist> <bibl id="bib8" idref="ref68" type="bt">8</bibl> <bibtext> Chalmers, R. P. (2012). mirt : A multidimensional item response theory package for the R environment. Journal of Statistical Software, 48 (6). doi: 10.18637/jss.v048.i06</bibtext> </blist> <blist> <bibl id="bib9" idref="ref1" type="bt">9</bibl> <bibtext> Christensen, P. R., Guilford, J. P., & Wilson, R. C. (1957). Relations of creative responses to working time and instructions. Journal of Experimental Psychology, 53 (2), 82 – 88. doi: 10.1037/h0045461</bibtext> </blist> <blist> <bibtext> Christensen, A. P., Silvia, P. J., Nusbaum, E. C., & Beaty, R. E. (2018). Clever people: Intelligence and humor production ability. Psychology of Aesthetics, Creativity, and the Arts, 12 (2), 136 – 143. doi: 10.1037/aca0000109</bibtext> </blist> <blist> <bibtext> Cicchetti, D. V. (2001). Methodological commentary the precision of reliability and validity estimates re-visited: Distinguishing between clinical and statistical significance of sample size requirements. Journal of Clinical and Experimental Neuropsychology, 23 (5), 695 – 700. doi: 10.1076/jcen.23.5.695.1249</bibtext> </blist> <blist> <bibtext> Dumas, D., Organisciak, P., & Doherty, M. (2020). Measuring divergent thinking originality with human raters and text-mining models: A psychometric comparison of methods. Psychology of Aesthetics, Creativity, and the Arts, 15 (4), 645 – 663. doi: 10.1037/aca0000319</bibtext> </blist> <blist> <bibtext> Ferrando, P. J., & Lorenzo-Seva, U. (2018). Assessing the quality and appropriateness of factor solutions and factor score estimates in exploratory item factor analysis. Educational and Psychological Measurement, 78 (5), 762 – 780. doi: 10.1177/0013164417719308</bibtext> </blist> <blist> <bibtext> Forthmann, B., Holling, H., Zandi, N., Gerwig, A., Çelik, P., Storme, M., & Lubart, T. (2017). Missing creativity: The effect of cognitive workload on rater (dis-)agreement in subjective divergent-thinking scores. Thinking Skills and Creativity, 23, 129 – 139. doi: 10.1016/j.tsc.2016.12.005</bibtext> </blist> <blist> <bibtext> Fürst, G. (2020). Measuring creativity with planned missing data. The Journal of Creative Behavior, 54 (1), 150 – 164. doi: 10.1002/jocb.352</bibtext> </blist> <blist> <bibtext> Graham, J. W. (2009). Missing data analysis: Making it work in the real world. Annual Review of Psychology, 60 (1), 549 – 576. doi: 10.1146/annurev.psych.58.110405.085530</bibtext> </blist> <blist> <bibtext> Hass, R. W., Rivera, M., & Silvia, P. J. (2018). On the dependability and feasibility of layperson ratings of divergent thinking. Frontiers in Psychology, 9, 9. doi: 10.3389/fpsyg.2018.01343</bibtext> </blist> <blist> <bibtext> Kaufman, J. C., Baer, J., Cropley, D. H., Reiter-Palmon, R., & Sinnett, S. (2013). Furious activity vs. Understanding: How much expertise is needed to evaluate creative work? Psychology of Aesthetics, Creativity, and the Arts, 7 (4), 332 – 340. doi: 10.1037/a0034809</bibtext> </blist> <blist> <bibtext> Kleinkorres, R., Forthmann, B., & Holling, H. (2021). An experimental approach to investigate the involvement of cognitive load in divergent thinking. Journal of Intelligence, 9 (1), 3. doi: 10.3390/jintelligence9010003</bibtext> </blist> <blist> <bibtext> Kornilov, S. A., Kornilova, T. V., & Grigorenko, E. L. (2016). The cross-cultural invariance of creative cognition: A case study of creative writing in U.S. and Russian college students. New Directions for Child and Adolescent Development, 2016 (151), 47 – 59. doi: 10.1002/cad.20149</bibtext> </blist> <blist> <bibtext> Long, H. (2014). More than appropriateness and novelty: Judges' criteria of assessing creative products in science tasks. Thinking Skills and Creativity, 13, 183 – 194. doi: 10.1016/j.tsc.2014.05.002</bibtext> </blist> <blist> <bibtext> Long, H., & Pang, W. (2015). Rater effects in creativity assessment: A mixed methods investigation. Thinking Skills and Creativity, 15, 13 – 25. doi: 10.1016/j.tsc.2014.10.004</bibtext> </blist> <blist> <bibtext> Matlock, K. L., Turner, R. C., & Gitchel, W. D. (2018). A study of reverse-worded matched item Pairs using the generalized partial credit and nominal response models. Educational and Psychological Measurement, 78 (1), 103 – 127. doi: 10.1177/0013164416670211</bibtext> </blist> <blist> <bibtext> Mouchiroud, C., & Lubart, T. (2001). Children's original thinking: An empirical examination of alternative measures derived from divergent thinking tasks. The Journal of Genetic Psychology, 162 (4), 382 – 401. doi: 10.1080/00221320109597491</bibtext> </blist> <blist> <bibtext> Muraki, E. (1992). A generalized partial credit model: Application of an EM algorithm. ETS Research Report Series, 1992 (1), i – 30. doi: 10.1002/j.2333-8504.1992.tb01436.x</bibtext> </blist> <blist> <bibtext> Myszkowski, N. (2021). Development of the R library "jrt": Automated item response theory procedures for judgment data and their application with the consensual assessment technique. Psychology of Aesthetics, Creativity, and the Arts, 15 (3), 426 – 438. doi: 10.1037/aca0000287</bibtext> </blist> <blist> <bibtext> Myszkowski, N., & Storme, M. (2019). Judge response theory? A call to upgrade our psychometrical account of creativity judgments. Psychology of Aesthetics, Creativity, and the Arts, 13 (2), 167 – 175. doi: 10.1037/aca0000225</bibtext> </blist> <blist> <bibtext> Nusbaum, E. C., Silvia, P. J., & Beaty, R. E. (2017). Ha ha? Assessing individual differences in humor production ability. Psychology of Aesthetics, Creativity, and the Arts, 11 (2), 231 – 241. doi: 10.1037/aca0000086</bibtext> </blist> <blist> <bibtext> Plucker, J. A., Beghetto, R. A., & Dow, G. T. (2004). Why isn't creativity more important to educational psychologists? Potentials, pitfalls, and future directions in creativity research. Educational Psychologist, 39 (2), 83 – 96. doi: 10.1207/s15326985ep3902_1</bibtext> </blist> <blist> <bibtext> Primi, R. (2014). Divergent productions of metaphors: Combining many-facet rasch measurement and cognitive psychology in the assessment of creativity. Psychology of Aesthetics, Creativity, and the Arts, 8 (4), 461 – 474. doi: 10.1037/a0038055</bibtext> </blist> <blist> <bibtext> Primi, R., Silvia, P. J., Jauk, E., & Benedek, M. (2019). Applying many-facet Rasch modeling in the assessment of creativity. Psychology of Aesthetics, Creativity, and the Arts, 13 (2), 176 – 186. doi: 10.1037/aca0000230</bibtext> </blist> <blist> <bibtext> R Core Team. (2021). R: A language and environment for statistical computing (4.1.2). Vienna, Austria : R Foundation for Statistical Computing.</bibtext> </blist> <blist> <bibtext> Reinig, B. A., & Briggs, R. O. (2013). Putting quality first in ideation research. Group Decision and Negotiation, 22 (5), 943 – 973. doi: 10.1007/s10726-012-9338-y</bibtext> </blist> <blist> <bibtext> Robitzsch, A., & Steinfeld, J. (2018). Item response models for human ratings: Overview, estimation methods, and implementation in R. Psychological Test and Assessment Modeling, 60 (1), 101 – 138.</bibtext> </blist> <blist> <bibtext> Runco, M. A., Plucker, J. A., & Lim, W. (2001). Development and psychometric integrity of a measure of ideational behavior. Creativity Research Journal, 13 (3–4), 393 – 400. doi: 10.1207/S15326934CRJ1334_16</bibtext> </blist> <blist> <bibtext> Samejima, F. (1969). Estimation of latent ability using a response pattern of graded scores. Psychometrika, 34 (S1), 1 – 97. doi: 10.1007/BF03372160</bibtext> </blist> <blist> <bibtext> Shaw, A. (2021). It works ... but can we make it easier? A comparison of three subjective scoring indexes in the assessment of divergent thinking. Thinking Skills and Creativity, 40, 100789. doi: 10.1016/j.tsc.2021.100789</bibtext> </blist> <blist> <bibtext> Signorell, A. (2021). DescTools: Tools for descriptive statistics (R package version 0.99.44).</bibtext> </blist> <blist> <bibtext> Silvia, P. J., Kwapil, T. R., Walsh, M. A., & Myin-Germeys, I. (2014). Planned missing-data designs in experience-sampling research: Monte Carlo simulations of efficient designs for assessing within-person constructs. Behavior Research Methods, 46 (1), 41 – 54. doi: 10.3758/s13428-013-0353-y</bibtext> </blist> <blist> <bibtext> Silvia, P. J., Martin, C., & Nusbaum, E. C. (2009). A snapshot of creativity: Evaluating a quick and simple method for assessing divergent thinking. Thinking Skills and Creativity, 4 (2), 79 – 85. doi: 10.1016/j.tsc.2009.06.005</bibtext> </blist> <blist> <bibtext> Silvia, P. J., Winterstein, B. P., Willse, J. T., Barona, C. M., Cram, J. T. ... Richard, C. A. (2008). Assessing creativity with divergent thinking tasks: Exploring the reliability and validity of new subjective scoring methods. Psychology of Aesthetics, Creativity, and the Arts, 2 (2), 68 – 85. doi: 10.1037/1931-3896.2.2.68</bibtext> </blist> <blist> <bibtext> Stevenson, C. E., Smal, I., Baas, M., Dahrendorf, M., Grasman, R. & van der Maas, H. (2020). Automated AUT scoring using a big data variant of the consensual assessment technique.</bibtext> </blist> <blist> <bibtext> Storme, M., Myszkowski, N., Çelik, P., & Lubart, T. (2014). Learning to judge creativity: The underlying mechanisms in creativity training for non-expert judges. Learning and Individual Differences, 32, 19 – 25. doi: 10.1016/j.lindif.2014.03.002</bibtext> </blist> <blist> <bibtext> Taylor, C. L., & Barbot, B. (2021). Dual pathways in creative writing processes. Psychology of Aesthetics, Creativity, and the Arts. doi: 10.1037/aca0000415</bibtext> </blist> <blist> <bibtext> Wagenmakers, E.-J., & Farrell, S. (2004). AIC model selection using Akaike weights. Psychonomic Bulletin & Review, 11 (1), 192 – 196. doi: 10.3758/BF03206482</bibtext> </blist> <blist> <bibtext> Wilson, R. C., Guilford, J. P., & Christensen, P. R. (1953). The measurement of individual differences in originality. Psychological Bulletin, 50 (5), 362 – 370. doi: 10.1037/h0060857</bibtext> </blist> </ref> <aug> <p>By Boris Forthmann; Benjamin Goecke and Roger E. Beaty</p> <p>Reported by Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib46" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib14" firstref="ref5"></nolink> <nolink nlid="nl3" bibid="bib41" firstref="ref6"></nolink> <nolink nlid="nl4" bibid="bib30" firstref="ref8"></nolink> <nolink nlid="nl5" bibid="bib10" firstref="ref9"></nolink> <nolink nlid="nl6" bibid="bib28" firstref="ref10"></nolink> <nolink nlid="nl7" bibid="bib12" firstref="ref13"></nolink> <nolink nlid="nl8" bibid="bib42" firstref="ref14"></nolink> <nolink nlid="nl9" bibid="bib24" firstref="ref15"></nolink> <nolink nlid="nl10" bibid="bib34" firstref="ref16"></nolink> <nolink nlid="nl11" bibid="bib37" firstref="ref19"></nolink> <nolink nlid="nl12" bibid="bib40" firstref="ref20"></nolink> <nolink nlid="nl13" bibid="bib26" firstref="ref21"></nolink> <nolink nlid="nl14" bibid="bib27" firstref="ref22"></nolink> <nolink nlid="nl15" bibid="bib29" firstref="ref24"></nolink> <nolink nlid="nl16" bibid="bib35" firstref="ref25"></nolink> <nolink nlid="nl17" bibid="bib33" firstref="ref26"></nolink> <nolink nlid="nl18" bibid="bib20" firstref="ref27"></nolink> <nolink nlid="nl19" bibid="bib44" firstref="ref28"></nolink> <nolink nlid="nl20" bibid="bib21" firstref="ref29"></nolink> <nolink nlid="nl21" bibid="bib22" firstref="ref30"></nolink> <nolink nlid="nl22" bibid="bib18" firstref="ref32"></nolink> <nolink nlid="nl23" bibid="bib17" firstref="ref33"></nolink> <nolink nlid="nl24" bibid="bib43" firstref="ref36"></nolink> <nolink nlid="nl25" bibid="bib31" firstref="ref38"></nolink> <nolink nlid="nl26" bibid="bib19" firstref="ref47"></nolink> <nolink nlid="nl27" bibid="bib16" firstref="ref52"></nolink> <nolink nlid="nl28" bibid="bib15" firstref="ref54"></nolink> <nolink nlid="nl29" bibid="bib36" firstref="ref61"></nolink> <nolink nlid="nl30" bibid="bib25" firstref="ref62"></nolink> <nolink nlid="nl31" bibid="bib13" firstref="ref71"></nolink> <nolink nlid="nl32" bibid="bib11" firstref="ref74"></nolink> <nolink nlid="nl33" bibid="bib32" firstref="ref76"></nolink> <nolink nlid="nl34" bibid="bib45" firstref="ref77"></nolink> <nolink nlid="nl35" bibid="bib23" firstref="ref81"></nolink> <nolink nlid="nl36" bibid="bib39" firstref="ref82"></nolink> <nolink nlid="nl37" bibid="bib38" firstref="ref83"></nolink>
Header DbId: eric
DbLabel: ERIC
An: EJ1458414
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Planning Missing Data Designs for Human Ratings in Creativity Research: A Practical Guide
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Boris+Forthmann%22">Boris Forthmann</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-9755-7304">0000-0001-9755-7304</externalLink>)<br /><searchLink fieldCode="AR" term="%22Benjamin+Goecke%22">Benjamin Goecke</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3050-1848">0000-0002-3050-1848</externalLink>)<br /><searchLink fieldCode="AR" term="%22Roger+E%2E+Beaty%22">Roger E. Beaty</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-6114-5973">0000-0001-6114-5973</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Creativity+Research+Journal%22"><i>Creativity Research Journal</i></searchLink>. 2025 37(1):167-178.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 12
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)<br />National Science Foundation (NSF), Division of Undergraduate Education (DUE)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: 1920653<br />2155070
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Creativity%22">Creativity</searchLink><br /><searchLink fieldCode="DE" term="%22Research%22">Research</searchLink><br /><searchLink fieldCode="DE" term="%22Researchers%22">Researchers</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Methodology%22">Research Methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Psychometrics%22">Psychometrics</searchLink><br /><searchLink fieldCode="DE" term="%22Simulation%22">Simulation</searchLink><br /><searchLink fieldCode="DE" term="%22Measurement+Techniques%22">Measurement Techniques</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Response+Theory%22">Item Response Theory</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluators%22">Evaluators</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Matrices%22">Matrices</searchLink><br /><searchLink fieldCode="DE" term="%22Data%22">Data</searchLink><br /><searchLink fieldCode="DE" term="%22Computation%22">Computation</searchLink><br /><searchLink fieldCode="DE" term="%22Cost+Effectiveness%22">Cost Effectiveness</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1080/10400419.2023.2250976
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1040-0419<br />1532-6934
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Human ratings are ubiquitous in creativity research. Yet, the process of rating responses to creativity tasks -- typically several hundred or thousands of responses, per rater -- is often time-consuming and expensive. Planned missing data designs, where raters only rate a subset of the total number of responses, have been recently proposed as one possible solution to decrease overall rating time and monetary costs. However, researchers also need ratings that adhere to psychometric standards, such as a certain degree of reliability, and psychometric work with planned missing designs is currently lacking in the literature. In this work, we introduce how judge response theory and simulations can be used to fine-tune planning of missing data designs. We provide open code for the community and illustrate our proposed approach by a cost-effectiveness calculation based on a realistic example. We clearly show that fine-tuning helps to save time (to perform the ratings) and monetary costs, while simultaneously targeting expected levels of reliability.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1458414
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1458414
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1080/10400419.2023.2250976
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 12
        StartPage: 167
    Subjects:
      – SubjectFull: Creativity
        Type: general
      – SubjectFull: Research
        Type: general
      – SubjectFull: Researchers
        Type: general
      – SubjectFull: Research Methodology
        Type: general
      – SubjectFull: Psychometrics
        Type: general
      – SubjectFull: Simulation
        Type: general
      – SubjectFull: Measurement Techniques
        Type: general
      – SubjectFull: Item Response Theory
        Type: general
      – SubjectFull: Evaluators
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Matrices
        Type: general
      – SubjectFull: Data
        Type: general
      – SubjectFull: Computation
        Type: general
      – SubjectFull: Cost Effectiveness
        Type: general
    Titles:
      – TitleFull: Planning Missing Data Designs for Human Ratings in Creativity Research: A Practical Guide
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Boris Forthmann
      – PersonEntity:
          Name:
            NameFull: Benjamin Goecke
      – PersonEntity:
          Name:
            NameFull: Roger E. Beaty
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-print
              Value: 1040-0419
            – Type: issn-electronic
              Value: 1532-6934
          Numbering:
            – Type: volume
              Value: 37
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Creativity Research Journal
              Type: main
ResultId 1