Revisiting the Usage of Alpha in Scale Evaluation: Effects of Scale Length and Sample Size
Saved in:
| Title: | Revisiting the Usage of Alpha in Scale Evaluation: Effects of Scale Length and Sample Size |
|---|---|
| Language: | English |
| Authors: | Leifeng Xiao (ORCID |
| Source: | Educational Measurement: Issues and Practice. 2024 43(2):74-81. |
| Availability: | Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us |
| Peer Reviewed: | Y |
| Page Count: | 8 |
| Publication Date: | 2024 |
| Document Type: | Journal Articles Reports - Research |
| Descriptors: | Measurement, Benchmarking, Item Sampling, Sample Size, Research Methodology, Guidelines, Test Construction |
| DOI: | 10.1111/emip.12604 |
| ISSN: | 0731-1745 1745-3992 |
| Abstract: | Short scales are time-efficient for participants and cost-effective in research. However, researchers often mistakenly expect short scales to have the same reliability as long ones without considering the effect of scale length. We argue that applying a universal benchmark for alpha is problematic as the impact of low-quality items is greater on shorter scales. In this study, we proposed simple guidelines for item reduction using the "alpha-if-item-deleted" procedure in scale construction. An item can be removed if alpha increases or decreases by less than 0.02, especially for short scales. Conversely, an item should be retained if alpha decreases by more than 0.04 upon its removal. For reliability benchmarks, 0.80 is relatively safe in most conditions, but higher benchmarks are recommended for longer scales and smaller sample sizes. Supplementary analyses, including item content, face validity, and content coverage, are critical to ensure scale quality. |
| Abstractor: | As Provided |
| Entry Date: | 2024 |
| Accession Number: | EJ1425083 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHerB91Ng78IFOIi2yVWtjwAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDElj1IX1cRFm2AtmgwIBEICBm4G9Y4VqbmucE5EVUmwEM7P9wrwL3erMqc5OdyX5VLRjVlxvzbQ_zfkNAcWGKDD8amS2NZUVnWmsoGczKIFZfd99XFv5QklMp1k9r_fwxd3M-bj1wX-YDeS7JqoIcANpTwRWJs61IN5EGJD5FpT8tByrlWJMAC9AFkM402ufTGW7ZGZKFb2VKEyTL_8aUrLVwDBhGoZJaxcX-XKr Text: Availability: 1 Value: <anid>AN0177337607;ems01jun.24;2024May22.04:42;v2.2.500</anid> <title id="AN0177337607-1">Revisiting the Usage of Alpha in Scale Evaluation: Effects of Scale Length and Sample Size </title> <p>Short scales are time‐efficient for participants and cost‐effective in research. However, researchers often mistakenly expect short scales to have the same reliability as long ones without considering the effect of scale length. We argue that applying a universal benchmark for alpha is problematic as the impact of low‐quality items is greater on shorter scales. In this study, we proposed simple guidelines for item reduction using the "alpha‐if‐item‐deleted" procedure in scale construction. An item can be removed if alpha increases or decreases by less than.02, especially for short scales. Conversely, an item should be retained if alpha decreases by more than.04 upon its removal. For reliability benchmarks,.80 is relatively safe in most conditions, but higher benchmarks are recommended for longer scales and smaller sample sizes. Supplementary analyses, including item content, face validity, and content coverage, are critical to ensure scale quality.</p> <p>Keywords: coefficient alpha; cutoff value; reliability; scale evaluation; scale length</p> <p>In behavioral and social science research, there is an increasing recognition of the benefits of using short scales to incorporate more dimensions while reducing administration time, thereby enhancing participant engagement (Schipolowski et al., [<reflink idref="bib32" id="ref1">32</reflink>]). The Organisation for Economic Cooperation and Development's (OECD) Programme for International Student Assessment (PISA) exemplifies this approach by including numerous short scales to assess 15‐year‐old students' abilities in reading, mathematics, and science (OECD, [<reflink idref="bib26" id="ref2">26</reflink>]). This widespread use of short scales, however, raises questions about the minimum number of items necessary for each scale and concerns over their seemingly lower reliability compared to longer scales (Ziegler et al., [<reflink idref="bib34" id="ref3">34</reflink>]). This paper addressed and clarified the misconceptions regarding the reliability differences between short and long scales and offered practical advice on comparing reliability across scale lengths and employing the alpha‐if‐item‐deleted criterion for scale revision.</p> <hd id="AN0177337607-2">Reliability and Coefficient Alpha</hd> <p>Reliability is a fundamental quality indicator of a scale. Defined by Cronbach ([<reflink idref="bib8" id="ref4">8</reflink>]) as the accuracy or dependability of measurements, it aims to ensure that all items consistently measure the same construct. Within the classical test theory framework, an observed score (<emph>X</emph>) comprises a true score (<emph>T</emph>) and measurement error (<emph>E</emph>), with <emph>X</emph> = <emph>T</emph> + <emph>E</emph>. Reliability is defined as the square of the correlation between <emph>X</emph> and <emph>T</emph> ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0001" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi mathvariant="normal"&gt;i&lt;/mi&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mi mathvariant="normal"&gt;e&lt;/mi&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;mspace width="0.33em" /&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#961;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;mrow&gt;&lt;mo&gt;)&lt;/mo&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${\mathrm{i}}.{\mathrm{e}}.,\ \rho &amp;#95;{TX}^2)$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ; operationally, it is defined as the ratio of the true score variance to the observed score variance ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0002" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$\sigma &amp;#95;T^2/\sigma &amp;#95;X^2$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ; Lord &amp; Novick, [<reflink idref="bib19" id="ref5">19</reflink>]). 1 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0003" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;Reliability&lt;/mi&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#961;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;/mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;mo linebreak="goodbreak"&gt;=&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mi&gt;T&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;msubsup&gt;&lt;mi&gt;&amp;#963;&lt;/mi&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}{\mathrm{Reliability}} = \rho &amp;#95;{TX}^2 = \sigma &amp;#95;T^2/\sigma &amp;#95;X^2.\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>Coefficient alpha, widely used and researched for over half a century, quantifies the proportion of random measurement error in the sum or average of a set of scale items (Cronbach, [<reflink idref="bib8" id="ref6">8</reflink>]; McNeish, [<reflink idref="bib20" id="ref7">20</reflink>]). Represented as 2 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0004" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mfrac&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;/mfenced&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;msubsup&gt;&lt;mo&gt;&amp;#8721;&lt;/mo&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msubsup&gt;&lt;msubsup&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;/mrow&gt;&lt;msubsup&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;,&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}\alpha = \left({\frac{k}{{k - 1}}} \right)\left({1 - \frac{{\sum\nolimits&amp;#95;i^k {S&amp;#95;i^2} }}{{S&amp;#95;X^2}}} \right),\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> where <emph>k</emph> is the number of test items, <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0005" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msubsup&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;annotation encoding="application/x-tex"&gt;$S&amp;#95;i^2$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the population variance of the <emph>i</emph>th item, and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0006" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msubsup&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mi&gt;x&lt;/mi&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/msubsup&gt;&lt;annotation encoding="application/x-tex"&gt;$S&amp;#95;x^2$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is the total variance of the test. Coefficient alpha generally ranges between 0 and 1, though it can be negative when item correlations are negative.</p> <hd id="AN0177337607-3">Scale Length and Coefficient Alpha</hd> <p>While alpha mathematically increases with scale length, several unresolved issues remain (Xiao &amp; Hau, [<reflink idref="bib33" id="ref8">33</reflink>]; Davenport et al., [<reflink idref="bib9" id="ref9">9</reflink>]; Edwards et al., [<reflink idref="bib10" id="ref10">10</reflink>]; Green &amp; Yang, [<reflink idref="bib13" id="ref11">13</reflink>]; Hoekstra et al., [<reflink idref="bib17" id="ref12">17</reflink>]). Empirical studies, such as meta‐analyses, have shown a relatively modest relationship (<emph>R</emph><sups>2</sups>) between reliability and scale length, suggesting interactive effects with other factors like scale strength (Churchill &amp; Peter, [<reflink idref="bib5" id="ref13">5</reflink>]; Peterson, [<reflink idref="bib28" id="ref14">28</reflink>]). This paper presented mathematical derivations and empirical examples to illustrate these interactions.</p> <p>Furthermore, the prevalent practice of adhering to an absolute reliability cutoff (e.g.,.7 or.8) without considering scale length may not be entirely appropriate (Nunnally, [<reflink idref="bib23" id="ref15">23</reflink>];Marsh et al.,[<reflink idref="bib21" id="ref16">21</reflink>]). This suggests the need to reconsider using the "alpha‐if‐item‐deleted" strategy for item retention or removal for scales with different scale lengths. Additionally, while sometimes necessary, item deletion can compromise scale quality, including reliability, content validity, and construct validity (Raykov, [<reflink idref="bib31" id="ref17">31</reflink>]). We aimed to balance brevity, reliability, and validity in scale construction.</p> <p>The paper is structured into three sections: (a) exploring factors that might affect the impact of scale length, (b) strategies for item retention or removal in scales of varying lengths, and (c) developing dynamic reliability benchmarks suitable for different scale lengths and qualities. These sections are interspersed with practical demonstrations and examples. We conclude the paper with practical advice for applied researchers and identify areas for future research.</p> <hd id="AN0177337607-4">Factors Potentially Affecting Scale Length Effects</hd> <p>In this section, we argued mathematically and demonstrated empirically how scale strength and sample size might interact with scale length in affecting reliability and precision.</p> <hd id="AN0177337607-5">Scale Length and Scale Strength</hd> <p>According to an alternative formula for alpha based on the average interitem correlation (Cronbach, [<reflink idref="bib8" id="ref18">8</reflink>]; Davenport et al., [<reflink idref="bib9" id="ref19">9</reflink>]), alpha is a function of scale length for any given average correlation: 3 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0007" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;mo linebreak="badbreak"&gt;=&lt;/mo&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mspace width="0.33em" /&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mspace width="0.33em" /&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;+&lt;/mo&gt;&lt;mspace width="0.33em" /&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mspace width="0.33em" /&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfrac&gt;&lt;mspace width="0.33em" /&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}\alpha \ = \frac{{\ k\ {{{\bar{r}}}&amp;#95;{{{X}&amp;#95;i}{{X}&amp;#95;j}}}\ }}{{1 + \ 1\left({k - 1} \right)\ {{{\bar{r}}}&amp;#95;{{{X}&amp;#95;i}{{X}&amp;#95;j}}}}}\.\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>Here, <emph>k</emph> is the number of items in a scale and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0008" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mo&gt;&amp;#175;&lt;/mo&gt;&lt;/mover&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\bar{r}}&amp;#95;{{{X}&amp;#95;i}{{X}&amp;#95;j}}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> denotes the average of all unique pairwise correlations of the <emph>i</emph>th item score ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0009" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{X}&amp;#95;i}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ) and the <emph>j</emph>th item score ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0010" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{X}&amp;#95;j}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ). Alpha increases with a greater number of items or stronger interitem correlations. Here higher interitem correlations indicate stronger connections between items, reflecting stronger scale strength. Equations (<reflink idref="bib2" id="ref20">2</reflink>) and (<reflink idref="bib3" id="ref21">3</reflink>) differ in that they are based on covariance and correlation matrices, respectively, which are identical when item scores are standardized. In a one‐factor Confirmatory Factor Analysis (CFA) model, the item correlations are products of their respective standardized factor loadings ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0011" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mspace width="0.33em" /&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{r}&amp;#95;{{{X}&amp;#95;i}{{X}&amp;#95;j}}} = {{\lambda }&amp;#95;{i\ }}\ \times {{\lambda }&amp;#95;j}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ), hence, factor loading is also a reflection of scale strength. Therefore, alpha could be calculated for any combination of item factor loadings.</p> <p>It is observed in various studies that increasing scale length did not automatically result in significant reliability improvement, suggesting potential interactions with scale strength (e.g., Xiao &amp; Hau,[<reflink idref="bib33" id="ref22">33</reflink>]). For an essential tau equivalent scale, the loadings of items are equal, the correlations among the items are relatively similar, and the items' variances are equal. We calculated and tabulated the alpha of scales varying in scale strength (loading ranged from.3 to.8) and the number of items (2 to 30 items) (details in Table S1 and Figure S1) for such conditions. Several trends were noted: (a) the relationship between scale length and alpha was not linear, with diminishing increase in alpha as scales lengthen; (b) alpha plateaued when the number of items approached 20; (c) for scales with strong loadings, the incremental increase in alpha per added item was minimal. Therefore, simply adding items did not guarantee increased reliability, especially for strong and lengthy scales. For example, an 8‐item scale with.6 loading (alpha =.818) was less reliable than a 5‐item scale with.7 loading (alpha =.828). That is, a scale with more items did not necessarily mean it was more reliable, and the scale strength also mattered in such a comparison.</p> <hd id="AN0177337607-6">Scale Length and Sample Size</hd> <p>Sample size has been an important concern in empirical work because insufficient statistical power is a persistent issue in behavioral sciences (Cohen, [<reflink idref="bib7" id="ref23">7</reflink>]). Hence, a sampling error of unknown magnitude and direction is expected in a typical calculation of alpha estimated with a random sample.</p> <p>For the precision of reliability estimation (confidence interval, CI), Bonett ([<reflink idref="bib3" id="ref24">3</reflink>]) has affirmed that both sample size and scale length jointly affected the precision of the reliability estimation, estimated as follows: 4 <ephtml> &lt;math display="block" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0012" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo linebreak="badbreak"&gt;&amp;#8722;&lt;/mo&gt;&lt;mspace width="0.33em" /&gt;&lt;mi&gt;exp&lt;/mi&gt;&lt;mfenced separators="" open="[" close="]"&gt;&lt;mrow&gt;&lt;mi&gt;In&lt;/mi&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#961;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;&amp;#177;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;Z&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;msup&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mfenced separators="" open="{" close="}"&gt;&lt;mrow&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mfenced separators="" open="(" close=")"&gt;&lt;mrow&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mrow&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msup&gt;&lt;/mrow&gt;&lt;/mfenced&gt;&lt;mo&gt;.&lt;/mo&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;$$\begin{equation}1 - \ {\mathrm{exp}}\left[ {{\mathrm{In}}\left({1 - {{{\hat{\rho }}}&amp;#95;k}} \right) \pm {{Z}&amp;#95;{\alpha /2}}{{{\left({2k/\left\{ {\left({k - 1} \right)\left({n - 2} \right)} \right\}} \right)}}^{1/2}}} \right].\end{equation}$$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml></p> <p>Here, <emph>k</emph> is the number of items, and <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0013" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mover accent="true"&gt;&lt;mi&gt;&amp;#961;&lt;/mi&gt;&lt;mo&gt;&amp;#770;&lt;/mo&gt;&lt;/mover&gt;&lt;mi&gt;k&lt;/mi&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{\hat{\rho }}&amp;#95;k}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> denotes coefficient alpha based on a scale having <emph>k</emph> items, where <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0014" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;Z&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{Z}&amp;#95;{\alpha /2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> is a point on the standard normal distribution with probability <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0015" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;msub&gt;&lt;mi&gt;Z&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;&amp;#945;&lt;/mi&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;mn&gt;2&lt;/mn&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;annotation encoding="application/x-tex"&gt;${{Z}&amp;#95;{\alpha /2}}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> .</p> <hd id="AN0177337607-7">Study 1: Precision of Reliability Alpha with Scale Length and Sample Size</hd> <p>Due to the complex relationship between scale length and sample size, we further investigated how the precision of alpha estimates varied with changes in scale length and sample size in a few actual conditions. As alpha is also affected by the scale strength, we set different levels of reliability to examine how the precision of alpha estimates changed.</p> <hd id="AN0177337607-8">Design</hd> <p>Using Equation (<reflink idref="bib4" id="ref25">4</reflink>), we computed the CI width for varying scale lengths (from 3 to 30 items), sample size (<emph>N</emph> = 50, 100, 200, 500, 1,000, and 2,000), and reliabilities (.70,.75,.80,.85, and.90) at.05 significant level Type I error. Consequently, the CI width is illustrated in Figure 1.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/EMS/01jun24/emip12604-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="emip12604-fig-0001.jpg" title="1 Confidence Intervals Width for Varying Scale Lengths, Sample Sizes, and Reliabilities" /> </p> <p></p> <hd id="AN0177337607-10">Results</hd> <p>The results showed that CI width was substantially larger for shorter scales or smaller sample sizes, specifically for scales with 3 to 6 items or <emph>N</emph><ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0016" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mo&gt;&amp;#60;&lt;/mo&gt;&lt;annotation encoding="application/x-tex"&gt;$ &lt; $&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> 500 (see Figure 1). However, as the scale length or sample size increased, the CI width decreased, suggesting greater precision in measurement for longer scales or larger sample sizes. This finding underscored the importance of using longer scales or larger sample sizes when precision was a research priority. Notably, when the sample size reached 500, the CI width reduced to less than.10 in most cases, even for short scales, implying minimal sampling errors and a deviation estimated smaller than.05 (half of the CI width). Thus, for short scales, employing larger sample sizes (e.g., <emph>N</emph> ≥ 500) could offset the accuracy loss due to scale brevity.</p> <hd id="AN0177337607-11">Item to Retain or Remove? Comparison Between Long and Short Scales</hd> <p>When constructing a test form using a set of items, some items may perform better than others, prompting researchers to select and retain the most effective ones while discarding those of lower quality. This process is nuanced, as removing a poorly performing item is expected to improve alpha; but simultaneously, alpha tends to decrease when the scale length is reduced. This section examined the combined effects of retaining or removing items of varying quality and the implications of changing scale length.</p> <hd id="AN0177337607-12">Retain or Remove Items in Scale Construction</hd> <p>Traditionally, researchers have adhered to a universal alpha benchmark, such as.70, removing items until this benchmark is met. However, expecting short scales to conform to the same benchmark as longer ones was a misconception (Ziegler et al., [<reflink idref="bib34" id="ref26">34</reflink>]). This section aimed to develop effective item selection strategies and appropriate criteria for comparing the reliability between long and short scales during the instrument construction process.</p> <p>The "alpha‐if‐item‐deleted" statistic is commonly used in this process. Through an iterative procedure, the item that, when removed, increases the alpha value most is identified and eliminated. This process continues until the scale meets the targeted alpha value (e.g.,.7). However, this decision‐making process is complex. Setting a universal alpha target for both long and short scales is inappropriate, as shown in our previous discussion. Moreover, a shorter scale tends to have a lower alpha following the removal of an item, and the impact of item removal on alpha varies significantly between long and short scales. While sophisticated methods like simulated annealing algorithms (Pereira et al., [<reflink idref="bib27" id="ref27">27</reflink>]) or HA algorithms (Hayes &amp; Coutts, [<reflink idref="bib15" id="ref28">15</reflink>]) have been proposed, this study explored more straightforward and broadly applicable guidelines through empirical analysis and examples.</p> <hd id="AN0177337607-13">Study 2: Decision Rules in Item Reduction: Strategy and Examples</hd> <p></p> <hd id="AN0177337607-14">Method: Design to develop item reduction rules</hd> <p>The decision to drop a poor item may not always lead to an increase in alpha, especially in shorter scales with inherently lower reliability. This section employed various scale quality conditions to establish benchmarks for alpha changes, useful for applied researchers working with shorter scales. We used high‐quality and average‐quality scales with average interitem correlations of.5 and.3, respectively. Subsequently, we calculated the number of items needed for a.1 alpha change at the shorter end of these scales using Equation (<reflink idref="bib3" id="ref29">3</reflink>). Since the effect of the number of items on alpha diminished with longer scales, a.1 alpha change in shorter scales required fewer additional items. Meanwhile, in standard statistical analyses, most results were in the middle and extremes were rare; we assumed that the results from extreme conditions could be applied as a relatively safety value to make a decision. Therefore, our results with the shortest end could be safely applied to the most common situations. Furthermore, we calculated the average reduction in alpha with one item removal, which could be applied as preliminary item reduction rules.</p> <hd id="AN0177337607-15">Results: Preliminary guidelines for item removal or retention</hd> <p>For high‐quality scales (average interitem correlation of.5), a.1 increase in alpha correlates with the addition of approximately five items (e.g., 4 items, alpha =.80; 9 items, alpha =.90), suggesting that for short versus long scales, the removal of an item corresponded to a change of approximately.02 in alpha (i.e., 1/5 of.1 change in alpha =.02). The implication is that if removing an item results in less than a.02 decrease in alpha, it is likely not a high‐quality item and should be removed. Thus, removing an item leading to an increase in alpha or a decrease of.02 or less should be considered desirable.</p> <p>Similarly, for average‐quality scales (.3 interitem correlation), the.1 alpha change is associated with approximately two or three items (<emph>k</emph> = 4, alpha =.632; <emph>k</emph> = 6, alpha =.720; <emph>k</emph> = 7, alpha =.750). That is, one item added would increase alpha by around.04 [i.e., (.720–.632)/(6–2) =.044; (.750–.632) / (7–4) =.04]. Therefore, an item should be retained if its removal would lead to an alpha drop of.04 or more.</p> <p>Combining the above two guidelines as extreme conditions in scale refinement, we concluded that an item causing more than a.04 alpha drop upon removal should be kept. In contrast, one leading to an alpha increase or a drop of less than.02 can be safely removed.</p> <p>Furthermore, we extended the application of our previously discussed rule through various practical scenarios to see whether our guidelines are practical in real‐world empirical applications. First, we examined the effectiveness of our strategy on several psychological and educational scales that possess both long and short versions. Specifically, we chose long and short versions Achievement Emotions Questionnaire, comprising a total of 24 subscales, which examined which examined eight kinds of achievement emotions from three educational settings (Bieleke et al., [<reflink idref="bib2" id="ref30">2</reflink>]). By incorporating the comprehensive scales, we were able to enhance the effectiveness of evaluating our preliminary guidelines. Second, we applied our rules to the large‐scale 2012 Programme for International Student Assessment (PISA), conducted by the Organisation for Economic Cooperation and Development (OECD, [<reflink idref="bib25" id="ref31">25</reflink>]). In this application, we intentionally incorporated contaminated items — items originating from a different scale — into the existing targeted scales. This methodology allowed us to assess the efficacy of our broad guidelines in identifying these mismatched items.</p> <hd id="AN0177337607-16">1 Example</hd> <hd1 id="AN0177337607-17">An application with Achievement Emotions Questionnaire scales</hd1> <p>Using data from the long and short versions of the Achievement Emotions Questionnaire by Bieleke et al. ([<reflink idref="bib2" id="ref32">2</reflink>]), we calculated the average change in alpha per removed item (see Table 1). The long scale consistently showed high reliability across different contexts. However, the average alpha change per item removed was less than.02, aligning with our preliminary criteria for item removal. This supported the development of a short version of the instrument.</p> <p>1 Table Comparisons of Reliability Alpha between Long and Short Versions of the Achievement Emotions Questionnaire (Bieleke et al., 2021)</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th&gt;Long&lt;/th&gt;&lt;th&gt;Short&lt;/th&gt;&lt;th /&gt;&lt;th /&gt;&lt;th /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th&gt;Scales&lt;/th&gt;&lt;th&gt;Items&lt;/th&gt;&lt;th&gt;Alpha&lt;/th&gt;&lt;th&gt;Items&lt;/th&gt;&lt;th&gt;Alpha&lt;/th&gt;&lt;th&gt;&amp;#916;Items&lt;/th&gt;&lt;th&gt;&amp;#916;Alpha&lt;/th&gt;&lt;th&gt;&amp;#916;Alpha/&amp;#916;Items&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Class Setting&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Enjoyment&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.75&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.10&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hope&lt;/td&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;.79&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.73&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.06&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Pride&lt;/td&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;.81&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.70&lt;/td&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;.11&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anger&lt;/td&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;.86&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.78&lt;/td&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;.08&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anxiety&lt;/td&gt;&lt;td&gt;12&lt;/td&gt;&lt;td&gt;.86&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.71&lt;/td&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;.15&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Shame&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.89&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.83&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.06&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hopelessness&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.81&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.09&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Boredom&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.93&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.88&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.05&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Learning Setting&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Enjoyment&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.78&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.64&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.14&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hope&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.77&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.76&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Pride&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.75&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.70&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;.05&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anger&lt;/td&gt;&lt;td&gt;9&lt;/td&gt;&lt;td&gt;.86&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.77&lt;/td&gt;&lt;td&gt;5&lt;/td&gt;&lt;td&gt;.09&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anxiety&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.84&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.69&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.15&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Shame&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.86&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.73&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.13&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hopelessness&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.79&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.11&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Boredom&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.91&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.06&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Test Setting&lt;/td&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;td /&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Enjoyment&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.78&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.66&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.12&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hope&lt;/td&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;.81&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.77&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.04&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Pride&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.86&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.74&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.12&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Relief&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.77&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.70&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;td&gt;.07&lt;/td&gt;&lt;td&gt;.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anger&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.86&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.74&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.12&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Anxiety&lt;/td&gt;&lt;td&gt;12&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.77&lt;/td&gt;&lt;td&gt;8&lt;/td&gt;&lt;td&gt;.13&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Shame&lt;/td&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;.87&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.83&lt;/td&gt;&lt;td&gt;6&lt;/td&gt;&lt;td&gt;.04&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Hopelessness&lt;/td&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;.92&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;7&lt;/td&gt;&lt;td&gt;.07&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Note</emph>. ΔItems and Δalpha are the differences in number of items and alphas, respectively, between long and short versions.</p> <hd id="AN0177337607-18">2 Example</hd> <hd1 id="AN0177337607-19">An application with large‐scale survey data</hd1> <p> <bold>Forming contaminated scales</bold>. The "Mathematics Work Ethic (MWE)" measure was chosen as the targeted scale, with two levels of length (9‐item: 6 from the target scale, 3 from a contaminated scale; 6‐item: 4 from the target scale, 2 from a contaminated scale). The contaminated items were of mild, moderate, and severe degrees of contamination by using items having high, moderate, and low correlations (hence, mild, moderate, and high degrees of mismatch/contamination) with the targeted scale (see Tables 2 and 3). They came from the "Mathematics Self‐Concept (MSC)," "Mathematics Self‐Efficacy (MSE)," and "Attributions to Failure (ATF)" scales. We examined the changes in reliability alpha and the usefulness of the strategies/guidelines developed from the above studies.</p> <p>2 Table Correlations and "Alpha‐if‐Item‐Deleted" for Items in Mild, Moderate, and Severe Contamination Scales (Six Items)</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th&gt;Mild Contamination (MWE + MSC)&lt;/th&gt;&lt;th align="center"&gt;Moderate Contamination (MWE + MSE)&lt;/th&gt;&lt;th align="center"&gt;Severe Contamination (MWE + ATF)&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th&gt;Items&lt;/th&gt;&lt;th&gt;Correlation&lt;/th&gt;&lt;th&gt;&amp;#916;Alpha&lt;/th&gt;&lt;th&gt;Correlation&lt;/th&gt;&lt;th align="center"&gt;&amp;#916;Alpha&lt;/th&gt;&lt;th&gt;Correlation&lt;/th&gt;&lt;th align="center"&gt;&amp;#916;Alpha&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Targeted item 1&lt;/td&gt;&lt;td&gt;.58&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;.57&lt;/td&gt;&lt;td&gt;.06&lt;/td&gt;&lt;td&gt;.43&lt;/td&gt;&lt;td&gt;.10&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 2&lt;/td&gt;&lt;td&gt;.63&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.62&lt;/td&gt;&lt;td&gt;.08&lt;/td&gt;&lt;td&gt;.50&lt;/td&gt;&lt;td&gt;.12&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 3&lt;/td&gt;&lt;td&gt;.70&lt;/td&gt;&lt;td&gt;.05&lt;/td&gt;&lt;td&gt;.57&lt;/td&gt;&lt;td&gt;.06&lt;/td&gt;&lt;td&gt;.42&lt;/td&gt;&lt;td&gt;.10&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 4&lt;/td&gt;&lt;td&gt;.59&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.58&lt;/td&gt;&lt;td&gt;.06&lt;/td&gt;&lt;td&gt;.47&lt;/td&gt;&lt;td&gt;.11&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mismatched item 1&lt;/td&gt;&lt;td&gt;.59&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.29&lt;/td&gt;&lt;td&gt;&amp;#8722;.02&lt;/td&gt;&lt;td&gt;.16&lt;/td&gt;&lt;td&gt;&amp;#8722;.02&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mismatched item 2&lt;/td&gt;&lt;td&gt;.50&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;td&gt;.29&lt;/td&gt;&lt;td&gt;&amp;#8722;.01&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;&amp;#8722;.09&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Initial alpha&lt;/td&gt;&lt;td&gt;.83&lt;/td&gt;&lt;td&gt;.75&lt;/td&gt;&lt;td&gt;.57&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>2 <emph>Note</emph>. Correlation: item‐total correlation; Δalpha: reduction of alpha after item removal; negative values indicate improvement of alpha after item removal.</item> <item>3 Targeted scale: MWE = Mathematics Work Ethic (targeted scale).</item> <item>4 Contaminated scales: MSC = Mathematics Self‐Concept; MSE = Mathematics Self‐Efficacy; ATF = Attributions to Failure.</item> <item>3 Table Correlations and "Alpha‐if‐Item‐Deleted" for Items in Mild, Moderate, and Severe Contamination Scales (Nine Items)</item> </ulist> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th /&gt;&lt;th&gt;Mild Contamination (MWE + MSC)&lt;/th&gt;&lt;th&gt;Moderate Contamination (MWE + MSE)&lt;/th&gt;&lt;th align="center"&gt;Severe Contamination (MWE + ATF)&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th&gt;Items&lt;/th&gt;&lt;th&gt;Correlation&lt;/th&gt;&lt;th&gt;&amp;#916;Alpha&lt;/th&gt;&lt;th&gt;Correlation&lt;/th&gt;&lt;th&gt;&amp;#916;Alpha&lt;/th&gt;&lt;th&gt;Correlation&lt;/th&gt;&lt;th align="center"&gt;&amp;#916;Alpha&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Targeted item 1&lt;/td&gt;&lt;td&gt;.57&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;td&gt;.57&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.47&lt;/td&gt;&lt;td&gt;.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 2&lt;/td&gt;&lt;td&gt;.64&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;.63&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.56&lt;/td&gt;&lt;td&gt;.09&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 3&lt;/td&gt;&lt;td&gt;.70&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;.60&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.46&lt;/td&gt;&lt;td&gt;.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 4&lt;/td&gt;&lt;td&gt;.60&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;td&gt;.60&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.52&lt;/td&gt;&lt;td&gt;.08&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 5&lt;/td&gt;&lt;td&gt;.63&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;.59&lt;/td&gt;&lt;td&gt;.03&lt;/td&gt;&lt;td&gt;.48&lt;/td&gt;&lt;td&gt;.07&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Targeted item 6&lt;/td&gt;&lt;td&gt;.49&lt;/td&gt;&lt;td&gt;.00&lt;/td&gt;&lt;td&gt;.49&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;.37&lt;/td&gt;&lt;td&gt;.04&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mismatched item 1&lt;/td&gt;&lt;td&gt;.66&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;.34&lt;/td&gt;&lt;td&gt;.00&lt;/td&gt;&lt;td&gt;.14&lt;/td&gt;&lt;td&gt;&amp;#8722;.01&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mismatched item 2&lt;/td&gt;&lt;td&gt;.57&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;td&gt;.36&lt;/td&gt;&lt;td&gt;.00&lt;/td&gt;&lt;td&gt;.02&lt;/td&gt;&lt;td&gt;&amp;#8722;.05&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mismatched item 3&lt;/td&gt;&lt;td&gt;.59&lt;/td&gt;&lt;td&gt;.01&lt;/td&gt;&lt;td&gt;.39&lt;/td&gt;&lt;td&gt;.00&lt;/td&gt;&lt;td&gt;.00&lt;/td&gt;&lt;td&gt;&amp;#8722;.06&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Initial alpha&lt;/td&gt;&lt;td&gt;.87&lt;/td&gt;&lt;td&gt;.81&lt;/td&gt;&lt;td&gt;.62&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <ulist> <item>5 <emph>Note</emph>. Correlation: item‐total correlation; Δalpha: reduction of alpha after item removal; negative values indicate improvement of alpha after item removal.</item> <item>6 Targeted scale: MWE = Mathematics Work Ethic (targeted scale).</item> <item>7 Contaminated scales: MSC = Mathematics Self‐Concept; MSE = Mathematics Self‐Efficacy; ATF = Attributions to Failure.</item> </ulist> <p> <bold>Results</bold>. The alpha values of the initial combined scales, along with the "alpha‐if‐item‐deleted" metric following the removal of each item in these contaminated scales, were analyzed (see Tables 2 and 3). The findings underscored the practicality and effectiveness of the guidelines we established.</p> <p>First, in most contaminated scales, the erroneous removal of targeted items consistently reduced alpha. Conversely, the correct elimination of mismatched items led to varied outcomes: an increase in alpha, no change, or a decrease, contingent upon the extent of contamination.</p> <p>Second, using the criterion that "an item should be removed if its removal leads to an increase in alpha or a decrease of.02 or less," we could easily and correctly remove severely and moderately contaminated items, especially in short scales. For example, in Table 2, for severely and moderately contaminated scales, removing mismatched items would decrease less than.02 in alpha. However, this criterion (a decrease of.02 in alpha) did not effectively identify and remove mildly contaminated items, especially for longer scales (see Table 2).</p> <p>Lastly, our analysis revealed that for scales with severe to moderate contamination, a threshold of.04 was a reliable and practical criterion to determine item retention (see Tables 2 and 3). In contrast, a lower threshold of.02 was more appropriate for longer scales when making similar decisions (refer to Table 3). These findings and decision rules are concisely summarized in Figure 2 for readers' convenience.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/EMS/01jun24/emip12604-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="emip12604-fig-0002.jpg" title="2 Broad Decision Rules in Retention or Removal of an Item" /> </p> <p></p> <hd id="AN0177337607-21">Trade‐Off Between Short Scale and Scale Quality</hd> <p></p> <hd id="AN0177337607-22">Dynamic Reliability Benchmarks</hd> <p>The prevalent practice of adhering to a singular reliability (alpha) standard regardless of varying conditions presents challenges. Diverse guidelines have emerged (Aiken, [<reflink idref="bib1" id="ref33">1</reflink>]; Clark &amp; Watson, [<reflink idref="bib6" id="ref34">6</reflink>]; Gregory, [<reflink idref="bib14" id="ref35">14</reflink>]; Nunnally, [[<reflink idref="bib22" id="ref36">22</reflink>]]). Nunnally initially considered a.60 or.50 alpha sufficient, later advocating for a.7 in exploratory studies (Nunnally, [<reflink idref="bib23" id="ref37">23</reflink>]). Studies indicated that coefficient alpha escalates with an increase in items (Xiao &amp; Hau,[<reflink idref="bib33" id="ref38">33</reflink>]; Hoekstra et al., [<reflink idref="bib17" id="ref39">17</reflink>]). Researchers like Cho and Kim ([<reflink idref="bib4" id="ref40">4</reflink>]) have warned against the rigid application of a universal cutoff criterion. Along these lines, Ponterotto and Ruckdeschel ([<reflink idref="bib29" id="ref41">29</reflink>]) introduced a matrix for estimating adequate internal consistency across various scale lengths and sample sizes, though its development methodology remained obscure.</p> <hd id="AN0177337607-23">Study 3: Relations among Benchmarks, Scale Length, and Scale Strength</hd> <p>Pertinent to this study, alpha is influenced by both scale length and scale strength. For instance, a 3‐item scale with an average correlation of.438 exhibits an alpha of.700, mirroring a 6‐item scale with an average correlation of.280. Building on this concept, we aimed to establish dynamic reliability benchmarks that adjust according to scale lengths.</p> <hd id="AN0177337607-24">Design</hd> <p>Employing Equation (<reflink idref="bib3" id="ref42">3</reflink>) and the foundational principles of Confirmatory Factor Analysis (CFA), where item correlations are the product of their respective standardized factor loadings ( <ephtml> &lt;math display="inline" altimg="urn:x-wiley:07311745:media:emip12604:emip12604-math-0017" xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;semantics&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;/msub&gt;&lt;msub&gt;&lt;mi&gt;X&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;i&lt;/mi&gt;&lt;mspace width="0.33em" /&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;mspace width="0.33em" /&gt;&lt;mo&gt;&amp;#215;&lt;/mo&gt;&lt;msub&gt;&lt;mi&gt;&amp;#955;&lt;/mi&gt;&lt;mi&gt;j&lt;/mi&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;annotation encoding="application/x-tex"&gt;${{r}&amp;#95;{{{X}&amp;#95;i}{{X}&amp;#95;j}}} = {{\lambda }&amp;#95;{i\ }}\ \times {{\lambda }&amp;#95;j}$&lt;/annotation&gt;&lt;/semantics&gt;&lt;/math&gt; </ephtml> ), we deduced that alpha can be predicted for any combination of item factor loadings for a specific number of items. Conversely, this relationship enables the direct inference of average factor loadings for any preestablished alpha values and number of items. Thus, we cataloged the corresponding factor loadings for each scale, considering specific benchmarks (alpha) and scale lengths.</p> <p>For benchmarks (alpha), we chose a spectrum ranging from.55 to.95 in increments of.05 (i.e.,.55,.60,.65,.70,.75,.80,.85,.90, and.95), encompassing a broad array of plausible values. Scale lengths varied from 3 to 20 items, in line with the typical range of subscales in psychological measures (Ponterotto &amp; Ruckdeschel, [<reflink idref="bib29" id="ref43">29</reflink>]). Applying these parameters during calculations, we derived the requisite average factor loadings for items. For detailed results, please refer to Table S2.</p> <hd id="AN0177337607-25">Results</hd> <p>Our design framework enabled us to tabulate the scale strength (as factor loading) for various combinations of benchmarks (alpha) and scale lengths (see Table S2). To enhance interpretability, we divided our scales into five categories: 6 items or fewer, 7 to 9 items, 10 to 12 items, 13 to 15 items, and 16 items or more. This categorization facilitated the determination of the appropriate alpha level for each scale length category (color highlighted in Table S2).</p> <p>Moreover, aligning with recent guidelines suggesting minimum factor loadings between.40 and.70 for scale quality (Knekta et al., [<reflink idref="bib18" id="ref44">18</reflink>]), we translated this into four validity levels: low, fair, good, and excellent, corresponding to factor loadings of.4,.5,.6, and.7, respectively. These levels then informed the alpha benchmarks for each scale length category (bolded in Table S2). Last, we summarized the four levels of alpha benchmarks for each scale length category in Table 4.</p> <p>4 Table Broad Dynamic Alpha Benchmarks for Varying Scale Quality and Lengths</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;th&gt;Scale Length&lt;/th&gt;&lt;th align="left"&gt;Scale Quality (Factor Loading)&lt;/th&gt;&lt;/tr&gt;&lt;tr&gt;&lt;th&gt;(No. of Items)&lt;/th&gt;&lt;th&gt;Excellent (.7)&lt;/th&gt;&lt;th&gt;Good (.6)&lt;/th&gt;&lt;th&gt;Fair (.5)&lt;/th&gt;&lt;th&gt;Low/Minimum (.4)&lt;/th&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;3&amp;#8211;6 items&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;.75&lt;/td&gt;&lt;td&gt;.65&lt;/td&gt;&lt;td&gt;.55&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;7&amp;#8211;9 items&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;.75&lt;/td&gt;&lt;td&gt;.65&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10&amp;#8211;12 items&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;.80&lt;/td&gt;&lt;td&gt;.70&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;13&amp;#8211;15 items&lt;/td&gt;&lt;td&gt;.95&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;.75&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;16&amp;#8211;20 items&lt;/td&gt;&lt;td&gt;.95&lt;/td&gt;&lt;td&gt;.90&lt;/td&gt;&lt;td&gt;.85&lt;/td&gt;&lt;td&gt;.80&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p> <emph>Note</emph>. "Low/Minimum": the minimal standard. Alpha values lower than the minimum can be considered "unsatisfactory."</p> <p>As Table 4 suggests, for scales with fewer than ten items, an alpha standard below.7 can be acceptable when supplemented with other validity data. We recommended a.8 minimum for most unidimensional scales, especially for longer scales (see lower right‐hand side value in Table 4), to guarantee satisfactory reliability. Higher reliability standards are advisable for extensive scales or larger sample sizes. Notably, alpha estimation deviation is minimal for sample sizes of 500 or more. Therefore, our proposed dynamic benchmarks are robust for large samples. For smaller samples, adopting a slightly higher standard than those in Table 4 is advisable, or alternatively, reporting the confidence interval can be beneficial.</p> <hd id="AN0177337607-26">Discussion and Conclusion</hd> <p>Researchers frequently question whether an alpha of.70 suffices for scale reliability, often comparing their findings against a widely accepted benchmark (Nunnally, [[<reflink idref="bib22" id="ref45">22</reflink>]]; Nunnally &amp; Bernstein, [<reflink idref="bib24" id="ref46">24</reflink>]). We argue for the need to adjust this alpha standard based on scale length, noting that most traditional guidelines fall short by not specifying the scale length they pertain to. Additionally, the impact of sample size on measurement precision is often under‐discussed.</p> <p>Our research not only questions the appropriateness of a singular reliability benchmark but also introduces and validates the use of "alpha‐if‐item‐deleted" information. We emphasize the compromise between the ease of using shorter scales and the ensuing reduction in scale quality. A dynamic benchmark table for various scale lengths is provided for practical application.</p> <p>A few implications are noted. First, our analyses suggest that scales, regardless of their length, with comparable alpha values do not necessarily share equivalent scale quality. The "alpha‐if‐item‐deleted" metric is similarly affected. We proposed a simple yet broad decision rule utilizing this metric for item retention or removal. Specifically, we recommend retaining an item if its removal decreases alpha by more than.04. In contrast, items that decrease alpha by less than.02 should be considered for removal, particularly in short scales; for alpha changes ranging from.02 to.04, retention is suggested for longer scales. However, decisions should also incorporate qualitative aspects such as item content coverage.</p> <p>Second, acknowledging the trade‐off between the administrative convenience of short scales and reliability, a broader benchmark like a.80 alpha might be safer, as suggested by Cho and Kim ([<reflink idref="bib4" id="ref47">4</reflink>]). For longer scales, we advise a higher reliability standard (see Table 4). It is essential to recognize that these benchmarks are not absolute; the research context and the nature of the construct must be considered. In certain contexts, a lower alpha might be acceptable, but reducing reliability standards solely due to scale brevity is not justifiable. While shorter scales may struggle to achieve high reliability, the focus should remain on the integrity and consistency of measurement.</p> <p>Additionally, high reliability does not equate to quality measurement alone (Flake &amp; Fried, [<reflink idref="bib11" id="ref48">11</reflink>]; Fried &amp; Flake, [<reflink idref="bib12" id="ref49">12</reflink>]). Elevated alpha values could result from including too many similar items. While removing weaker items is necessary, care must be taken not to compromise item content coverage (validity; Raykov, [[<reflink idref="bib30" id="ref50">30</reflink>]]). Helms et al. ([<reflink idref="bib16" id="ref51">16</reflink>]) suggested reporting the number of items in both original and revised measures and providing indices of sample variances and subscale intercorrelations when removing items. Our research aids in elucidating some complexities in using coefficient alpha.</p> <p>Notably, our dynamic reliability benchmark is based on simplified analytical reasoning rather than empirical meta‐analysis. Therefore, these standards should be judiciously applied, considering the specific constructs and research context. Higher standards are likely necessary for larger sample sizes and longer scales, which afford greater reliability precision. A sample size exceeding 500 is generally advisable. Moreover, factors such as item content, sample characteristics, construct nature, dimensionality, and administration context, among others, may limit the applicability of our guidelines for practitioners. We recommend further exploration and consideration of these practical factors in applying our study's findings.</p> <hd id="AN0177337607-27">CONFLICT OF INTEREST STATEMENT</hd> <p>We have no conflicts of interest to declare. This research was not funded by any external sources.</p> <p>GRAPH: Supplementary Material</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/EMS/01jun24/emip12604-sup-0002-FigureS1.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="emip12604-sup-0002-FigureS1.jpg" title="Supplementary Material" /> </p> <p></p> <ref id="AN0177337607-29"> <title> References </title> <blist> <bibl id="bib1" idref="ref33" type="bt">1</bibl> <bibtext> Aiken, L. R. (2000). Psychological testing and assessment (10th ed.). Boston, MA : Allyn &amp; Bacon.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref20" type="bt">2</bibl> <bibtext> Bieleke, M., Gogol, K., Goetz, T., Daniels, L., &amp; Pekrun, R. (2021). The AEQ‐S: A short version of the Achievement Emotions Questionnaire. Contemporary Educational Psychology, 65, Article 101940.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref21" type="bt">3</bibl> <bibtext> Bonett, D. G. (2002). Sample size requirements for testing and estimating coefficient alpha. Journal of Educational and Behavioral Statistics, 27 (4), pp. 335 – 340.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref25" type="bt">4</bibl> <bibtext> Cho, E., &amp; Kim, S. (2015). Cronbach's coefficient alpha: Well known but poorly understood. Organizational Research Methods, 18 (2), pp. 207 – 230.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref13" type="bt">5</bibl> <bibtext> Churchill, G. A., Jr., &amp; Peter, J. P. (1984). Research design effects on the reliability of rating scales: A meta‐analysis. Journal of Marketing Research, 21 (4), pp. 360 – 375.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref34" type="bt">6</bibl> <bibtext> Clark, L. A., &amp; Watson, D. (1995). Constructing validity: Basic issues in objective scale development. Psychological Assessment, 7, pp. 309 – 319.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref23" type="bt">7</bibl> <bibtext> Cohen, J. (1988). Statistical power for the behavioral sciences. Hillsdale, NJ : Erlbaum.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref4" type="bt">8</bibl> <bibtext> Cronbach, L. J. (1951). Coefficient alpha and the internal structure of tests. Psychometrika, 16 (3), pp. 297 – 334.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref9" type="bt">9</bibl> <bibtext> Davenport, E. C., Davison, M. L., Liou, P. Y., &amp; Love, Q. U. (2015). Reliability, dimensionality, and internal consistency as defined by Cronbach: Distinct albeit related concepts. Educational Measurement: Issues and Practice, 34 (4), pp. 4 – 9.</bibtext> </blist> <blist> <bibtext> Edwards, A. A., Joyner, K. J., &amp; Schatschneider, C. (2021). A simulation study on the performance of different reliability estimation methods. Educational and Psychological Measurement, 81 (6), pp. 1089 – 1117.</bibtext> </blist> <blist> <bibtext> Flake, J. K., &amp; Fried, E. I. (2020). Measurement schmeasurement: Questionable measurement practices and how to avoid them. Advances in Methods and Practices in Psychological Science, 3 (4), pp. 456 – 465.</bibtext> </blist> <blist> <bibtext> Fried, E. I., &amp; Flake, J. K. (2018). Measurement matters. APS Observer, p. 31.</bibtext> </blist> <blist> <bibtext> Green, S. B., &amp; Yang, Y. (2009). Commentary on coefficient alpha: A cautionary tale. Psychometrika, 74 (1), pp. 121 – 135.</bibtext> </blist> <blist> <bibtext> Gregory, R. J. (2000). Psychological testing: History, principles, and applications (3rd ed.). Boston, MA : Allyn &amp; Bacon.</bibtext> </blist> <blist> <bibtext> Hayes, A. F., &amp; Coutts, J. J. (2020). Use omega rather than Cronbach's alpha for estimating reliability. But.... Communication Methods and Measures, 14 (1), pp. 1 – 24.</bibtext> </blist> <blist> <bibtext> Helms, J. E., Henze, K. T., Sass, T. L., &amp; Mifsud, V. A. (2006). Treating Cronbach's alpha reliability coefficients as data in counseling research. The Counseling Psychologist, 34 (5), pp. 630 – 660.</bibtext> </blist> <blist> <bibtext> Hoekstra, R., Vugteveen, J., Warrens, M. J., &amp; Kruyen, P. M. (2019). An empirical analysis of alleged misunderstandings of coefficient alpha. International Journal of Social Research Methodology, 22 (4), pp. 351 – 364.</bibtext> </blist> <blist> <bibtext> Knekta, E., Runyon, C., &amp; Eddy, S. (2019). One size doesn't fit all: Using factor analysis to gather validity evidence when using surveys in your research. CBE—Life Sciences Education, 18 (1), pp. 1 – 17.</bibtext> </blist> <blist> <bibtext> Lord, F. M., &amp; Novick, M. R. (1968). Statistical theories of mental test scores. Reading, MA : Addison Wesley.</bibtext> </blist> <blist> <bibtext> McNeish, D. (2018). Thanks coefficient alpha, we'll take it from here. Psychological Methods, 23 (3), pp. 412 – 433.</bibtext> </blist> <blist> <bibtext> Marsh, H. W., Hau, K. T., &amp; Wen, Z. (2004). In search of golden rules: Comment on hypothesis‐testing approaches to setting cutoff values for fit indexes and dangers in over generalizing Hu and Bentler's (1999) findings. Structural Equation Modeling, 11 (3), 320 – 341.</bibtext> </blist> <blist> <bibtext> Nunnally, J. C. (1967). Psychometric theory. New York : McGraw‐Hill.</bibtext> </blist> <blist> <bibtext> Nunnally, J. C. (1978). Psychometric theory (2nd ed.). New York : McGraw‐Hill.</bibtext> </blist> <blist> <bibtext> Nunnally, J. C., &amp; Bernstein, I. H. (1994). Psychometric theory (3rd ed.). New York : McGraw‐Hill.</bibtext> </blist> <blist> <bibtext> OECD. (2013). PISA 2012 assessment and analytical framework. OECD Publishing.</bibtext> </blist> <blist> <bibtext> OECD. (2019). PISA 2018 assessment and analytical framework. OECD Publishing.</bibtext> </blist> <blist> <bibtext> Pereira, V., da Hora, H. R. M., Costa, H. G., &amp; de Oliveira Nepomuceno, L. D. (2014). An if‐item‐deleted sensitive analysis of Cronbach's alpha technique using simulated anneling algorithm. Cadernos Do IME‐Série Estatística, 36 (1), pp. 29 – 37.</bibtext> </blist> <blist> <bibtext> Peterson, R. A. (1994). A meta‐analysis of Cronbach's coefficient alpha. Journal of Consumer Research, 21 (9), pp. 381 – 391.</bibtext> </blist> <blist> <bibtext> Ponterotto, J. G., &amp; Ruckdeschel, D. E. (2007). An overview of coefficient alpha and a reliability matrix for estimating adequacy of internal consistency coefficients with psychological research measures. Perceptual and Motor Skills, 105 (3), pp. 997 – 1014.</bibtext> </blist> <blist> <bibtext> Raykov, T. (2007). Reliability if deleted, not "alpha if deleted": Evaluation of scale reliability following component deletion. British Journal of Mathematical and Statistical Psychology, 60 (2), pp. 201 – 216.</bibtext> </blist> <blist> <bibtext> Raykov, T. (2008). Alpha if item deleted: A note on criterion validity loss in scale revision if maximizing coefficient alpha. British Journal of Mathematical and Statistical Psychology, 61 (2), pp. 275 – 285.</bibtext> </blist> <blist> <bibtext> Schipolowski, S., Schroeders, U., &amp; Wilhelm, O. (2014). Pitfalls and challenges in constructing short forms of cognitive ability measures. Journal of Individual Differences, 35 (4), pp. 190 – 200.</bibtext> </blist> <blist> <bibtext> Xiao, L., &amp; Hau, K. T. (2023). Accuracy and sensitivity of coefficient alpha and its alternatives with unidimensional and contaminated scales. Applied Measurement in Education, 36 (1), 31 – 44.</bibtext> </blist> <blist> <bibtext> Ziegler, M., Kemper, C. J., &amp; Kruyen, P. (2014). Short scales – Five misunderstandings and ways to overcome them. Journal of Individual Differences, 35 (4), pp. 185 – 189.</bibtext> </blist> </ref> <aug> <p>By Leifeng Xiao; Kit‐Tai Hau and Melissa Dan Wang</p> <p>Reported by Author; Author; Author</p> <p></p> <p>Dr. Leifeng XIAO is currently a specially appointed associate professor in the School of Psychology, Shanghai Normal University; meanwhile, she is also a member of the Lab for Educational Big Data and Policy Making. Her research interests involve educational methodology research. In addition, she is interested in issues of diversity within the field of educational psychology.</p> <p>Dr. Kit‐Tai Hau is currently a full professor in the Department of Educational Psychology, at the Chinese University of Hong Kong. His research interests involve research methodology, structural equation modeling, and motivation.</p> <p>Melissa Dan WANG is PhD in the Department of Educational Psychology, at the Chinese University of Hong Kong, researching insufficient effortful response. She has published in computerized instructional technology.</p> </aug> <nolink nlid="nl1" bibid="bib32" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib26" firstref="ref2"></nolink> <nolink nlid="nl3" bibid="bib34" firstref="ref3"></nolink> <nolink nlid="nl4" bibid="bib19" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib20" firstref="ref7"></nolink> <nolink nlid="nl6" bibid="bib33" firstref="ref8"></nolink> <nolink nlid="nl7" bibid="bib10" firstref="ref10"></nolink> <nolink nlid="nl8" bibid="bib13" firstref="ref11"></nolink> <nolink nlid="nl9" bibid="bib17" firstref="ref12"></nolink> <nolink nlid="nl10" bibid="bib28" firstref="ref14"></nolink> <nolink nlid="nl11" bibid="bib23" firstref="ref15"></nolink> <nolink nlid="nl12" bibid="bib21" firstref="ref16"></nolink> <nolink nlid="nl13" bibid="bib31" firstref="ref17"></nolink> <nolink nlid="nl14" bibid="bib27" firstref="ref27"></nolink> <nolink nlid="nl15" bibid="bib15" firstref="ref28"></nolink> <nolink nlid="nl16" bibid="bib25" firstref="ref31"></nolink> <nolink nlid="nl17" bibid="bib14" firstref="ref35"></nolink> <nolink nlid="nl18" bibid="bib22" firstref="ref36"></nolink> <nolink nlid="nl19" bibid="bib29" firstref="ref41"></nolink> <nolink nlid="nl20" bibid="bib18" firstref="ref44"></nolink> <nolink nlid="nl21" bibid="bib24" firstref="ref46"></nolink> <nolink nlid="nl22" bibid="bib11" firstref="ref48"></nolink> <nolink nlid="nl23" bibid="bib12" firstref="ref49"></nolink> <nolink nlid="nl24" bibid="bib30" firstref="ref50"></nolink> <nolink nlid="nl25" bibid="bib16" firstref="ref51"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1425083 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Revisiting the Usage of Alpha in Scale Evaluation: Effects of Scale Length and Sample Size – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Leifeng+Xiao%22">Leifeng Xiao</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0001-7125-1067">0000-0001-7125-1067</externalLink>)<br /><searchLink fieldCode="AR" term="%22Kit-Tai+Hau%22">Kit-Tai Hau</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-4065-1855">0000-0003-4065-1855</externalLink>)<br /><searchLink fieldCode="AR" term="%22Melissa+Dan+Wang%22">Melissa Dan Wang</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Measurement%3A+Issues+and+Practice%22"><i>Educational Measurement: Issues and Practice</i></searchLink>. 2024 43(2):74-81. – Name: Avail Label: Availability Group: Avail Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 8 – Name: DatePubCY Label: Publication Date Group: Date Data: 2024 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Research – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Measurement%22">Measurement</searchLink><br /><searchLink fieldCode="DE" term="%22Benchmarking%22">Benchmarking</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Sampling%22">Item Sampling</searchLink><br /><searchLink fieldCode="DE" term="%22Sample+Size%22">Sample Size</searchLink><br /><searchLink fieldCode="DE" term="%22Research+Methodology%22">Research Methodology</searchLink><br /><searchLink fieldCode="DE" term="%22Guidelines%22">Guidelines</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Construction%22">Test Construction</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1111/emip.12604 – Name: ISSN Label: ISSN Group: ISSN Data: 0731-1745<br />1745-3992 – Name: Abstract Label: Abstract Group: Ab Data: Short scales are time-efficient for participants and cost-effective in research. However, researchers often mistakenly expect short scales to have the same reliability as long ones without considering the effect of scale length. We argue that applying a universal benchmark for alpha is problematic as the impact of low-quality items is greater on shorter scales. In this study, we proposed simple guidelines for item reduction using the "alpha-if-item-deleted" procedure in scale construction. An item can be removed if alpha increases or decreases by less than 0.02, especially for short scales. Conversely, an item should be retained if alpha decreases by more than 0.04 upon its removal. For reliability benchmarks, 0.80 is relatively safe in most conditions, but higher benchmarks are recommended for longer scales and smaller sample sizes. Supplementary analyses, including item content, face validity, and content coverage, are critical to ensure scale quality. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2024 – Name: AN Label: Accession Number Group: ID Data: EJ1425083 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1425083 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1111/emip.12604 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 8 StartPage: 74 Subjects: – SubjectFull: Measurement Type: general – SubjectFull: Benchmarking Type: general – SubjectFull: Item Sampling Type: general – SubjectFull: Sample Size Type: general – SubjectFull: Research Methodology Type: general – SubjectFull: Guidelines Type: general – SubjectFull: Test Construction Type: general Titles: – TitleFull: Revisiting the Usage of Alpha in Scale Evaluation: Effects of Scale Length and Sample Size Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Leifeng Xiao – PersonEntity: Name: NameFull: Kit-Tai Hau – PersonEntity: Name: NameFull: Melissa Dan Wang IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 06 Type: published Y: 2024 Identifiers: – Type: issn-print Value: 0731-1745 – Type: issn-electronic Value: 1745-3992 Numbering: – Type: volume Value: 43 – Type: issue Value: 2 Titles: – TitleFull: Educational Measurement: Issues and Practice Type: main |
| ResultId | 1 |