Using Full-Information Item Analysis to Improve Item Quality
Saved in:
| Title: | Using Full-Information Item Analysis to Improve Item Quality |
|---|---|
| Language: | English |
| Authors: | Haladyna, Thomas M., Rodriguez, Michael C. |
| Source: | Educational Assessment. 2021 26(3):198-211. |
| Availability: | Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals |
| Peer Reviewed: | Y |
| Page Count: | 14 |
| Publication Date: | 2021 |
| Document Type: | Journal Articles Reports - Descriptive |
| Education Level: | Elementary Education Grade 6 Intermediate Grades Middle Schools |
| Descriptors: | Test Items, Item Analysis, Reading Tests, Mathematics Tests, Grade 6, Difficulty Level |
| DOI: | 10.1080/10627197.2021.1946390 |
| ISSN: | 1062-7197 |
| Abstract: | Full-information item analysis provides item developers and reviewers comprehensive empirical evidence of item quality, including option response frequency, point-biserial index (PBI) for distractors, mean-scores of respondents selecting each option, and option trace lines. The multi-serial index (MSI) is introduced as a more informative item-total correlation, accounting for variable distractor performance. The overall item PBI is empirically compared to the MSI. For items from an operational mathematics and reading test, poorly performing distractors are systematically removed to recompute the MSI, indicating improvements in item quality. Case studies for specific items with different characteristics are described to illustrate a variety of outcomes, focused on improving item discrimination. Full-information item analyses are presented for each case study item, providing clear examples of interpretation and use of item analyses. A summary of recommendations for item analysts is provided. |
| Abstractor: | As Provided |
| Entry Date: | 2021 |
| Accession Number: | EJ1309544 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwEqn-PH4D1QjVPb0Xo1a_VAAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDILNJhe2ng4rapmL2wIBEICBm-g7N-YaEIEK6ZqlshqBRkiTShTkD0KDvSDZRgOR8pP4_iCX3m2AehqcplSdPOZVrOPHD4d6TvooetaVi_G5VQbSt_If2MYdvG7nkhmufennL_9Ra3yXWFSUV74HVlTrDPxfa-t8MWrPDuxvZFvUbCpdgJmw-HvjCTHUxWbOHTUWtt0oPcVEkntz9NEDKmwJ2Z8E3PI4btLyiTlE Text: Availability: 1 Value: <anid>AN0152008618;7ls01jul.21;2021Aug23.04:04;v2.2.500</anid> <title id="AN0152008618-1">Using Full-information Item Analysis to Improve Item Quality </title> <sbt id="AN0152008618-2">Introduction</sbt> <p>Full-information item analysis provides item developers and reviewers comprehensive empirical evidence of item quality, including option response frequency, point-biserial index (PBI) for distractors, mean-scores of respondents selecting each option, and option trace lines. The multi-serial index (MSI) is introduced as a more informative item-total correlation, accounting for variable distractor performance. The overall item PBI is empirically compared to the MSI. For items from an operational mathematics and reading test, poorly performing distractors are systematically removed to recompute the MSI, indicating improvements in item quality. Case studies for specific items with different characteristics are described to illustrate a variety of outcomes, focused on improving item discrimination. Full-information item analyses are presented for each case study item, providing clear examples of interpretation and use of item analyses. A summary of recommendations for item analysts is provided.</p> <p>The evaluation of any newly written multiple-choice (MC) item begins with the considerable experience and skill of subject-matter experts (SMEs) and editors (Rodriguez, [<reflink idref="bib17" id="ref1">17</reflink>]). One of the most important later steps in item development and validation is item analysis. Item analysis has many objectives (Haladyna, [<reflink idref="bib6" id="ref2">6</reflink>], p. 392–393). One is to provide evidence for the dimensionality for what is being measured. A second is to determine which items will be retained in the item bank for use, revised, or discarded. A third is to provide information regarding specific content to guide test design and construction. A fourth is to supply information for the development of parallel test forms. A fifth objective is key validation after an operational test is administered to ensure that correct options are actually correct. To accomplish these objectives, we need information concerning items, not test takers (Livingston, [<reflink idref="bib11" id="ref3">11</reflink>], p. 421).</p> <p>We focus on the second objective. To determine if an item should be used, revised, or discarded, the item analyst must have good information. We support a full-information item analysis that introduces a new discrimination index that responds to the differential information that exists in MC distractors.</p> <p>The item analyst needs to know the difficulty and discrimination of each item. Also, it is important to know how well MC distractors are performing (Gierl et al., [<reflink idref="bib4" id="ref4">4</reflink>]; Haladyna &amp; Rodriguez, [<reflink idref="bib8" id="ref5">8</reflink>]; Livingston, [<reflink idref="bib11" id="ref6">11</reflink>]; Thissen, [<reflink idref="bib20" id="ref7">20</reflink>]; Thissen, Steinberg, &amp; Fitzpatrick, [<reflink idref="bib19" id="ref8">19</reflink>]). A poorly performing distractor can affect how well the item measures what it is intended to measure. The quest is to evaluate and, if possible, improve item quality in the quest for valid test score interpretation and use.</p> <p>First, the concept of reciprocity is presented as it bears on the study of distractor performance. Second, full-information item analysis is described. Third, a new item discrimination index, the multiserial (MSI), is introduced and described. The MSI is contrasted with the traditional point-biserial discrimination index. Fourth, two studies are reported bearing on the usefulness of the MSI as part of full-information item analysis. Finally, the use of full-information item analysis with the MSI enables the item analyst to evaluate item performance and, in many instances, improve item discrimination. With improved item discrimination, test score reliability is increased; similarly, with better estimates of item discrimination, we obtain better estimates of reliability.</p> <hd id="AN0152008618-3">Reciprocity</hd> <p>An important concept for understanding distractor discrimination is <emph>reciprocity</emph>. For any MC test item, reciprocity exists for responses to the correct option and responses to the team of distractors (Haladyna &amp; Rodriguez, [<reflink idref="bib8" id="ref9">8</reflink>], p. 346; Haladyna, Rodriguez, &amp; Stevens, [<reflink idref="bib9" id="ref10">9</reflink>]). A trace line is a graphical representation of test taker responses to a MC item as a function of ordered-score groups (Haladyna, [<reflink idref="bib6" id="ref11">6</reflink>]; Livingston, [<reflink idref="bib11" id="ref12">11</reflink>]). Figure 1 contains a trace line for the correct option and a trace line for the team of distractors. If the correlation between the correct option and total score is.40, then the correlation between the team of distractors and total score will be −.40. This reciprocity between the correct option and team of distractors is true for any MC item.</p> <p>PHOTO (COLOR): Figure 1. Trace lines illustrating the principle of reciprocity between the correct response and collective incorrect responses (team of distractors)</p> <p>If one distractor is not contributing to the team effort, does the overall item discrimination suffer? Would removing a poorly performing distractor improve item discrimination? If so, how much? What dynamics are involved?</p> <hd id="AN0152008618-4">Item and distractor evaluation</hd> <p>The evaluation of MC test item quality requires the frequency of responses to options for the item, item difficulty and discrimination, and the discrimination of each distractor.</p> <hd id="AN0152008618-5">Frequency of response</hd> <p>The tabulation of responses to each test item is a first step in evaluating discrimination. A low frequency of response to a distractor often yields a recommendation for revision or removal of the distractor because it indicates implausibility (Haladyna &amp; Downing, [<reflink idref="bib7" id="ref13">7</reflink>]; Raymond, Stevens, &amp; Bucak, [<reflink idref="bib16" id="ref14">16</reflink>]). Removal of a low-frequency distractor can shorten the test or create space on the test to add additional items that may increase content coverage and score reliability.</p> <hd id="AN0152008618-6">Item difficulty</hd> <p>Generally, the proportion correct (<emph>p</emph>-value = number of test takers with correct response/total number of test takers) of each item is computed, or IRT can be used to compute an item difficulty parameter. Although on different scales, the item <emph>p</emph>-value and the item difficulty parameter are very highly correlated; it matters little which is used when we evaluate item performance. For completeness, for one-parameter IRT models, the item difficulty is a function of the log odds of correct response, <emph>ln</emph>([1-p]/p). The item <emph>p</emph>-value as a measure of item difficulty is the simplest and easiest to interpret.</p> <hd id="AN0152008618-7">Item discrimination</hd> <p>Item discrimination is a measure of the association between the correct option (selected = 1, not selected = 0) and total score. Three item discrimination indexes are contrasted here: the point-biserial index (PBI), the biserial index (BIS), and the multiserial index (MSI).</p> <hd id="AN0152008618-8">PBI</hd> <p>The PBI is a product-moment correlation that is simple to compute and interpret. The formula for the PBI is</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;r&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;=&lt;/mo&gt;&lt;mrow&gt;&lt;mfrac&gt;&lt;mrow&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mn&gt;1&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mo&gt;&amp;#8722;&lt;/mo&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;M&lt;/mi&gt;&lt;mn&gt;0&lt;/mn&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/mrow&gt;&lt;mi&gt;s&lt;/mi&gt;&lt;/mfrac&gt;&lt;/mrow&gt;&lt;msqrt&gt;&lt;mi&gt;p&lt;/mi&gt;&lt;mi&gt;q&lt;/mi&gt;&lt;/msqrt&gt;&lt;/math&gt; </ephtml> . <emph>M</emph><subs>1</subs> is the mean test score of individuals correctly responding to the item; <emph>M</emph><subs>0</subs> is the mean test score of individuals incorrectly responding to the item; <emph>s</emph> is the standard deviation of total test scores; <emph>p</emph> is the proportion who chose the correct response; and <emph>q</emph> is 1 – <emph>p</emph>. PBI<sups>2</sups> is the proportion of test score variance accounted for by the choice mean keyed (correct) response as contrasted with the choice mean of the combined distractors (the mean of test takers with an incorrect response). Item reliability is the product of the item score variance (<emph>pq</emph>) and PBI, acknowledging that the contribution of an item to the total test score reliability is a function of both the item variance and discrimination (Crocker &amp; Algina, [<reflink idref="bib2" id="ref15">2</reflink>]). Score reliability for the entire test can be estimated using a formula provided by Ebel ([<reflink idref="bib3" id="ref16">3</reflink>]) based on item variances and item discrimination values. Thus, there is a direct functional association between item discrimination and score reliability.</p> <hd id="AN0152008618-9">BIS</hd> <p>The BIS is computationally complex. Theoretically, a BIS correlation estimates the magnitude of association between an artificially dichotomized variable (item response score) and a continuous variable. If you correlate a set of PBIs and BISs for any set of test item data, you will find these two indexes provide essentially the same information but on different scales. The BIS scale is higher (closer to 1.0) and is less sensitive to extreme item <emph>p</emph>-values.</p> <hd id="AN0152008618-10">MSI</hd> <p>The MSI is simply a multiple correlation involving the item option response (specifically the item nominal response data; typically A, B, C, and D for a four option item) as the independent variable (fixed factor) and the total test score as the dependent variable. A one-way analysis of variance (ANOVA) is used to compute MSI. Simply, it is the square root of the between sum of squares divided by the total sum of squares: MSI =</p> <p>Graph</p> <p> <ephtml> &lt;math xmlns="http://www.w3.org/1998/Math/MathML"&gt;&lt;msqrt&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;b&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;w&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;e&lt;/mi&gt;&lt;mi&gt;n&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;mrow&gt;&lt;mo&gt;/&lt;/mo&gt;&lt;/mrow&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mrow&gt;&lt;msub&gt;&lt;mi&gt;S&lt;/mi&gt;&lt;mrow&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;o&lt;/mi&gt;&lt;mi&gt;t&lt;/mi&gt;&lt;mi&gt;a&lt;/mi&gt;&lt;mi&gt;l&lt;/mi&gt;&lt;/mrow&gt;&lt;/msub&gt;&lt;/mrow&gt;&lt;/msqrt&gt;&lt;/math&gt; </ephtml> . Thus, MIS<sups>2</sups> is then the multiple R-squared, the amount of test score variance accounted for by the item response across all options. The PBI is much like a <emph>t</emph>-test, where the mean of those choosing the correct option is compared to the mean of those choosing any distractor. Stated as a null hypothesis, the PBI H<subs>0</subs>: μ<subs>correct</subs> = μ<subs>incorrect</subs>. For MSI the null hypothesis is H<subs>0</subs>: μ<subs>a</subs> = μ<subs>b</subs> = μ<subs>c</subs> = ... = μ<emph><subs>n</subs></emph>. In typical MC test items, the number of options is four or five.</p> <p>Because MSI is sensitive to variation in distractor discrimination (which is rarely, if ever, uniform), MSI should exceed PBI. By accounting for more test score variance using MSI, reliability will be more accurately estimated (accounting for more variance in total test scores), which is beneficial especially with high-stakes test score interpretation and use. Moreover, Metsämuuronen ([<reflink idref="bib12" id="ref17">12</reflink>]) argued that using PBI underestimates reliability by as much as 13%. The MSI is sensitive to differences in distractor choice means, whereas the PBI uses the combined choice means of all distractors.</p> <hd id="AN0152008618-11">Alternatives</hd> <p>Alternatively, many IRT and other methods for estimating discrimination exist (Gierl et al., [<reflink idref="bib4" id="ref18">4</reflink>]; Metsämuuronen, [<reflink idref="bib13" id="ref19">13</reflink>]). The discrimination parameter in the two-parameter IRT model is nearly perfectly correlated with classical discrimination. For the three-parameter IRT model, discrimination is linearly related to classical discrimination except in the lower part of the scale where the lower-asymptote (the c-parameter) affects the estimation of discrimination. The PBI is the simplest, mainstream index of item discrimination suitable for this study. The polyserial and polychoric correlation coefficients have been shown to be related to the PBI and BIS (Olsson, Drasgow, &amp; Dorans, [<reflink idref="bib15" id="ref20">15</reflink>]). These alternative procedures for estimating item discrimination have not yet been shown to be advantageous.</p> <hd id="AN0152008618-12">Distractor discrimination</hd> <p>We know from experience that distractors vary considerably in their associations with total scores. Moreover, researchers have reported associations between distractors and total scores that may improve the test score (Haladyna, [<reflink idref="bib6" id="ref21">6</reflink>]; Levine &amp; Drasgow, [<reflink idref="bib10" id="ref22">10</reflink>]; Thissen, [<reflink idref="bib20" id="ref23">20</reflink>]; Thissen et al., [<reflink idref="bib19" id="ref24">19</reflink>]). Samejima ([<reflink idref="bib18" id="ref25">18</reflink>]) proposed models to investigate differential distractor functioning. Wang ([<reflink idref="bib22" id="ref26">22</reflink>]) also proposed a factorial model in a series of studies. These procedures have not found their way to standard item analysis perhaps due to their complexity and lack of evidence that they provide new and different useful information regarding how well a distractor functions. For the purpose of measuring distractor discrimination, the PBI seems appropriate, (Attali &amp; Fraenkel, [<reflink idref="bib1" id="ref27">1</reflink>], provided a correction to focus the PBI on the comparison between the distractor and correct option, rather than between the distractor and the set of all other options).</p> <hd id="AN0152008618-13">Full-information item analysis</hd> <p>In <emph>The Future of Item Analysis</emph>, Wainer ([<reflink idref="bib21" id="ref28">21</reflink>]) provided examples of a more comprehensive examination of item responses. He provided additional rationale for full-information item analysis. Metsämuuronen ([<reflink idref="bib14" id="ref29">14</reflink>]) featured trace lines as an essential analytical tool for item and distractor discrimination analysis. Full information includes statistical and visual aids toward understanding how well an item functions and discriminates.</p> <p>Figure 2 contains an example of full-information item analysis for a single item. This item is highly discriminating (on all metrics) and requires no further attention. The correct choice frequency is monotonically increasing as a function of the ordered quintile groups. Distractor choice frequencies are monotonically decreasing as a function of the ordered quintile groups. These trace lines are what one expects of a well performing item with high discrimination. The upper table contains information about the <emph>p</emph>-value (response proportions) for each option, the PBI, the MSI, and the test score mean of test takers choosing each option (referred to as <emph>choice mean</emph>). In this example, the item difficulty is.737; the PBI is.466; the MSI is.487; and the choice mean for each option is listed. As expected, the correct option has the highest choice mean, and distractors have lower choice means. The lower table in Figure 2 contains the percentage of test takers choosing each option as a function of five (or ten if preferred) equal, ordered score groups. As previously noted, the frequencies for the correct option in the ordered score groups increase monotonically:.363,.589,.789,.927, and.976. Only the lowest scoring test takers are inclined to choose distractors. Researchers often advocate dropping low frequency distractors (Gierl et al., [<reflink idref="bib4" id="ref30">4</reflink>]; Haladyna &amp; Rodriguez, [<reflink idref="bib8" id="ref31">8</reflink>]).</p> <p>PHOTO (COLOR): Figure 2. A highly effective item (mathematics item 36)</p> <p>The importance of full-information item analysis is that it also describes the performance of the <emph>team</emph> of distractors. That information helps analysts better understand the dynamism among distractors and guides them in determining how to improve an item's capacity to measure more accurately and precisely the construct it represents. Two studies reported here reveal the benefits of full-information item analysis.</p> <hd id="AN0152008618-14">Two empirical studies</hd> <p>As a result of the foregoing discussion, two research questions were posed. Which index (PBI or MSI) accounts for the most test score variance and is more informative about distractor discrimination? What happens to item discrimination when a poorly performing distractor is removed from the test item and does it affect MSI?</p> <p>Test results from statewide sixth grade Reading and Mathematics tests were used (with strong validity evidence and alignment to state content standards). To effect a useful examination of item quality, a sample of 1,000 student item responses was employed. For the purpose of this study, the sample did not contain omitted or not-reached item responses. This was done to assure that information reported was not influenced by aberrant or careless student responses. As stated previously, item analysis focuses on item performance not test-taker performance. Descriptive statistics for the Reading and Mathematics tests were presented in Table 1. These test results had characteristics we would expect from well-validated, professionally-developed, large-scale achievement tests. The Mathematics test was more difficult than the Reading test (both were challenging to test takers). Both tests had high coefficient alpha reliability estimates.</p> <p>Table 1. Descriptive statistics for the reading and mathematics tests</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;Subject&lt;/td&gt;&lt;td&gt;n of items&lt;/td&gt;&lt;td&gt;&lt;italic&gt;M&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;SD&lt;/italic&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;Alpha&lt;/italic&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Reading&lt;/td&gt;&lt;td&gt;112&lt;/td&gt;&lt;td&gt;62.1%&lt;/td&gt;&lt;td&gt;17.5%&lt;/td&gt;&lt;td&gt;.947&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mathematics&lt;/td&gt;&lt;td&gt;88&lt;/td&gt;&lt;td&gt;46.6%&lt;/td&gt;&lt;td&gt;16.8%&lt;/td&gt;&lt;td&gt;.919&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0152008618-15">Study 1: Comparing PBI and MSI</hd> <p>In the first study, we compared PBI and MSI indexes for all Reading and Mathematic items. The hypothesis was that the MSI should exceed PBI in all instances except when the distractor choice means are equivalent. Because PBI<sups>2</sups> and MSI<sups>2</sups> represent the percentage of test score variance accounted for by the item response, analysts should view the higher coefficient as a more accurate and precise measure of item discrimination (accounting for more variance in test scores). The difference between MSI and PBI offers clues as to the performance of each distractor as part of the team of distractors.</p> <hd id="AN0152008618-16">Method</hd> <p>PBI and MSI were computed for all 112 Reading and 88 Mathematics items. The differences between the respective sets of discriminations (MSI-PBI) were evaluated using a one-way ANOVA with repeated measures. The highest MSI-PBI differences were investigated for clues as to why these differences were large.</p> <hd id="AN0152008618-17">Results and discussion</hd> <p>Table 2 contains the means of the test difficulty (proportion correct), PBI, MSI, and MSI-PBI for the 112 Reading and 88 Mathematics items. For Reading, the mean difference between MSI and PBI was.048. MSI exceeded PBI for all 112 items. The highest difference was.167, and the lowest difference was.010. The repeated measures ANOVA was statistically significant, (<emph>F</emph> = 203.4, <emph>p</emph> &lt;.0001). This accounted for 64.4% of variance, a very large effect size. For Mathematics, the mean difference between MSI and PBI was.047. As with Reading, MSI exceeded PBI for all 88 items. The highest difference was.141, and the lowest difference was.013. The repeated measures ANOVA was significant (<emph>F</emph> = 302.0, <emph>p</emph> &lt;.0001). This accounted for 77.8% of variance, a very large effect size.</p> <p>Table 2. Mean PBI, MSI, and item difficulty for the 112 reading and 88 mathematics items</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;Subject&lt;/td&gt;&lt;td&gt;Test difficulty&lt;/td&gt;&lt;td&gt;PBI&lt;/td&gt;&lt;td&gt;MSI&lt;/td&gt;&lt;td&gt;MSI-PBI&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Reading&lt;/td&gt;&lt;td&gt;.643&lt;/td&gt;&lt;td&gt;.365&lt;/td&gt;&lt;td&gt;.413&lt;/td&gt;&lt;td&gt;.048&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mathematics&lt;/td&gt;&lt;td&gt;.470&lt;/td&gt;&lt;td&gt;.322&lt;/td&gt;&lt;td&gt;.369&lt;/td&gt;&lt;td&gt;.047&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Correlations among item difficulty, PBI, MSI, and MSI-PBI are presented in Table 3. The upper diagonal of coefficients is for the Reading test; the lower diagonal of coefficients is for the Mathematics test. Higher PBI and MSIs were associated with easier items (with even higher correlations in Mathematics compared to Reading). In addition, the negative association between PBI and MSI with the difference MSI-PBI indicated that the larger the PBI and MSI, the smaller the difference; as PBI and MSI increased, they became more similar (we note that this was expected due to the ceiling value of 1.0). Moreover, we noted the moderate negative correlation between item difficulty and MSI-PBI difference (−.532 for Reading, −.506 for Mathematics); when item difficulty was higher, the difference between MSI and PBI was smaller. Where MSI-PBI was small, the choice means of the three distractors were approximately equal. In these instances, MSI represented a small improvement in the accuracy of a discrimination estimate. Where MSI-PBI was largest, the full-information item analysis offered clues about why PBI did not perform as well, as a discrimination index.</p> <p>Table 3. Product-moment correlations among item difficulty, PBI, MSI, and MSI-PBI</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;Difficulty&lt;/td&gt;&lt;td&gt;PBI&lt;/td&gt;&lt;td&gt;MSI&lt;/td&gt;&lt;td&gt;MSI-PBI&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Difficulty&lt;/td&gt;&lt;td /&gt;&lt;td&gt;.459&lt;/td&gt;&lt;td&gt;.326&lt;/td&gt;&lt;td&gt;&amp;#8722;.532&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PBI&lt;/td&gt;&lt;td&gt;.603&lt;/td&gt;&lt;td /&gt;&lt;td&gt;.938&lt;/td&gt;&lt;td&gt;&amp;#8722;.670&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MSI&lt;/td&gt;&lt;td&gt;.578&lt;/td&gt;&lt;td&gt;.986&lt;/td&gt;&lt;td /&gt;&lt;td&gt;&amp;#8722;.371&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MSI-PBI&lt;/td&gt;&lt;td&gt;&amp;#8722;.506&lt;/td&gt;&lt;td&gt;&amp;#8722;.725&lt;/td&gt;&lt;td&gt;&amp;#8722;.602&lt;/td&gt;&lt;td /&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>1 <emph>Note</emph>. Upper diagonal is reading; lower diagonal is mathematics. Difficulty = average item difficulty, PBI = point-biserial coefficients, MSI = multi-serial coefficients.</p> <p>To probe more deeply into MSI-PBI differences, 12 items with the highest MSI-PBI differences were identified and investigated for Reading and Mathematics. For Reading, the items with the greatest MSI-PBI differences were the lowest discriminating. For 12 items with the greatest MSI-PBI difference, PBIs varied between.169 and.327. MSIs varied between.296 and.434. These items had MSI-PBI differences of.100 to.218. Reading items 32 and 97 had distractors with a high choice mean and a low PBI. The trace lines visually identified the problem. The high choice mean mimicked the correct option. That typically would indicate removing or replacing it. Reading item 36 was simply a poor item that should not have been in the operational test. Nonetheless, the full-information analysis revealed that one distractor mimicked the correct option. Its removal may reconstitute the item as more discriminating. Reading item 13 had a weak distractor that also had a high choice mean.</p> <p>For the 12 Mathematics items with a large MSI-PBI difference, item difficulty ranged from.147 to.655 with a mean of.343. The mean of the PBI was.147, ranging from.004 to.246. Clearly, these were not very discriminating items. MSI had a mean of.246, ranging from.139 to.347. The mean difference between MSI and PBI was.099, ranging from.065 to.141. Item 8 was a very weak discriminator with a distractor that resembled a correct option. This item did not look salvageable. Item 11 had two statistically discriminating options. Item 55 had a very low discriminating distractor. Item 64 appeared to have to multiple correct choices. Item 78 was simply a very difficult and non-discriminating item.</p> <p>The MSI-PBI difference identified items where the PBI seemed to greatly underestimate item discrimination. In some instances, such items should be retired or revised. In other instances, a single distractor was not discriminating. More infrequently, a distractor competed with the correct option by being selected by many test takers. These items require further scrutiny by the SME committee.</p> <p>Concerning the practical significance of the difference between PBI and MSI, Ebel ([<reflink idref="bib3" id="ref32">3</reflink>]) developed a formula for estimating reliability from a discrimination index. The formula used the sum of discrimination indexes and the sum of item variance (<emph>pq</emph>). Table 4 contains reliability estimates using Ebel's formula for hypothetical tests of three test lengths (<reflink idref="bib50" id="ref33">50</reflink>, 75, and 100 items). For the shortest test, MSI estimates exceeded the PBI estimates sizably. For longer tests, the margin of difference shrank. Nonetheless, reliability estimates using MSI were always higher, and thus more precise.</p> <p>Table 4. Reliability estimates based on PBI and MSI discrimination for three test lengths</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;Reading&lt;/td&gt;&lt;td&gt;50 Items&lt;/td&gt;&lt;td&gt;75 Items&lt;/td&gt;&lt;td&gt;100 Items&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;PBI&lt;/td&gt;&lt;td&gt;.806&lt;/td&gt;&lt;td&gt;.871&lt;/td&gt;&lt;td&gt;.904&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MSI&lt;/td&gt;&lt;td&gt;.854&lt;/td&gt;&lt;td&gt;.903&lt;/td&gt;&lt;td&gt;.928&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Mathematics&lt;/td&gt;&lt;td&gt;50 Items&lt;/td&gt;&lt;td&gt;75 Items&lt;/td&gt;&lt;td&gt;100 Items&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;PBI&lt;/td&gt;&lt;td&gt;.772&lt;/td&gt;&lt;td&gt;.849&lt;/td&gt;&lt;td&gt;.887&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;MSI&lt;/td&gt;&lt;td&gt;.828&lt;/td&gt;&lt;td&gt;.886&lt;/td&gt;&lt;td&gt;.915&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <hd id="AN0152008618-18">Conclusions</hd> <p>MSI accounts for more test score variance than PBI, as expected since every item has options with different selection rates and discrimination. MSI is a more precise estimate of item discrimination because it takes into account the differential performance of the team of distractors (accounts for more variance). As noted previously, the PBI is a simplification of item discrimination that obscures the important information contained in the team of distractors.</p> <p>The recommendation is that the MSI be used to evaluate how well an item discriminates. Where the MSI-PBI difference is great, this clue should lead to further understanding about the dynamics of distractor performance. Distractor/total test score PBI, choice mean, and trace line provide a useful combination of information for evaluating how well each distractor functions.</p> <hd id="AN0152008618-19">Study 2: Removing poorly performing distractors</hd> <p>In study 1, we examined the differences between MSI and PBI for 200 test items. In study 2, we investigated the effects of eliminating distractors from scoring. Would item discrimination improve? If so, under what circumstances?</p> <hd id="AN0152008618-20">Method</hd> <p>To assist in completing this study, a typology was created that would classify types of distractor problems. Table 5 contains the typology, with associated figures containing examples.</p> <p></p> <ulist> <item> <emph>Highly discriminating distractors</emph> have a high negative PBI with the total score. Generally, a family of three highly discriminating distractors have similar PBIs. Typically, the criterion is a PBI below −.10 (subjectively determined). These are well-performing distractors (Figure 2).</item> <p></p> <item> <emph>Low frequency distractors</emph> have often been targeted as less useful due to the fact that the presence of the distractor increases the length of the item and reading time and, therefore, administration time, with little information relevant to score interpretation and use. In this study, we used the response rate of.05 as a criterion for identifying low frequency distractors.</item> <p></p> <item> <emph>Low PBI with the total test score</emph> generally results in a relatively flat trace line. Generally, these distractors had PBIs that were less than −.10. Figure 3 contains a distractor D that fails to discriminate with a relatively flat trace line. Its removal was hypothesized to improve item discrimination.</item> <p></p> <item> <emph>A distractor with a non-monotonic trace line</emph> often emerges as an inverted U-shape. This would appear typically as monotonically increasing in the lower quintiles and decreasing in upper quintiles. Like a flat trace line, this inverted U also has low distractor discrimination. The choice means of the quintiles verify this kind of trace line. Figure 4 contains a good example of this kind of distractor. Distractor A rises in the lower quintiles as if it was a correct option and then falls in the upper quintiles as if it was a plausible distractor.</item> <p></p> <item> <emph>A distractor with a high choice mean</emph> may or may not be discriminating. If it was positively correlated with the total score, it would appear to be a key error or item-writing flaw. This is a rare type of distractor (in test development programs with structured item development and review procedures). Figure 5 contains an example of this kind of distractor (see distractors A and C). Its removal was hypothesized to improve MSI.</item> </ulist> <p>Table 5. A typology for classifying distractor performance</p> <p> <ephtml> &lt;table&gt;&lt;thead&gt;&lt;tr&gt;&lt;td&gt;Type&lt;/td&gt;&lt;td&gt;Figure&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;High negative PBI with total test score&lt;/td&gt;&lt;td&gt;2&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Low frequency of selection&lt;/td&gt;&lt;td&gt;3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Low PBI with total test score&lt;/td&gt;&lt;td&gt;3&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Non-monotonic PBI &amp;#8211; differential discrimination&lt;/td&gt;&lt;td&gt;4&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;High choice mean &amp;#8211; competing with the correct option&lt;/td&gt;&lt;td&gt;5&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>PHOTO (COLOR): Figure 3. A flat trace line with poorly performing distractor (B) (reading item 92)</p> <p>PHOTO (COLOR): Figure 4. Non-monotonic distractor with inverted U-shape (mathematics item 25)</p> <p>PHOTO (COLOR): Figure 5. High choice mean distractors (A, C) (reading item 10)</p> <p>Items were classified using the typology. For each poorly performing distractor, responses to that distractor were coded as missing. This experimental simulation used a method described by Guo, Zu, and Kyllonen ([<reflink idref="bib5" id="ref34">5</reflink>]). After the poorly performing distractor had its responses omitted, MSI was recomputed and labeled MSIR. The difference between MSIR and MSI represented an estimate of the effect of removing the poorly performing distractor. This recoding was not ideal. Experimental removal of a distractor and retesting the items in an operational test or field test would be ideal.</p> <p>One threat to generalizability of the findings is that the sub-sample containing removed student responses to poorly performing distractors might bias the results due to a change in the character of the sample. MSIR was computed on a smaller sample of item responses where the omitted distractor was not selected by test takers. However, the total score of the remaining test takers of the sub-sample was based on all items and all item responses. Results should reflect a variety of circumstances and small and large differences between MSI and MSIR. The criterion measure for this study was MSIR-MSI.</p> <hd id="AN0152008618-21">Results and discussion</hd> <p>Although these two tests were operational, four Reading and 16 Mathematics items had low MSIs. These 20 items were removed from further consideration because none of these items seemed reparable by removing a poorly performing distractor.</p> <hd id="AN0152008618-22">High-functioning set of distractors</hd> <p>On the Reading test for 78 of the remaining 108 items (72%), all three distractors discriminated as intended. These items had a mean MSI of.446 with a standard deviation of.063. The Mathematics test had 44 of 64 remaining items (69%) with all three distractors discriminating as intended. The mean MSI was.432 (<emph>SD</emph> =.070). Ideally, an operational test should have all items high functioning. From these results, there was room for improvement.</p> <hd id="AN0152008618-23">Low-frequency distractors</hd> <p>Using a conventional.05 value for identifying low-frequency distractors, Reading had 42 items and Mathematics had only eight items. These findings were interesting. Why would the Mathematics items have so few low-frequency distractors? For the most part, low-frequency distractors had monotonically decreasing trace lines, high discrimination, and low choice means. In other words, low-frequency distractors modeled the characteristics of high-functioning distractors. The difference being the frequency with which the distractor was chosen. Removing a low-frequency, highly discriminating distractor had a slight negative effect on MSI. For example, Reading item 5 had one distractor with low frequency, where MSI was.354 and MSIR was.335. This distractor behaved like other highly discriminating distractors; therefore, its removal had a slightly negative effect on discrimination.</p> <p>For a low-frequency distractor that was not discriminating very well, its removal had a small effect because the number of test takers for that distractor was very small. This was an important finding, because low-frequency distractors are often advocated for removal. This recommendation would be appropriate if the objective was to remove distractors that increase reading and reduce administration time. On the other hand, the removal of a discriminating low-frequency distractor would not be justified if maintaining item discrimination was the objective.</p> <hd id="AN0152008618-24">Flat trace-line distractors</hd> <p>For Reading, 21 items had low-discriminating distractors (flat trace lines). The mean MSI was.362 and the mean MSIR was.399. The difference (.037) represented the improvement in discrimination as a consequence of removing a low discriminating distractor. For Mathematics, 16 items had low-discriminating distractors. The mean MSI was.389 and the mean MSIR was.419. The difference of.020 was less than the MSIR-MSI difference for Reading low-discriminating distractors. With low-discriminating distractors, removing one poorly performing distractors improved item discrimination.</p> <hd id="AN0152008618-25">Non-monotonic distractors</hd> <p>The five Reading items identified as having distractors with non-monotonic trace lines had a mean MSI of.371 and a mean MSIR of.420. The MSIR-MSI difference was.049. For Mathematics, 11 items were identified with non-monotonic distractors. The mean MSI was.332 and the mean MSIR was.363. The MSIR-MSI difference of.031 was less than the MSIR-MSI difference for Reading. These were very small samples. However, their identification and replacement made a marked difference for Reading and a somewhat smaller difference for Mathematics.</p> <hd id="AN0152008618-26">High choice-mean distractors</hd> <p>Only four Reading items had distractors with high choice means. These items typically had low discrimination and the high choice-mean distractors tended to have non-monotonic trace lines (as seen in Figure 5). High choice means were easily identified by viewing trace lines (see option A). For the four Reading items, the mean MSI was.366 and the mean MSIR was.455. The difference of.089 was the largest observed. Only one Mathematics items had a high choice mean distractor. MSI was.314 and MSIR was.363, with a difference of.049.</p> <p>In this experiment, the low-frequency distractors were discovered to have reasons to retain them and also reasons to reject them, above and beyond the low frequency of selection. For low discriminating distractors, the shape and location of trace lines were informative, as was examination of the choice means. In this experiment, items with distractors with flat trace lines, non-monotonic trace lines, or high choice means, benefited from removal of such distractors.</p> <hd id="AN0152008618-27">Summary and recommendations</hd> <p>Our analysis of item discrimination started with the concept of reciprocity between the discrimination of the correct option and the discrimination of a team of distractors. Drawing from the concept of reciprocity, MSI is more sensitive than PBI as a discrimination index. Trace lines and choice means provide additional insight into how well a distractor functions. In the second study, a typology of distractor performance was created for classifying distractors. Through an experiment removing poorly performing distractors, removal had positive effects on discrimination. The most interesting finding was the role of low-frequency distractors. Generally, if low-frequency distractors have high discrimination, their removal is not warranted. If shortening the test is more important, at the risk of a slight decrease in item discrimination, the traditional advice to remove low-frequency distractors is justifiable.</p> <p>A reviewer asked if criteria were available for PBI or MSI regarding item quality. This is a complex question and we know of no item discrimination values that justify decisions across all types of tests. Such decisions are often made based on relative information across items within the test. For example, a highly homogenous test of restricted content or cognitive skills, such as a mathematics test of algebra alone, item discrimination (PBI) values should be high and relatively uniform, such as around.40. For a certification test covering broad content standards and multiple cognitive skills, item discrimination values could be much lower, such as.20 or lower.</p> <p>If we remove distractors, test length and administration time could be shortened. Or, with a shorter test, we might add new streamlined items that increase content-related validity evidence. These actions potentially improve the validity of score interpretations and reduce random error. In all testing contexts, this is a positive outcome, as one common threat to validity is limited content coverage.</p> <p>A SME committee can use full-information item analysis to study items that are not performing as well as expected. Standard statistical criteria are not sufficient to reject an item or replace a distractor. However, the SME committee should consider all available information before deciding what should be done to an item that fails to perform as expected. Full-information item analysis recommendations include:</p> <p></p> <ulist> <item> Use MSI to measure item discrimination. It is more sensitive and comprehensive than PBI.</item> <p></p> <item> Use PBI to measure distractor discrimination.</item> <p></p> <item> Use trace lines to augment the study of distractor discrimination.</item> <p></p> <item> Use choice means to further study distractor discrimination.</item> <p></p> <item> Use the distractor performance typology (Table 5) to organize results of the item analysis for poorly performing distractors.</item> <p></p> <item> Do not remove low-frequency distractors if their trace lines and PBIs are negative. However, removing low-frequency distractors does shorten a test and reduce administration time without significantly lowering discrimination.</item> <p></p> <item> After a thorough evaluation of distractor performance for an item, the SME committee can decide whether to keep the set of distractors, revise and replace poorly performing distractors, or remove them from the item.</item> </ulist> <p>Item analysis is an essential practice in test development. Full-information item analysis provides a complete picture of the performance of test items. Furthermore, our experience with full-information item analysis reminds us of the role of options in a MC item, as each option should contribute to item quality. Furthermore, the number of distractors should be tailored to secure item quality rather than be fixed for each item as a rule (based on item-specifications or item-writer instructions). Full-information item analysis both equips the test developer with a total view of item performance and frees them from arbitrary rules of standardized item-writing guidelines.</p> <hd id="AN0152008618-28">Disclosure statement</hd> <p>No potential conflict of interest was reported by the author(s).</p> <ref id="AN0152008618-29"> <title> References </title> <blist> <bibl id="bib1" idref="ref27" type="bt">1</bibl> <bibtext> Attali, Y., &amp; Fraenkel, T. (2000). The point-biserial as a discrimination index for distractors in multiple-choice items: Deficiencies in usage and an alternative. Journal of Educational Measurement, 37 (1), 77 – 86. doi: 10.1111/j.1745-3984.2000.tb01077.x</bibtext> </blist> <blist> <bibl id="bib2" idref="ref15" type="bt">2</bibl> <bibtext> Crocker, L., &amp; Algina, J. (1986). Introduction to classical and modern test theory. New York, NY: Harcourt Brace Jovanovich.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref16" type="bt">3</bibl> <bibtext> Ebel, R. L. (1967). The relationship of item discrimination to test reliability. Journal of Educational Measurement, 4, 125 – 128. doi: 10.1111/j.1745-3984.1967.tb00579.x</bibtext> </blist> <blist> <bibl id="bib4" idref="ref4" type="bt">4</bibl> <bibtext> Gierl, M. J., Bulut, O., Guo, Q., &amp; Zhang, X. (2017). Developing, analyzing, and using distractors for multiple-choice tests in education: A comprehensive review. Review of Educational Research, 87 (6), 1082 – 1116. doi: 10.3102/0034654317726529</bibtext> </blist> <blist> <bibl id="bib5" idref="ref34" type="bt">5</bibl> <bibtext> Guo, H., Zu, J., &amp; Kyllonen, P. (2018). A simulation-based method for finding the optimal number of options for multiple-choice items on a test (Research Report No. RR-18-22). Educational Testing Service. doi: 10.1002/ets2.12209</bibtext> </blist> <blist> <bibl id="bib6" idref="ref2" type="bt">6</bibl> <bibtext> Haladyna, T. M. (2016). Item analysis for selected-response test items. In S. Lane, M. R. Raymond, &amp; T. M. Haladyna (Eds.), Handbook of test development (pp. 392 – 409). New York, NY: Routledge.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref13" type="bt">7</bibl> <bibtext> Haladyna, T. M., &amp; Downing, S. M. (1993). How many options is enough for a multiple-choice test item. Educational and Psychological Measurement, 53 (4), 999 – 1010. doi: 10.1177/0013164493053004013</bibtext> </blist> <blist> <bibl id="bib8" idref="ref5" type="bt">8</bibl> <bibtext> Haladyna, T. M., &amp; Rodriguez, M. C. (2013). Developing and validating test items. New York, NY: Routledge.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref10" type="bt">9</bibl> <bibtext> Haladyna, T. M., Rodriguez, M. C., &amp; Stevens, C. (2019). Are multiple-choice items too fat? Applied Measurement in Education, 32 (4), 350 – 364. doi: 10.1080/08957347.2019.1660348</bibtext> </blist> <blist> <bibtext> Levine, M. V., &amp; Drasgow, F. (1983). The relation between incorrect option choice and estimated ability. Educational and Psychological Measurement, 43 (675), 685. doi: 10.1177/001316448304300301</bibtext> </blist> <blist> <bibtext> Livingston, S. (2006). Item analysis. In S. M. Downing &amp; T. M. Haladyna (Eds.), Handbook of test development (1st ed., pp. 421 – 441). New York, NY: Routledge.</bibtext> </blist> <blist> <bibtext> Metsämuuronen, J. (2016). Item-total correlation as the cause for the underestimation of the alpha estimate for the reliability of the scale. Global Journal for Research Analysis, 5, 471 – 477.</bibtext> </blist> <blist> <bibtext> Metsämuuronen, J. (2018a). Generalized discrimination index and its connection to the latent item difficulty—Some impurities in proportion of correct answers (p) as an estimator of the latent item difficulty.</bibtext> </blist> <blist> <bibtext> Metsämuuronen, J. (2018b). Essentials of visual diagnosis of test items—Logical and pathological patterns in items to be detected. doi: 10.13140/RG.2.2.13950.23364</bibtext> </blist> <blist> <bibtext> Olsson, U., Drasgow, F., &amp; Dorans, N. J. (1982). The polyserial correlation coefficient. Psychometrika, 47 (3), 337 – 347. doi: 10.1007/BF02294164</bibtext> </blist> <blist> <bibtext> Raymond, M. R., Stevens, C., &amp; Bucak, S. D. (2019). The optimal number of options for multiple-choice questions on high-stakes tests: Application of a revised index for detecting nonfunctional distractors. Advances in Health Sciences Education, 24 (1), 141 – 150. doi: 10.1007/s10459-018-9855-9</bibtext> </blist> <blist> <bibtext> Rodriguez, M. C. (2016). Selected-response item development. In S. Lane, M. Raymond, &amp; T. M. Haladyna (Eds.), Handbook of test development (2nd ed., pp. 259 – 273). New York, NY: Routledge.</bibtext> </blist> <blist> <bibtext> Samejima, F. (1994). Non parametric estimation of plausibility functions of distractors of vocabulary test items. Applied Psychological Measurement, 18, 35 – 51. doi: 10.1177/014662169401800104</bibtext> </blist> <blist> <bibtext> Thissen, D., Steinberg, L., &amp; Fitzpatrick, A. R. (1989). Multiple-choice models: The distractors are also part of the item. Journal of Educational Measurement, 26 (2), 161 – 176. doi: 10.1111/j.1745-3984.1989.tb00326.x</bibtext> </blist> <blist> <bibtext> Thissen, D. M. (1976). Information in wrong responses to the Raven progressive matrices. Journal of Educational Measurement, 14, 201 – 214. doi: 10.1111/j.1745-3984.1976.tb00011.x</bibtext> </blist> <blist> <bibtext> Wainer, H. (1989). The future of item analysis. Journal of Educational Measurement, 26 (2), 191 – 208. doi: 10.1111/j.1745-3984.1989.tb00328.x</bibtext> </blist> <blist> <bibtext> Wang, W.-C. (2000). Factorial modeling of differential distractor functioning in multiple-choice test items. Journal of Applied Measurement, 1 (3), 238 – 256.</bibtext> </blist> </ref> <aug> <p>By Thomas M. Haladyna and Michael C. Rodriguez</p> <p>Reported by Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib17" firstref="ref1"></nolink> <nolink nlid="nl2" bibid="bib11" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib20" firstref="ref7"></nolink> <nolink nlid="nl4" bibid="bib19" firstref="ref8"></nolink> <nolink nlid="nl5" bibid="bib16" firstref="ref14"></nolink> <nolink nlid="nl6" bibid="bib12" firstref="ref17"></nolink> <nolink nlid="nl7" bibid="bib13" firstref="ref19"></nolink> <nolink nlid="nl8" bibid="bib15" firstref="ref20"></nolink> <nolink nlid="nl9" bibid="bib10" firstref="ref22"></nolink> <nolink nlid="nl10" bibid="bib18" firstref="ref25"></nolink> <nolink nlid="nl11" bibid="bib22" firstref="ref26"></nolink> <nolink nlid="nl12" bibid="bib21" firstref="ref28"></nolink> <nolink nlid="nl13" bibid="bib14" firstref="ref29"></nolink> <nolink nlid="nl14" bibid="bib50" firstref="ref33"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ1309544 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Using Full-Information Item Analysis to Improve Item Quality – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Haladyna%2C+Thomas+M%2E%22">Haladyna, Thomas M.</searchLink><br /><searchLink fieldCode="AR" term="%22Rodriguez%2C+Michael+C%2E%22">Rodriguez, Michael C.</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22Educational+Assessment%22"><i>Educational Assessment</i></searchLink>. 2021 26(3):198-211. – Name: Avail Label: Availability Group: Avail Data: Routledge. Available from: Taylor & Francis, Ltd. 530 Walnut Street Suite 850, Philadelphia, PA 19106. Tel: 800-354-1420; Tel: 215-625-8900; Fax: 215-207-0050; Web site: http://www.tandf.co.uk/journals – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 14 – Name: DatePubCY Label: Publication Date Group: Date Data: 2021 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Descriptive – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Grade+6%22">Grade 6</searchLink><br /><searchLink fieldCode="EL" term="%22Intermediate+Grades%22">Intermediate Grades</searchLink><br /><searchLink fieldCode="EL" term="%22Middle+Schools%22">Middle Schools</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Analysis%22">Item Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Reading+Tests%22">Reading Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Mathematics+Tests%22">Mathematics Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Grade+6%22">Grade 6</searchLink><br /><searchLink fieldCode="DE" term="%22Difficulty+Level%22">Difficulty Level</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1080/10627197.2021.1946390 – Name: ISSN Label: ISSN Group: ISSN Data: 1062-7197 – Name: Abstract Label: Abstract Group: Ab Data: Full-information item analysis provides item developers and reviewers comprehensive empirical evidence of item quality, including option response frequency, point-biserial index (PBI) for distractors, mean-scores of respondents selecting each option, and option trace lines. The multi-serial index (MSI) is introduced as a more informative item-total correlation, accounting for variable distractor performance. The overall item PBI is empirically compared to the MSI. For items from an operational mathematics and reading test, poorly performing distractors are systematically removed to recompute the MSI, indicating improvements in item quality. Case studies for specific items with different characteristics are described to illustrate a variety of outcomes, focused on improving item discrimination. Full-information item analyses are presented for each case study item, providing clear examples of interpretation and use of item analyses. A summary of recommendations for item analysts is provided. – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: DateEntry Label: Entry Date Group: Date Data: 2021 – Name: AN Label: Accession Number Group: ID Data: EJ1309544 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1309544 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/10627197.2021.1946390 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 14 StartPage: 198 Subjects: – SubjectFull: Test Items Type: general – SubjectFull: Item Analysis Type: general – SubjectFull: Reading Tests Type: general – SubjectFull: Mathematics Tests Type: general – SubjectFull: Grade 6 Type: general – SubjectFull: Difficulty Level Type: general Titles: – TitleFull: Using Full-Information Item Analysis to Improve Item Quality Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Haladyna, Thomas M. – PersonEntity: Name: NameFull: Rodriguez, Michael C. IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2021 Identifiers: – Type: issn-print Value: 1062-7197 Numbering: – Type: volume Value: 26 – Type: issue Value: 3 Titles: – TitleFull: Educational Assessment Type: main |
| ResultId | 1 |