A Generalized Logistic Regression Procedure to Detect Differential Item Functioning among Multiple Groups
Saved in:
| Title: | A Generalized Logistic Regression Procedure to Detect Differential Item Functioning among Multiple Groups |
|---|---|
| Language: | English |
| Authors: | Magis, David, Raiche, Gilles, Beland, Sebastien, Gerard, Paul |
| Source: | International Journal of Testing. 2011 11(4):365-386. |
| Availability: | Routledge. Available from: Taylor & Francis, Ltd. 325 Chestnut Street Suite 800, Philadelphia, PA 19106. Tel: 800-354-1420; Fax: 215-625-2940; Web site: http://www.tandf.co.uk/journals |
| Peer Reviewed: | Y |
| Page Count: | 22 |
| Publication Date: | 2011 |
| Document Type: | Journal Articles Reports - Descriptive |
| Education Level: | Higher Education |
| Descriptors: | Language Skills, Identification, Foreign Countries, Evaluation Methods, Comparative Analysis, English (Second Language), French, Models, Tests, Computation, Test Bias, College Students, Colleges |
| Geographic Terms: | Canada |
| DOI: | 10.1080/15305058.2011.602810 |
| ISSN: | 1530-5058 |
| Abstract: | We present an extension of the logistic regression procedure to identify dichotomous differential item functioning (DIF) in the presence of more than two groups of respondents. Starting from the usual framework of a single focal group, we propose a general approach to estimate the item response functions in each group and to test for the presence of uniform DIF, nonuniform DIF, or both. This generalized procedure is compared to other existing DIF methods for multiple groups with a real data set on language skill assessment. Emphasis is put on the flexibility, completeness, and computational easiness of the generalized method. (Contains 5 tables and 2 figures.) |
| Abstractor: | As Provided |
| Number of References: | 52 |
| Entry Date: | 2011 |
| Accession Number: | EJ946932 |
| Database: | ERIC |
|
Full text is not displayed to guests.
Login for full access.
|
|
| FullText | Links: – Type: pdflink Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwHVUk4Pa9WKieCgKkvEyVi1AAAA4TCB3gYJKoZIhvcNAQcGoIHQMIHNAgEAMIHHBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDJNH1Q2b27UXDo7OtAIBEICBmcbNZS0eNwsOkN62Vxv_rdZ7d7acs7aa81y_ZOvJ3Fddd7u_ijPblmnLu5FSoYoouN74RTnhKcU8Hx15bW8Y4o-m6Chx-elKl0uSQU3E9mwfq7EeKvQLmmVu0knDU8lbNHJkUQqKsnrd2S6XfgAcBfQYXtgRNy7I3Wk_0GA3e372xNwb5QkNBv4uHUOjt-hIK2TidFxFoVL7aA== Text: Availability: 1 Value: <anid>AN0067043770;k0301oct.11;2019Mar11.13:18;v2.2.500</anid> <title id="AN0067043770-1">A Generalized Logistic Regression Procedure to Detect Differential Item Functioning Among Multiple Groups. </title> <p>We present an extension of the logistic regression procedure to identify dichotomous differential item functioning (DIF) in the presence of more than two groups of respondents. Starting from the usual framework of a single focal group, we propose a general approach to estimate the item response functions in each group and to test for the presence of uniform DIF, nonuniform DIF, or both. This generalized procedure is compared to other existing DIF methods for multiple groups with a real data set on language skill assessment. Emphasis is put on the flexibility, completeness, and computational easiness of the generalized method.</p> <p>Keywords: differential item functioning; logistic regression; multiple groups</p> <hd id="AN0067043770-2">INTRODUCTION</hd> <p>An important research area in psychometrics is the identification of differential item functioning (DIF). In the context of dichotomous responses, an item is said to function differently (or shortly, to be DIF) if respondents with the same ability level but from different groups of examinees have different probabilities of answering this item correctly. DIF is an unwilling phenomenon that can lead to biased measurements of ability (Ackerman, [<reflink idref="bib2" id="ref1">2</reflink>]). The identification of DIF items is therefore a crucial issue for assessing valid psychometric and educational measurement. In this article, we restrict our purpose to dichotomously scored items.</p> <p>The DIF detection can be classified according to two main factors: the methodological approach—either based on item response theory (IRT) models or not—and the type of DIF effect—uniform DIF or nonuniform DIF. Methods based on IRT require the fitting of item response models, for instance the logistic model with one or more parameters. Non-IRT methods, on the other hand, are based on statistical methods for assessing the presence of DIF based on observed test scores and without requiring an IRT solution. Moreover, an item exhibits uniform DIF if the interaction between the item responses and the groups of respondents is independent of the ability level. Nonuniform (or crossing) DIF is characterized by an item-group interaction, which can vary along the ability scale (Clauser &amp; Mazor, [<reflink idref="bib10" id="ref2">10</reflink>]; Hanson, [<reflink idref="bib16" id="ref3">16</reflink>]).</p> <p>The most known IRT-based methods are the Lord's χ<sups>2</sups> test (Lord, [<reflink idref="bib27" id="ref4">27</reflink>]), the Raju's area method (Raju, [<reflink idref="bib38" id="ref5">38</reflink>]) and the likelihood-ratio test (Thissen, Steinberg, &amp; Wainer, [<reflink idref="bib46" id="ref6">46</reflink>]). Among the non-IRT methods, the Mantel-Haenszel method (Holland &amp; Thayer, [<reflink idref="bib20" id="ref7">20</reflink>]), the SIBTEST method (Shealy &amp; Stout, [<reflink idref="bib43" id="ref8">43</reflink>]), and the logistic regression procedure (Swaminathan &amp; Rogers, [<reflink idref="bib45" id="ref9">45</reflink>]) are the most commonly used methods. For a recent review of the DIF detection methods, see Clauser and Mazor ([<reflink idref="bib10" id="ref10">10</reflink>]), Osterlind and Everson ([<reflink idref="bib33" id="ref11">33</reflink>]), Penfield and Camilli ([<reflink idref="bib35" id="ref12">35</reflink>]) and Penfield and Lam ([<reflink idref="bib34" id="ref13">34</reflink>]).</p> <p>The aforementioned methods are designed to compare two groups of respondents: the reference group and the focal group. However, practical situations might require the detection of items that function differently between three or more groups. For instance, one could be interested in comparing item performance between different classrooms within a school or between schools within a country. At an international level, DIF investigation could be performed in international studies or surveys such as the Program for International Student Assessment (PISA) or Trends in International Mathematics and Science Study (TIMSS)—large-scale assessment studies involving much more than two groups (countries) of respondents. Another possible application of multiple groups DIF detection methods is the comparison of item performance from a common test, administered repeatedly across years. For example, entrance or admission tests that are administered every year to students entering into higher studies could be screened for DIF across the years of administration. The practical example analyzed in this article is from that kind.</p> <p>Until now, however, the extension of DIF detection methods to multiple groups has received very little attention. When dealing with more than one focal group, the common practical approach is to perform pairwise comparisons between the reference group and each focal group (see e.g., Angoff &amp; Sharon, [<reflink idref="bib6" id="ref14">6</reflink>]; Ellis &amp; Kimmel, [<reflink idref="bib12" id="ref15">12</reflink>]; Schmitt &amp; Dorans, [<reflink idref="bib41" id="ref16">41</reflink>]; Zwick &amp; Ercikan, [<reflink idref="bib52" id="ref17">52</reflink>]). But because of the simultaneous testing of multiple hypotheses, this approach provokes a Type I error inflation, and ad-hoc methods must be used such as Bonferroni correction of the nominal significance level.</p> <p>To date, two methods have been specifically extended for testing DIF among multiple groups: the Mantel-Haenzel method (Penfield, [<reflink idref="bib34" id="ref18">34</reflink>]; see also Fidalgo &amp; Madeira, [<reflink idref="bib13" id="ref19">13</reflink>]; Fidalgo &amp; Scalon, [<reflink idref="bib14" id="ref20">14</reflink>]) and Lord's χ<sups>2</sups> test (Kim, Cohen, &amp; Park, [<reflink idref="bib25" id="ref21">25</reflink>]). The so-called generalized Mantel-Haenszel method focuses on uniform DIF among multiple groups, while the generalized Lord's test can detect both types of effects. To our knowledge, the logistic regression procedure has not been formally extended to more than two groups. Millsap and Everson (1993, p. 305) mentioned that this method can be generalized to more than two groups, and Van den Noortgate and De Boeck (2005) introduced logistic mixed models that can be considered with several groups of respondents. Kanjee ([<reflink idref="bib24" id="ref22">24</reflink>]) recommended merging all focal groups in a single one before applying the usual logistic regression procedure. This approach avoids pairwise comparisons and subsequently controls for Type I error inflation. In addition, the power of the method increases as the "pooled" focal group has larger size. However, by merging all focal groups into a single one, one distorts the structure of the data set, which can lead to the misidentification of items as DIF or non-DIF. In particular, items that exhibit some DIF effect between one focal group and the reference group could become undetected when all focal groups are merged together, especially if the sample size of the focal group is small regarding other focal groups.</p> <p>The purpose of this article is to present an extension of the logistic regression procedure in the case of multiple focal groups. Emphasis is put on the usefulness and easiness of the method as well as its flexibility in modelling DIF effects and testing within subgroups of respondents. For each group of examinees a response probability curve is fitted, and DIF can be statistically assessed by an appropriate comparison of group-specific parameters. No merging of the focal groups is required, and both types of DIF effect can be tested. This constitutes an improvement of the generalized Mantel-Haenszel method, and it can be seen as the non-IRT counterpart of the generalized Lord's test with the two-parameter logistic (2PL) model. To keep the DIF terminology consistent, we further refer to this extension as the generalized logistic regression procedure for DIF detection.</p> <p>The article is organized as follows. In Section 2 we present the generalized logistic regression procedure by introducing the notations of the logistic model and highlighting the relationships between the model parameters and the DIF effects. Two methods of null hypothesis testing (i.e., non-DIF hypothesis) are discussed in Section 3: the Wald test and the likelihood-ratio tests. These are well-established methods (Agresti, [<reflink idref="bib5" id="ref23">5</reflink>]) and their application in this DIF framework is displayed in detail. The generalized logistic regression procedure is then illustrated in Section 4 by analyzing a real example about language skill assessment in Quebec colleges. The results from this method are presented and compared to those from the generalized Mantel-Haenszel method and generalized Lord's test.</p> <hd id="AN0067043770-3">GENERALIZED LOGISTIC REGRESSION</hd> <p>The generalized logistic regression is an extension of Swaminathan and Rogers' ([<reflink idref="bib45" id="ref24">45</reflink>]) approach in the usual case of DIF between two groups. We focus on one item of interest, and all other tests items are assumed to be DIF free (this is sometimes referred to as the all-other anchor items setting).</p> <p>Let <emph>π<subs>ig</subs></emph> be the probability that respondent <emph>i</emph> from group <emph>g</emph> answers the item correctly, where <emph>g</emph> = 0 for the reference group, and <emph>g</emph> = 1, 2, ..., <emph>F</emph> respectively for the first, second, ..., last focal group. Set moreover <emph>S<subs>i</subs></emph> as the test score of respondent <emph>i</emph>, which acts as the matching variable and as a proxy for respondent's ability. With these notations, the generalized logistic regression DIF model takes the following form:</p> <p>Graph</p> <p>where <emph>α</emph><subs>g</subs> and <emph>β</emph><subs>g</subs> are the intercept and the slope parameters of group <emph>g</emph> (<emph>g</emph> = 0, 1, ..., <emph>F</emph>), respectively.</p> <p>A completely equivalent form of model (<reflink idref="bib1" id="ref25">1</reflink>) is obtained by setting common, group-independent intercept and slope parameters, so that <emph>α</emph><subs>g</subs> and <emph>β</emph><subs>g</subs> are group-specific parameters:</p> <p>Graph</p> <p>where <emph>α</emph> and <emph>β</emph> are the common intercept and slope parameters for all groups. Because model (<reflink idref="bib2" id="ref26">2</reflink>) is overparametrized with respect to model (<reflink idref="bib1" id="ref27">1</reflink>), group-specific parameters must be constrained to avoid identification issues. Setting all reference group parameters to zero is the most obvious constraint; that is, <emph>α</emph><subs>0</subs> = <emph>β</emph><subs>0</subs> = 0. Hence, model (<reflink idref="bib2" id="ref28">2</reflink>) can be rewritten as</p> <p>Graph</p> <p>The intercept and the slope parameters are respectively equal to <emph>α</emph> and <emph>β</emph> in the reference group and to (<emph>α</emph> + <emph>α<subs>g</subs></emph>) and (<emph>β</emph> +<emph>β<subs>g</subs></emph>) in the focal group <emph>g</emph> (<emph>g</emph> = 1, ..., <emph>F</emph>). In the following we make use of the parameterization of model (<reflink idref="bib3" id="ref29">3</reflink>).</p> <p>The tested item exhibits DIF if the response probability <emph>π<subs>ig</subs></emph> varies across the groups of examinees, or equivalently, if there is some interaction between the item responses and the group membership. According to model (<reflink idref="bib3" id="ref30">3</reflink>), this occurs if at least one of the group parameters <emph>α<subs>g</subs></emph> and <emph>β<subs>g</subs></emph> is different from zero. If all group-specific parameters equal zero, no DIF effect is present. Furthermore, the presence of nonuniform DIF is assessed by a significant difference in the slopes of the logistic response curves—when at least one slope parameter <emph>β<subs>g</subs></emph> is different from zero, whatever the values of the intercept parameters. Finally, uniform DIF is present if the item-group interaction does not depend on the matching variable—if at least one intercept parameter <emph>α<subs>g</subs></emph> is different from zero, given that all slope parameters <emph>β<subs>g</subs></emph> are equal to zero. In sum, three types of DIF effects can be tested: uniform DIF (<emph>UDIF</emph>), nonuniform DIF (<emph>NUDIF</emph>) and both types of DIF effects altogether (<emph>DIF</emph>). Each framework is characterized by the following null hypotheses:</p> <p>Graph</p> <p>Graph</p> <p>Graph</p> <p>The alternative hypotheses are such that at least one of the tested parameters in the null hypothesis is different from zero. The null hypothesis (<reflink idref="bib6" id="ref31">6</reflink>) states that the focal group-specific intercept parameters <emph>α</emph><subs>1</subs>, ..., <emph>α<subs>F</subs></emph> are all equal to zero, given that all group-specific slope parameters <emph>β</emph><subs>1</subs>, ..., <emph>β<subs>F</subs></emph> are equal to zero. This implies that the hypothesis of absence of uniform DIF effect is tested on the basis of the simpler logistic model</p> <p>Graph</p> <p>The null hypotheses (<reflink idref="bib4" id="ref32">4</reflink>) and (<reflink idref="bib5" id="ref33">5</reflink>), however, must be assessed with the full model (<reflink idref="bib3" id="ref34">3</reflink>) since the group-specific slope parameters <emph>β<subs>g</subs></emph> must also be tested.</p> <p>As the number of groups increases, small discrepancies between groups' logistic curves would probably lead to flagging the item as DIF, so it is expected that at least one item will often be identified as DIF using this method. This would be argument against the usual the standard approach of testing the usual null hypotheses of absence of DIF. One solution consists in switching to the Bayesian paradigm, allowing for prior information regarding the potential DIF status of the items and computing posterior probabilities and credible intervals for the model parameters. This approach is nevertheless skipped from this article, as emphasis is put on a direct extension of the usual logistic regression method based on maximum likelihood estimation and testing of the model parameters, as explained in the next section.</p> <hd id="AN0067043770-4">DIF IDENTIFICATION</hd> <p>Statistical assessment of DIF is done as follows. Set first as the vector of model parameters:</p> <p>Graph</p> <p>The vector can be estimated by maximum likelihood estimation (Agresti, [<reflink idref="bib3" id="ref35">3</reflink>]). Let be the maximum likelihood estimate of . There are two reasons for focusing on simple maximum likelihood estimation. First, the estimates are unique (Wedderburn, [<reflink idref="bib49" id="ref36">49</reflink>]), asymptotically multivariate normally distributed (Bock, [<reflink idref="bib7" id="ref37">7</reflink>]; Rao, [<reflink idref="bib39" id="ref38">39</reflink>]) with mean vector and covariance matrix , where is the inverse of Fisher's information matrix (Nelder &amp; Wedderburn, [<reflink idref="bib32" id="ref39">32</reflink>]). In short, . Second, the likelihood equations can be solved with an iterative optimization routine, for instance the Newton-Raphson or Fisher scoring method (Agresti, [<reflink idref="bib4" id="ref40">4</reflink>], p. 94).</p> <p>Under the maximum likelihood framework, the null hypotheses (<reflink idref="bib4" id="ref41">4</reflink>) to (<reflink idref="bib6" id="ref42">6</reflink>) can be tested by several methods. This article restricts to two well-known approaches: the Wald test and the likelihood ratio test.</p> <hd id="AN0067043770-5">Wald Test</hd> <p>The null hypotheses (<reflink idref="bib4" id="ref43">4</reflink>) to (<reflink idref="bib6" id="ref44">6</reflink>) can be written in a common matrix form <emph>H</emph><subs>0</subs>: , where is given by (<reflink idref="bib8" id="ref45">8</reflink>), <bold>0</bold> is a vector of zeros, and <emph><bold>C</bold></emph> is an appropriate contrast matrix. The alternative hypothesis also takes a simple form: <emph>H<subs>A</subs></emph>: . To write the matrix <bold><emph>C</emph></bold> properly in each framework, set <bold>0</bold><emph><subs>n × m</subs></emph> as the <emph>n</emph>-by-<emph>m</emph> matrix of zeros and <bold><emph>I</emph></bold><emph><subs>n</subs></emph> as the identity matrix of dimension <emph>n</emph>. Then, in the <emph>DIF</emph> framework, <bold><emph>C</emph></bold> is the (2<emph>F</emph>)-by-(2<emph>F</emph> + 2) matrix</p> <p>Graph</p> <p>In the <emph>NUDIF</emph> framework, <bold><emph>C</emph></bold> is the <emph>F</emph>-by-(2<emph>F</emph> + 2) matrix</p> <p>Graph</p> <p>and in the <emph>UDIF</emph> framework, <bold><emph>C</emph></bold> is the <emph>F</emph>-by-(<emph>F</emph> + 2) matrix</p> <p>Graph</p> <p>The forms (<reflink idref="bib9" id="ref46">9</reflink>) to (<reflink idref="bib11" id="ref47">11</reflink>) of the matrix <bold><emph>C</emph></bold> are straightforward extensions of the contrast matrices used in the simple situation of a single focal group. For instance, in the <emph>DIF</emph> framework with one focal group, reduces to (<emph>α</emph>, <emph>α</emph><subs>1</subs>, <emph>β</emph>, <emph>β</emph><subs>1</subs>)<emph><sups>T</sups></emph> and the null hypothesis (<reflink idref="bib4" id="ref48">4</reflink>) to <emph>H</emph><subs>0</subs>: <emph>α</emph><subs>1</subs> = <emph>β</emph><subs>1</subs> = 0. The appropriate contrast matrix <bold><emph>C</emph></bold> is then equal to</p> <p>Graph</p> <p>(Swaminathan &amp; Rogers, [<reflink idref="bib45" id="ref49">45</reflink>], p. 365), which is the particular case of (<reflink idref="bib9" id="ref50">9</reflink>) when <emph>F</emph> equals one.</p> <p>The rank of <bold><emph>C</emph></bold> is equal to <emph>p</emph> = <emph>2F</emph> in the <emph>DIF</emph> framework and <emph>p</emph> = <emph>F</emph> in the <emph>UDIF</emph> and <emph>NUDIF</emph> frameworks. Since is asymptotically multivariate normally distributed, the <emph>p</emph>-dimensional vector is also asymptotically multivariate normally distributed, with mean vector and covariance matrix (Johnson &amp; Wichern, [<reflink idref="bib23" id="ref51">23</reflink>], p. 165). It follows that the one-dimensional variable</p> <p>Graph</p> <p>has an asymptotic chi-squared distribution with <emph>p</emph> degrees of freedom (Rao, [<reflink idref="bib39" id="ref52">39</reflink>], p. 188). Thus, under the null hypothesis <emph>H</emph><subs>0</subs>: , the test statistic</p> <p>Graph</p> <p>has an asymptotic chi-squared distribution with <emph>p</emph> degrees of freedom. When the value of <emph>Q</emph> exceeds the corresponding cut-score from the chi-squared distribution, the null hypothesis is rejected and the presence of DIF is statistically assessed. Swaminathan and Rogers ([<reflink idref="bib45" id="ref53">45</reflink>]) referred to the statistic (<reflink idref="bib14" id="ref54">14</reflink>) as <emph>χ</emph><sups>2</sups> (equation 14, p. 365), but we use <emph>Q</emph> instead to avoid confusion with the chi-squared distribution. This method is referred to as the <emph>Wald test</emph> since it is derived form the asymptotic normality of the maximum likelihood estimates of the model parameters (Wald, [<reflink idref="bib48" id="ref55">48</reflink>]). For this reason, <emph>Q</emph> is further referred to as the <emph>Wald statistic</emph>.</p> <p>Another asset of the Wald test is that it can be used to test for DIF among a subset of groups of examinees. This is particularly useful when one wants to determine where the differential functioning comes from. Subtests of groups of examinees can be specified with an appropriate contrast matrix and by using the output of the logistic regression model fitted to all groups of examinees. For instance, under the <emph>DIF</emph> framework, the 4 × (2<emph>F</emph> + 2) contrast matrix</p> <p>Graph</p> <p>can be used to test whether the item functions differently between the reference group and the first two focal groups (assuming that the number of focal groups <emph>F</emph> is at least equal to three).</p> <hd id="AN0067043770-6">Likelihood Ratio Test</hd> <p>The likelihood ratio test compares two nested models: one referring to the null hypothesis and one to the alternative hypothesis. The most suitable model is retained by comparing their maximized likelihood values. This method was introduced by Wilks ([<reflink idref="bib50" id="ref56">50</reflink>]) in the general context of comparing composite hypotheses (see also Agresti, [<reflink idref="bib3" id="ref57">3</reflink>]; McCullagh &amp; Nelder, [<reflink idref="bib29" id="ref58">29</reflink>]).</p> <p>More precisely, let <emph>M</emph><subs>0</subs> and <emph>M</emph><subs>1</subs> be the logistic models, which are used to represent the null and the alternative hypotheses respectively. The model <emph>M</emph><subs>0</subs>, referred to as the <emph>null model</emph>, is equal to</p> <p>Graph</p> <p>while the model <emph>M</emph><subs>1</subs>, called the <emph>alternative model</emph>, is given by</p> <p>Graph</p> <p>The null hypothesis of absence of the tested DIF effect is assessed when the model <emph>M</emph><subs>0</subs> is preferred to the model <emph>M</emph><subs>1</subs>, so the tested hypotheses can be rewritten in this context as: <emph>H</emph><subs>0</subs>: model <emph>M</emph><subs>0</subs> is preferred to model <emph>M</emph><subs>1</subs> vs <emph>H</emph><subs>1</subs>: model <emph>M</emph><subs>1</subs> is preferred to model <emph>M</emph><subs>0</subs>.</p> <p>Once the maximum likelihood parameters estimates available, the corresponding maximized likelihoods are computed, say <emph>L</emph><subs>0</subs> for model <emph>M</emph><subs>0</subs> and <emph>L</emph><subs>1</subs> for model <emph>M</emph><subs>1</subs>. Wilks ([<reflink idref="bib50" id="ref59">50</reflink>]) introduced the lambda statistic :</p> <p>Graph</p> <p>and showed that, under the null hypothesis, this statistic has an asymptotic chi-squared distribution with as many degrees of freedom as the difference in the number of parameters between the models <emph>M</emph><subs>0</subs> and <emph>M</emph><subs>1</subs> (see also Agresti, [<reflink idref="bib3" id="ref60">3</reflink>]). The statistic is often called the <emph>likelihood ratio statistic</emph>. In this framework, large values of indicate the presence of the tested DIF effect.</p> <p>Because there are <emph>F</emph> group-specific intercept and <emph>F</emph> group-specific slope parameters, the degrees of freedom of the asymptotic null distribution of equal 2<emph>F</emph> in the <emph>DIF</emph> framework and <emph>F</emph> in both the <emph>UDIF</emph> and <emph>NUDIF</emph> frameworks. Thus, both the Wald statistic <emph>Q</emph> and the likelihood ratio statistic share the same asymptotic distribution. The tests are therefore asymptotically equivalent for detecting DIF effects among the items and would return the same results with sufficiently large sample sizes.</p> <p>It is important to notice that this test is not directly related to the so-called likelihood ratio test (LRT) method for DIF identification introduced by Thissen, Steinberg, and Wainer (1988). Although the basic idea is common to both approaches (that is, the statistical comparison of two nested models by means of the likelihood ratio statistic), the latter is used with nested IRT models, while in this context we focus on logistic models. The present LRT method serves as a statistical tool for testing for the presence of DIF with multiple-groups logistic models.</p> <hd id="AN0067043770-7">Wald or LRT?</hd> <p>Two methods are available to test for the presence of DIF among multiple groups. Applying both will generally return similar results, especially when the samples are large. This is because of the asymptotic equivalence between the two statistics (Cox &amp; Hinkley, [<reflink idref="bib11" id="ref61">11</reflink>]). However, the likelihood ratio test is more reliable than the Wald test with smaller samples (Agresti, [<reflink idref="bib5" id="ref62">5</reflink>]). This is notably because the Wald test suffers from poor estimation of the standard errors of the parameters with smaller samples, while the likelihood ratio test is unaffected. Note, however, that discrepancy between both tests could also be due to a bad fit of the logistic model to the data, as the consequence of an improper model selection for data analysis, for instance.</p> <p>On the other hand, the Wald test is suitable for performing subtesting between some groups of respondents, as explained previously. The likelihood ratio test, however, cannot perform such specific comparisons. One may therefore recommend to start by comparing the results of both statistical tests, and in case of acceptable agreement between their results, subsequent comparisons could be performed with the Wald test. Great discrepancy between the outputs of the methods indicates that the samples are not large enough to ensure asymptotic validity of the conclusions, and great care should be taken with their interpretation.</p> <hd id="AN0067043770-8">Practical Implementation</hd> <p>The generalized logistic regression procedure can be implemented with any suitable software that performs logistic regression modelling, such as SAS or STATISTICA. For the purpose of this article, however, the method has been implemented within the <emph>R</emph> package <emph>difR</emph> (Magis et al., [<reflink idref="bib28" id="ref63">28</reflink>]) as well as many other DIF methods for two or more than two groups. The results of the data set analysis in the next section were obtained from this implementation. The R code can be obtained freely from the first author.</p> <hd id="AN0067043770-9">AN EXAMPLE</hd> <p>We illustrate the usefulness and flexibility of the generalized logistic regression procedure by analyzing an example about assessment of English, as a second language, skills. After a short description of the data set, we proceed to a complete DIF analysis, first with this method, then with other multiple-groups DIF methods. Although the data set restricts to assessment testing of Quebec students entering into college, this example illustrates how the method can be easily adapted to international testing studies, across times of administration, countries, or both.</p> <hd id="AN0067043770-10">The TCALS-II Data Set</hd> <p>The TCALS-II test is administered to Canadian French-speaking students as prior to entering into college education in Quebec province, to assess their aptitude in English as a second language. The primary goal of this test is to evaluate the English level of the students in order to assign them into classes of appropriate difficulty level (Laurier et al., [<reflink idref="bib26" id="ref64">26</reflink>]; Raîche, [<reflink idref="bib37" id="ref65">37</reflink>]). The test consists of 85 multiple-choice items, divided in eight subgroups, and is identical for all French-speaking colleges of the Quebec province (Canada). In this study we focus on the items for which the students have to answer questions related to the reading of short English texts. There are 15 such items (referred to as items 1 to 15) and we consider this subset of items for further analysis.</p> <p>In order to define the different groups of examinees, we focus on the results of the TCALS-II test from students entering into the College of Outaouais (Gatineau, Quebec, Canada). The groups of examinees are defined by the different years the test was assigned. Four years were selected, respectively 1998, 2000, 2002 and 2004. The year 1998 corresponds to the very first year the test was administered, and is selected as the reference year. The other years of administration (2000, 2002, and 2004) are the focal groups. Although the TCALS-II test was assigned every year from 1998, the data from several years of administration are not available anymore (in particular, years 1999 and 2001), so that we restricted to focal groups made by every two years of administration. Moreover, the generalized logistic regression can handle any number of groups of respondents, but for practical illustrative purposes we restricted to four groups of TCALS-II administration.</p> <p>The sample sizes range from 1277 to 1547 and are relatively large. The average scores to the 15 items range from 10.14 to 10.65, and the standard deviations of these scores range from 3.36 to 3.68. Moreover, the skewness coefficients of the scores are all negative, ranging from –0.60 to –0.67, which indicates an asymmetric distribution of the scores with a larger proportion of high scores than low scores (Raîche, [<reflink idref="bib37" id="ref66">37</reflink>]). This is partly due to the fact that students from Outaouais are more often bilingual than in the others regions of the Quebec province (mainly French speaking), because of the closeness of the Ontario province (mainly English speaking). Finally, Laurier and colleagues (1998) established that the full TCALS-II questionnaire exhibits a high fidelity level, with Cronbach's <emph>α</emph> of 0.96, as well as the unidimensionality of the questionnaire (see also Raîche, [<reflink idref="bib37" id="ref67">37</reflink>], for another dimensionality analysis but with similar conclusions).</p> <hd id="AN0067043770-11">DIF Analysis</hd> <p>Our interest is to discover whether some items perform differently over the successive years of administration. Because the test is identical from year to year, this problem can be investigated by a DIF analysis of these items.</p> <p>We start by an analysis of both types of DIF effect on the whole set of 15 items, using separately the Wald test and the likelihood ratio test. For the Wald test, the model (<reflink idref="bib3" id="ref68">3</reflink>) is fitted and the null hypothesis (<reflink idref="bib4" id="ref69">4</reflink>) is tested by means of the contrast matrix</p> <p>Graph</p> <p>that is, the matrix <bold><emph>C</emph></bold> given by (<reflink idref="bib9" id="ref70">9</reflink>) and with <emph>F</emph> = 4. For the likelihood ratio test, both models (<reflink idref="bib16" id="ref71">16</reflink>) and (<reflink idref="bib17" id="ref72">17</reflink>) are fitted and compared by means of the statistic (<reflink idref="bib18" id="ref73">18</reflink>). To reduce the impact of DIF items into the set of anchor (DIF-free) items, the process of item purification described by Candell and Drasgow ([<reflink idref="bib9" id="ref74">9</reflink>]) is performed within each test. That is, items flagged as DIF are removed from the test score computation and the DIF process is re-run; this step is repeated until two successive iterations return the same classification of items as DIF or non-DIF (see also Clauser &amp; Mazor, [<reflink idref="bib10" id="ref75">10</reflink>]). Significance level was set to 5% and was not adjusted (by means of Bonferroni correction) at that step. This avoids missing some items that are potentially functioning differently and is in line with other approaches for extending DIF methods to more than two groups (Kim et al., [<reflink idref="bib25" id="ref76">25</reflink>]; Penfield, [<reflink idref="bib34" id="ref77">34</reflink>]).</p> <p>The Wald test and the likelihood ratio test required respectively five and four iterations of the item purification process to reach convergence of the results. Table 1 summarizes the test statistics and related <emph>p</emph>-values for the 15 items. The <emph>p</emph>-values are computed on the basis of the chi-squared distribution with six degrees of freedom—the rank of the matrix (<reflink idref="bib19" id="ref78">19</reflink>).</p> <p>TABLE 1 Wald and Likelihood Ratio DIF Statistics, TCALS-II Data Set</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;Wald Test&lt;/td&gt;&lt;td&gt;Likelihood Ratio Test&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Item&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;1&lt;/td&gt;&lt;td&gt;18.054&lt;/td&gt;&lt;td&gt;0.006*&lt;/td&gt;&lt;td&gt;18.594&lt;/td&gt;&lt;td&gt;0.005*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;2&lt;/td&gt;&lt;td&gt;9.627&lt;/td&gt;&lt;td&gt;0.141&lt;/td&gt;&lt;td&gt;8.147&lt;/td&gt;&lt;td&gt;0.228&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;3&lt;/td&gt;&lt;td&gt;7.612&lt;/td&gt;&lt;td&gt;0.268&lt;/td&gt;&lt;td&gt;6.757&lt;/td&gt;&lt;td&gt;0.344&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;4&lt;/td&gt;&lt;td&gt;10.765&lt;/td&gt;&lt;td&gt;0.096&lt;/td&gt;&lt;td&gt;10.740&lt;/td&gt;&lt;td&gt;0.097&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;5&lt;/td&gt;&lt;td&gt;2.808&lt;/td&gt;&lt;td&gt;0.832&lt;/td&gt;&lt;td&gt;2.526&lt;/td&gt;&lt;td&gt;0.866&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;6&lt;/td&gt;&lt;td&gt;11.861&lt;/td&gt;&lt;td&gt;0.065&lt;/td&gt;&lt;td&gt;12.555&lt;/td&gt;&lt;td&gt;0.051&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;7&lt;/td&gt;&lt;td&gt;3.678&lt;/td&gt;&lt;td&gt;0.720&lt;/td&gt;&lt;td&gt;3.690&lt;/td&gt;&lt;td&gt;0.719&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;8&lt;/td&gt;&lt;td&gt;35.126&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;36.518&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;9&lt;/td&gt;&lt;td&gt;1.343&lt;/td&gt;&lt;td&gt;0.969&lt;/td&gt;&lt;td&gt;2.231&lt;/td&gt;&lt;td&gt;0.897&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;26.345&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;29.500&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;21.654&lt;/td&gt;&lt;td&gt;0.001*&lt;/td&gt;&lt;td&gt;22.495&lt;/td&gt;&lt;td&gt;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;12&lt;/td&gt;&lt;td&gt;12.443&lt;/td&gt;&lt;td&gt;0.053&lt;/td&gt;&lt;td&gt;12.656&lt;/td&gt;&lt;td&gt;0.049*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;13&lt;/td&gt;&lt;td&gt;10.049&lt;/td&gt;&lt;td&gt;0.123&lt;/td&gt;&lt;td&gt;9.694&lt;/td&gt;&lt;td&gt;0.138&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;14&lt;/td&gt;&lt;td&gt;3.043&lt;/td&gt;&lt;td&gt;0.803&lt;/td&gt;&lt;td&gt;2.800&lt;/td&gt;&lt;td&gt;0.833&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;15&lt;/td&gt;&lt;td&gt;2.830&lt;/td&gt;&lt;td&gt;0.830&lt;/td&gt;&lt;td&gt;3.251&lt;/td&gt;&lt;td&gt;0.777&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;*Item is flagged as DIF at significance level 0.05.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Four items are flagged as DIF with very high confidence: items 1, 8, 10, and 11. Items 6 and 12 are borderline. The remaining nine items are not detected as functioning differently. Both tests provide the same classification of the items as DIF or non DIF, except for item 12 for which the <emph>p</emph>-values are equal to 0.049 (Wald test) and 0.053 (likelihood ratio test). Figure 1 displays the fitted response probabilities for the four items which exhibit a highly significant DIF effect. These curves are also displayed on a logit scale in Figure 2, to better distinguish between each group. The fitted curves are based on the results of the Wald test. The likelihood ratio test provides very similar curves and the corresponding curves are not displayed here.</p> <p>Graph: FIGURE 1 Item probability curves for items 1 (top left), 8 (top right), 10 (bottom left), and 11 (bottom right), based on the results of the Wald test.</p> <p>Graph: FIGURE 2 Logits of item probability curves for items 1 (top left), 8 (top right), 10 (bottom left), and 11 (bottom right), based on the results of the Wald test.</p> <p>For items 1, 8, and 11, the DIF effect mainly occurs between the reference group (year 1998) on the one hand, and the three focal groups on the other hand. The focal groups have very close probability curves, while the reference group has larger probabilities almost overall. For item 10, the focal group 2000 has slightly larger probabilities and the focal group 2002 has slightly lower probabilities overall. Both the reference group and the focal group 2004 have similar probability curves. In addition, the response probability curves are somewhat parallel for the items 1 and 11, which is an indicator of the presence of uniform DIF only. On the opposite, one can observe different slopes in the probability curves for items 8 and 10, so that nonuniform DIF is most probably present for those items. These findings are identical whenever either the Wald test or the likelihood ratio test is considered.</p> <p>Table 2 provides the model parameter estimates and the standard errors of for the four items flagged as DIF. The parameter estimates permit to describe the graphical display of the response probabilities in Figures 1 and 2. For items 1 and 11, the slope parameters <emph>β</emph><subs>00</subs>, <emph>β</emph><subs>02</subs>, and <emph>β</emph><subs>04</subs> for respectively the focal groups 2000, 2002, and 2004 are close to zero, while the intercept parameters <emph>α</emph><subs>00</subs>, <emph>α</emph><subs>02</subs>, and <emph>α</emph><subs>04</subs> take negative and very close values. For item 8, the intercepts are all close and negative but the slopes are positive so the three probability curves are steeper and right-shifted with respect to the reference group curve. Finally, the intercept and slope parameters for item 10 are either positive or negative and with opposite signs. Consequently, and as noted from Figure 1, nonuniform DIF is most probably present for these two items. Note, however, that some of these parameters are not statistically significant whenever taken individually. The previous discussion is therefore purely descriptive and in line with Figures 1 and 2, as expected. A more formal investigation is provided.</p> <p>TABLE 2 Group-Specific Parameter Estimates and Standard Errors (in Parentheses) for the Four Items with Significant DIF Effect and Both Tests. Subscripts 00, 02, and 04 Refer Respectively to the Years 2000, 2002, and 2004 of the Focal Groups</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;td&gt;Test&lt;/td&gt;&lt;td&gt;Item&lt;/td&gt;&lt;td&gt;&lt;italic&gt;&amp;#945;&lt;/italic&gt;&lt;sub&gt;00&lt;/sub&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;&amp;#945;&lt;/italic&gt;&lt;sub&gt;02&lt;/sub&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;&amp;#945;&lt;/italic&gt;&lt;sub&gt;04&lt;/sub&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;&amp;#946;&lt;/italic&gt;&lt;sub&gt;00&lt;/sub&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;&amp;#946;&lt;/italic&gt;&lt;sub&gt;02&lt;/sub&gt;&lt;/td&gt;&lt;td&gt;&lt;italic&gt;&amp;#946;&lt;/italic&gt;&lt;sub&gt;04&lt;/sub&gt;&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Wald&lt;/td&gt;&lt;td&gt;&amp;#8199;1&lt;/td&gt;&lt;td&gt;&amp;#8722;0.51&amp;#160; (0.35)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.59&amp;#160; (0.37)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.47&amp;#160; (0.37)&lt;/td&gt;&lt;td&gt;0.02&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.01&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.02&amp;#160; (0.05)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;&amp;#8199;8&lt;/td&gt;&lt;td&gt;&amp;#8722;1.20&amp;#160; (0.31)&lt;/td&gt;&lt;td&gt;&amp;#8722;1.25&amp;#160; (0.34)&lt;/td&gt;&lt;td&gt;&amp;#8722;1.16&amp;#160; (0.33)&lt;/td&gt;&lt;td&gt;0.14&amp;#160; (0.04)&lt;/td&gt;&lt;td&gt;0.13&amp;#160; (0.04)&lt;/td&gt;&lt;td&gt;0.09&amp;#160; (0.04)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;1.75&amp;#160; (0.69)&lt;/td&gt;&lt;td&gt;0.47&amp;#160; (0.76)&lt;/td&gt;&lt;td&gt;&amp;#8722;1.37&amp;#160; (0.85)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.16&amp;#160; (0.07)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.05&amp;#160; (0.08)&lt;/td&gt;&lt;td&gt;0.12&amp;#160; (0.09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;&amp;#8722;0.31&amp;#160; (0.40)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.69&amp;#160; (0.43)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.81&amp;#160; (0.44)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.01&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.04&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.04&amp;#160; (0.05)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;LRT&lt;/td&gt;&lt;td&gt;&amp;#8199;1&lt;/td&gt;&lt;td&gt;&amp;#8722;0.49&amp;#160; (0.34)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.72&amp;#160; (0.37)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.59&amp;#160; (0.37)&lt;/td&gt;&lt;td&gt;0.03&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.03&amp;#160; (0.06)&lt;/td&gt;&lt;td&gt;0.03&amp;#160; (0.06)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;&amp;#8199;8&lt;/td&gt;&lt;td&gt;&amp;#8722;1.19&amp;#160; (0.32)&lt;/td&gt;&lt;td&gt;&amp;#8722;1.31&amp;#160; (0.34)&lt;/td&gt;&lt;td&gt;&amp;#8722;1.19&amp;#160; (0.33)&lt;/td&gt;&lt;td&gt;0.15&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.15&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.10&amp;#160; (0.05)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;1.74&amp;#160; (0.66)&lt;/td&gt;&lt;td&gt;0.26&amp;#160; (0.73)&lt;/td&gt;&lt;td&gt;&amp;#8722;1.51&amp;#160; (0.83)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.18&amp;#160; (0.07)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.03&amp;#160; (0.08)&lt;/td&gt;&lt;td&gt;0.15&amp;#160; (0.09)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;&amp;#8722;0.22&amp;#160; (0.40)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.64&amp;#160; (0.43)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.91&amp;#160; (0.45)&lt;/td&gt;&lt;td&gt;&amp;#8722;0.02&amp;#160; (0.05)&lt;/td&gt;&lt;td&gt;0.04&amp;#160; (0.06)&lt;/td&gt;&lt;td&gt;0.06&amp;#160; (0.06)&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>The slight differences in model parameters between the two tests are due to different final results after item purification because of the different classification of item 12 as DIF or non-DIF. Recall however that item 12 returns very borderline <emph>p</emph>-values, which explains this apparent contradiction.</p> <hd id="AN0067043770-12">Subset Comparisons</hd> <p>In order to illustrate how the Wald test can perform subset comparisons, all possible triplets of groups of examinees were compared for the four items with significant DIF effect. For instance, the comparison between the reference group and the focal groups 2000 and 2004 was achieved with the contrast matrix</p> <p>Graph</p> <p>As four triplets of groups of examinees are considered in a multiple comparison scheme, adjustment of the significance level was performed. Three methods were considered: the usual Bonferroni adjustment; Šidák correction (Šidák, 1967); and the Holm-Bonferroni method (Holm, [<reflink idref="bib21" id="ref79">21</reflink>]; see also Abdi, [<reflink idref="bib1" id="ref80">1</reflink>] and Shaffer, 1995). With four comparisons, the adjusted Bonferroni and Šidák significance levels equal respectively 0.0125 and 0.0127. With Holm-Bonferroni method, the four <emph>p</emph>-values are sorted by increasing order and compared respectively to the increasing significance levels 0.0125, 0.0250, 0.0375, and 0.050.</p> <p>Table 3 contains the Wald statistics and the associated <emph>p</emph>-values for each item and each triplet of groups of examinees. First, the adjusted Bonferroni and Šidák significance levels are so close that they lead to the same conclusions, so they will not be distinguished further. Moreover, the Holm-Bonferroni method returns the same significant <emph>p</emph>-values in any case, except for item 10 for which the triplet of years 1998, 2000, and 2002 has significant <emph>p</emph>-value according to this method.</p> <p>TABLE 3 Subtests of DIF Among All Possible Triples of Groups of Examinees, for the Four Items with Significant DIF Effect</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;td&gt;Item&lt;/td&gt;&lt;td&gt;Groups&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;td&gt;Item&lt;/td&gt;&lt;td&gt;Groups&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;Item 1&lt;/td&gt;&lt;td&gt;(98-00-02)&lt;/td&gt;&lt;td&gt;17.770&lt;/td&gt;&lt;td&gt;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;td&gt;Item 8&lt;/td&gt;&lt;td&gt;(98-00-02)&lt;/td&gt;&lt;td&gt;21.497&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(98-00-04)&lt;/td&gt;&lt;td&gt;10.466&lt;/td&gt;&lt;td&gt;0.033&lt;/td&gt;&lt;td /&gt;&lt;td&gt;(98-00-04)&lt;/td&gt;&lt;td&gt;32.284&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(98-02-04)&lt;/td&gt;&lt;td&gt;17.549&lt;/td&gt;&lt;td&gt;0.002*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;td /&gt;&lt;td&gt;(98-02-04)&lt;/td&gt;&lt;td&gt;30.938&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(00-02-04)&lt;/td&gt;&lt;td&gt;2.746&lt;/td&gt;&lt;td&gt;0.601&lt;/td&gt;&lt;td&gt;&amp;#160;&lt;/td&gt;&lt;td&gt;(00-02-04)&lt;/td&gt;&lt;td&gt;9.102&lt;/td&gt;&lt;td&gt;0.059&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Item 10&lt;/td&gt;&lt;td&gt;(98-00-02)&lt;/td&gt;&lt;td&gt;11.444&lt;/td&gt;&lt;td&gt;0.022&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;td&gt;Item 11&lt;/td&gt;&lt;td&gt;(98-00-02)&lt;/td&gt;&lt;td&gt;15.833&lt;/td&gt;&lt;td&gt;0.003*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(98-00-04)&lt;/td&gt;&lt;td&gt;26.081&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;td /&gt;&lt;td&gt;(98-00-04)&lt;/td&gt;&lt;td&gt;20.228&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(98-02-04)&lt;/td&gt;&lt;td&gt;6.775&lt;/td&gt;&lt;td&gt;0.148&lt;/td&gt;&lt;td /&gt;&lt;td&gt;(98-02-04)&lt;/td&gt;&lt;td&gt;19.622&lt;/td&gt;&lt;td&gt;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;(00-02-04)&lt;/td&gt;&lt;td&gt;24.569&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;sup&gt;&amp;#9633;&lt;/sup&gt;&lt;/td&gt;&lt;td /&gt;&lt;td&gt;(00-02-04)&lt;/td&gt;&lt;td&gt;2.390&lt;/td&gt;&lt;td&gt;0.664&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;* Significant at &lt;italic&gt;&amp;#945;&lt;/italic&gt; level 0.0125 (Bonferroni) and 0.0127 (&amp;#352;id&amp;#225;k).&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&lt;sup&gt;&amp;#9633;&lt;/sup&gt; Significant with Holm-Bonferroni method.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>For each item, at least one triplet of groups yields a nonsignificant <emph>p</emph>-value, which indicates that the significant DIF effect mostly occurs from the fourth group not included in the nonsignificant triplet. For items 1, 10, and 11, the nonsignificant triplet is made of the three focal groups. That is, the presence of the reference group in any triplet leads to a significant difference between the probability curves. For item 8, however, the <emph>p</emph>-value is somewhat borderline (<emph>p</emph> = 0.059) and for item 1, another triplet (without the focal group 2002) is also nonsignificant, but the <emph>p</emph>-value (0.033) is also borderline. For item 10, the triplet of groups with a nonsignificant result is made of the reference group and the focal groups 2002 and 2004. That is, the focal group 2000 behaves differently than the three other groups. Finally, all other triplets provide very significant results.</p> <p>These results are in agreement with Figures 1 and 2. It was noticed that for items 1, 8, and 11, the reference group tends to have higher probabilities of answering the item correctly than the other groups, while the three focal groups have quite close probability curves. For item 8, however, the discrepancy between the probability curves of the focal groups is much more visible, which explains the borderline <emph>p</emph>-value for the corresponding Wald test. Finally, the focal group 2000 has larger probabilities with item 10, while the three other groups are much closer in terms of probability curves.</p> <hd id="AN0067043770-13">Uniform and Nonuniform DIF Effects</hd> <p>Till now, both DIF effects have been tested simultaneously as an omnibus approach. For the four items with established significant DIF effect, a subsequent distinct analysis of both effects is conducted. Nonuniform DIF is tested first, and uniform DIF is investigated in a second step among those items without nonuniform DIF. The results of this two-step analysis, both with the Wald test and the likelihood ratio test, are summarized in Table 4.</p> <p>TABLE 4 Results of the Tests of Nonuniform and Uniform DIF for the Four Items with Significant DIF Effect</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;td /&gt;&lt;td /&gt;&lt;td&gt;Wald Test&lt;/td&gt;&lt;td&gt;Likelihood Ratio Test&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Effect&lt;/td&gt;&lt;td&gt;Item&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;NUDIF&lt;/td&gt;&lt;td&gt;&amp;#8199;1&lt;/td&gt;&lt;td&gt;0.537&lt;/td&gt;&lt;td&gt;0.911&lt;/td&gt;&lt;td&gt;0.531&lt;/td&gt;&lt;td&gt;0.912&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;&amp;#8199;8&lt;/td&gt;&lt;td&gt;8.736&lt;/td&gt;&lt;td&gt;0.033*&lt;/td&gt;&lt;td&gt;8.658&lt;/td&gt;&lt;td&gt;0.034*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;23.673&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;24.590&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;2.129&lt;/td&gt;&lt;td&gt;0.546&lt;/td&gt;&lt;td&gt;2.121&lt;/td&gt;&lt;td&gt;0.548&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;UDIF&lt;/td&gt;&lt;td&gt;&amp;#8199;1&lt;/td&gt;&lt;td&gt;17.628&lt;/td&gt;&lt;td&gt;0.001*&lt;/td&gt;&lt;td&gt;18.076&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;19.810&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;20.164&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;* Item is flagged as DIF at significance level 0.05.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Both tests provide very close results so they will not be distinguished further. It appears from the top part of Table 4 that the items 8 and 10 exhibit nonuniform DIF, with significant <emph>p</emph>-values for the former and highly significant <emph>p</emph>-values for the latter. Items 1 and 11, on the other hand, are not affected by nonuniform DIF and are therefore retained for an investigation of uniform DIF. From the bottom part of Table 4, one concludes that these two items have a highly significant uniform DIF effect. Thus, and as expected, each item has a significant DIF effect, either uniform or nonuniform. Moreover, it was observed in Figure 1 that the probability curves are rather parallel for items 1 and 11, while there is some variability in the slopes for items 8 and 10.</p> <hd id="AN0067043770-14">Other Methods</hd> <p>Finally, an investigation of DIF was performed by using the two usual methods for DIF identification in multiple groups: the generalized Mantel-Haenszel method (further referred to as the GMH method) and the generalized Lord's χ<sups>2</sups> test. Lord's test was performed using both the 1PL model and the 2PL model. Item purification was performed with each method. Table 5 lists the DIF statistics and related <emph>p</emph>-values.</p> <p>TABLE 5 Identification of DIF Items Using the Generalized Mantel-Haenszel (GMH) Method and the Generalized Lord's Chi-square Test Under the 1PL and the 2PL Models</p> <p> <ephtml> &lt;table&gt;&lt;thead valign="bottom"&gt;&lt;tr&gt;&lt;td /&gt;&lt;td&gt;GMH&lt;/td&gt;&lt;td&gt;Lord's Test (1PL)&lt;/td&gt;&lt;td&gt;Lord's Test (2PL)&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;Item&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;td&gt;Statistic&lt;/td&gt;&lt;td&gt;&lt;italic&gt;p&lt;/italic&gt;-value&lt;/td&gt;&lt;/tr&gt;&lt;/thead&gt;&lt;tbody&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;1&lt;/td&gt;&lt;td&gt;17.233&lt;/td&gt;&lt;td&gt;0.001*&lt;/td&gt;&lt;td&gt;14.402&lt;/td&gt;&lt;td&gt;0.002*&lt;/td&gt;&lt;td&gt;3.528&lt;/td&gt;&lt;td&gt;0.740&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;2&lt;/td&gt;&lt;td&gt;3.281&lt;/td&gt;&lt;td&gt;0.350&lt;/td&gt;&lt;td&gt;3.480&lt;/td&gt;&lt;td&gt;0.323&lt;/td&gt;&lt;td&gt;26.247&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;3&lt;/td&gt;&lt;td&gt;6.810&lt;/td&gt;&lt;td&gt;0.078&lt;/td&gt;&lt;td&gt;5.133&lt;/td&gt;&lt;td&gt;0.162&lt;/td&gt;&lt;td&gt;16.526&lt;/td&gt;&lt;td&gt;0.011*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;4&lt;/td&gt;&lt;td&gt;5.032&lt;/td&gt;&lt;td&gt;0.169&lt;/td&gt;&lt;td&gt;4.527&lt;/td&gt;&lt;td&gt;0.210&lt;/td&gt;&lt;td&gt;12.403&lt;/td&gt;&lt;td&gt;0.054&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;5&lt;/td&gt;&lt;td&gt;2.403&lt;/td&gt;&lt;td&gt;0.493&lt;/td&gt;&lt;td&gt;1.503&lt;/td&gt;&lt;td&gt;0.682&lt;/td&gt;&lt;td&gt;2.779&lt;/td&gt;&lt;td&gt;0.836&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;6&lt;/td&gt;&lt;td&gt;4.436&lt;/td&gt;&lt;td&gt;0.218&lt;/td&gt;&lt;td&gt;3.527&lt;/td&gt;&lt;td&gt;0.317&lt;/td&gt;&lt;td&gt;9.929&lt;/td&gt;&lt;td&gt;0.128&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;7&lt;/td&gt;&lt;td&gt;0.797&lt;/td&gt;&lt;td&gt;0.850&lt;/td&gt;&lt;td&gt;2.010&lt;/td&gt;&lt;td&gt;0.570&lt;/td&gt;&lt;td&gt;1.579&lt;/td&gt;&lt;td&gt;0.954&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;8&lt;/td&gt;&lt;td&gt;20.865&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;19.616&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;24.379&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;&amp;#8199;9&lt;/td&gt;&lt;td&gt;1.010&lt;/td&gt;&lt;td&gt;0.799&lt;/td&gt;&lt;td&gt;0.977&lt;/td&gt;&lt;td&gt;0.807&lt;/td&gt;&lt;td&gt;1.531&lt;/td&gt;&lt;td&gt;0.957&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;10&lt;/td&gt;&lt;td&gt;13.326&lt;/td&gt;&lt;td&gt;0.004*&lt;/td&gt;&lt;td&gt;10.809&lt;/td&gt;&lt;td&gt;0.013*&lt;/td&gt;&lt;td&gt;17.913&lt;/td&gt;&lt;td&gt;0.006*&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;11&lt;/td&gt;&lt;td&gt;18.280&lt;/td&gt;&lt;td&gt;&amp;#60;0.001*&lt;/td&gt;&lt;td&gt;14.510&lt;/td&gt;&lt;td&gt;0.002*&lt;/td&gt;&lt;td&gt;6.410&lt;/td&gt;&lt;td&gt;0.379&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;12&lt;/td&gt;&lt;td&gt;6.429&lt;/td&gt;&lt;td&gt;0.093&lt;/td&gt;&lt;td&gt;3.887&lt;/td&gt;&lt;td&gt;0.274&lt;/td&gt;&lt;td&gt;3.184&lt;/td&gt;&lt;td&gt;0.785&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;13&lt;/td&gt;&lt;td&gt;7.375&lt;/td&gt;&lt;td&gt;0.061&lt;/td&gt;&lt;td&gt;7.153&lt;/td&gt;&lt;td&gt;0.067&lt;/td&gt;&lt;td&gt;6.834&lt;/td&gt;&lt;td&gt;0.336&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;14&lt;/td&gt;&lt;td&gt;0.279&lt;/td&gt;&lt;td&gt;0.964&lt;/td&gt;&lt;td&gt;0.269&lt;/td&gt;&lt;td&gt;0.966&lt;/td&gt;&lt;td&gt;3.367&lt;/td&gt;&lt;td&gt;0.762&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;15&lt;/td&gt;&lt;td&gt;1.251&lt;/td&gt;&lt;td&gt;0.741&lt;/td&gt;&lt;td&gt;0.237&lt;/td&gt;&lt;td&gt;0.971&lt;/td&gt;&lt;td&gt;4.999&lt;/td&gt;&lt;td&gt;0.544&lt;/td&gt;&lt;/tr&gt;&lt;tr&gt;&lt;td&gt;*Item is flagged as DIF at significance level 0.05.&lt;/td&gt;&lt;/tr&gt;&lt;/tbody&gt;&lt;/table&gt; </ephtml> </p> <p>Both the GMH method and Lord's test with the 1PL model focus on uniform DIF detection. It appears from Table 5 that they identify items 1, 8, 10, and 11 with a significant uniform DIF effect. It is expected that these methods return very close results as they are designed to detect uniform DIF, assuming that nonuniform DIF is absent (Holland and Thayer, [<reflink idref="bib20" id="ref81">20</reflink>], highlighted this particular relationship). In addition, items flagged as DIF are identical to those obtained from the logistic regression procedure.</p> <p>With Lord's test and the 2PL model, however, items 2, 3, 8, and 10 are flagged as DIF. Items 8 and 10 are still detected as DIF items, but items 2 and 3 are now flagged as having a significant DIF effect while items 1 and 11 are not identified as such anymore. These differences in conclusions can be explained as follows. First, items 2 and 3 exhibit statistically significant differences in item discriminations under the 2PL model, which is impossible to detect with the 1PL model (or the GMH method). Note that these differences in item discriminations were also detected by the generalized logistic regression method, although judged as not statistically significant. Second, for items 1 and 11 the DIF effect is mostly uniform. When switching from the 1PL model to the 2PL model, however, the item discriminations are much larger than with the 1PL model—although there is no statistical difference between the groups of respondents—and simultaneously the item difficulties get more closer so that no DIF effect is detected.</p> <p>In sum, items 8 and 10 tend exhibit nonuniform DIF, while items 1 and 11 are more affected by uniform DIF. Items 2 and 3 were not identified by the logistic regression procedure, although the generalized Lord's test with the 2PL model indicates a possible presence of nonuniform DIF effect.</p> <hd id="AN0067043770-15">DISCUSSION</hd> <p>This article focused on the identification of DIF in the presence of more than two groups of respondents. The logistic regression procedure finds a natural extension into this multiple group framework. Simultaneous statistical inference permits to identify the items that exhibit a significant DIF effect. Either uniform, nonuniform, or both effects can be tested. In addition, two statistical approaches for testing for DIF, the Wald test and the likelihood ratio test, are available. The Wald test is also appropriate to perform subset comparisons between some groups of examinees by using appropriate contrast matrices. Finally, the practical working of the method was illustrated by the analysis of a real example on language skill assessment. It was highlighted that several conclusions can be drawn from the fitted logistic curves in each group, which reinforces the usefulness of the generalized logistic regression procedure in this context. In addition, the conclusions were in line with other methods for multiple groups DIF detection, which strengthens the usefulness of the present approach.</p> <p>The generalized logistic regression procedure was presented under the usual approach of several focal groups that are compared to a single reference group. To this end, the identification constraints of the logistic regression model parameters were naturally defined with respect to this reference group. One main asset of the method, however, is that it can also apply when no clear reference group can be set up. For example, in international studies such as PISA or TIMMS, it is not straightforward to set up the reference country for DIF investigations. This approach allows for comparisons of countries altogether, by specific comparisons of model parameters, and the way these parameters are constrained does not influence the conclusions themselves but only the parameterization of the logistic models. In addition, the larger the group sizes the better the estimation of model parameters, and hence the better the statistical inference for DIF identification. This method is therefore naturally suitable for large-scale assessment studies or international surveys.</p> <p>The main goal of this article was to highlight the fact that the logistic regression procedure can be easily extended to multiple-groups DIF testing. Although sometimes suggested, it has apparently never been practically achieved. However, this is only the first step in the process of proposing a novel methodology for multiple-groups DIF. Indeed, the practical efficiency of this method should be carefully checked, notably with respect to the control of Type I error and by evaluating the empirical power of detecting DIF items. In addition, Monte-Carlo comparisons with other multiple-groups DIF methods should be performed. The practical example pointed out some similarities between the generalized versions of Mantel-Haenszel and logistic regression methods but also some discrepancy with generalized Lord's test and the 2PL model. It makes sense therefore to investigate further in this way. For example, from previous studies in the case of a single focal group (e.g., Rogers &amp; Swaminathan, [<reflink idref="bib40" id="ref82">40</reflink>]), one might expect that the generalized logistic regression and the generalized Mantel-Haenszel methods will perform similarly in the presence of uniform DIF, while the former will be more adequate for nonuniform DIF identification.</p> <p>Finally, identifying the items that may exhibit DIF between several groups is an important step, but providing some measure of effect size to evaluate the magnitude of DIF is another important task. In the usual framework of a single focal group, Jodoin and Gierl ([<reflink idref="bib22" id="ref83">22</reflink>]) and Zumbo and Thomas ([<reflink idref="bib51" id="ref84">51</reflink>]) proposed to consider the difference Δ<emph>R</emph><sups>2</sups> between Nagelkerke's <emph>R</emph><sups>2</sups> coefficients (Nagelkerke, [<reflink idref="bib31" id="ref85">31</reflink>]) of the models <emph>M</emph><subs>0</subs> and <emph>M</emph><subs>1</subs>. This measure of effect size could directly be extended in this multiple-groups framework. Its relative usefulness, however, should be assessed since it is known (Hidalgo &amp; Lopez-Pina, [<reflink idref="bib17" id="ref86">17</reflink>]) that Nagelkerke's <emph>R</emph><sups>2</sups> coefficient tends to underestimate the true DIF effect size. Other approaches might involve the computation of the area between the logistic curves, although it is not easy to apply it with more than two groups of respondents.</p> <p>To our opinion, selecting an appropriate measure of DIF effect size is a central issue and is worth being studied carefully. This could actually counterbalance the issue mentioned at the end of Section 2—items could be more often flagged as DIF as more than groups are compared simultaneously. Flagging more items is a drawback, but computing then the effect size measures could correct for potential increase of Type I error (that is, flagging as DIF some items that are not functioning differently). In addition, as suggested earlier, this method would certainly benefit from a Bayesian approach by computing posterior DIF probabilities and related effect size estimates. One avenue of potential interest is the evaluation of informative hypotheses using either hypothesis testing (Hoijtink, [<reflink idref="bib18" id="ref87">18</reflink>]) or the Bayes factor (Hoijtink, Klugkist, &amp; Boelen, [<reflink idref="bib19" id="ref88">19</reflink>]). Although the Bayesian paradigm was recently introduced in DIF research (Bolt &amp; Cohen, [<reflink idref="bib8" id="ref89">8</reflink>]; Frederickx, Tuerlinckx, De Boeck, &amp; Magis, [<reflink idref="bib15" id="ref90">15</reflink>]), it has apparently not yet been applied in this multiple groups context, but would be of interest for the future.</p> <hd id="AN0067043770-16">Acknowledgments</hd> <p>The authors wish to thank Prof. Stephen G. Sireci, editor, Prof. Rob R. Meijer, co-editor, and two anonymous reviewers for their helpful comments. This research was funded by a grant "Chargé de recherches" of the National Funds for Scientific Research (FNRS), Belgium, the Research Funds of the K. U. Leuven, and a grant from the Social Sciences Research Council of Canada (SSRC).</p> <ref id="AN0067043770-17"> <title> REFERENCES </title> <blist> <bibl id="bib1" idref="ref25" type="bt">1</bibl> <bibtext> Abdi, H.2007. "Bonferroni and Šidák corrections for multiple comparisons". In Encyclopedia of measurement and statistics, Edited by: Salkind, N.J.Thousand Oaks, CA: Sage.</bibtext> </blist> <blist> <bibl id="bib2" idref="ref1" type="bt">2</bibl> <bibtext> Ackerman, T.A.1992. A didactic explanation of item bias, item impact, and item validity from a multidimensional perspective. Journal of Educational Measurement, 29: 67–91.</bibtext> </blist> <blist> <bibl id="bib3" idref="ref29" type="bt">3</bibl> <bibtext> Agresti, A.1990. Categorical data analysis, New York: Wiley.</bibtext> </blist> <blist> <bibl id="bib4" idref="ref32" type="bt">4</bibl> <bibtext> Agresti, A.1996. An introduction to categorical data analysis, New York: Wiley.</bibtext> </blist> <blist> <bibl id="bib5" idref="ref23" type="bt">5</bibl> <bibtext> Agresti, A.2002. Categorical data analysis (2nd ed.), New York: Wiley.</bibtext> </blist> <blist> <bibl id="bib6" idref="ref14" type="bt">6</bibl> <bibtext> Angoff, W.H. and Sharon, A.T.1974. The evaluation of differences in test performance of two or more groups. Educational and Psychological Measurement, 34: 807–816.</bibtext> </blist> <blist> <bibl id="bib7" idref="ref37" type="bt">7</bibl> <bibtext> Bock, R.D.1975. Multivariate statistical methods, New York: McGraw-Hill.</bibtext> </blist> <blist> <bibl id="bib8" idref="ref45" type="bt">8</bibl> <bibtext> Bolt, D.M. and Cohen, A.S.2005. A mixture model analysis of differential item functioning. Journal of Educational Measurement, 42: 133–148.</bibtext> </blist> <blist> <bibl id="bib9" idref="ref46" type="bt">9</bibl> <bibtext> Candell, G.L. and Drasgow, F.1988. An iterative procedure for linking metrics and assessing item bias in item response theory. Applied Psychological Measurement, 12: 253–260.</bibtext> </blist> <blist> <bibtext> Clauser, B.E. and Mazor, K.M.1998. Using statistical procedures to identify differential item functioning test items. Educational Measurement: Issues and Practice, 17: 31–44.</bibtext> </blist> <blist> <bibtext> Cox, D.R. and Hinkley, D.V.1974. Theoretical statistics, London: Chapman and Hall.</bibtext> </blist> <blist> <bibtext> Ellis, B.B. and Kimmel, H.D.1992. Identification of unique cultural response patterns by means of item response theory. Journal of Applied Psychology, 77: 177–184.</bibtext> </blist> <blist> <bibtext> Fidalgo, A.M. and Madeira, J.M.2008. Generalized Mantel-Haenszel methods for differential item functioning detection. Educational and Psychological Measurement, 68: 940–958.</bibtext> </blist> <blist> <bibtext> Fidalgo, A.M. and Scalon, J.D.2010. Using generalized Mantel-Haenszel statistics to assess DIF among multiple groups. Journal of Psychoeducational Assessment, 28: 60–69.</bibtext> </blist> <blist> <bibtext> Frederickx, S., Tuerlinckx, F., De Boeck, P. and Magis, D.2010. RIM: A random item mixture model to detect differential item functioning. Journal of Educational Measurement, 47: 432–457.</bibtext> </blist> <blist> <bibtext> Hanson, B.A.1998. Uniform DIF and DIF defined by differences in item response functions. Journal of Educational and Behavioral Statistics, 23: 244–253.</bibtext> </blist> <blist> <bibtext> Hidalgo, M.D. and Lopez-Pina, J.A.2004. Differential item functioning detection and effect size: a comparison between logistic regression and Mantel-Haenszel procedures. Educational and Psychological Measurement, 64: 903–915.</bibtext> </blist> <blist> <bibtext> Hoijtink, H.1998. Constrained latent class analysis using the Gibbs sampler and posterior predictive p-values: Applications to educational testing. Statistica Sinica, 8: 691–712.</bibtext> </blist> <blist> <bibtext> Hoijtink, H., Klugkist, I. and Boelen, P.A.2008. Bayesian information of informative hypotheses, New York: Springer.</bibtext> </blist> <blist> <bibtext> Holland, P.W. and Thayer, D.T.1988. "Differential item performance and the Mantel-Haenszel procedure". In Test validity, Edited by: Wainer, H. and Braun, H.I.129–145. Hillsdale, NJ: Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibtext> Holm, S.1979. A simple sequentially rejective multiple testing procedure. Scandinavian Journal of Statistics, 6: 65–70.</bibtext> </blist> <blist> <bibtext> Jodoin, M.G. and Gierl, M.J.2001. Evaluating type I error and power rates using an effect size measure with the logistic regression procedure for DIF detection. Applied Measurement in Education, 14: 329–349.</bibtext> </blist> <blist> <bibtext> Johnson, R.A. and Wichern, D.W.1998. Applied multivariate statistical analysis (4th ed.), Upper Saddle River, NJ: Prentice-Hall.</bibtext> </blist> <blist> <bibtext> Kanjee, A.2007. Using logistic regression to detect bias when multiple groups are tested. South African Journal of Psychology, 37: 47–61.</bibtext> </blist> <blist> <bibtext> Kim, S.-H., Cohen, A.S. and Park, T.-H.1995. Detection of differential item functioning in multiple groups. Journal of Educational Measurement, 32: 261–276.</bibtext> </blist> <blist> <bibtext> Laurier, M.D., Froio, L., Paero, C. and Fournier, M.1998. L'élaboration d'un test provincial pour le classement des étudiants en anglais langue seconde au collégial [The elaboration of a provincial test to classify students in English, as a second language, in colleges], Québec, QC: Direction générale de l'enseignement collégial, ministère de l'Education du Québec.</bibtext> </blist> <blist> <bibtext> Lord, F.M.1980. Applications of item response theory to practical testing problems, Hillsdale, NJ: Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibtext> Magis, D., Béland, S., Tuerlinckx, F. and De Boeck, P.2010. A general framework and an R package for the detection of dichotomous differential item functioning. Behavior Research Methods, 42: 847–862.</bibtext> </blist> <blist> <bibtext> McCullagh, P. and Nelder, J.1989. Generalized linear models (2nd ed.), London: Chapman &amp; Hall.</bibtext> </blist> <blist> <bibtext> Millsap, R.E. and Everson, H.T.1993. Methodology review: statistical approaches for assessing measurement bias. Applied Psychological Measurement, 17: 297–334.</bibtext> </blist> <blist> <bibtext> Nagelkerke, N.J. D.1991. A note on a general definition of the coefficient of determination. Biometrika, 78: 691–692.</bibtext> </blist> <blist> <bibtext> Nelder, J. and Wedderburn, R.W. M.1972. Generalized linear models. Journal of the Royal Statistical Society (Series A), 135: 370–384.</bibtext> </blist> <blist> <bibtext> Osterlind, S.J. and Everson, H.T.2009. Differential item functioning (2nd ed.), Thousand Oakes, CA: Sage.</bibtext> </blist> <blist> <bibtext> Penfield, R.D.2001. Assessing differential item functioning among multiple groups: a comparison of three Mantel-Haenszel procedures. Applied Measurement in Education, 14: 235–259.</bibtext> </blist> <blist> <bibtext> Penfield, R.D. and Camilli, G.2007. "Differential item functioning and item bias". In Handbook of statistics 26: psychometrics, Edited by: Rao, C.R. and Sinharray, S.125–167. Amsterdam, , The Netherlands: Elsevier.</bibtext> </blist> <blist> <bibtext> Penfield, R.D. and Lam, T.C. M.2001. Assessing differential item functioning in performance assessment: Review and recommendations. Educational Measurement: Issues and Practice, 19: 5–15.</bibtext> </blist> <blist> <bibtext> Raîche, G.2002. Le dépistage du sous-classement aux tests de classement en anglais, langue seconde, au collégial [The detection of under-classification at English, as a second language, test in college], Gatineau, QC: Collège de l'Outaouais.</bibtext> </blist> <blist> <bibtext> Raju, N.S.1990. Determining the significance of estimated signed and unsigned areas between two item response functions. Applied Psychological Measurement, 14: 197–207.</bibtext> </blist> <blist> <bibtext> Rao, C.R.1973. Linear statistical inference and its applications (second edition), New York: Wiley.</bibtext> </blist> <blist> <bibtext> Rogers, H.J. and Swaminathan, H.1993. A comparison of logistic regression and Mantel-Haenszel procedures for detecting differential item functioning. Applied Psychological Measurement, 17: 105–116.</bibtext> </blist> <blist> <bibtext> Schmitt, A.P. and Dorans, N.J.1990. Differential item functioning for minority examinees on the SAT. Journal of Educational Measurement, 27: 67–81.</bibtext> </blist> <blist> <bibtext> Shaffer, J.P.1995. Multiple hypothesis testing. Annual Review of Psychology, 46: 561–584.</bibtext> </blist> <blist> <bibtext> Shealy, R.T. and Stout, W.1993. A model based standardization approach that separates true bias/DIF from group ability differences and detects test bias/DIF as well as item bias/DIF. Psychometrika, 58: 159–194.</bibtext> </blist> <blist> <bibtext> Šidák, Z.1967. Rectangular confidence region for the means of multivariate normal distributions. Journal of the American Statistical Association, 62: 626–633.</bibtext> </blist> <blist> <bibtext> Swaminathan, H. and Rogers, H.J.1990. Detecting differential item functioning using logistic regression procedures. Journal of Educational Measurement, 27: 361–370.</bibtext> </blist> <blist> <bibtext> Thissen, D., Steinberg, L. and Wainer, H.1988. "Use of item response theory in the study of group difference in trace lines". In Test validity, Edited by: Wainer, H. and Braun, H.147–170. Hillsdale, NJ: Lawrence Erlbaum Associates.</bibtext> </blist> <blist> <bibtext> Van den Noortgate, W. and De Boeck, P.2005. Assessing and explaining differential item functioning using logistic mixed models. Journal of Educational and Behavioral Statistics, 30: 443–464.</bibtext> </blist> <blist> <bibtext> Wald, A.1939. Contributions to the theory of statistical estimation and testing hypotheses. Annals of Mathematical Statistics, 10: 299–326.</bibtext> </blist> <blist> <bibtext> Wedderburn, R.W. M.1976. On the existence and uniqueness of the maximum likelihood estimates for certain generalized linear models. Biometrika, 63: 27–32.</bibtext> </blist> <blist> <bibtext> Wilks, S.S.1938. The large-sample distribution of the likelihood ratio for testing composite hypotheses. Annals of Mathematical Statistics, 9: 60–62.</bibtext> </blist> <blist> <bibtext> Zumbo, B.D. and Thomas, D.R.1997. "A measure of effect size for a model-based approach for studying DIF". Prince George, , BC, Canada: University of Northern British Columbia, Edgeworth Laboratory for Quantitative Behavioral Science.</bibtext> </blist> <blist> <bibtext> Zwick, R. and Ercikan, K.1989. Analysis of differential item functioning in the NAEP history assessment. Journal of Educational Measurement, 26: 55–66.</bibtext> </blist> </ref> <aug> <p>By David Magis; Gilles Raîche; Sébastien Béland and Paul Gérard</p> <p>Reported by Author; Author; Author; Author</p> </aug> <nolink nlid="nl1" bibid="bib10" firstref="ref2"></nolink> <nolink nlid="nl2" bibid="bib16" firstref="ref3"></nolink> <nolink nlid="nl3" bibid="bib27" firstref="ref4"></nolink> <nolink nlid="nl4" bibid="bib38" firstref="ref5"></nolink> <nolink nlid="nl5" bibid="bib46" firstref="ref6"></nolink> <nolink nlid="nl6" bibid="bib20" firstref="ref7"></nolink> <nolink nlid="nl7" bibid="bib43" firstref="ref8"></nolink> <nolink nlid="nl8" bibid="bib45" firstref="ref9"></nolink> <nolink nlid="nl9" bibid="bib33" firstref="ref11"></nolink> <nolink nlid="nl10" bibid="bib35" firstref="ref12"></nolink> <nolink nlid="nl11" bibid="bib34" firstref="ref13"></nolink> <nolink nlid="nl12" bibid="bib12" firstref="ref15"></nolink> <nolink nlid="nl13" bibid="bib41" firstref="ref16"></nolink> <nolink nlid="nl14" bibid="bib52" firstref="ref17"></nolink> <nolink nlid="nl15" bibid="bib13" firstref="ref19"></nolink> <nolink nlid="nl16" bibid="bib14" firstref="ref20"></nolink> <nolink nlid="nl17" bibid="bib25" firstref="ref21"></nolink> <nolink nlid="nl18" bibid="bib24" firstref="ref22"></nolink> <nolink nlid="nl19" bibid="bib49" firstref="ref36"></nolink> <nolink nlid="nl20" bibid="bib39" firstref="ref38"></nolink> <nolink nlid="nl21" bibid="bib32" firstref="ref39"></nolink> <nolink nlid="nl22" bibid="bib11" firstref="ref47"></nolink> <nolink nlid="nl23" bibid="bib23" firstref="ref51"></nolink> <nolink nlid="nl24" bibid="bib48" firstref="ref55"></nolink> <nolink nlid="nl25" bibid="bib50" firstref="ref56"></nolink> <nolink nlid="nl26" bibid="bib29" firstref="ref58"></nolink> <nolink nlid="nl27" bibid="bib28" firstref="ref63"></nolink> <nolink nlid="nl28" bibid="bib26" firstref="ref64"></nolink> <nolink nlid="nl29" bibid="bib37" firstref="ref65"></nolink> <nolink nlid="nl30" bibid="bib17" firstref="ref72"></nolink> <nolink nlid="nl31" bibid="bib18" firstref="ref73"></nolink> <nolink nlid="nl32" bibid="bib19" firstref="ref78"></nolink> <nolink nlid="nl33" bibid="bib21" firstref="ref79"></nolink> <nolink nlid="nl34" bibid="bib40" firstref="ref82"></nolink> <nolink nlid="nl35" bibid="bib22" firstref="ref83"></nolink> <nolink nlid="nl36" bibid="bib51" firstref="ref84"></nolink> <nolink nlid="nl37" bibid="bib31" firstref="ref85"></nolink> <nolink nlid="nl38" bibid="bib15" firstref="ref90"></nolink> |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: EJ946932 AccessLevel: 3 PubType: Academic Journal PubTypeId: academicJournal PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: A Generalized Logistic Regression Procedure to Detect Differential Item Functioning among Multiple Groups – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Magis%2C+David%22">Magis, David</searchLink><br /><searchLink fieldCode="AR" term="%22Raiche%2C+Gilles%22">Raiche, Gilles</searchLink><br /><searchLink fieldCode="AR" term="%22Beland%2C+Sebastien%22">Beland, Sebastien</searchLink><br /><searchLink fieldCode="AR" term="%22Gerard%2C+Paul%22">Gerard, Paul</searchLink> – Name: TitleSource Label: Source Group: Src Data: <searchLink fieldCode="SO" term="%22International+Journal+of+Testing%22"><i>International Journal of Testing</i></searchLink>. 2011 11(4):365-386. – Name: Avail Label: Availability Group: Avail Data: Routledge. Available from: Taylor & Francis, Ltd. 325 Chestnut Street Suite 800, Philadelphia, PA 19106. Tel: 800-354-1420; Fax: 215-625-2940; Web site: http://www.tandf.co.uk/journals – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: Y – Name: Pages Label: Page Count Group: Src Data: 22 – Name: DatePubCY Label: Publication Date Group: Date Data: 2011 – Name: TypeDocument Label: Document Type Group: TypDoc Data: Journal Articles<br />Reports - Descriptive – Name: Audience Label: Education Level Group: Audnce Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink> – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Language+Skills%22">Language Skills</searchLink><br /><searchLink fieldCode="DE" term="%22Identification%22">Identification</searchLink><br /><searchLink fieldCode="DE" term="%22Foreign+Countries%22">Foreign Countries</searchLink><br /><searchLink fieldCode="DE" term="%22Evaluation+Methods%22">Evaluation Methods</searchLink><br /><searchLink fieldCode="DE" term="%22Comparative+Analysis%22">Comparative Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22English+%28Second+Language%29%22">English (Second Language)</searchLink><br /><searchLink fieldCode="DE" term="%22French%22">French</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Tests%22">Tests</searchLink><br /><searchLink fieldCode="DE" term="%22Computation%22">Computation</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Bias%22">Test Bias</searchLink><br /><searchLink fieldCode="DE" term="%22College+Students%22">College Students</searchLink><br /><searchLink fieldCode="DE" term="%22Colleges%22">Colleges</searchLink> – Name: Subject Label: Geographic Terms Group: Su Data: <searchLink fieldCode="DE" term="%22Canada%22">Canada</searchLink> – Name: DOI Label: DOI Group: ID Data: 10.1080/15305058.2011.602810 – Name: ISSN Label: ISSN Group: ISSN Data: 1530-5058 – Name: Abstract Label: Abstract Group: Ab Data: We present an extension of the logistic regression procedure to identify dichotomous differential item functioning (DIF) in the presence of more than two groups of respondents. Starting from the usual framework of a single focal group, we propose a general approach to estimate the item response functions in each group and to test for the presence of uniform DIF, nonuniform DIF, or both. This generalized procedure is compared to other existing DIF methods for multiple groups with a real data set on language skill assessment. Emphasis is put on the flexibility, completeness, and computational easiness of the generalized method. (Contains 5 tables and 2 figures.) – Name: AbstractInfo Label: Abstractor Group: Ab Data: As Provided – Name: Ref Label: Number of References Group: RefInfo Data: 52 – Name: DateEntry Label: Entry Date Group: Date Data: 2011 – Name: AN Label: Accession Number Group: ID Data: EJ946932 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ946932 |
| RecordInfo | BibRecord: BibEntity: Identifiers: – Type: doi Value: 10.1080/15305058.2011.602810 Languages: – Text: English PhysicalDescription: Pagination: PageCount: 22 StartPage: 365 Subjects: – SubjectFull: Language Skills Type: general – SubjectFull: Identification Type: general – SubjectFull: Foreign Countries Type: general – SubjectFull: Evaluation Methods Type: general – SubjectFull: Comparative Analysis Type: general – SubjectFull: English (Second Language) Type: general – SubjectFull: French Type: general – SubjectFull: Models Type: general – SubjectFull: Tests Type: general – SubjectFull: Computation Type: general – SubjectFull: Test Bias Type: general – SubjectFull: College Students Type: general – SubjectFull: Colleges Type: general – SubjectFull: Canada Type: general Titles: – TitleFull: A Generalized Logistic Regression Procedure to Detect Differential Item Functioning among Multiple Groups Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Magis, David – PersonEntity: Name: NameFull: Raiche, Gilles – PersonEntity: Name: NameFull: Beland, Sebastien – PersonEntity: Name: NameFull: Gerard, Paul IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 01 Type: published Y: 2011 Identifiers: – Type: issn-print Value: 1530-5058 Numbering: – Type: volume Value: 11 – Type: issue Value: 4 Titles: – TitleFull: International Journal of Testing Type: main |
| ResultId | 1 |