Differential Item Functioning of the Scales for Assessing Emotional Disturbance-3 for White and African American Students

Saved in:
Bibliographic Details
Title: Differential Item Functioning of the Scales for Assessing Emotional Disturbance-3 for White and African American Students
Language: English
Authors: Lambert, Matthew C. (ORCID 0000-0002-7387-3780), Martin, Jodie, Epstein, Michael H., Cullinan, Douglas, Katsiyannis, Antonis
Source: Psychology in the Schools. Mar 2021 58(3):553-568.
Availability: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
Peer Reviewed: Y
Page Count: 16
Publication Date: 2021
Document Type: Journal Articles
Reports - Research
Education Level: Elementary Education
Secondary Education
Descriptors: Disability Identification, Rating Scales, Test Items, Item Response Theory, Psychometrics, Emotional Disturbances, Behavior Disorders, Special Education, Scores, Test Bias, Racial Differences, African American Students, White Students, Elementary School Students, Secondary School Students
DOI: 10.1002/pits.22463
ISSN: 0033-3085
Abstract: The present study investigated the psychometric properties of the "Scales for Assessing Emotional Disturbance -- Third Edition: Rating Scale" (SAED-3 RS), which is designed for use in identifying students with emotional disturbance for special education services. The purposes of this study were to evaluate (a) the measurement invariance of SAED-3 RS scores between White and African American students and (b) the impact of differential item functioning (DIF) on test scores from the SAED-3 RS. The sample consisted of 855 K-12 students from throughout the United States. The findings suggested that SAED-3 RS items exhibited small to negligible levels of DIF and that DIF did not significantly impact scores. The results supported the SAED-3 RS, a teacher-completed rating scale, as relatively consistent in measuring the emotional and behavioral status of school-age students from different racial backgrounds. Researchers and practitioners can have confidence that scores from the SAED-3 RS are not substantially affected by DIF when assessing the emotional and behavioral functioning of African American and White school-age students. Research limitations, future research, and implications for school professionals are discussed.
Abstractor: As Provided
Entry Date: 2021
Accession Number: EJ1284319
Database: ERIC
Full text is not displayed to guests.
FullText Links:
  – Type: pdflink
    Url: https://content.ebscohost.com/cds/retrieve?content=AQICAHj0k_4E0hTGH8RJwT4gCJyBsGNe_WN95AvKlDbXJGqwxwEqjbP1NTObzBeVNdR1VnWVAAAA4zCB4AYJKoZIhvcNAQcGoIHSMIHPAgEAMIHJBgkqhkiG9w0BBwEwHgYJYIZIAWUDBAEuMBEEDMXNOFL75ShBpvdaBAIBEICBm-vyZU3faK2SE12DYgkctL0KN299cqqSNzk5BuQ3y9_VXosnJuh2p2QsW9z39XHYOvhbubanjEsra1mmabRpDdZYesmAkQyh0a7cwDKU75SHBwWafI1KwRvsmXq22nG7fEWBBC600rYkSMxlSnyfULCjYVrFF3_6ZhbFRcM6SOv1LvdpDOo2pzqOlpBnWmsUAinT3sA_BO4YoSbX
Text:
  Availability: 1
  Value: <anid>AN0148430754;pis01mar.21;2021Feb03.03:30;v2.2.500</anid> <title id="AN0148430754-1">Differential item functioning of the Scales for Assessing Emotional Disturbance‐3 for White and African American students </title> <p>The present study investigated the psychometric properties of the Scales for Assessing Emotional Disturbance – Third Edition: Rating Scale (SAED‐3 RS), which is designed for use in identifying students with emotional disturbance for special education services. The purposes of this study were to evaluate (a) the measurement invariance of SAED‐3 RS scores between White and African American students and (b) the impact of differential item functioning (DIF) on test scores from the SAED‐3 RS. The sample consisted of 855 K‐12 students from throughout the United States. The findings suggested that SAED‐3 RS items exhibited small to negligible levels of DIF and that DIF did not significantly impact scores. The results supported the SAED‐3 RS, a teacher‐completed rating scale, as relatively consistent in measuring the emotional and behavioral status of school‐age students from different racial backgrounds. Researchers and practitioners can have confidence that scores from the SAED‐3 RS are not substantially affected by DIF when assessing the emotional and behavioral functioning of African American and White school‐age students. Research limitations, future research, and implications for school professionals are discussed.</p> <p>Keywords: behavior disorders; differential item functioning; Scales for Assessing Emotional Disturbance</p> <p>A significant number of school‐age children in the United States demonstrate emotional or behavioral challenges, including difficulties that reflect clinical criteria for a mental disorder. Numerous research reports and government policy papers state that about 14%–20% of school‐age children have at one time a mental health problem (e.g., Ghandour et al., 2019; Jaffee et al., 2005; Merikangas et al., 2010; National Research Council and Institute of Medicine [NRC and IoM], 2009). Such problems obviously affect children, families, and peers in and out of school. Unfortunately, a very small number of these students receive mental health services, and an even smaller number of students receive special education services for emotional and behavioral disorders. Specifically, less than 1% of school‐age children are school identified with an emotional disturbance (ED) and are afforded special education or related services under the Individuals with Disabilities Education Improvement Act of 2004 (IDEA, 2004; Mitchell et al., 2019; United States Department of Education, Office of Special Education Programs, 2020). Over the past several years, this rate has remained remarkably stable (Kauffman & Landrum, 2018; United States Department of Education, Office of Special Education Programs, 2020). The discrepancy between the number of school‐age children with mental health challenges and the number who are school identified with ED suggests that more efforts and improvements need to focus on accurately screening, identifying, and serving students with or at risk of ED.</p> <p>Students with ED tend to demonstrate extremely poor educational and life outcomes (Bradley et al., 2008). For instance, students with ED demonstrate significant levels of academic underperformance in all academic content areas, and this underperformance increases with age (Trout et al., 2003). In addition, these students tend to underperform their same‐age peers with and without disabilities. In school, students with ED receive poor grades, numerous course, and grade failures, and experience high levels of school behavior referrals, absenteeism, suspensions, and expulsions (United States Department of Education, Office of Special Education Programs, 2020; Wagner et al., 2005). Students with ED are much more likely than their peers to drop out of school (United States Department of Education, Office of Special Education Programs, 2020; Wagner et al., 2003). Postschool students with ED are likely to demonstrate heightened levels of unemployment, involvement with the criminal justice system, and substance dependency and abuse (Kauffman & Landrum, 2018; Sanford et al., 2011; Wagner et al., 2005).</p> <p>Test scores that are psychometrically sound enable school professionals to begin assessment such as screening and identification efforts as early as possible. When students with or at risk of ED go unidentified and underserved, there is a greater probability that their initial emotional and behavioral problems will persist throughout school into adulthood and result in greater mental health challenges (Costello et al., 2003; Essex et al., 2009). On the other hand, researchers have reported that proper identification and intervention for school‐age children who demonstrate at‐risk problems can prevent or lessen the level of emotional and behavioral problems (Conroy et al., 2004; Mrazek & Mrazek, 2005; NRC & loM, 2009). In light of the research indicating the importance of school‐based identification and intervention, school administrators and policymakers have advocated for psychometrically sound assessment tools to identify students who are at risk for ED (Mitchell et al., 2019). To this end, assessment experts have identified a number of essential elements for identification instruments: (a) adequacy of the intended use, (b) acceptable psychometric functioning, and (c) consumer usefulness and acceptability (Glover & Albers, 2007).</p> <p>Along with assessment instruments meeting these criteria, tests used in the identification of students with ED need to align with the criteria included in federal legislation, specifically in the IDEA. The federal definition for ED was written almost 50 years ago and has received considerable criticism over the years. The primary criticisms leveled at the definition are that it includes vague and ambiguous terms, lacks research support, includes arbitrary exclusionary clauses, and is outdated (e.g., Florell, 2018; Forness & Knitzer, 1992; Hanchon & Allen, 2013, 2018; Merrell & Walker, 2004; Skiba & Grizzle, 1992). Nonetheless, the federal definition of ED in IDEA has not been significantly altered over the past 50 years and yet is to be used by school personnel in the identification of students with ED (Becker et al., 2011). In light of these concerns with the definition, determining whether a specific student meets the definitional criteria is a challenge to educators, both philosophically and practically. While there are many adequate assessment instruments available to measure emotional and behavior problems of school‐aged students (see Achenbach & Edelbrock, 1983; Goodman, 2001; Reynolds & Kamphus, 2015), few are directly linked to the criteria of ED as defined in the IDEA. The failure to directly link instrument structure and IDEA definition characteristics and criteria make the identification process of students with ED especially challenging (e.g., Algozzine, 2017; Becker et al., 2011; Mitchell et al., 2019).</p> <p>One assessment instrument designed specifically to assist school professionals in the screening and identification of students with ED is the <emph>Scales for Assessing Emotional Disturbance‐3</emph> (SAED‐3; Epstein et al., 2020). The SAED‐3 consists of several assessment components including the 45‐item Rating Scale (SAED‐RS), which could be used as <emph>part</emph> of a school‐based system to identify students with ED. The SAED‐3 RS is a standardized, norm‐referenced instrument developed to operationalize the five primary characteristics and other essential components of the federal definition of ED. The SAED‐3 RS contains 45 rated items that measure emotional, social, and behavior problems. The five characteristics (a)–(e) in the ED definition are reflected in five of the SAED‐3 RS's subscales: <emph>Inability to Learn, Relationship Problems, Inappropriate Behavior, Unhappiness or Depression</emph>, and <emph>Physical Symptoms or Fears</emph>. A sixth RS subscale (<emph>Socially Maladjusted</emph>) helps to determine whether the student manifests social maladjustment and a further single item assists in making the "adversely affects educational performance" decision called for.</p> <p>There is evidence that the scores from the SAED‐3 RS have strong internal consistency, test–retest reliability, interrater reliability, criterion validity, convergent validity, and factor structure (Cullinan et al., 2002; Epstein et al., 1999; Epstein, Cullinan, et al., 2002; Epstein, Nordness, et al., 2002). Furthermore, the RS scores appear accurate in discriminating students at risk of having ED across subgroups of students based on age, sex, and student ability group (Epstein & Cullinan, 1998, 2010). These psychometric findings were replicated in the initial research of the SAED‐3 RS (Epstein et al., 2020). Together these findings suggest that the SAED‐3 RS scores demonstrate acceptable psychometric properties with the targeted population of students; however, researchers have not yet evaluated whether the SAED‐3 RS scores are biased in assessing the emotional and behavioral risk for students of different racial backgrounds.</p> <p>Disproportionality in special education, in general, and in ED, in particular, has been a persistent concern (see Zhang et al., 2014). Disproportionality occurs when members of a particular demographic group are identified in a proportion that is substantially less or greater than what should be expected (National Association of School Psychologists, 2013). For example, according to a "risk ratio" method for calculating disproportionality, <emph>Black or African American</emph> students are two times as likely to have been identified with ED (United States Department of Education, Office of Special Education Programs, 2020).</p> <p>Some authorities state or imply that disproportional identification of African American students as ED demonstrates that the process of identification, and/or the professionals who carry out that process, are racially biased (Harry & Klingner, 2006; Mitchell et al., 2019; National Association of School Psychologists, 2013). An important way to address racial bias in the process of identification is to ensure the assessment tools and procedures that underlie decision‐making contain little to no racial bias or other discriminatory bias. Specifically, instruments that contribute to decisions about students identified with ED should contain few, if any, racially‐biased items. Whether or not scores from an assessment instrument demonstrate consistency across different groups of students is referred to as psychometric measurement invariance (Cole & Zieky, 2001), which is essential for comparing students from diverse backgrounds to the same set of normative standards. Selecting an assessment that yields scores that have demonstrated measurement invariance is one important component, but not the only component (e.g., testing in primary language) necessary for nondiscriminatory evaluations.</p> <p>Differential item functioning (DIF), one way to quantify a lack of measurement invariance, occurs when two students with the same underlying emotional and behavioral problems receive different ratings for an item. DIF may suggest that items are measuring a different trait across demographic subgroups, perhaps because the items have different meanings for each subgroup or that raters (i.e., teachers) perceive certain behaviors differently across subgroups. The cumulative result of an assessment with items demonstrating DIF could be that there are systematically lower or higher scores for specific subgroups leading to overidentification or underidentification of certain subgroups (American Educational Research Association, American Psychological Association, & National Council on Measurement in Education, 2014). It is important to highlight that DIF does not reflect "true" group differences but reflects confounded measurement (i.e., construct‐irrelevant variance). Therefore, DIF methods are appropriate for use to assess whether the underlying traits of emotional and behavioral problems are measured consistently across students from different racial or ethnic backgrounds. To build on the previous studies of psychometric functioning of the SAED‐3 RS scores, the purposes of this study were to evaluate (a) the measurement invariance of SAED‐3 RS items between White and African American students and (b) the impact of DIF on test scores from the RS.</p> <hd id="AN0148430754-2">METHOD</hd> <p></p> <hd id="AN0148430754-3">Participants</hd> <p>Participants were drawn from the normative sample for the development of the SAED‐3 (Epstein et al., 2020). The normative sample included a total of 1430 students (53% males, 47% females) ranging from 5 to 18 years of age (<emph>M</emph> = 11.58, <emph>SD</emph> = 4.03). Other important characteristics of the sample indicated that it was nationally representative to a substantial degree in terms of race–ethnic status (White = 72%; African American = 17%; Asian = 3%; Native American = 2%; and two or more = 6%). Hispanic students comprised 22% of the sample. Students attended schools in 23 states representing all US geographic regions (Northeast = 17%; South = 40%, Midwest = 21; and West = 22%). A comparison of the sample percentages to those reported in the ProQuest Statistical Abstract of the United States, 2019 (ProQuest LLC, 2018) demonstrated that the normative sample was representative of the nation as a whole regarding the geographic region, gender, race, Hispanic status, exceptionality status, parent education level, and household income (see Epstein et al., 2020).</p> <p>The analytic sample included 855 students without exceptionalities who were identified as White/Non‐Hispanic (78.5%; <emph>n</emph> = 671) and African American/Non‐Hispanic (21.5%; <emph>n</emph> = 184). By and large, the analytic sample was similar to the overall normative sample. The mean age of students for the analytic sample was 11.45 years (<emph>SD</emph> = 4.09), and the sample was fairly evenly split between male (52.2%; <emph>n</emph> = 446) and female students (47.8%; <emph>n</emph> = 409). The majority of students in the analytic sample attended public schools (77.0%; <emph>n</emph> = 658), and fewer than 2% of the students (<emph>n</emph> = 17) reported a primary language other than English.</p> <p>We excluded students with exceptionalities (including students who had been school‐identified with ED) because we were concerned that some degree of bias might have contributed to the school‐identification of students thus confounding our analysis of DIF. It is important to note that scores for students included in the analytic sample spanned the entire range of the SAED‐3 RS scale. As would be expected for a representative sample, approximately 14% of students scored at least one <emph>SD</emph> above the population mean, so the analytic sample included a sufficient number of students who had high scores on the SAED‐3 RS.</p> <hd id="AN0148430754-4">Measures</hd> <p>The SAED‐3 RS is a 45‐item measure of the federal definition of Emotional Disturbance as defined by IDEA. The SAED‐3 RS consists of five core subscales (Inability to Learn, Relationship Problems, Inappropriate Behavior, Unhappiness or Depression, and Physical Symptoms or Fears) for ages 5 through 18, and one supplemental subscale (Socially Maladjusted) for ages 12 through 18. Each item is rated on a four‐point Likert scale by an individual familiar with the student's behavior for at least 2 months. Items are rated 0 = "not a problem," 1 = "mild problem," 2 = "considerable problem," and 3 = "severe problem." Items composing each of the subscales are summed to obtain a raw score which is transformed to a scaled score. Subscale scaled scores 13 or below are described as "not indicative of ED," subscale scaled scores of 14 to 16 are described as "indicative of ED," and subscale scaled scores 17 or above are described as "highly indicative of ED." A composite rating scale index combines the core subscale results to provide a global measure of general behavioral functioning. Finally, a scaled score of 14 or greater on the SM subscale (equal to the 91st percentile) indicates antisocial and delinquent behavior in the community, thereby indicating a need for services beyond those provided by IDEA. For this study, the items from the supplemental scale were not included in the analysis.</p> <hd id="AN0148430754-5">Data collection</hd> <p>Data were collected as part of the norming study of the SAED‐3 (Epstein et al., 2020). Before data being collected, two university Internal Review Boards (University of Nebraska‐Lincoln and Elon University) approved recruitment and data collection protocols. Data were collected from Fall 2015 through Spring 2018. The instrument's authors recruited teachers and other school personnel by mail or telephone. Those who volunteered to participate were instructed on how to complete the rating scale. They were asked to rate students whom they had in their class for at least 2 months and to provide student demographic information. They were asked to complete the rating scale on all of their students, or else to select an unbiased sample of their students using the following simple procedure: (a) decide how many students they wished to rate, (b) start either at the top or bottom of their class roster and select every other student to rate, and (c) stop selecting students when they reached the number they had decided to rate.</p> <hd id="AN0148430754-6">Data analysis plan</hd> <p>The data used in this study were the item‐level responses to each of the 39 individual items from five core subscales of the SAED‐3 RS that align with the federal definition of ED. While the primary purpose of this study was to examine the DIF of the RS item ratings, score dimensionality needed to be established before evaluating DIF. Score dimensionality was evaluated using <emph>Mplus</emph> v8 software (Muthén & Muthén, 2017). After establishing score dimensionality, DIF was examined through an iterative process combining item response theory modeling and logistic ordinal regression using the <emph>lordif</emph> package (Choi et al., 2011) in <emph>R</emph> statistical software.</p> <hd id="AN0148430754-7">Score dimensionality</hd> <p>The dimensionality of the data was assessed through a series of confirmatory factor analysis models: (a) single‐factor model, (b) correlated five‐factor model, and (c) bifactor version of the five‐factor model. For each model, the ratings were treated as categorical indicators and the WLSMV estimator was used. The fit of the models was evaluated primarily based on the alternative fit indices (AFIs): comparative fit index (CFI), Tucker‐Lewis index (TLI), and the root mean squared error of approximation (RMSEA). Models with CFI and TLI values greater than 0.95 (Browne & Cudeck, 1993) and RMSEA values less than 0.06 (Hu & Bentler, 1999) were considered to represent an acceptable fit to the data.</p> <p>The single‐factor model was used to examine the tenability of a strictly unidimensional interpretation of the SAED‐3 RS data. The correlated five‐factor model was used to assess the multidimensionality of the SAED‐3 RS data. Finally, the bifactor model was used to assess the tenability of a unidimensional interpretation of the data while allowing for content heterogeneity across items (i.e., multidimensionality).</p> <p>The bifactor model decomposes the item‐level variance into two sources: variance attributable to a <emph>general</emph> factor and (residual) variance attributable to a set of <emph>group</emph> factors (representing content heterogeneity or subscales). This model helps to evaluate whether the data are "unidimensional enough" to be used in an item response theory (IRT; Lord, 1980) modeling framework (Reise, 2012). To assess whether the data were unidimensional enough for IRT and DIF analyses, we computed the explained common variance (ECV) index and the omega hierarchical estimate for the <emph>general</emph> factor (see Reise, 2012 for an in‐depth description of bifactor models). The ECV index provides information on the ratio of explained variance attributable to the general factor compared to the group factors of the bifactor model—higher values indicate a greater proportion of variance attributable to the general factor. The interpretation of ECV relies on the percentage of uncontaminated correlations (PUC; Reise et al., 2013), the percentage of correlations between items that are due to <emph>only</emph> the general factor. Rodriguez et al. (2016) suggest that scores are "essentially unidimensional" when ECV > 0.70 and PUC > 0.70, and Reise et al. (2013) suggest that when PUC is greater than 0.80 then an ECV index less than 0.70 indicates limited bias when estimating a model as unidimensional. Finally, the omega hierarchical (<emph>ω</emph><subs>H</subs>) estimate provides information about how much <emph>raw</emph> score variance is attributable to the general factor—that is to say, omega hierarchical represents the proportion of the observed score that reflects a single, common factor (i.e., unidimensionality).</p> <hd id="AN0148430754-8">Detecting differential item functioning</hd> <p>To detect potential differential item functioning, we used a combination of item response theory (Lord, 1980) and logistic regression approaches built upon the framework developed by Swaminathan and Rogers (1990), which utilizes a comparison of logistic ordinal regression (LOR) models. While Swaminathan and Rogers (1990) used observed scores as the "matching" criterion (i.e., the variable used to match on the underlying degree of emotional and behavioral problems), we used IRT <emph>person</emph> parameters from a two‐parameter graded response model (Samejima, 1968) as the matching criterion.</p> <p>After obtaining the IRT person parameters (i.e., theta), we developed two LOR models to evaluate the presence of DIF. In the first model (Equation 1), the IRT person parameter was used to predict the rating of each item. In the second model (Equation 2), the IRT person parameter, race, and the interaction between the IRT person parameter and the race was used to predict the rating of each item.</p> <p>1 <ephtml> <math altimg="urn:x-wiley:00333085:media:pits22463:pits22463-math-0001" display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mtext>logit</mtext><mi>P</mi><mo>(</mo><msub><mi>u</mi><mi>i</mi></msub><mo>≥</mo><mi>k</mi><mo>)</mo><mo>=</mo><msub><mi>α</mi><mi>k</mi></msub><mo>+</mo><msub><mi>β</mi><mn>1</mn></msub><mo>*</mo><mtext>theta</mtext></math> </ephtml></p> <p>2 <ephtml> <math altimg="urn:x-wiley:00333085:media:pits22463:pits22463-math-0002" display="block" xmlns="http://www.w3.org/1998/Math/MathML"><mtext>logit</mtext><mi>P</mi><mo>(</mo><msub><mi>u</mi><mi>i</mi></msub><mo>≥</mo><mi>k</mi><mo>)</mo><mo>=</mo><msub><mi>α</mi><mi>k</mi></msub><mo>+</mo><msub><mi>β</mi><mn>1</mn></msub><mo>*</mo><mtext>theta</mtext><mo>+</mo><msub><mi>β</mi><mn>2</mn></msub><mo>*</mo><mtext>race</mtext><mo>+</mo><msub><mi>β</mi><mn>3</mn></msub><mo>*</mo><mtext>theta</mtext><mo>*</mo><mtext>race</mtext></math> </ephtml></p> <p>The fit statistics for the two models were then compared using a <emph>Χ</emph><sups>2</sups> difference test to determine if the addition of race to the model significantly improved the ability to predict item‐level ratings. If the addition of race significantly improved the model's ability to predict the rating for an item, after controlling for the IRT person parameter (i.e., the underlying degree of emotional and behavioral problems), it would indicate that the item demonstrated potential DIF. If the addition of race did not significantly improve the model's ability to predict ratings, then this indicated that the ratings were not influenced by race after controlling for the IRT person parameter.</p> <p>This study required a total of 39 statistical comparisons (i.e., comparing two LOR models for each core subscale item). We adopted two per‐test significance levels to evaluate the statistical significance of the DIF analyses to strike a balance between Type I (i.e., false positives) and Type II errors (i.e., false negatives). We used the unadjusted per‐test significance levels of 0.01 and a family‐wise adjusted level of 0.0013 to evaluate items. When using the significance level of 0.01, we would expect to have a 32% chance of making one Type I error across the study (and a 6% chance of making two Type I errors), but a lower Type II error rate. When using the adjusted significance level of 0.0013, we would have the nominal 5% chance of making one Type I error across the study, but a higher Type II error rate.</p> <hd id="AN0148430754-9">Distinguishing meaningfulness of DIF</hd> <p>Once items were flagged for potential DIF based on the statistical significance of the <emph>Χ</emph><sups>2</sups> difference between LOR models, the practical significance was evaluated by examining the change in McFadden <emph>R</emph><sups>2</sups>, which has been suggested as an effect size measure (Zumbo, 1999). The change in <emph>R</emph><sups>2</sups> (denoted as Δ<emph>R</emph><sups>2</sups>) represents the difference between the <emph>R</emph><sups>2</sups> values for each of the models described above. Our interpretation of Δ<emph>R</emph><sups>2</sups> was guided by Jodoin and Gierl (2001) who suggested that a difference of less than 0.035 indicates negligible DIF, a difference between 0.035 and 0.069 indicates moderate DIF, and a difference of greater than 0.069 indicates large DIF.</p> <p>The meaningfulness of DIF was also examined by plotting the item characteristic curves for DIF items. These plots display the difference in the expected rating for White and African American students across the range of possible rating scale scores. These plots help to identify for which group the differential item functioning "favors"—the group that received the lower than expected score (lower scores represent less likelihood of ED). Because one DIF item might "favor" White students and another item might "favor" African American students, we also plotted the test characteristic curves (TCCs) for White and African American students across all of the DIF items to understand the cumulative effect of DIF in terms of which group was "favored"—the group for which the TCC indicates a lower than expected score (across DIF items) when matched on their underlying level of emotional and behavioral difficulties.</p> <hd id="AN0148430754-10">Evaluating individual‐level impact of DIF</hd> <p>We also evaluated the degree to which DIF impacted the overall RS score (scaled as an IRT person parameter) using the framework developed by Choi et al. (2011). This required a two‐step process: (a) estimating "naïve" IRT person parameters using all of the SAED‐3 RS core subscale items, and (b) estimating "purified" IRT person parameters accounting for DIF items by using group‐specific item characteristic curves for the DIF items. The two sets of IRT person parameters were equated (i.e., scaled in a common metric) using the Stocking‐Lord approach (Stocking & Lord, 1983) so that the two sets of parameters could be directly compared. Then the difference between the unadjusted (naïve) and adjusted (purified) parameter for each student was computed and compared to the median standard error (<emph>SE</emph>) of the naïve scores as well as to the individual students' naïve standard error. If the difference in a student's IRT person parameter was greater than the median naïve <emph>SE</emph> or their own naïve <emph>SE</emph> (i.e., the uncertainty of their naïve score), then the impact of DIF was considered salient.</p> <hd id="AN0148430754-11">Missing data</hd> <p>Missing data were minimal (<1%). Missing values were accounted for differently depending on the specific set of analyses. For the CFA models, a pairwise‐present approach was used as is the default in Mplus when estimating models with WLSMV. For the IRT modeling analysis, missing data were included through maximum likelihood estimation. For the ordinal regression analyses, missing data were excluded using a pairwise‐present approach because the dependent variables in the analyses were observed item ratings.</p> <hd id="AN0148430754-12">RESULTS</hd> <p>The primary purpose of this study was to examine each of the 39 SAED‐3 RS core subscale items to understand the degree of DIF in the assessment. However, before DIF was examined, the dimensionality of the scores was established using confirmatory factor analysis to aid in the interpretation of the DIF findings. Table 1 presents the results of the confirmatory factor analysis models including the <emph>Χ</emph><sups>2</sups> goodness‐of‐fit test and the alternative fit indices. Note that the <emph>Χ</emph><sups>2</sups> difference tests were not reported because the focus of this series of CFA models was not on the fit of nested models.</p> <p>1 TableConfirmatory factor analysis model fit indicators</p> <p> <ephtml> <table><thead valign="bottom"><tr valign="bottom"><th /><th align="left"><italic>χ</italic><sup>2</sup><sub><italic>(df)</italic></sub></th><th align="left">CFI</th><th align="left">TLI</th><th align="left">RMSEA [90% CI]</th></tr></thead><tbody valign="top"><tr><td align="left">Single‐factor</td><td align="left">3874.92<sub>(702)</sub></td><td char="." align="char">0.893</td><td char="." align="char">0.887</td><td char="(" align="char">0.073 [0.070, 0.075]</td></tr><tr><td align="left">Five‐factor</td><td align="left">1981.67<sub>(692)</sub></td><td char="." align="char">0.956</td><td char="." align="char">0.953</td><td char="(" align="char">0.047 [0.044, 0.049]</td></tr><tr><td align="left">Bifactor version</td><td align="left">1731.26<sub>(663)</sub></td><td char="." align="char">0.964</td><td char="." align="char">0.960</td><td char="(" align="char">0.043 [0.041, 0.046]</td></tr></tbody></table> </ephtml> </p> <p>1 Abbreviations: CFI, comparative fit index; RMSEA, root mean square error of approximation; TLI, Tucker‐Lewis index.</p> <p>The single‐factor model, which evaluates the strict unidimensionality of the RS data, fit the data poorly according to the CFI and TLI measures (CFI = 0.893, TLI = 0.887, and RMSEA = 0.073). The correlated five‐factor model fit the data well according to all of the alternative fit indices (CFI = 0.956, TLI = 0.953, and RMSEA = 0.047). The bifactor version of the correlated five‐factor model indicated the closest fit to the data across each of the AFIs (CFI = 0.964, TLI = 0.960, and RMSEA = 0.043).</p> <p>The bifactor version of the five‐factor model indicated that the explained common variance (ECV) index was 0.69—that is, 69% of the variance that the model can explain was attributable to the <emph>general</emph> factor. Furthermore, when the <emph>general</emph> factor parameters (i.e., factor loadings) were compared to the single‐factor parameters, the similarity between parameters revealed that the general factor was the <emph>same</emph> as the single‐factor. This makes sense because the percentage of uncontaminated correlations (PUC) was 0.82 indicating that 82% of correlations between RS items were due to only the general factor. In addition, the omega hierarchical reliability coefficient for the general factor was 0.907—that is, nearly 91% of the <emph>raw</emph> score variance was attributable to the general factor. Taken together, these pieces of evidence suggest that the SAED‐3 RS data are "unidimensional enough" to be analyzed in a unidimensional item response theory model for the purpose of examining DIF (Reise, 2012; Reise et al., 2013; Rodriguez et al., 2016).</p> <hd id="AN0148430754-13">Differential item functioning</hd> <p>Four of the items met the <emph>p</emph> < 0.01 criteria for potential DIF: Item 7: <emph>Anxious, worried, tense</emph> (<emph>p</emph> = 0.0013; Δ<emph>R</emph><sups>2</sups> = 0.011), Item 8: <emph>Verbally abuses, teases, taunts people</emph> (<emph>p</emph> = 0.0009; Δ<emph>R</emph><sups>2</sups> = 0.018), Item 16: <emph>Has feelings of worthlessness</emph> (<emph>p</emph> = 0.0020; Δ<emph>R</emph><sups>2</sups> = 0.0005), and Item 18: <emph>Makes threats to others</emph> (<emph>p</emph> = 0.0001; Δ<emph>R</emph><sups>2</sups> =0 .032). When using the family‐wise adjusted significance level (0.0013), three items would have been identified (<reflink idref="bib7" id="ref1">7</reflink>, 8, and 18). However, all of the items demonstrated effect sizes (Δ<emph>R</emph><sups>2</sups>) indicating a negligible degree of DIF (Δ<emph>R</emph><sups>2</sups> < 0.035; Jodoin & Gierl, 2001).</p> <p>For a visual representation of DIF at the item‐level, the item characteristic curves of these four items were plotted in Figure 1. White students were given higher than expected ratings on Items 7 and 16 while African American students were given higher than expected ratings on Items 8 and 18. Note that nonuniform DIF was observed, so there are some minor exceptions to those patterns (e.g., for Item 26, African American students received a slightly higher than expected rating for individuals with an IRT score between approximately −1.00 and 0.75). Remembering that higher ratings represent a greater likelihood of identification as ED, two items "favor" White and two items "favor" African American students.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/PIS/01mar21/pits22463-fig-0001.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pits22463-fig-0001.jpg" title="1 Item characteristic curves for items demonstrating DIF. Note: (a) Item 8, (b) Item 18, (c) Item 16, (d) Item 7. DIF, differential item functioning" /> </p> <p></p> <p>For a visual representation of DIF at the test‐level, the TCCs for the DIF items were plotted in Figure 2. As can be seen in the plot, there was virtually no difference between the two TCCs except for students with higher ratings between the IRT score of approximately 1.75 and 4.00. For individuals in this range, African American students were slightly "favored" by the differential item functioning—they received marginally lower than expected scores.</p> <p> <img src="https://imageserver.ebscohost.com/img/embimages/rdk/PIS/01mar21/pits22463-fig-0002.jpg?ephost1=dGJyMNXb4kSepq84yOvqOLCmsE6epq5Srqa4SK6WxWXS" alt="pits22463-fig-0002.jpg" title="2 Test characteristic curves for DIF items. DIF, differential item functioning" /> </p> <p></p> <hd id="AN0148430754-16">Individual‐level impact of DIF</hd> <p>The IRT person parameter estimates from the purified model showed only minor differences when compared to the estimates from the naïve model. The differences between IRT estimates at the individual level ranged from −0.15 to 0.07 with a mean difference of −0.01. None of the individual differences were larger than the median standard error of the naïve estimates (<emph>Mdn</emph> = 0.281; range: 0.12 to 0.56) nor were any of the differences larger than an individual's naïve standard error (the individual with the −0.15 difference in IRT estimates had a naïve standard error of 0.30). Therefore, none of the differences in IRT estimates from the purified and naïve models reached the thresholds for indicating a salient impact of DIF on individual scores.</p> <hd id="AN0148430754-17">DISCUSSION</hd> <p>Given the poor educational and life outcomes for school‐age students identified with ED (Bradley et al., 2004; Wagner et al., 2005), it is essential that psychometrically sound assessment instruments are used to screen and identify students as eligible for special education and support services to provide timely and appropriate intervention. The primary purpose of this study was to examine whether the SAED‐3 RS items demonstrated measurement invariance across White and African American students. The analyses demonstrated that the items on the SAED‐3 RS for students were largely invariant and that teachers generally rated White and African American students as expected given the students' underlying emotional and behavioral difficulties. Overall, there were only four items that demonstrated DIF between White and African American students. Moreover, the degree of DIF reported was small and considered trivial. Thus, the overall DIF findings indicate that African American and White students with comparable SAED‐3 RS scores were not rated significantly differently on most assessment items (e.g., items that did not demonstrate significant DIF), and that the cumulative effect of certain items that demonstrated DIF did not saliently impact the SAED‐3 RS scale scores. Both of these findings support the psychometric quality of the SAED‐3 scores, the comparability of scores across White and African American students, and the appropriateness of the normative standards provided in the SAED‐3 Examiner's Manual (Epstein et al., 2020) which is an important aspect of non‐discriminatory evaluation.</p> <p>Research conducted on the previous editions of the SAED‐RS (Epstein & Cullinan, 1998, 2010) demonstrated the evidence of reliability (Epstein et al., 1999), validity (Cullinan et al., 2002; Epstein, Cullinan, et al., 2002) and an underlying theoretical structure for measuring ED (Epstein, Nordness, et al., 2002). However, in the previous studies and in earlier editions of the RS, no research was reported on whether the RS items were biased in measuring the emotional and behavioral problems of students exhibited by students from diverse racial and ethnic backgrounds. The results from this study identified that a few RS items resulted in significant DIF yet had a small or trivial impact. The present study filled an important gap in the psychometric research on the SAED‐3 RS by determining the measurement invariance of the SAED‐3 RS items for White and African American students from a large nationally representative sample. The findings afford an important first step in indicating that the SAED‐3 RS items appear to be relatively non‐biased in measuring the emotional and behavioral problems of students, particularly African American students.</p> <p>Examining differential item functioning is a useful approach for understanding construct‐irrelevant variance (i.e., bias) related to racial, ethnic, or gender groups when using assessments for making placement decisions. DIF and other measures of bias are particularly important when items are rated by individuals who have limited experience with the behavioral norms of a subgroup. When considering rating scales, it is imperative that test users remember that ratings are the interpretations of the rater and <emph>not</emph> a direct measure of the behavior being considered. Therefore, this study is best understood as an evaluation of bias in the items when rated by teachers. Whether items demonstrate DIF, it should be understood that the bias lies in the rating process, and not necessarily in the student's behavior. In the context of this framework, the findings of this study are positive and encouraging—only four items out of the 39 SAED‐3 RS items were flagged for potential bias, and the degrees of bias for those items were negligible.</p> <hd id="AN0148430754-18">Limitations and future research</hd> <p>Several limitations of this study should be noted. First, a noteworthy limitation of this study related to sample size—in particular, the sample of African American students. A general guideline for the minimal sample size for estimating reliable parameters in the grade response IRT model is 250 participants (Reeve & Fayer, 2005), and a minimal sample size guideline for DIF analyses is 200 participants per group (Zieky, 1993). The current sample of African American students did not satisfy either threshold. However, in light of our sample approaching both of these guidelines, we weighed the benefit of investigating DIF against the potential consequence of not, and decided it was better to investigate DIF (with a smaller sample size) than to ignore the possibility of DIF given the reality of overidentification (of African American students) and the potential role of bias in assessing ED. With that said, the results of the DIF analysis should be considered with caution and future research should seek to replicate the findings with larger samples.</p> <p>Second, the present study only investigated African American and White students. Thus, the findings may not be applicable to students from other diverse backgrounds including Hispanic, Asian/Native Hawaiian/Other Pacific Islander, American Indian/Alaska Native, and multiracial groups. In light of the estimates that the US student population is likely to become more racially and ethnically diverse (Hussar & Bailey, 2016), standardized assessment instruments must demonstrate that they are acceptable for use with students of various multicultural backgrounds. To that end, future researchers need to assess the measurement invariance of the SAED‐3 RS across other racial and ethnic groups. Third, this study only examined the measurement invariance of the SAED‐3 RS with African American and White students with no analysis of gender. Future investigators need to replicate this study focusing separately on female and male students. Fourth, the socioeconomic status (SES) of students is an important demographic variable that was not evaluated, which is shortsighted as SES is frequently associated with the school behavior of students (e.g., Cholewa et al., 2018). Fifth, no information on the demographic characteristics of the teachers who provided the student ratings was collected. Future investigators need to collect teacher information such as gender, age, race, terminal degree, and years taught, and determine how these variables may be related to teacher ratings. Finally, while this study evaluated racial bias in items on SAED‐3 RS, it did not measure the potential bias of the scores when used at different phases of assessment, specifically the referral, identification, progress monitoring, and program evaluation phases of assessment. Future researchers need to extend this study in these areas.</p> <p>While the outcome of this study contributed to understanding the psychometric properties of items and scores from the SAED‐3 RS, further research is needed. Future researchers need to determine the diagnostic utility of the SAED‐3 RS scores—how well do the scores discriminate between students with and without ED? Similarly, future studies might investigate the predictive validity of the SAED‐3 RS scores with important criterion‐related outcomes, such as student engagement, suspensions, office discipline referrals, grades, to name a few. Also, young students without an educational diagnosis of ED could be evaluated using the SAED‐3 RS; then after a significant time period such as 1 or 2 years, the test could be reapplied to see how well the SAED‐3 RS predicted which students were identified with ED or developed other educational issues. If a student's score on the SAED‐3 RS was able to predict his or her classification in later years, the SAED‐3 RS's value in school screening and identification efforts would be determined. Other investigators may wish to examine the cross‐informant reliability of the SAED‐3 RS across different groups of respondents such as teachers and parents. This type of research would measure how similarly adults with different roles and experiences rate the emotional and behavioral problems of students as measured by the SAED‐3 RS.</p> <p>The findings from the factor analysis highlight another area for future research and an issue that the developers of the SAED‐3 RS may need to address. The developers of the SAED‐3 RS suggest that the ratings form five subscale scores and an overall global score; however, the developers only suggest that schools use the subscale scores when identifying students. The results of the factor analysis do not necessarily support that approach, but also do not directly refute that approach. The factor analysis results do, however, suggest that the subscale scores are moderate to highly correlated with one another and that a substantial proportion of explained variance is attributable to a general factor rather than subscale factors (i.e., the subscale score may lack discriminant validity). Future research should replicate the factor analysis reported in this study and, if the findings are replicable, the developers of the SAED‐3 RS should reconsider how best to use the subscale and global scores.</p> <p>Finally, as prior research has documented that teachers are more likely to judge African American students as demonstrating greater levels of externalizing behaviors than White students (e.g., DuPaul et al., 2016), one area for investigators to examine is related to the possible interactive effects of student race and teacher race upon teacher ratings of student emotion and behavior. Specifically, researchers have reported that White teachers are more likely than African American teachers to rate African American students as having greater externalizing problems (Bates & Glick, 2013). Therefore, future researchers need to assess the potential impact of teachers' race and ethnicity on RS scores.</p> <hd id="AN0148430754-19">Implications for practice</hd> <p>The outcomes of this study add to the existing research of the SAED‐3 RS and provide important data that the SAED‐3 RS appears to be an assessment instrument with acceptable psychometric properties. In this investigation, it was documented that the SAED‐3 RS appears to be consistent in assessing the emotional and behavioral functioning of school students of White and African American backgrounds. We concur with other researchers (e.g., Bruhn et al., 2014) who recommend that school psychologists, teachers, and other school personnel use assessments like the SAED‐3 RS within systematic identification programs so that students at risk of ED are identified at the appropriate time.</p> <p>For many reasons, the SAED‐3 RS may be a useful tool to facilitate determinations of eligibility for special education under the category of ED. First, the SAED‐3 RS is a reliable and valid measure, which meets the key safeguard under IDEA that calls for the use of technically sound instruments in eligibility determinations. Second, this study provides support for the measurement invariance of the SAED‐3 RS scores across White and African American students. In light of persistent concerns regarding disproportionality in the identification of students for special education, the use of an instrument without item bias is a critical consideration. The SAED‐3 RS scores for African American and White students were comparable and were not rated differentially on specific test items thus representing a non‐biased measure. Though we recognize the complexity of disproportionality and the multitude of factors involved, having instruments such as SAED‐RS with items free of bias in light of racial background is a significant, positive contribution. African American students represent a particularly vulnerable group facing academic (Hussar et al., 2020; National Assessment of Educational Progress, 2020a, 2020b) and behavioral challenges including disproportionate rates of exclusionary discipline and school arrests (Gage et al., 2019; United States Department of Education, Office for Civil Rights, 2019). Finally, this SAED‐3 is aligned with the current IDEA definition of and primary characteristics of emotional disturbance—Inability to Learn, Relationship Problems, Inappropriate Behavior, Unhappiness or Depression, and Physical Symptoms or Fears—which addresses limitations of other popular measures. Thus, the results may be compatible with a comprehensive evaluation to determine whether a referred student meets the eligibility criteria for special education services for ED.</p> <p>A final implication for the SAED‐3 RS is its relevance to the tiered prevention and intervention model, which has been recently advocated by numerous professional organizations, policymakers, and researchers to address the emotional and behavioral problems of students. The frameworks have been referred to as Positive Behavior Intervention and Support (PBIS), Schoolwide Discipline, or Response to Intervention (RTI; e.g., Jimerson et al., 2016; Kerr & Nelson, 2010; Lane et al., 2009; Stoiber & Gettinger, 2016; Sugai & Horner, 2006). Perhaps the most widely used term is Multitiered System of Support (MTSS) which advocates for the use of multiple levels of instruction for academic and behavioral development where students considered to be at low risk receive universal supports, students considered to have some risk receive targeted supports, and students considered high risk receive intensive, individualized supports.</p> <p>The National Association of School Psychologists (2010, 2013, 2016) has identified the basic guidelines for MTSS including the use of non‐biased, empirically based, psychometrically sound instruments in the identification, and decision‐making process. A key criterion of MTSS models is that in the typical school environment, students generally fit one of three levels of risk for emotional and behavioral problems: Tier 1—Universal support for all students; Tier 2—Targeted support for at‐risk students; and Tier 3—Intensive support for identified students. Moreover, each tier or level requires a different degree or approach to prevention or intervention. Thus, to change a student's designation from Tier 2 for at‐risk students to Tier 3 for students identified as needing intensive and individualized supports is a significant undertaking. Clearly, there can be major negative consequences for students who are incorrectly identified and placed in different behavioral or instructional programs or in more restrictive settings. Therefore, based on the accumulating research, it appears that the SAED‐3 RS can be part of a comprehensive assessment battery, along with other psychometrically sound instruments, interviews, and direct observation measures, in the decision‐making process by providing a psychometrically sound assessment based on teacher ratings.</p> <hd id="AN0148430754-20">CONFLICT OF INTERESTS</hd> <p>It is important to note that the third and fourth authors are developers of the SAED‐3, and receive royalties from sales of the assessment. The second author is currently employed by the SAED‐3 publisher and therefore has an indirect financial interest in the assessment. It is also important to note that the data were analyzed and interpreted by the lead author independently from the other authors. The lead author has neither a financial interest related to the SAED‐3 nor any other conflict of interest related to this study.</p> <ref id="AN0148430754-21"> <title> REFERENCES </title> <blist> <bibl id="bib1" type="bt">1</bibl> <bibtext> Achenbach, T. M., & Edelbrock, C. (1983). Child behavior checklist. University Associates in Psychiatry.</bibtext> </blist> <blist> <bibl id="bib2" type="bt">2</bibl> <bibtext> Algozzine, B. (2017). Toward an acceptable definition of emotional disturbance: Waiting for change. Behavioral Disorders, 42 (3), 136 – 144. https://doi.org/10.1177/0198742917702117</bibtext> </blist> <blist> <bibl id="bib3" type="bt">3</bibl> <bibtext> American Educational Research Association, American Psychological Association, & National Council on Measurement in Education. (2014). The standards for educational and psychological testing. AERA.</bibtext> </blist> <blist> <bibl id="bib4" type="bt">4</bibl> <bibtext> Bates, L. A., & Glick, J. E. (2013). Does it matter if teachers and schools match the student? Racial and ethnic disparities in problem behaviors. Social Science Research, 42 (5), 1180 – 1190. https://doi.org/10.1016/j.ssresearch.2013.04.005</bibtext> </blist> <blist> <bibl id="bib5" type="bt">5</bibl> <bibtext> Becker, D. V., Anderson, U. S., Mortensen, C. R., Neufeld, S. L., & Neel, R. (2011). The face in the crowd effect unconfounded: Happy faces, not angry faces, are more efficiently detected in single‐ and multiple‐target visual search tasks. Journal of Experimental Psychology: General, 140 (4), 637 – 659. https://doi.org/10.1037/a0024060</bibtext> </blist> <blist> <bibl id="bib6" type="bt">6</bibl> <bibtext> Bradley, R., Doolittle, J., & Bartolotta, R. (2008). Building on the data and adding to the discussion: The experiences and outcomes of students with emotional disturbance. Journal of Behavioral Education, 17, 4 – 23. https://doi.org/10.1007/s10864-007-9058-6</bibtext> </blist> <blist> <bibl id="bib7" idref="ref1" type="bt">7</bibl> <bibtext> Bradley, R., Henderson, K., & Monfore, D. A. (2004). A national perspective on children with emotional disorders. Behavioral Disorders, 29 (3), 211 – 223. https://<ulink href="http://www.jstor.org/stable/23889470">www.jstor.org/stable/23889470</ulink></bibtext> </blist> <blist> <bibl id="bib8" type="bt">8</bibl> <bibtext> Browne, M. W., & Cudeck, R. (1993). Alternative ways of assessing model fit. In K. A. Bollen, & J. S. Long (Eds.), Testing structural equation models (pp. 136 – 162). Sage.</bibtext> </blist> <blist> <bibl id="bib9" type="bt">9</bibl> <bibtext> Bruhn, A. L., Wood‐Groves, S., & Huddle, S. (2014). A preliminary investigation of emotional and behavioral screening practices in K‐12 schools. Education and Treatment of Children, 37 (4), 611 – 634. https://doi.org/10.1353/etc.2014.0039</bibtext> </blist> <blist> <bibtext> Choi, S. W., Gibbons, L. E., & Crane, P. K. (2011). lordif: An R package for detecting differential item functioning using iterative hybrid ordinal logistic regression/item response theory and Monte Carlo simulations. Journal of Statistical Software, 39 (8), 1 – 30. https://doi.org/10.18637/jss.v039.i08</bibtext> </blist> <blist> <bibtext> Cholewa, B., Hull, M. F., Babcock, C. R., & Smith, A. D. (2018). Predictors and academic outcomes associated with in‐school suspension. School Psychology Quarterly, 33 (2), 191 – 199. https://doi.org/10.1037/spq0000213</bibtext> </blist> <blist> <bibtext> Cole, N. S., & Zieky, M. J. (2001). The new faces of fairness. Journal of Educational Measurement, 38 (4), 369 – 382. https://doi.org/10.1111/j.1745-3984.2001.tb01132.x</bibtext> </blist> <blist> <bibtext> Conroy, M. A., Hendrickson, J. M., & Sester, P. P. (2004). Early identification and prevention of emotional and behavioral disorders. In R. B. Rutherford, M. M. Quinn, & S. R. Mathur (Eds.), Handbook of research in emotional and behavioral disorders (pp. 199 – 215). Guilford Press.</bibtext> </blist> <blist> <bibtext> Costello, E. J., Mustillo, S., Erkanli, A., Keeler, G., & Angold, A. (2003). Prevalence and development of psychiatric disorders in childhood and adolescence. Archives of General Psychiatry, 60 (8), 837 – 844. https://doi.org/10.1001/archpsyc.60.8.837</bibtext> </blist> <blist> <bibtext> Cullinan, D., Harniss, M. K., Epstein, M. H., & Ryser, G. (2002). The scale for assessing emotional disturbance: Concurrent validity. Journal of Child and Family Studies, 10, 449 – 466. https://doi.org/10.1023/A:1016709407756</bibtext> </blist> <blist> <bibtext> DuPaul, G., Reid, R., Anastopoulos, A., Lambert, M. C., Watkins, M., & Power, T. (2016). Parent and teacher ratings of attention‐deficit/hyperactivity disorder symptoms: Factor structure and normative data. Psychological Assessment, 28 (2), 214 – 225. https://doi.org/10.1037/pas0000166</bibtext> </blist> <blist> <bibtext> Epstein, M. H., & Cullinan, D. (1998). Scale for Assessing Emotional Disturbance (SAED). PRO‐ED.</bibtext> </blist> <blist> <bibtext> Epstein, M. H., & Cullinan, D. (2010). Scales for Assessing Emotional Disturbance (2nd ed.). PRO‐ED.</bibtext> </blist> <blist> <bibtext> Epstein, M. H., Cullinan, D., Harniss, M. K., & Ryser, G. (1999). The scale for assessing emotional disturbance: Test‐retest and interrater reliability. Behavioral Disorders, 24 (3), 222 – 230. https://doi.org/10.1177/019874299902400301</bibtext> </blist> <blist> <bibtext> Epstein, M. H., Cullinan, D., Pierce, C., Huscroft‐D'Angelo, J., & Wery, J. (2020). Scales for Assessing Emotional Disturbance (3rd ed.). PRO‐ED.</bibtext> </blist> <blist> <bibtext> Epstein, M. H., Cullinan, D., Ryser, G., & Pearson, N. (2002). Development of a scale to assess emotional disturbance. Behavioral Disorders, 28 (1), 5 – 22. https://doi.org/10.1177/019874290202800101</bibtext> </blist> <blist> <bibtext> Epstein, M. H., Nordness, P. D., Cullinan, D., & Hertzog, M. (2002). Scale for assessing emotional disturbance: Long term test–retest reliability and convergent validity with kindergarten and first grade students. Remedial and Special Education, 23 (3), 141 – 148. https://doi.org/10.1177/07419325020230030201</bibtext> </blist> <blist> <bibtext> Essex, M., Kraemer, H. C., Slattery, M. J., Burk, L. R., Boyce, W. T., Woodward, H. R., & Kupfer, D. J. (2009). Screening for childhood mental health problems: Outcomes and early identification. Journal of Child Psychology and Psychiatry, 50, 562 – 570. https://doi.org/10.1111/j.1469-7610.2008.02015.x</bibtext> </blist> <blist> <bibtext> Florell, D. (2018). Historical foundations. In S. L. Grapin, & J. H. Kranzler (Eds.), School psychology: Professional issues and practices (pp. 21 – 41). Springer Publishing Co.</bibtext> </blist> <blist> <bibtext> Forness, S. R., & Knitzer, J. (1992). A new proposed definition and terminology to replace "serious emotional disturbance" in individuals with disabilities education act. School Psychology Review, 21 (1), 12 – 20. https://doi.org/10.1080/02796015.1992.12085587</bibtext> </blist> <blist> <bibtext> Gage, N. A., Whitford, D. K., Katsiyannis, A., Adams, S., & Jasper, A. (2019). National analysis of the disciplinary exclusion of black students with and without disabilities. Journal of Child and Family Studies, 28, 1754 – 1764. https://doi.org/10.1007/s10826-019-01407-7</bibtext> </blist> <blist> <bibtext> Ghandour, R. M., Sherman, L. J., Vladutiu, C. J., Ali, M. M., Lynch, S. E., Bitsko, R. H., & Blumberg, S. J. (2019). Prevalence and treatment of depression, anxiety, and conduct problems in U.S. children. The Journal of Pediatrics, 206, 256 – 267.</bibtext> </blist> <blist> <bibtext> Glover, T. A., & Albers, C. A. (2007). Considerations for evaluating universal screening assessments. Journal of School Psychology, 45, 117 – 135. https://doi.org/10.1016/j.jpeds.2018.09.021</bibtext> </blist> <blist> <bibtext> Goodman, R. (2001). Psychometric properties of the Strengths and Difficulties Questionnaire (SDQ). Journal of the American Academy of Child and Adolescent Psychiatry, 40 (11), 1337 – 1345. https://doi.org/10.1097/00004583-200111000-00015</bibtext> </blist> <blist> <bibtext> Hanchon, T. A., & Allen, R. A. (2013). Identifying students with emotional disturbance: School psychologists' practices and perceptions. Psychology in the Schools, 50 (2), 193 – 208. https://doi.org/10.1002/pits.21668</bibtext> </blist> <blist> <bibtext> Hanchon, T. A., & Allen, R. A. (2018). The identification of students with emotional disturbance: Moving the field toward responsible assessment practices. Psychology in the Schools, 55 (2), 176 – 189. https://doi.org/10.1002/pits.22099</bibtext> </blist> <blist> <bibtext> Harry, B., & Klingner, J. (2006). Why are so many minority students in special education? Understanding race and disability in schools. Teachers College Press.</bibtext> </blist> <blist> <bibtext> Hu, L., & Bentler, P. M. (1999). Cutoff criteria for fit indexes in covariance structure analysis: Conventional criteria versus new alternatives. Structural Equation Modeling, 6, 1 – 55. https://doi.org/10.1080/10705519909540118</bibtext> </blist> <blist> <bibtext> Hussar, W. J., & Bailey, T. M. (2016). Projections of Education Statistics to 2024: Forty‐third edition (NCES 2016‐013). US Department of Education, National Center for Educational Statistics, US Government Printing Office. https://nces.ed.gov/pubsearch/pubsinfo.asp?pubid=2016013</bibtext> </blist> <blist> <bibtext> Hussar, B., Zhang, J., Hein, S., Wang, K., Roberts, A., Cui, J., Smith, M., Bullock Mann, F., Barmer, A., & Dilig, R. (2020). The Condition of Education 2020 (NCES 2020‐144). National Center for Education Statistics. https://nces.ed.gov/pubsearch/pubsinfo.asp?pubid=2020144</bibtext> </blist> <blist> <bibtext> Individuals with Disabilities Education Improvement Act of 2004, Pub. L. No. 108‐446, 20 U.S.C. §1400 et seq. (2004).</bibtext> </blist> <blist> <bibtext> Jaffee, S. R., Harrington, H., Cohen, P., & Moffitt, T. E. (2005). Cumulative prevalence of psychiatric disorder in youths. Journal of the American Academy of Child and Adolescent Psychiatry, 44 (5), 406 – 407. https://doi.org/10.1097/01.chi.0000155317.38265.61</bibtext> </blist> <blist> <bibtext> Jimerson, S. R., Burns, M. K., & VanDerHeyden, A. M. (2016). Handbook of response to intervention: The science and practice of multi‐tiered systems of support (2nd ed.). Springer.</bibtext> </blist> <blist> <bibtext> Jodoin, M. G., & Gierl, M. J. (2001). Evaluating Type I error and power rates using an effect size measure with the logistic regressions procedure for DIF detection. Applied Measurement in Education, 14 (4), 329 – 349. https://doi.org/10.1207/S15324818AME1404_2</bibtext> </blist> <blist> <bibtext> Kauffman, J. M., & Landrum, T. J. (2018). Characteristics of emotional and behavioral disorders of children and youth (11th ed.). Pearson.</bibtext> </blist> <blist> <bibtext> Kerr, M. M., & Nelson, C. M. (2010). Strategies for addressing behavior problems in the classroom (6th ed.). Pearson Merrill.</bibtext> </blist> <blist> <bibtext> Lane, K. L., Kalberg, J. R., & Menzies, H. M. (2009). Developing schoolwide programs to prevent and manage problem behaviors: A step‐by‐step approach. Guilford Press.</bibtext> </blist> <blist> <bibtext> Lord, F. M. (1980). Applications of item response theory to practical testing problems. Erlbaum.</bibtext> </blist> <blist> <bibtext> Merikangas, K. R., He, J., Brody, D., Fisher, P. W., Bourdon, K., & Koretz, D. S. (2010). Prevalence and treatment of mental disorders among U.S. children in the 2001–2004 NHANES. Pediatrics, 125 (1), 75 – 81. https://doi.org/10.1542/peds.2008-2598</bibtext> </blist> <blist> <bibtext> Merrell, K. W., & Walker, H. M. (2004). Deconstructing a definition: Social maladjustment versus emotional disturbance and moving the EBD field forward. Psychology in the Schools, 41 (8), 899 – 910. https://doi.org/10.1002/pits.20046</bibtext> </blist> <blist> <bibtext> Mitchell, B. S., Kern, L., & Conroy, M. A. (2019). Supporting students with emotional or behavioral disorders: State of the field. Behavioral Disorders, 44 (2), 70 – 84. https://doi.org/10.1177/0198742918816518</bibtext> </blist> <blist> <bibtext> Mrazek, D., & Mrazek, P. J. (2005). Prevention of psychiatric disorders in children and adolescents. In B. J. Sadock, & V. A. Sadock (Eds.), Kaplan & Sadock's comprehensive textbook of psychiatry (Vol. 2, pp. 3513 – 3518). Lippincott Williams & Wilkins.</bibtext> </blist> <blist> <bibtext> Muthén, L. K., & Muthén, B. O. (2017). Mplus: Statistical analysis with latent variables: User's guide (Version 8). Muthén & Muthén.</bibtext> </blist> <blist> <bibtext> National Assessment of Educational Progress. (2020a). NAEP report card: Mathematics. https://<ulink href="http://www.nationsreportcard.gov/mathematics/nation/achievement?grade=4">www.nationsreportcard.gov/mathematics/nation/achievement?grade=4</ulink></bibtext> </blist> <blist> <bibtext> National Assessment of Educational Progress. (2020b). NAEP report card: Reading. https://<ulink href="http://www.nationsreportcard.gov/reading/nation/achievement?grade=4">www.nationsreportcard.gov/reading/nation/achievement?grade=4</ulink></bibtext> </blist> <blist> <bibtext> National Association of School Psychologists. (2010). Model for comprehensive and integrated school psychological services. National Association of School Psychologists.</bibtext> </blist> <blist> <bibtext> National Association of School Psychologists. (2013). Racial and ethnic disproportionality in education [Position statement]. National Association of School Psychologists.</bibtext> </blist> <blist> <bibtext> National Association of School Psychologists. (2016). Integrated model of academic and behavioral supports [Position statement]. National Association of School Psychologists.</bibtext> </blist> <blist> <bibtext> National Research Council and Institute of Medicine. (2009). Preventing mental, emotional, and behavioral disorders among young people: Progress and possibilities. National Academies Press.</bibtext> </blist> <blist> <bibtext> ProQuest LLC. (2018). ProQuest statistical abstracts of the United States, 2019 (7th ed.). Bernan.</bibtext> </blist> <blist> <bibtext> Reeve, B. B., & Fayer, P. (2005). Applying item response theory modeling for evaluating questionnaire item and scale properties. In P. M. Fayers, & R. D. Hays (Eds.), Assessing quality of life in clinical trials: Methods and practice (2nd ed., pp. 55 – 73). Oxford University Press.</bibtext> </blist> <blist> <bibtext> Reise, S. P. (2012). Invited paper: The rediscovery of bifactor measurement models. Multivariate Behavioral Research, 47 (5), 667 – 696. https://doi.org/10.1080/00273171.2012.715555</bibtext> </blist> <blist> <bibtext> Reise, S. P., Scheines, R., Widaman, K. F., & Haviland, M. G. (2013). Multidimensionality and structural coefficient bias in structural equation modeling: A bifactor perspective. Educational and Psychological Measurement, 73 (1), 5 – 26. https://doi.org/10.1177/0013164412449831</bibtext> </blist> <blist> <bibtext> Reynolds, C. R., & Kamphus, R. W. (2015). Behavior assessment system for children (3rd ed.). Pearson.</bibtext> </blist> <blist> <bibtext> Rodriguez, A., Reise, S. P., & Haviland, M. G. (2016). Applying bifactor statistical indices in the evaluation of psychological measures. Journal of Personality Assessment, 98 (3), 223 – 237. https://doi.org/10.1080/00223891.2015.1089249</bibtext> </blist> <blist> <bibtext> Samejima, F. (1968). Estimation of latent ability using a response pattern of graded scores. ETS Research Bulletin Series, 1968 (1), 1 – 169. https://doi.org/10.1002/j.2333-8504.1968.tb00153.x</bibtext> </blist> <blist> <bibtext> Sanford, C., Newman, L., Wagner, M., Cameto, R., Knokey, A. M., & Shaver, D. (2011). The post‐high school outcomes of young adults with disabilities up to 6 years after high school: Key findings from the National Longitudinal Transition Study‐2 (NLTS2) (NCSER 2011‐3004). SRI International.</bibtext> </blist> <blist> <bibtext> Skiba, R., & Grizzle, K. L. (1992). Qualifications v. logic and data: Excluding conduct disorders from the SED definition. School Psychology Review, 21 (1), 23 – 28. https://doi.org/10.1080/02796015.1992.12085589</bibtext> </blist> <blist> <bibtext> Stocking, M. L., & Lord, F. M. (1983). Developing a common metric in item response theory. Applied Psychological Measurement, 7 (2), 201 – 210. https://doi.org/10.1177/014662168300700208</bibtext> </blist> <blist> <bibtext> Stoiber, K. C., & Gettinger, M. (2016). Multi‐tiered systems of support and evidence‐based practices. In S. R. Jimerson, M. K. Burns, & A. M. VanDerHeyden (Eds.), Handbook of response to intervention : The science and practice of multi‐tiered systems of support (2nd ed., pp. 121 – 141). Springer.</bibtext> </blist> <blist> <bibtext> Sugai, G., & Horner, R. R. (2006). A promising approach for expanding and sustaining school‐wide positive behavior support. School Psychology Review, 35 (2), 245 – 259. https://doi.org/10.1080/02796015.2006.12087989</bibtext> </blist> <blist> <bibtext> Swaminathan, H., & Rogers, H. J. (1990). Detecting differential item functioning using logistic regression procedures. Journal of Educational Measurement, 27 (4), 361 – 370. https://doi.org/10.1111/j.1745-3984.1990.tb00754.x</bibtext> </blist> <blist> <bibtext> Trout, A. L., Nordness, P. D., Pierce, C. D., & Epstein, M. H. (2003). Research on the academic status of children with emotional and behavioral disorders: A review of the literature from 1961 to 2000. Journal of Emotional & Behavioral Disorders, 11 (4), 198 – 210. https://doi.org/10.1177/10634266030110040201</bibtext> </blist> <blist> <bibtext> United States Department of Education, Office for Civil Rights. (2019). 2015–16 Civil Rights Data Collection: School climate and safety. United States Department of Education, Office for Civil Rights.</bibtext> </blist> <blist> <bibtext> United States Department of Education, Office of Special Education Programs. (2020). 41st Annual Report to Congress on the Implementation of the Individuals with Disabilities Education Act, 2019. United States Department of Education, Office of Special Education Programs.</bibtext> </blist> <blist> <bibtext> Wagner, M., Cameto, R., & Newman, L. (2003). Youth with disabilities: A changing population. SRI International.</bibtext> </blist> <blist> <bibtext> Wagner, M., Kutash, K., Duchnowski, A. J., Epstein, M. H., & Sumi, W. C. (2005). The Special Education Elementary Longitudinal Study (SEELS) and the National Longitudinal Transition Study (NLTS2): Study designs and implications for children and youth with emotional disturbance. Journal of Emotional and Behavioral Disorders, 13, 25 – 41. https://doi.org/10.1177/10634266050130010301</bibtext> </blist> <blist> <bibtext> Wagner, M., Newman, L., Cameto, R., Garza, N., & Levin, P. (2005). After high school: A first look at the postschool experiences of youth with disabilities. A report from the National Longitudinal Transition Study‐2 (NLTS2). SRI International.</bibtext> </blist> <blist> <bibtext> Zhang, D., Katsiyannis, A., Ju, S., & Roberts, E. L. (2014). Minority representation in special education: Five‐year trends. Journal of Child and Family Studies, 23, 118 – 127. https://doi.org/10.1007/s10826-012-9698-6</bibtext> </blist> <blist> <bibtext> Zieky, M. (1993). Practical questions in the use of DIF statistics in test development. In P. W. Holland, & H. Wainer (Eds.), Differential item functioning (pp. 337 – 347). Lawrence Erlbaum Associates, Inc.</bibtext> </blist> <blist> <bibtext> Zumbo, B. D. (1999). A handbook on the theory and methods of differential item functioning (DIF). Directorate of Human Resources Research, Department of National Defense.</bibtext> </blist> </ref> <aug> <p>By Matthew C. Lambert; Jodie Martin; Michael H. Epstein; Douglas Cullinan and Antonis Katsiyannis</p> <p>Reported by Author; Author; Author; Author; Author</p> </aug>
Header DbId: eric
DbLabel: ERIC
An: EJ1284319
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Differential Item Functioning of the Scales for Assessing Emotional Disturbance-3 for White and African American Students
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Lambert%2C+Matthew+C%2E%22">Lambert, Matthew C.</searchLink> (ORCID <externalLink term="http://orcid.org/0000-0002-7387-3780">0000-0002-7387-3780</externalLink>)<br /><searchLink fieldCode="AR" term="%22Martin%2C+Jodie%22">Martin, Jodie</searchLink><br /><searchLink fieldCode="AR" term="%22Epstein%2C+Michael+H%2E%22">Epstein, Michael H.</searchLink><br /><searchLink fieldCode="AR" term="%22Cullinan%2C+Douglas%22">Cullinan, Douglas</searchLink><br /><searchLink fieldCode="AR" term="%22Katsiyannis%2C+Antonis%22">Katsiyannis, Antonis</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Psychology+in+the+Schools%22"><i>Psychology in the Schools</i></searchLink>. Mar 2021 58(3):553-568.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Wiley. Available from: John Wiley & Sons, Inc. 111 River Street, Hoboken, NJ 07030. Tel: 800-835-6770; e-mail: cs-journals@wiley.com; Web site: https://www.wiley.com/en-us
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 16
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2021
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Elementary+Education%22">Elementary Education</searchLink><br /><searchLink fieldCode="EL" term="%22Secondary+Education%22">Secondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Disability+Identification%22">Disability Identification</searchLink><br /><searchLink fieldCode="DE" term="%22Rating+Scales%22">Rating Scales</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Items%22">Test Items</searchLink><br /><searchLink fieldCode="DE" term="%22Item+Response+Theory%22">Item Response Theory</searchLink><br /><searchLink fieldCode="DE" term="%22Psychometrics%22">Psychometrics</searchLink><br /><searchLink fieldCode="DE" term="%22Emotional+Disturbances%22">Emotional Disturbances</searchLink><br /><searchLink fieldCode="DE" term="%22Behavior+Disorders%22">Behavior Disorders</searchLink><br /><searchLink fieldCode="DE" term="%22Special+Education%22">Special Education</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Bias%22">Test Bias</searchLink><br /><searchLink fieldCode="DE" term="%22Racial+Differences%22">Racial Differences</searchLink><br /><searchLink fieldCode="DE" term="%22African+American+Students%22">African American Students</searchLink><br /><searchLink fieldCode="DE" term="%22White+Students%22">White Students</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+School+Students%22">Elementary School Students</searchLink><br /><searchLink fieldCode="DE" term="%22Secondary+School+Students%22">Secondary School Students</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1002/pits.22463
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 0033-3085
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The present study investigated the psychometric properties of the "Scales for Assessing Emotional Disturbance -- Third Edition: Rating Scale" (SAED-3 RS), which is designed for use in identifying students with emotional disturbance for special education services. The purposes of this study were to evaluate (a) the measurement invariance of SAED-3 RS scores between White and African American students and (b) the impact of differential item functioning (DIF) on test scores from the SAED-3 RS. The sample consisted of 855 K-12 students from throughout the United States. The findings suggested that SAED-3 RS items exhibited small to negligible levels of DIF and that DIF did not significantly impact scores. The results supported the SAED-3 RS, a teacher-completed rating scale, as relatively consistent in measuring the emotional and behavioral status of school-age students from different racial backgrounds. Researchers and practitioners can have confidence that scores from the SAED-3 RS are not substantially affected by DIF when assessing the emotional and behavioral functioning of African American and White school-age students. Research limitations, future research, and implications for school professionals are discussed.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2021
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1284319
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1284319
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1002/pits.22463
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 16
        StartPage: 553
    Subjects:
      – SubjectFull: Disability Identification
        Type: general
      – SubjectFull: Rating Scales
        Type: general
      – SubjectFull: Test Items
        Type: general
      – SubjectFull: Item Response Theory
        Type: general
      – SubjectFull: Psychometrics
        Type: general
      – SubjectFull: Emotional Disturbances
        Type: general
      – SubjectFull: Behavior Disorders
        Type: general
      – SubjectFull: Special Education
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Test Bias
        Type: general
      – SubjectFull: Racial Differences
        Type: general
      – SubjectFull: African American Students
        Type: general
      – SubjectFull: White Students
        Type: general
      – SubjectFull: Elementary School Students
        Type: general
      – SubjectFull: Secondary School Students
        Type: general
    Titles:
      – TitleFull: Differential Item Functioning of the Scales for Assessing Emotional Disturbance-3 for White and African American Students
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Lambert, Matthew C.
      – PersonEntity:
          Name:
            NameFull: Martin, Jodie
      – PersonEntity:
          Name:
            NameFull: Epstein, Michael H.
      – PersonEntity:
          Name:
            NameFull: Cullinan, Douglas
      – PersonEntity:
          Name:
            NameFull: Katsiyannis, Antonis
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 03
              Type: published
              Y: 2021
          Identifiers:
            – Type: issn-print
              Value: 0033-3085
          Numbering:
            – Type: volume
              Value: 58
            – Type: issue
              Value: 3
          Titles:
            – TitleFull: Psychology in the Schools
              Type: main
ResultId 1