Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity.
Saved in:
| Title: | Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity. |
|---|---|
| Language: | English |
| Authors: | Kelly, P. Adam |
| Peer Reviewed: | N |
| Page Count: | 43 |
| Publication Date: | 2001 |
| Sponsoring Agency: | Educational Testing Service, Princeton, NJ. Graduate Record Examination Board Program. |
| Document Type: | Reports - Research Speeches/Meeting Papers |
| Descriptors: | Computer Assisted Testing, Elementary Secondary Education, Essays, Higher Education, Scores, Scoring, Test Scoring Machines, Validity, Writing Evaluation |
| Abstract: | The purpose of this research was to establish, within the constraints of the methods presented, whether the computer is capable of scoring essays in much the same way that human experts rate essays. The investigation attempted to establish what was actually going on within the computer and within the mind of the rater and to describe the degree to which these processes equated. Revealing this parallelism depended on the careful assessment of the "intrinsic" aspects of validity as proposed by S. Messick (1995) of computerized essay scores. The focus was "e-rater" (TM), a computer-based essay scoring system developed by the Educational Testing Service. The study used 1,794 existing Graduate Record Examination Writing Assessment essays written and scored during recent test administration. Factor analysis and the advice of expert raters were used to guide the deconstruction of essay scoring models into subscore models corresponding to writing characteristics within the essay. The writing characteristics identified in this process were used as the basis for developing characteristic-specific scoring rubrics to be used by expert raters. Fresh essay samples were scored by expert rates, both holistically and characteristic-by-characteristic. The same essay samples were assigned both holistic scores and character-wise subscores by the computer scoring models. The degree of convergent validity of scores was evidenced by the proportion of agreement and strength of pairwise correlation among scores and subscores. The statistics derived in this study suggest that simpler e-rater models might do just as well at agreeing with the scores of expert rates, although the proportion of total variance in the expert rater scores explained by the e-rater scores might decrease from an already modest level. (Contains 7 tables and 30 references.) (Author/SLD) |
| Entry Date: | 2002 |
| Accession Number: | ED458296 |
| Database: | ERIC |
| FullText | Text: Availability: 0 CustomLinks: – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED458296 Name: ERIC Full Text Category: fullText Text: Full Text from ERIC |
|---|---|
| Header | DbId: eric DbLabel: ERIC An: ED458296 AccessLevel: 3 PubType: Report PubTypeId: report PreciseRelevancyScore: 0 |
| IllustrationInfo | |
| Items | – Name: Title Label: Title Group: Ti Data: Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity. – Name: Language Label: Language Group: Lang Data: English – Name: Author Label: Authors Group: Au Data: <searchLink fieldCode="AR" term="%22Kelly%2C+P%2E+Adam%22">Kelly, P. Adam</searchLink> – Name: PeerReviewed Label: Peer Reviewed Group: SrcInfo Data: N – Name: Pages Label: Page Count Group: Src Data: 43 – Name: DatePubCY Label: Publication Date Group: Date Data: 2001 – Name: SourceSuprt Label: Sponsoring Agency Group: SrcSuprt Data: Educational Testing Service, Princeton, NJ. Graduate Record Examination Board Program. – Name: TypeDocument Label: Document Type Group: TypDoc Data: Reports - Research<br />Speeches/Meeting Papers – Name: Subject Label: Descriptors Group: Su Data: <searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+Secondary+Education%22">Elementary Secondary Education</searchLink><br /><searchLink fieldCode="DE" term="%22Essays%22">Essays</searchLink><br /><searchLink fieldCode="DE" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Scoring+Machines%22">Test Scoring Machines</searchLink><br /><searchLink fieldCode="DE" term="%22Validity%22">Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink> – Name: Abstract Label: Abstract Group: Ab Data: The purpose of this research was to establish, within the constraints of the methods presented, whether the computer is capable of scoring essays in much the same way that human experts rate essays. The investigation attempted to establish what was actually going on within the computer and within the mind of the rater and to describe the degree to which these processes equated. Revealing this parallelism depended on the careful assessment of the "intrinsic" aspects of validity as proposed by S. Messick (1995) of computerized essay scores. The focus was "e-rater" (TM), a computer-based essay scoring system developed by the Educational Testing Service. The study used 1,794 existing Graduate Record Examination Writing Assessment essays written and scored during recent test administration. Factor analysis and the advice of expert raters were used to guide the deconstruction of essay scoring models into subscore models corresponding to writing characteristics within the essay. The writing characteristics identified in this process were used as the basis for developing characteristic-specific scoring rubrics to be used by expert raters. Fresh essay samples were scored by expert rates, both holistically and characteristic-by-characteristic. The same essay samples were assigned both holistic scores and character-wise subscores by the computer scoring models. The degree of convergent validity of scores was evidenced by the proportion of agreement and strength of pairwise correlation among scores and subscores. The statistics derived in this study suggest that simpler e-rater models might do just as well at agreeing with the scores of expert rates, although the proportion of total variance in the expert rater scores explained by the e-rater scores might decrease from an already modest level. (Contains 7 tables and 30 references.) (Author/SLD) – Name: DateEntry Label: Entry Date Group: Date Data: 2002 – Name: AN Label: Accession Number Group: ID Data: ED458296 |
| PLink | https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED458296 |
| RecordInfo | BibRecord: BibEntity: Languages: – Text: English PhysicalDescription: Pagination: PageCount: 43 Subjects: – SubjectFull: Computer Assisted Testing Type: general – SubjectFull: Elementary Secondary Education Type: general – SubjectFull: Essays Type: general – SubjectFull: Higher Education Type: general – SubjectFull: Scores Type: general – SubjectFull: Scoring Type: general – SubjectFull: Test Scoring Machines Type: general – SubjectFull: Validity Type: general – SubjectFull: Writing Evaluation Type: general Titles: – TitleFull: Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity. Type: main BibRelationships: HasContributorRelationships: – PersonEntity: Name: NameFull: Kelly, P. Adam IsPartOfRelationships: – BibEntity: Dates: – D: 01 M: 04 Type: published Y: 2001 |
| ResultId | 1 |