Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity.

Saved in:
Bibliographic Details
Title: Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity.
Language: English
Authors: Kelly, P. Adam
Peer Reviewed: N
Page Count: 43
Publication Date: 2001
Sponsoring Agency: Educational Testing Service, Princeton, NJ. Graduate Record Examination Board Program.
Document Type: Reports - Research
Speeches/Meeting Papers
Descriptors: Computer Assisted Testing, Elementary Secondary Education, Essays, Higher Education, Scores, Scoring, Test Scoring Machines, Validity, Writing Evaluation
Abstract: The purpose of this research was to establish, within the constraints of the methods presented, whether the computer is capable of scoring essays in much the same way that human experts rate essays. The investigation attempted to establish what was actually going on within the computer and within the mind of the rater and to describe the degree to which these processes equated. Revealing this parallelism depended on the careful assessment of the "intrinsic" aspects of validity as proposed by S. Messick (1995) of computerized essay scores. The focus was "e-rater" (TM), a computer-based essay scoring system developed by the Educational Testing Service. The study used 1,794 existing Graduate Record Examination Writing Assessment essays written and scored during recent test administration. Factor analysis and the advice of expert raters were used to guide the deconstruction of essay scoring models into subscore models corresponding to writing characteristics within the essay. The writing characteristics identified in this process were used as the basis for developing characteristic-specific scoring rubrics to be used by expert raters. Fresh essay samples were scored by expert rates, both holistically and characteristic-by-characteristic. The same essay samples were assigned both holistic scores and character-wise subscores by the computer scoring models. The degree of convergent validity of scores was evidenced by the proportion of agreement and strength of pairwise correlation among scores and subscores. The statistics derived in this study suggest that simpler e-rater models might do just as well at agreeing with the scores of expert rates, although the proportion of total variance in the expert rater scores explained by the e-rater scores might decrease from an already modest level. (Contains 7 tables and 30 references.) (Author/SLD)
Entry Date: 2002
Accession Number: ED458296
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED458296
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: ED458296
AccessLevel: 3
PubType: Report
PubTypeId: report
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity.
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Kelly%2C+P%2E+Adam%22">Kelly, P. Adam</searchLink>
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: N
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 43
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2001
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Educational Testing Service, Princeton, NJ. Graduate Record Examination Board Program.
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Reports - Research<br />Speeches/Meeting Papers
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Elementary+Secondary+Education%22">Elementary Secondary Education</searchLink><br /><searchLink fieldCode="DE" term="%22Essays%22">Essays</searchLink><br /><searchLink fieldCode="DE" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="DE" term="%22Scores%22">Scores</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Test+Scoring+Machines%22">Test Scoring Machines</searchLink><br /><searchLink fieldCode="DE" term="%22Validity%22">Validity</searchLink><br /><searchLink fieldCode="DE" term="%22Writing+Evaluation%22">Writing Evaluation</searchLink>
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: The purpose of this research was to establish, within the constraints of the methods presented, whether the computer is capable of scoring essays in much the same way that human experts rate essays. The investigation attempted to establish what was actually going on within the computer and within the mind of the rater and to describe the degree to which these processes equated. Revealing this parallelism depended on the careful assessment of the "intrinsic" aspects of validity as proposed by S. Messick (1995) of computerized essay scores. The focus was "e-rater" (TM), a computer-based essay scoring system developed by the Educational Testing Service. The study used 1,794 existing Graduate Record Examination Writing Assessment essays written and scored during recent test administration. Factor analysis and the advice of expert raters were used to guide the deconstruction of essay scoring models into subscore models corresponding to writing characteristics within the essay. The writing characteristics identified in this process were used as the basis for developing characteristic-specific scoring rubrics to be used by expert raters. Fresh essay samples were scored by expert rates, both holistically and characteristic-by-characteristic. The same essay samples were assigned both holistic scores and character-wise subscores by the computer scoring models. The degree of convergent validity of scores was evidenced by the proportion of agreement and strength of pairwise correlation among scores and subscores. The statistics derived in this study suggest that simpler e-rater models might do just as well at agreeing with the scores of expert rates, although the proportion of total variance in the expert rater scores explained by the e-rater scores might decrease from an already modest level. (Contains 7 tables and 30 references.) (Author/SLD)
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2002
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED458296
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED458296
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 43
    Subjects:
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Elementary Secondary Education
        Type: general
      – SubjectFull: Essays
        Type: general
      – SubjectFull: Higher Education
        Type: general
      – SubjectFull: Scores
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Test Scoring Machines
        Type: general
      – SubjectFull: Validity
        Type: general
      – SubjectFull: Writing Evaluation
        Type: general
    Titles:
      – TitleFull: Computerized Scoring of Essays for Analytical Writing Assessments: Evaluating Score Validity.
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Kelly, P. Adam
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 04
              Type: published
              Y: 2001
ResultId 1