Multitask Summary Scoring with Longformers

Saved in:
Bibliographic Details
Title: Multitask Summary Scoring with Longformers
Language: English
Authors: Botarleanu, Robert-Mihai, Dascalu, Mihai, Allen, Laura K., Crossley, Scott Andrew, McNamara, Danielle S.
Source: Grantee Submission. 2022Paper presented at International Conference on Artificial Intelligence in Education (AIED) (2022).
Peer Reviewed: Y
Page Count: 7
Publication Date: 2022
Sponsoring Agency: Institute of Education Sciences (ED)
Office of Naval Research (ONR) (DOD)
Contract Number: R305A180144
R305A180261
N000141712300
N000142012623
N000141912424
N000142012627
Document Type: Reports - Research
Speeches/Meeting Papers
Descriptors: Automation, Scoring, Documentation, Likert Scales, Artificial Intelligence, Task Analysis, Natural Language Processing, Learning Processes, Regression (Statistics), Error Patterns, Models, Prediction
DOI: 10.1007/978-3-031-11644-5_79
Abstract: Automated scoring of student language is a complex task that requires systems to emulate complex and multi-faceted human evaluation criteria. Summary scoring brings an additional layer of complexity to automated scoring because it involves two texts of differing lengths that must be compared. In this study, we present our approach to automate summary scoring by evaluating a corpus of approximately 5,000 summaries based on 103 source texts, each summary being scored on a 4-point Likert scale for seven different evaluation criteria. We train and evaluate a series of Machine Learning models that use a combination of independent textual complexity indices from the ReaderBench framework and Deep Learning models based on the Transformer architecture in a multitask setup to predict concurrently all criteria. Our models achieve significantly lower errors than previous work using a similar dataset, with MAE ranging from 0.10-0.16 and corresponding R[superscript 2] values of up to 0.64. Our findings indicate that Longformer-based models are adequate for contextualizing longer text sequences and effectively scoring summaries according to a variety of human-defined evaluation criteria using a single Neural Network. [This paper was published in: "AIED 2022, LNCS 13355," edited by M. M. Rodrigo et al., Springer Nature Switzerland, 2022, pp. 756-761.]
Abstractor: As Provided
IES Funded: Yes
Entry Date: 2023
Accession Number: ED629735
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=ED629735
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: ED629735
AccessLevel: 3
PubType: Report
PubTypeId: report
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Multitask Summary Scoring with Longformers
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Botarleanu%2C+Robert-Mihai%22">Botarleanu, Robert-Mihai</searchLink><br /><searchLink fieldCode="AR" term="%22Dascalu%2C+Mihai%22">Dascalu, Mihai</searchLink><br /><searchLink fieldCode="AR" term="%22Allen%2C+Laura+K%2E%22">Allen, Laura K.</searchLink><br /><searchLink fieldCode="AR" term="%22Crossley%2C+Scott+Andrew%22">Crossley, Scott Andrew</searchLink><br /><searchLink fieldCode="AR" term="%22McNamara%2C+Danielle+S%2E%22">McNamara, Danielle S.</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Grantee+Submission%22"><i>Grantee Submission</i></searchLink>. 2022Paper presented at International Conference on Artificial Intelligence in Education (AIED) (2022).
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 7
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2022
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: Institute of Education Sciences (ED)<br />Office of Naval Research (ONR) (DOD)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: R305A180144<br />R305A180261<br />N000141712300<br />N000142012623<br />N000141912424<br />N000142012627
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Reports - Research<br />Speeches/Meeting Papers
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Scoring%22">Scoring</searchLink><br /><searchLink fieldCode="DE" term="%22Documentation%22">Documentation</searchLink><br /><searchLink fieldCode="DE" term="%22Likert+Scales%22">Likert Scales</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Task+Analysis%22">Task Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Natural+Language+Processing%22">Natural Language Processing</searchLink><br /><searchLink fieldCode="DE" term="%22Learning+Processes%22">Learning Processes</searchLink><br /><searchLink fieldCode="DE" term="%22Regression+%28Statistics%29%22">Regression (Statistics)</searchLink><br /><searchLink fieldCode="DE" term="%22Error+Patterns%22">Error Patterns</searchLink><br /><searchLink fieldCode="DE" term="%22Models%22">Models</searchLink><br /><searchLink fieldCode="DE" term="%22Prediction%22">Prediction</searchLink>
– Name: DOI
  Label: DOI
  Group: ID
  Data: 10.1007/978-3-031-11644-5_79
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Automated scoring of student language is a complex task that requires systems to emulate complex and multi-faceted human evaluation criteria. Summary scoring brings an additional layer of complexity to automated scoring because it involves two texts of differing lengths that must be compared. In this study, we present our approach to automate summary scoring by evaluating a corpus of approximately 5,000 summaries based on 103 source texts, each summary being scored on a 4-point Likert scale for seven different evaluation criteria. We train and evaluate a series of Machine Learning models that use a combination of independent textual complexity indices from the ReaderBench framework and Deep Learning models based on the Transformer architecture in a multitask setup to predict concurrently all criteria. Our models achieve significantly lower errors than previous work using a similar dataset, with MAE ranging from 0.10-0.16 and corresponding R[superscript 2] values of up to 0.64. Our findings indicate that Longformer-based models are adequate for contextualizing longer text sequences and effectively scoring summaries according to a variety of human-defined evaluation criteria using a single Neural Network. [This paper was published in: "AIED 2022, LNCS 13355," edited by M. M. Rodrigo et al., Springer Nature Switzerland, 2022, pp. 756-761.]
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: CodeSource
  Label: IES Funded
  Group: SrcInfo
  Data: Yes
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2023
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED629735
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED629735
RecordInfo BibRecord:
  BibEntity:
    Identifiers:
      – Type: doi
        Value: 10.1007/978-3-031-11644-5_79
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 7
    Subjects:
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Scoring
        Type: general
      – SubjectFull: Documentation
        Type: general
      – SubjectFull: Likert Scales
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Task Analysis
        Type: general
      – SubjectFull: Natural Language Processing
        Type: general
      – SubjectFull: Learning Processes
        Type: general
      – SubjectFull: Regression (Statistics)
        Type: general
      – SubjectFull: Error Patterns
        Type: general
      – SubjectFull: Models
        Type: general
      – SubjectFull: Prediction
        Type: general
    Titles:
      – TitleFull: Multitask Summary Scoring with Longformers
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Botarleanu, Robert-Mihai
      – PersonEntity:
          Name:
            NameFull: Dascalu, Mihai
      – PersonEntity:
          Name:
            NameFull: Allen, Laura K.
      – PersonEntity:
          Name:
            NameFull: Crossley, Scott Andrew
      – PersonEntity:
          Name:
            NameFull: McNamara, Danielle S.
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2022
          Titles:
            – TitleFull: Grantee Submission
              Type: main
ResultId 1