CorreGram: Using Corpus Data to Develop Student-Adaptable Automated Corrective Feedback for the L2 Spanish Language Classroom

Saved in:
Bibliographic Details
Title: CorreGram: Using Corpus Data to Develop Student-Adaptable Automated Corrective Feedback for the L2 Spanish Language Classroom
Language: English
Authors: Samuel S. Davidson
Source: ProQuest LLC. 2024Ph.D. Dissertation, University of California, Davis.
Availability: ProQuest LLC. 789 East Eisenhower Parkway, P.O. Box 1346, Ann Arbor, MI 48106. Tel: 800-521-0600; Web site: http://www.proquest.com/en-US/products/dissertations/individuals.shtml
Peer Reviewed: N
Page Count: 162
Publication Date: 2024
Document Type: Dissertations/Theses - Doctoral Dissertations
Education Level: Higher Education
Postsecondary Education
Descriptors: Computer Assisted Testing, Automation, Student Evaluation, Feedback (Response), Second Language Learning, Spanish, Error Correction, College Students, Computer Software, Artificial Intelligence, Grammar
Geographic Terms: California
ISBN: 979-83-8448-361-8
Abstract: Automated corrective feedback (ACF), in which a computer system helps language learners identify and correct errors in their writing or speech, is considered an important tool for language instruction by many researchers. Such systems allow learners to correct their own mistakes, thereby reducing teacher workload and potentially preventing issues related to grammatical error fossilization. Research in this area has led to the development and widespread adoption of tools such as Grammarly for English learners. However, research in grammatical error correction (GEC) and other forms of ACF in languages other than English has been much more limited. This dearth of research is in part due to the large demand for English instruction, but is also driven by the limited training data available for non-English languages. However, a new corpus of learner Spanish collected at UC Davis, COWS-L2H, provided me with an opportunity to explore development of ACF for students studying Spanish. In my dissertation work, I explore the error patterns present in writing by students of Spanish in COWS-L2H, and use this information to inform a novel data augmentation technique to generate synthetic data for training language models capable of correcting learner errors in Spanish text. I then use this synthetic data, along with learner data from COWS-L2H, to train an AI-based GEC model for Spanish learners that is adaptable to learner L1 and proficiency level. Finally, I explore how this automatically corrected writing can be used to present feedback to learners in a pedagogically motivated way. To that end, I combine the GEC model trained using data from COWS-L2H with hand-written templates and feedback produced by generative LLMs to craft appropriate feedback for learners using the system. The end goal is a grammar-checker that is able to not only explain why something a student wrote is potentially incorrect, but is also able to guide the student to make the correction themselves. I demonstrate this novel system, CorreGram, and further discuss details of its implementation and proposals for how the system may be effectively utilized in the language classroom. [The dissertation citations contained here are published with the permission of ProQuest LLC. Further reproduction is prohibited without permission. Copies of dissertations may be obtained by Telephone (800) 1-800-521-0600. Web page: http://www.proquest.com/en-US/products/dissertations/individuals.shtml.]
Abstractor: As Provided
Entry Date: 2024
Access URL: https://gateway.proquest.com/openurl?url_ver=Z39.88-2004&rft_val_fmt=info:ofi/fmt:kev:mtx:dissertation&res_dat=xri:pqm&rft_dat=xri:pqdiss:31561972
Accession Number: ED663221
Database: ERIC
FullText Text:
  Availability: 0
Header DbId: eric
DbLabel: ERIC
An: ED663221
AccessLevel: 3
PubType: Dissertation/ Thesis
PubTypeId: dissertation
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: CorreGram: Using Corpus Data to Develop Student-Adaptable Automated Corrective Feedback for the L2 Spanish Language Classroom
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Samuel+S%2E+Davidson%22">Samuel S. Davidson</searchLink>
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22ProQuest+LLC%22"><i>ProQuest LLC</i></searchLink>. 2024Ph.D. Dissertation, University of California, Davis.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: ProQuest LLC. 789 East Eisenhower Parkway, P.O. Box 1346, Ann Arbor, MI 48106. Tel: 800-521-0600; Web site: http://www.proquest.com/en-US/products/dissertations/individuals.shtml
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: N
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 162
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2024
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Dissertations/Theses - Doctoral Dissertations
– Name: Audience
  Label: Education Level
  Group: Audnce
  Data: <searchLink fieldCode="EL" term="%22Higher+Education%22">Higher Education</searchLink><br /><searchLink fieldCode="EL" term="%22Postsecondary+Education%22">Postsecondary Education</searchLink>
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Computer+Assisted+Testing%22">Computer Assisted Testing</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Student+Evaluation%22">Student Evaluation</searchLink><br /><searchLink fieldCode="DE" term="%22Feedback+%28Response%29%22">Feedback (Response)</searchLink><br /><searchLink fieldCode="DE" term="%22Second+Language+Learning%22">Second Language Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Spanish%22">Spanish</searchLink><br /><searchLink fieldCode="DE" term="%22Error+Correction%22">Error Correction</searchLink><br /><searchLink fieldCode="DE" term="%22College+Students%22">College Students</searchLink><br /><searchLink fieldCode="DE" term="%22Computer+Software%22">Computer Software</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Grammar%22">Grammar</searchLink>
– Name: Subject
  Label: Geographic Terms
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22California%22">California</searchLink>
– Name: ISBN
  Label: ISBN
  Group: ISBN
  Data: 979-83-8448-361-8
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: Automated corrective feedback (ACF), in which a computer system helps language learners identify and correct errors in their writing or speech, is considered an important tool for language instruction by many researchers. Such systems allow learners to correct their own mistakes, thereby reducing teacher workload and potentially preventing issues related to grammatical error fossilization. Research in this area has led to the development and widespread adoption of tools such as Grammarly for English learners. However, research in grammatical error correction (GEC) and other forms of ACF in languages other than English has been much more limited. This dearth of research is in part due to the large demand for English instruction, but is also driven by the limited training data available for non-English languages. However, a new corpus of learner Spanish collected at UC Davis, COWS-L2H, provided me with an opportunity to explore development of ACF for students studying Spanish. In my dissertation work, I explore the error patterns present in writing by students of Spanish in COWS-L2H, and use this information to inform a novel data augmentation technique to generate synthetic data for training language models capable of correcting learner errors in Spanish text. I then use this synthetic data, along with learner data from COWS-L2H, to train an AI-based GEC model for Spanish learners that is adaptable to learner L1 and proficiency level. Finally, I explore how this automatically corrected writing can be used to present feedback to learners in a pedagogically motivated way. To that end, I combine the GEC model trained using data from COWS-L2H with hand-written templates and feedback produced by generative LLMs to craft appropriate feedback for learners using the system. The end goal is a grammar-checker that is able to not only explain why something a student wrote is potentially incorrect, but is also able to guide the student to make the correction themselves. I demonstrate this novel system, CorreGram, and further discuss details of its implementation and proposals for how the system may be effectively utilized in the language classroom. [The dissertation citations contained here are published with the permission of ProQuest LLC. Further reproduction is prohibited without permission. Copies of dissertations may be obtained by Telephone (800) 1-800-521-0600. Web page: http://www.proquest.com/en-US/products/dissertations/individuals.shtml.]
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2024
– Name: URL
  Label: Access URL
  Group: URL
  Data: <link linkTarget="URL" linkTerm="https://gateway.proquest.com/openurl?url_ver=Z39.88-2004&rft_val_fmt=info:ofi/fmt:kev:mtx:dissertation&res_dat=xri:pqm&rft_dat=xri:pqdiss:31561972" linkWindow="_blank">https://gateway.proquest.com/openurl?url_ver=Z39.88-2004&rft_val_fmt=info:ofi/fmt:kev:mtx:dissertation&res_dat=xri:pqm&rft_dat=xri:pqdiss:31561972</link>
– Name: AN
  Label: Accession Number
  Group: ID
  Data: ED663221
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=ED663221
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 162
    Subjects:
      – SubjectFull: Computer Assisted Testing
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Student Evaluation
        Type: general
      – SubjectFull: Feedback (Response)
        Type: general
      – SubjectFull: Second Language Learning
        Type: general
      – SubjectFull: Spanish
        Type: general
      – SubjectFull: Error Correction
        Type: general
      – SubjectFull: College Students
        Type: general
      – SubjectFull: Computer Software
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Grammar
        Type: general
      – SubjectFull: California
        Type: general
    Titles:
      – TitleFull: CorreGram: Using Corpus Data to Develop Student-Adaptable Automated Corrective Feedback for the L2 Spanish Language Classroom
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Samuel S. Davidson
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2024
          Identifiers:
            – Type: isbn-print
              Value: 979-83-8448-361-8
          Titles:
            – TitleFull: ProQuest LLC
              Type: main
ResultId 1