Qualitative Coding with GPT-4: Where It Works Better

Saved in:
Bibliographic Details
Title: Qualitative Coding with GPT-4: Where It Works Better
Language: English
Authors: Xiner Liu (ORCID 0009-0004-3796-2251), Andres Felipe Zambrano (ORCID 0000-0003-0692-1209), Ryan S. Baker (ORCID 0000-0002-3051-3232), Amanda Barany (ORCID 0000-0003-2239-2271), Jaclyn Ocumpaugh (ORCID 0000-0002-9667-8523), Jiayi Zhang (ORCID 0000-0002-7334-4256), Maciej Pankiewicz (ORCID 0000-0002-6945-0523), Nidhi Nasiar (ORCID 0009-0006-7063-5433), Zhanlan Wei (ORCID 0009-0002-3931-6398)
Source: Journal of Learning Analytics. 2025 12(1):169-185.
Availability: Society for Learning Analytics Research. 121 Pointe Marsan, Beaumont, AB T4X 0A2, Canada. Tel: +61-429-920-838; e-mail: info@solaresearch.org; Web site: https://learning-analytics.info/index.php/JLA/index
Peer Reviewed: Y
Page Count: 17
Publication Date: 2025
Sponsoring Agency: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
Contract Number: 2301173
Document Type: Journal Articles
Reports - Research
Descriptors: Coding, Artificial Intelligence, Automation, Data Analysis, Educational Research, Engineering, Man Machine Systems, Algebra, Tutoring, Game Based Learning, Troubleshooting, Introductory Courses, Programming, Interrater Reliability, Program Effectiveness, Prompting
ISSN: 1929-7750
Abstract: This study explores the potential of the large language model GPT-4 as an automated tool for qualitative data analysis by educational researchers, exploring which techniques are most successful for different types of constructs. Specifically, we assess three different prompt engineering strategies -- Zero-shot, Few-shot, and Fewshot with contextual information -- as well as the use of embeddings. We do so in the context of qualitatively coding three distinct educational datasets: Algebra I semi-personalized tutoring session transcripts, student observations in a game-based learning environment, and debugging behaviours in an introductory programming course. We evaluated the performance of each approach based on its inter-rater agreement with human coders and explored how different methods vary in effectiveness depending on a construct's degree of clarity, concreteness, objectivity, granularity, and specificity. Our findings suggest that while GPT-4 can code a broad range of constructs, no single method consistently outperforms the others, and the selection of a particular method should be tailored to the specific properties of the construct and context being analyzed. We also found that GPT-4 has the most difficulty with the same constructs than human coders find more difficult to reach inter-rater reliability on.
Abstractor: As Provided
Entry Date: 2025
Accession Number: EJ1465623
Database: ERIC
FullText Text:
  Availability: 0
CustomLinks:
  – Url: https://eric.ed.gov/contentdelivery/servlet/ERICServlet?accno=EJ1465623
    Name: ERIC Full Text
    Category: fullText
    Text: Full Text from ERIC
Header DbId: eric
DbLabel: ERIC
An: EJ1465623
AccessLevel: 3
PubType: Academic Journal
PubTypeId: academicJournal
PreciseRelevancyScore: 0
IllustrationInfo
Items – Name: Title
  Label: Title
  Group: Ti
  Data: Qualitative Coding with GPT-4: Where It Works Better
– Name: Language
  Label: Language
  Group: Lang
  Data: English
– Name: Author
  Label: Authors
  Group: Au
  Data: <searchLink fieldCode="AR" term="%22Xiner+Liu%22">Xiner Liu</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0004-3796-2251">0009-0004-3796-2251</externalLink>)<br /><searchLink fieldCode="AR" term="%22Andres+Felipe+Zambrano%22">Andres Felipe Zambrano</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-0692-1209">0000-0003-0692-1209</externalLink>)<br /><searchLink fieldCode="AR" term="%22Ryan+S%2E+Baker%22">Ryan S. Baker</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-3051-3232">0000-0002-3051-3232</externalLink>)<br /><searchLink fieldCode="AR" term="%22Amanda+Barany%22">Amanda Barany</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0003-2239-2271">0000-0003-2239-2271</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jaclyn+Ocumpaugh%22">Jaclyn Ocumpaugh</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-9667-8523">0000-0002-9667-8523</externalLink>)<br /><searchLink fieldCode="AR" term="%22Jiayi+Zhang%22">Jiayi Zhang</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-7334-4256">0000-0002-7334-4256</externalLink>)<br /><searchLink fieldCode="AR" term="%22Maciej+Pankiewicz%22">Maciej Pankiewicz</searchLink> (ORCID <externalLink term="https://orcid.org/0000-0002-6945-0523">0000-0002-6945-0523</externalLink>)<br /><searchLink fieldCode="AR" term="%22Nidhi+Nasiar%22">Nidhi Nasiar</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0006-7063-5433">0009-0006-7063-5433</externalLink>)<br /><searchLink fieldCode="AR" term="%22Zhanlan+Wei%22">Zhanlan Wei</searchLink> (ORCID <externalLink term="https://orcid.org/0009-0002-3931-6398">0009-0002-3931-6398</externalLink>)
– Name: TitleSource
  Label: Source
  Group: Src
  Data: <searchLink fieldCode="SO" term="%22Journal+of+Learning+Analytics%22"><i>Journal of Learning Analytics</i></searchLink>. 2025 12(1):169-185.
– Name: Avail
  Label: Availability
  Group: Avail
  Data: Society for Learning Analytics Research. 121 Pointe Marsan, Beaumont, AB T4X 0A2, Canada. Tel: +61-429-920-838; e-mail: info@solaresearch.org; Web site: https://learning-analytics.info/index.php/JLA/index
– Name: PeerReviewed
  Label: Peer Reviewed
  Group: SrcInfo
  Data: Y
– Name: Pages
  Label: Page Count
  Group: Src
  Data: 17
– Name: DatePubCY
  Label: Publication Date
  Group: Date
  Data: 2025
– Name: SourceSuprt
  Label: Sponsoring Agency
  Group: SrcSuprt
  Data: National Science Foundation (NSF), Division of Research on Learning in Formal and Informal Settings (DRL)
– Name: NumberContract
  Label: Contract Number
  Group: NumCntrct
  Data: 2301173
– Name: TypeDocument
  Label: Document Type
  Group: TypDoc
  Data: Journal Articles<br />Reports - Research
– Name: Subject
  Label: Descriptors
  Group: Su
  Data: <searchLink fieldCode="DE" term="%22Coding%22">Coding</searchLink><br /><searchLink fieldCode="DE" term="%22Artificial+Intelligence%22">Artificial Intelligence</searchLink><br /><searchLink fieldCode="DE" term="%22Automation%22">Automation</searchLink><br /><searchLink fieldCode="DE" term="%22Data+Analysis%22">Data Analysis</searchLink><br /><searchLink fieldCode="DE" term="%22Educational+Research%22">Educational Research</searchLink><br /><searchLink fieldCode="DE" term="%22Engineering%22">Engineering</searchLink><br /><searchLink fieldCode="DE" term="%22Man+Machine+Systems%22">Man Machine Systems</searchLink><br /><searchLink fieldCode="DE" term="%22Algebra%22">Algebra</searchLink><br /><searchLink fieldCode="DE" term="%22Tutoring%22">Tutoring</searchLink><br /><searchLink fieldCode="DE" term="%22Game+Based+Learning%22">Game Based Learning</searchLink><br /><searchLink fieldCode="DE" term="%22Troubleshooting%22">Troubleshooting</searchLink><br /><searchLink fieldCode="DE" term="%22Introductory+Courses%22">Introductory Courses</searchLink><br /><searchLink fieldCode="DE" term="%22Programming%22">Programming</searchLink><br /><searchLink fieldCode="DE" term="%22Interrater+Reliability%22">Interrater Reliability</searchLink><br /><searchLink fieldCode="DE" term="%22Program+Effectiveness%22">Program Effectiveness</searchLink><br /><searchLink fieldCode="DE" term="%22Prompting%22">Prompting</searchLink>
– Name: ISSN
  Label: ISSN
  Group: ISSN
  Data: 1929-7750
– Name: Abstract
  Label: Abstract
  Group: Ab
  Data: This study explores the potential of the large language model GPT-4 as an automated tool for qualitative data analysis by educational researchers, exploring which techniques are most successful for different types of constructs. Specifically, we assess three different prompt engineering strategies -- Zero-shot, Few-shot, and Fewshot with contextual information -- as well as the use of embeddings. We do so in the context of qualitatively coding three distinct educational datasets: Algebra I semi-personalized tutoring session transcripts, student observations in a game-based learning environment, and debugging behaviours in an introductory programming course. We evaluated the performance of each approach based on its inter-rater agreement with human coders and explored how different methods vary in effectiveness depending on a construct's degree of clarity, concreteness, objectivity, granularity, and specificity. Our findings suggest that while GPT-4 can code a broad range of constructs, no single method consistently outperforms the others, and the selection of a particular method should be tailored to the specific properties of the construct and context being analyzed. We also found that GPT-4 has the most difficulty with the same constructs than human coders find more difficult to reach inter-rater reliability on.
– Name: AbstractInfo
  Label: Abstractor
  Group: Ab
  Data: As Provided
– Name: DateEntry
  Label: Entry Date
  Group: Date
  Data: 2025
– Name: AN
  Label: Accession Number
  Group: ID
  Data: EJ1465623
PLink https://search.ebscohost.com/login.aspx?direct=true&site=eds-live&db=eric&AN=EJ1465623
RecordInfo BibRecord:
  BibEntity:
    Languages:
      – Text: English
    PhysicalDescription:
      Pagination:
        PageCount: 17
        StartPage: 169
    Subjects:
      – SubjectFull: Coding
        Type: general
      – SubjectFull: Artificial Intelligence
        Type: general
      – SubjectFull: Automation
        Type: general
      – SubjectFull: Data Analysis
        Type: general
      – SubjectFull: Educational Research
        Type: general
      – SubjectFull: Engineering
        Type: general
      – SubjectFull: Man Machine Systems
        Type: general
      – SubjectFull: Algebra
        Type: general
      – SubjectFull: Tutoring
        Type: general
      – SubjectFull: Game Based Learning
        Type: general
      – SubjectFull: Troubleshooting
        Type: general
      – SubjectFull: Introductory Courses
        Type: general
      – SubjectFull: Programming
        Type: general
      – SubjectFull: Interrater Reliability
        Type: general
      – SubjectFull: Program Effectiveness
        Type: general
      – SubjectFull: Prompting
        Type: general
    Titles:
      – TitleFull: Qualitative Coding with GPT-4: Where It Works Better
        Type: main
  BibRelationships:
    HasContributorRelationships:
      – PersonEntity:
          Name:
            NameFull: Xiner Liu
      – PersonEntity:
          Name:
            NameFull: Andres Felipe Zambrano
      – PersonEntity:
          Name:
            NameFull: Ryan S. Baker
      – PersonEntity:
          Name:
            NameFull: Amanda Barany
      – PersonEntity:
          Name:
            NameFull: Jaclyn Ocumpaugh
      – PersonEntity:
          Name:
            NameFull: Jiayi Zhang
      – PersonEntity:
          Name:
            NameFull: Maciej Pankiewicz
      – PersonEntity:
          Name:
            NameFull: Nidhi Nasiar
      – PersonEntity:
          Name:
            NameFull: Zhanlan Wei
    IsPartOfRelationships:
      – BibEntity:
          Dates:
            – D: 01
              M: 01
              Type: published
              Y: 2025
          Identifiers:
            – Type: issn-electronic
              Value: 1929-7750
          Numbering:
            – Type: volume
              Value: 12
            – Type: issue
              Value: 1
          Titles:
            – TitleFull: Journal of Learning Analytics
              Type: main
ResultId 1